Cellar
cellar
CEL — agent-agnostic infrastructure for computer use.
CEL (Context Execution Layer) is an open-source platform that fuses accessibility trees, CDP, vision, network, and app-specific adapters into one structured device understanding — and exposes stable execution primitives over MCP, CLI, SDK, and N-API. The planner is pluggable: use LangGraph, Mastra, Claude Code, Cursor, Codex, GPT, Gemini, n8n, a raw MCP client, or CEL's built-in cel_think fallback. CEL owns the device; you bring the agent.
Status: Active development. Core runtime fully functional on macOS. MCP server with 4 composable tools. Linux support available. Windows planned.
Three-layer architecture
+------------------------------------------------------------+
| Agents LangGraph | Mastra | Claude Code | Cursor |
| Codex | GPT | Gemini | n8n | MCP clients |
+------------------------------------------------------------+
| CEL / crates context fusion, stream normalization, |
| canonical execution, adapter dispatch, |
| stable MCP / CLI / SDK / N-API surfaces |
+------------------------------------------------------------+
| Adapters browser | Numbers | Excel | Figma | Slack |
| Cursor | Docker Desktop | ... |
+------------------------------------------------------------+Adapters — where app-specific structured truth lives. Third-party extensible.
CEL (this repo) — the durable core: fused context, execution, adapter routing, tool surfaces.
Agents — planners/orchestrators. Every framework is a first-class client; none of them defines the platform.
See docs/what-cel-is.md for the full platform boundary and docs/adapters-cel-agents.md for the north-star design doc.
Related MCP server: native-devtools-mcp
What CEL owns vs. what's pluggable
CEL owns (durable) | Pluggable (agent's choice) |
Fused context: AX, CDP, vision, network, audio, adapters | Which agent framework plans / orchestrates |
Freshness, anomaly, and state tracking (Cortex) | Retry / branching / checkpoint policy |
Canonical | Which LLM(s) back each role |
Adapter lifecycle, dispatch, and the | Human-approval / done-policy |
Stable MCP, CLI ( | App-specific intelligence (lives in adapters) |
Supported agents
First-class integrations. See docs/agents/README.md for the full matrix.
Agent | Cookbook | Transport |
Claude Code | MCP | |
Cursor | MCP | |
LangGraph | MCP / SDK | |
Mastra | MCP / SDK | |
Codex | MCP | |
n8n | MCP / HTTP | |
Raw MCP client | MCP | |
Built-in | In-process |
Adapters
First-party and community adapters. Full catalog in docs/adapter-catalog.md; build your own with docs/adapter-sdk.md.
Adapter | Status | Notes |
Browser (CDP + DOM fusion) | Stable | Primary runtime today |
Numbers | In progress | Spreadsheet truth via app model, not AX guesswork |
Excel | Planned | COM bridge; roadmap in docs/adapter-roadmap.md |
Slack | Planned | Workspace-aware messaging/context |
Figma | Planned | Design-file structured operations |
Cursor (IDE adapter) | Planned | IDE-specific code/editor operations |
Docker Desktop | Planned | Container lifecycle + logs |
Legend: Stable = shipping; In progress = active dev; Planned = on the roadmap.
Hybrid Runtime: What It Handles That Screenshots Can't
Scenario | Screenshot Agents | CEL |
Browser → Desktop handoff | Lose track when focus leaves the browser | Cortex detects context shift via a11y, continues in native app |
Stale state (dynamic content changes between read and act) | Act on where the button was | Freshness model detects staleness, re-reads before acting |
Ambiguous targets (8 identical "Delete" buttons) | ~12.5% chance of clicking the right one | a11y tree resolves by label, role, and structural context |
Unintended side effects (unexpected modal/popup) | Get stuck or blindly click through | Cortex catches the side effect, records it, agent recovers |
Impossible actions (auth-blocked, disabled) | Loop forever or timeout | Escalation ceiling: structured → semantic → vision → terminal stop |
Run these scenarios yourself: ./scripts/demo.sh — see DEMO.md for the full walkthrough.
What Makes CEL Different
Structure-first perception — reads what's actually on screen through OS-level APIs, not what pixels look like. Vision is the fallback, not the foundation.
Hybrid runtime with strategy router — per-action routing: structured → semantic → vision → refresh → terminal failure. Escalation ceiling prevents infinite loops.
Continuous awareness — Cortex tracks what changed, not just what's there now. Freshness model (fresh / soft-stale / hard-stale) prevents acting on stale state.
Works everywhere — browsers, desktop apps, terminals, legacy software. One runtime, not separate products for browser vs. desktop.
Model-agnostic — works with any LLM. Sends structured text, not screenshots. A local 7B model works for most workflows.
Agent-agnostic — LangGraph, Mastra, Claude Code, Codex, GPT, Gemini, Cursor, n8n, or future runtimes should all be able to use CEL.
200x cheaper — structured context extraction eliminates expensive vision model inference on every step.
The Problem
Agentic computer use — AI that operates software through the UI — is the defining trend in AI. But it does not work reliably yet.
In browsers, agents have the DOM but still produce unstable results because they depend entirely on LLM interpretation. Outside the browser — on desktop apps, terminals, native software — it's far worse. Agents rely on screenshots alone, feeding pixels to vision models and hoping they correctly identify buttons, fields, and values.
Meanwhile, rich structured information already exists on every computer: accessibility trees, native application APIs, network traffic, input events. No tool combines these signals into a standard format that any agent can consume.
MCP solved this problem for tool access. CEL solves it for computer use.
The Solution: CEL
CEL (Context Execution Layer) is both a context extraction and execution layer. It fuses five streams into a single structured JSON output with per-element confidence scoring:
Stream | What it provides |
Vision | Screen capture + vision model analysis |
Accessibility tree | Platform APIs (AT-SPI2, AXUIElement, UIA) |
Native API bridge | App-specific adapters (Excel COM, SAP Scripting, etc.) |
Input layer | Mouse/keyboard — injected, intercepted, logged, replayable |
Network layer | Traffic monitoring for state change detection |
The agent calls getContext() and gets structured JSON with confidence scores — regardless of which source provided the data. Then it executes actions through CEL using the same multi-source approach. Workflows become replayable sequences of structured contexts and actions, not brittle screenshot-to-click chains.
Works on any interface: browser, terminal, Finder, Excel, SAP, Bloomberg — any OS, any application.
Unlike screenshot-only approaches that route every action through expensive LLM inference, CEL uses structured sources (accessibility tree, native APIs) first and escalates to vision models only when needed. Faster, cheaper, more predictable — and capable of running fully offline.
Use CEL with Claude Code (MCP)
CEL ships as an MCP server with 4 tools. Connect it to Claude Code, Cursor, or any MCP client:
# Build everything
pnpm install && pnpm -r build
# Build native module (macOS)
cargo build --release -p cel-napi
cp target/release/libcel_napi.dylib cel/cel-napi/cel-napi.darwin-arm64.node
codesign -fs - cel/cel-napi/cel-napi.darwin-arm64.nodePick an LLM provider — the fastest path is the interactive setup (writes ~/.cellar/config.toml):
cellar initOptions: paste a Gemini / Anthropic / OpenAI API key, or install Gemma 4 E4B locally via Ollama for fully-private runs. If you'd rather configure via .mcp.json directly (see below), skip init.
Configuration hierarchy
Environment variables override ~/.cellar/config.toml, which overrides compiled defaults.
# ~/.cellar/config.toml
[llm]
provider = "gemini" # openai | anthropic | gemini | ollama | compatible
api_key = "your-key"
model = "gemini-2.0-flash"
[audio] # optional — enables audio transcription in the Cortex
whisper_endpoint = "https://api.openai.com/v1/audio/transcriptions"
whisper_api_key = "sk-..."
whisper_model = "whisper-1"
# whisper_language = "en" # ISO 639-1 hint — improves accuracyFull variable list: docs/api-reference.md.
Add to .mcp.json in your project root:
{
"mcpServers": {
"cellar": {
"command": "node",
"args": ["/path/to/cellar/mcp-server/dist/index.js"],
"env": {
"CEL_LLM_PROVIDER": "gemini",
"CEL_LLM_API_KEY": "your-api-key",
"CEL_LLM_MODEL": "gemini-2.0-flash"
}
}
}
}Restart Claude Code and you'll have four tools:
Tool | What it does | Modes/Actions |
cel_see | Read the screen — structured elements with types, labels, bounds, confidence scores | 14 modes |
cel_act | Click, type, scroll, drag — by coordinates, element ID, or accessibility API | 11 actions + CDP eval |
cel_think | Plan, remember, track runs, autonomous execution (run_goal) | 16 modes |
cel_perceive | Always-on perception engine (Cortex) — continuous screen awareness | 7 modes |
On startup, the Cortex boots automatically (screen model is warm before your first call) and Chrome CDP is auto-detected.
See docs/quickstart.md for the full setup guide and docs/mcp-server.md for the complete tool reference.
Current State
Cellar is in prototype phase on macOS. The bar for exit is defined in docs/PROTOTYPE_EXIT_CRITERIA.md; the curated regression suite that gates it lives in eval/prototype-subset/.
Gated today (macOS local):
Local execution on macOS via AX + CDP + screen capture + input injection
MCP server with 4 composable tools:
cel_see/cel_act/cel_think/cel_perceiveCortex — always-on perception with background event streams
Autonomous execution (
run_goal) over the prototype scenario suite: browser happy-paths, grounding, ambiguity, recovery, browser-to-desktop handoffCLI entry points:
cellar init(setup) andcellar run-goal "<goal>"BYOK providers (OpenAI, Anthropic, Gemini) and local Ollama (Gemma 4 E4B default)
Per-role LLM routing — Planner / Observer / Vision / Validator
Built but outside the prototype exit bar:
Audio capture + Whisper transcription fused into the Cortex world model
Embedded SQLite + FTS5 for memory / semantic search
First-party adapters — Excel, SAP GUI, Bloomberg, MetaTrader
Recorder, live-view, and the wider benchmarks/ suite (50+ tasks + hybrid scenarios)
napi-rs Rust ↔ Node.js bridge
Later phase — explicitly not prototype work (see docs/ROADMAP.md):
Linux accessibility (AT-SPI2) and Windows UI Automation bridges
Remote worker / Docker image / managed VMs (
cellar-worker/exists in-tree as a preview; not wired into prototype gates)Managed cloud, control plane, billing
Production confidence calibration
Portable context maps, community workflow registry
Architecture
cellar/
cel/ ← Cortex + perception layer (Rust, Apache 2.0)
cel-accessibility/ ← accessibility bridge (AXUIElement, AT-SPI2)
cel-context/ ← unified context API + multi-source fusion + references
cel-display/ ← screen capture (xcap)
cel-input/ ← input injection (enigo)
cel-vision/ ← vision model integration (multi-provider)
cel-network/ ← traffic monitoring + idle detection
cel-store/ ← embedded SQLite + FTS5 (memory, knowledge)
cel-llm/ ← LLM provider abstraction
cel-planner/ ← built-in planner / runner code (useful, but not the repo's main value)
cel-napi/ ← Node.js native bindings (napi-rs)
agent/ ← agent integrations and runtime experiments
mcp-server/ ← generic tool surface for external agents
adapters/ ← app-specific adapters (browser, Excel, SAP)
benchmarks/ ← eval harness (50+ tasks + 5 hybrid scenarios)
live-view/ ← real-time debug surface (screen + runtime decisions)
cli/ ← `cellar` CLIGetting Started
Quickstart — Claude Code (recommended)
See docs/quickstart.md for the full step-by-step guide. The short version:
# 1. Build
pnpm install && pnpm -r build
cargo build --release -p cel-napi
cp target/release/libcel_napi.dylib cel/cel-napi/cel-napi.darwin-arm64.node
codesign -fs - cel/cel-napi/cel-napi.darwin-arm64.node
# 2. Configure .mcp.json (see quickstart for full config)
# 3. Grant Accessibility permissions in System Settings
# 4. Restart Claude Code — tools are readyQuickstart — see what the agent sees
No Rust build needed. Just Node.js 20+ and pnpm:
pnpm install && pnpm -r build
npx tsx examples/quickstart.ts https://github.com/loginThis launches a browser, extracts DOM elements as structured ContextElements with confidence scores, and shows the kind of context any external agent runtime would receive.
Prerequisites
Node.js 20+ and pnpm 9+ (TypeScript packages)
Rust 1.75+ (CEL core, accessibility bridge, native bindings)
macOS 13+ with Accessibility permissions
Chrome (optional, for CDP features)
Build
# Build everything
make build
# Or separately
make build-rust # cargo build --workspace
make build-ts # pnpm install && pnpm build
# Run tests
make testCLI
cellar init # Interactive first-run setup (pick LLM provider or install Gemma 4)
cellar setup # Configure AX + CDP permissions on this machine
cellar context # Show unified context with confidence scores
cellar context --json # Output raw JSON
cellar context --watch # Live-update context in terminal
cellar capture # Capture screenshot to file
cellar action click 500 300 # Click at coordinates
cellar action type "Hello" # Type text
cellar action key Enter # Press a key
cellar action combo Ctrl C # Key combination
cellar mcp # Start MCP server (stdio)
cellar mcp install # Print Claude Desktop config
cellar run <workflow> # Execute a saved workflow
cellar train # Enter training modeBenchmarks
Hybrid Runtime Scenarios (CEL advantage)
5 scenarios designed to test where multi-source perception matters. Run them: ./scripts/demo.sh
Scenario | What breaks screenshot agents | CEL metric |
Browser → Desktop handoff | Lose context across app boundary |
|
Stale state (2s shuffle) | Click where button was |
|
Ambiguous targets (8 similar names) | Can't distinguish identical buttons |
|
Side-effect detection (unexpected modal) | Stuck or blindly proceed |
|
Terminal failure (auth-blocked) | Loop forever |
|
General Web Tasks
We also benchmark on 50+ general web tasks against other tools:
Tool | Approach |
Cellar | Multi-source fusion (DOM + a11y + vision + network), confidence scoring, incremental updates |
Anthropic Computer Use | Screenshot-only, pixel-coordinate actions via API |
Browser-Use (OSS) | Hybrid screenshot + DOM (Python) |
Browserbase + Stagehand | Cloud CDP + AI SDK |
Browser-Use Cloud | Managed browser-use + custom model |
Measured on Apple M-series (arm64, 12 cores, 18GB RAM), April 2026. Hybrid suite: 5 tasks testing browser-desktop handoff, stale state recovery, ambiguous targets, side-effect detection, terminal failure. All local tools use Gemini 2.5 Flash. Computer Use locked to Claude Sonnet.
Benchmark results (April 2026 — Hybrid Suite, 5 tasks)
Tool | Avg Time | LLM Calls | Cost/Task | Success |
CEL | 20.8s | 1.4 | $0.0005 | 100% |
Browser-Use OSS | 23.4s | 3.0 | $0.001 | 100% |
Stagehand v3 | 35.6s | 18.2 | $0.005 | 20% |
Computer Use | 36.2s | 6.2 | $0.155 | 100% |
Browser-Use Cloud | 46.5s | 5.6 | $0.003 | 100% |
CEL vs the field:
vs | Speed | Cost | Accuracy |
Computer Use (Anthropic) | 1.7x faster | 310x cheaper | Same (100%) |
Browser-Use Cloud | 2.2x faster | 6x cheaper | Same (100%) |
Stagehand v3 | 1.7x faster | 10x cheaper | 5x better (100% vs 20%) |
Browser-Use OSS | 1.1x faster | 2x cheaper | Same (100%) |
Why CEL wins:
1 LLM call per task — structured context means most tasks extract data in a single pass. Competitors need 3-18 calls.
$0.0005/task — Gemini Flash + context distillation. At 1000 tasks: CEL $0.50 vs Computer Use $155.
Structured context is free — 500+ elements extracted in 100-400ms via Rust-native DOM fusion, no LLM required.
Full Rust execution loop — perceive, plan, execute, verify all in Rust. No FFI in the hot path.
For building reliable automation (not one-off tasks), structured context is the foundation
See benchmarks/README.md for full methodology, per-task breakdown, and how to reproduce.
Roadmap
The forward plan lives in docs/ROADMAP.md. Two related references:
docs/deployment.md — topology: local / remote worker / managed cloud, Docker scope, model backends.
docs/oss-boundary.md — what's OSS vs commercial, mirror strategy.
Contributing
See CONTRIBUTING.md for how to get started, and DEVELOPMENT.md for build instructions and conventions.
We welcome contributions — especially:
Accessibility bridges (macOS AXUIElement, Windows UI Automation)
New application adapters — see docs/building-adapters.md
MCP tool improvements
Test coverage for platform-specific code
Documentation and examples
Platform Support
Platform | Status |
macOS | Primary platform. AXUIElement bridge, Cortex, MCP server — all fully functional. |
Linux | AT-SPI2 accessibility bridge working |
Windows | Planned (UI Automation bridge designed, not yet implemented) |
License
Everything OSS-destined (
cel/,agent/,cli/,mcp-server/,cellar-worker/,live-view/,recorder/,registry/,docs/,benchmarks/,examples/,e2e/,tests/,box/): Apache License 2.0.Community adapters (
adapters/): MIT.Commercial-only (
app/, futurecontrol-plane/,cloud/,billing/): proprietary — not covered by this license.
See docs/oss-boundary.md for the full license map and what stays private.
Available Tools
4 toolscel_actCEL ActA
Execute actions on the screen: mouse clicks, keyboard input, accessibility actions, drag & drop, and direct value setting. Always use cel_see first to understand the screen.
For click/move: provide (x, y) coordinates or a target_ref from cel_see make_reference. For form filling: prefer set_value over type — faster and more reliable. For buttons/checkboxes: prefer ax_action over click — uses native accessibility API.
Coordinate Actions (x,y or target_ref): click, right_click, double_click, mouse_move.
Keyboard: type (text string), key_press (single key: Enter, Tab, Escape, etc.), key_combo (modifier combinations: ['Ctrl','C'], ['Cmd','Shift','S']).
Accessibility API (preferred for reliability): ax_action — native a11y actions on element_id: click, activate, press, increment, decrement, cancel, show_menu, scroll_to_visible, raise, pick, delete. set_value — direct value injection on element_id: text for fields, 'true'/'false' for checkboxes.
Deterministic spreadsheet actions: write_cells (atomic Numbers cell writes with optional readback verification), read_cells (read Numbers cell values from the document model instead of guessing from AX text).
Other: scroll (dx,dy at optional x,y), drag (from_x,from_y to to_x,to_y), cdp_eval (execute JavaScript in browser via CDP — best for cookie banners, iframes, overlays, and elements invisible to the accessibility tree).
Batching: pass array of 1-4 actions for sequential execution (100ms default delay). Re-observe with cel_see after each batch to avoid stale-state cascading failures.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries full burden. It describes each action type, mentions deterministic spreadsheet actions, batching with default delay, and warns about stale-state cascading failures. Side effects (UI mutation) are implied, and no contradictions exist.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is lengthy but well-structured with bullet points and sections. It front-loads purpose and general guidance. Some redundancy exists (e.g., repeating 'prefer'), but overall it is organized and earn its detail for the variety of actions.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema is provided, and the description does not explain what the tool returns. Additionally, the input schema is empty, creating a mismatch with the description that implies parameters. The missing return value and schema inconsistency reduce completeness.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Input schema has zero parameters, so baseline is 4 per instructions. The description adds substantial meaning by detailing all action types and their required coordinates, target_ref, element_id, etc., far beyond the empty schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states that the tool executes actions on the screen including mouse clicks, keyboard input, accessibility actions, drag & drop, and direct value setting. It also distinguishes itself from siblings by advising to use cel_see first, making its purpose distinct and specific.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Extensive guidelines are provided: always use cel_see first, prefer set_value over type for form filling, prefer ax_action over click for buttons/checkboxes, and detailed recommendations for each action type. Batching and re-observing instructions are also given, offering clear when-to-use and when-not-to-use guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
cel_perceiveCEL PerceiveA
Always-on perception engine (Cortex). Maintains a continuously-updated mental model via background event streams with periodic accessibility tree refreshes on significant changes, and vision/screenshots when flagged as needed.
IMPORTANT: Singleton — only one perception session can be active at a time. cel_see 'watch' mode is unavailable during an active session.
Modes:
start: Boot the cortex with a goal. Set enable_suggestions=true (default) for LLM-powered next-action recommendations on each read.
read: Get the mental model snapshot (instant — model is kept warm by background events).
feed: Report an action you took (action, target, expected outcome). Cortex waits for screen to settle, diffs against current model, returns verification.
checkpoint: Summarize completed work and reset action history. Use between phases of multi-step tasks.
configure: Update goal or enable_suggestions mid-session.
status: Cortex health — confidence score, uptime, cycle count, element counts (stable vs volatile), temporal state (loading, errors, focus trail).
stop: Shutdown the cortex and get a summary.
The model includes temporal awareness (loading states, error persistence, focus trail) and element stability classification (stable vs volatile targets).
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are present, so the description carries full burden. It discloses that the tool is always-on, maintains a background mental model, uses event streams, accessibility refreshes, and optional screenshots. It explains each mode's behavior and side effects (e.g., feed waits for screen settle, diffs model). Minor ambiguity about whether feedback modifies state, but overall highly transparent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is somewhat lengthy but well-organized with a clear mode list and important constraints upfront. Every sentence adds information, though some details could be tightened. Front-loading the singleton note is effective.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no parameters, no output schema, and no annotations, the description covers all essential information: purpose, modes, constraints, sibling differentiation, and behavioral model. It is fully adequate for an agent to select and invoke the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has zero parameters (100% documented by schema), so baseline is 4. The description adds value by explaining the modes which act as sub-operations, but no parameter details are needed.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool is an 'always-on perception engine' that maintains a mental model, and explicitly lists all modes (start, read, feed, etc.) with specific verbs and resources. It effectively distinguishes from siblings like cel_see by noting that 'watch' mode is unavailable during active session.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides explicit guidance on when to use each mode, including when not to use certain modes (e.g., 'cel_see watch mode is unavailable during an active session'). It also highlights the singleton constraint, aiding the agent in choosing this tool appropriately.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
cel_seeCEL SeeA
Read and observe the current screen state. Returns structured UI elements, window lists, screenshots, CDP page content, accessibility element details, and screen change events. Always use this BEFORE acting to understand what's on screen.
Screen Context: context (elements with filter/compression — use detail 'compact' to save tokens), screenshot (PNG capture), windows (visible window list), monitors (display list).
Element Inspection: focused (high-fidelity detail for one element_id), element_at (hit-test x,y coordinates), is_settable (check if set_value works), make_reference (resilient ref that survives across snapshots), cursor_position.
Browser (CDP): cdp_status (debug targets & connection state), cdp_page (full page content as text).
Observation Recall: observation (load a persisted context snapshot by observation_id).
Waiting & Watching: wait_for_element (poll for element by type/label, default 10s timeout), wait_for_idle (poll until screen stabilizes — requires 2 consecutive stable polls), watch (event-driven — 18 event types: tree_changed, network_idle, focus_changed, value_changed, window_created, menu_opened, menu_closed, sheet_created, layout_changed, title_changed, app_activated, app_deactivated, window_moved, window_resized, window_minimized, window_restored, selection_changed, row_count_changed). Note: watch is unavailable during an active cel_perceive session.
Limits: CDP enrichment caps at 50 text_blocks, 50 interactive_elements, 3000 char body_text.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Despite no annotations, the description provides rich behavioral details: default timeout for wait_for_element (10s), requirement for wait_for_idle (2 consecutive stable polls), 18 event types for watch, CDP limits (50 text_blocks, etc.), and conflict note about cel_perceive. This goes far beyond simple annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-organized with clear sections (Screen Context, Element Inspection, Browser, Observation Recall, Waiting & Watching, Limits). Each sentence adds value, providing necessary detail without redundancy. Front-loaded with purpose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no output schema, the description lists categories of returned data but does not fully specify output structure. However, it covers key aspects like limits and sub-function behaviors. It feels complete for a read tool, though a more structured output spec would be even better.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There are zero parameters in the input schema, and the description compensates by thoroughly explaining all the tool's sub-functions (Screen Context, Element Inspection, etc.). According to guidelines, 0 params = baseline 4; this description exceeds that with detailed breakdown of capabilities.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description explicitly states 'Read and observe the current screen state' and lists many capabilities. It distinguishes from siblings by saying 'Always use this BEFORE acting', making clear this is the observation tool while cel_act is for actions and cel_perceive for perception.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear usage advice: 'Always use this BEFORE acting'. It also notes a limitation (watch unavailable during cel_perceive session). However, it does not explicitly state when not to use or provide direct comparison with cel_perceive.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
cel_thinkCEL ThinkC
CEL's cognitive layer: delegated autonomy, planning, knowledge, run tracking, and LLM passthrough.
Efficiency rule: if the MCP host already reasons well step-by-step, prefer cel_see + cel_act and keep planning in the host. Use run_goal only when you intentionally want CEL to take over the control loop.
Delegated Autonomous Execution: run_goal — give a natural language goal, CEL runs a full internal see→plan→act loop autonomously. This can be convenient, but it adds an internal planner loop and may be slower or more expensive than host-driven execution. Only goal, max_steps (default 80), and timeout_ms (default 900_000) are tunable — vision, self-healing, decomposition, and notebook are implicit in the canonical loop and no longer per-invocation knobs (see docs/canonical-agent-plan.md).
Planning: plan (LLM-powered step planning with optional history for multi-step context), plan_with_vision (plan with screenshot — use for visual/spatial tasks).
Knowledge Store (persisted to ~/.cellar/cel-store.db): store_knowledge (save facts with source and optional tags), search_knowledge (FTS5 full-text search, default 10 results, scope by workflow).
Working Memory: memory_get, memory_set (per-workflow scratchpad, not persisted across sessions).
Observations: observe (record insight with priority high/medium/low), get_observations (retrieve, default 50).
Run Tracking: run_start, run_finish, run_log_step (per-step with confidence score), run_history, run_steps.
LLM Passthrough: llm_complete (text, 4096 tokens default), llm_complete_with_image (vision, 4096 tokens default).
Maintenance: eviction (TTL cleanup — default 90 days runs, 365 days knowledge).
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description details many capabilities but fails to disclose what happens on invocation without arguments. No annotations are provided to clarify behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is excessively long and poorly structured, lacking front-loading. It lists many sub-functions without clear organization.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the complexity of the described features and the lack of parameters or output schema, the description is incomplete for an agent to know how to effectively use the tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has no parameters, so schema-description coverage is 100%. No parameter semantics are needed, baseline score of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description lists many sub-operations but does not state what the tool does when invoked with no parameters. The purpose is vague and ambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description includes efficiency guidance preferring cel_see+cel_act in some cases, but does not clarify how to invoke any of the listed sub-operations since the tool takes no parameters.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
4 tool updates
v0.1.0- First observed
cel_act - First observed
cel_perceive - First observed
cel_see - First observed
cel_think
TDQS
Scored across 4 tools
Each tool has a distinct role: cel_act for executing actions, cel_perceive for continuous perception, cel_see for reading screen state, and cel_think for cognitive planning. Descriptions clarify boundaries despite some perceptual overlap.
All tool names follow the consistent pattern 'cel_verb' (act, perceive, see, think), using lowercase with underscores throughout. No deviations or mixed conventions.
With only 4 tools, the set is compact but each encapsulates many sub-operations via parameters and modes. The count is slightly low but appropriate for the server's focused domain of screen automation and perception.
The tools cover perception, action, and cognitive planning comprehensively for UI automation. Minor gaps exist (e.g., no explicit system-level operations), but core workflows are well supported and no obvious dead ends.
Maintenance
Related MCP Connectors
Melaya is a remote MCP server. It gives an assistant hands on your own Android phone and browser: it reads the screen through the accessibility tree, then taps, types and navigates inside the apps and sites you allow-list, with no per-app API. It also builds, schedules and runs agent pipelines across 6k+ connected tools. OAuth 2.1, nothing to install.
A comprehensive Model Context Protocol (MCP) server that enables AI assistants to interact with yo…
Remote streamable-HTTP MCP server running on a single Cloudflare Worker. Your assistant gets live Airbnb, Amazon, Booking.com, Google Flights, Maps and Reddit data, social search on X, Instagram and TikTok, the Meta Ad Library, and image/video generation without any keys. Connect your own accounts to let it send WhatsApp or Telegram messages, work an IMAP inbox, manage Meta Ads campaigns and publish to X and LinkedIn. OAuth 2.1 with PKCE; stored credentials are AES-256-GCM encrypted.
Operate Linux, macOS and Windows from your LLM. Every action runs through an auditable allowlist.
Related MCP Servers
- AlicenseNot gradedqualityFmaintenanceAn open-source MCP server for macOS and Windows that provides native desktop control via Accessibility APIs, OCR, and Chrome CDP. It enables AI agents to interact with applications, manage browser sessions, and automate workflows with high-speed native UI actions.91 npm15AGPL 3.0
- AlicenseNot gradedqualityAmaintenanceGives AI agents and MCP clients direct control over native desktop apps, Chrome/Electron browsers, and Android devices with screenshots, OCR, accessibility-based element lookup, input simulation, window management, CDP, and ADB in one local server.133MIT
- AlicenseNot gradedqualityBmaintenanceMCP server that gives AI agents hands and eyes on macOS, enabling them to see and operate native and Electron applications via structured accessibility queries, screenshots, OCR, and application-specific skills.4MIT
- FlicenseAqualityAmaintenanceCross-platform desktop automation MCP server that lets AI agents capture screenshots, run OCR with UI-element classification, control mouse/keyboard, and launch programs on Linux, macOS, and Windows.201-