Skip to main content
Glama

cellar

Glama MCP server

CEL — agent-agnostic infrastructure for computer use.

CEL (Context Execution Layer) is an open-source platform that fuses accessibility trees, CDP, vision, network, and app-specific adapters into one structured device understanding — and exposes stable execution primitives over MCP, CLI, SDK, and N-API. The planner is pluggable: use LangGraph, Mastra, Claude Code, Cursor, Codex, GPT, Gemini, n8n, a raw MCP client, or CEL's built-in cel_think fallback. CEL owns the device; you bring the agent.

Status: Active development. Core runtime fully functional on macOS. MCP server with 4 composable tools. Linux support available. Windows planned.

Three-layer architecture

+------------------------------------------------------------+
|  Agents       LangGraph | Mastra | Claude Code | Cursor    |
|               Codex | GPT | Gemini | n8n | MCP clients      |
+------------------------------------------------------------+
|  CEL / crates context fusion, stream normalization,         |
|               canonical execution, adapter dispatch,        |
|               stable MCP / CLI / SDK / N-API surfaces       |
+------------------------------------------------------------+
|  Adapters     browser | Numbers | Excel | Figma | Slack    |
|               Cursor | Docker Desktop | ...                 |
+------------------------------------------------------------+
  • Adapters — where app-specific structured truth lives. Third-party extensible.

  • CEL (this repo) — the durable core: fused context, execution, adapter routing, tool surfaces.

  • Agents — planners/orchestrators. Every framework is a first-class client; none of them defines the platform.

See docs/what-cel-is.md for the full platform boundary and docs/adapters-cel-agents.md for the north-star design doc.

Related MCP server: native-devtools-mcp

What CEL owns vs. what's pluggable

CEL owns (durable)

Pluggable (agent's choice)

Fused context: AX, CDP, vision, network, audio, adapters

Which agent framework plans / orchestrates

Freshness, anomaly, and state tracking (Cortex)

Retry / branching / checkpoint policy

Canonical cel_see / cel_act / cel_perceive / cel_think tool surface

Which LLM(s) back each role

Adapter lifecycle, dispatch, and the AdapterDriver trait

Human-approval / done-policy

Stable MCP, CLI (cellar), SDK, and N-API bindings

App-specific intelligence (lives in adapters)

Supported agents

First-class integrations. See docs/agents/README.md for the full matrix.

Agent

Cookbook

Transport

Claude Code

docs/agents/claude-code.md

MCP

Cursor

docs/agents/cursor.md

MCP

LangGraph

docs/agents/README.md

MCP / SDK

Mastra

docs/agents/mastra.md

MCP / SDK

Codex

docs/agents/codex.md

MCP

n8n

docs/agents/README.md

MCP / HTTP

Raw MCP client

docs/agents/README.md

MCP

Built-in cel_think (fallback)

docs/mcp-server.md

In-process

Adapters

First-party and community adapters. Full catalog in docs/adapter-catalog.md; build your own with docs/adapter-sdk.md.

Adapter

Status

Notes

Browser (CDP + DOM fusion)

Stable

Primary runtime today

Numbers

In progress

Spreadsheet truth via app model, not AX guesswork

Excel

Planned

COM bridge; roadmap in docs/adapter-roadmap.md

Slack

Planned

Workspace-aware messaging/context

Figma

Planned

Design-file structured operations

Cursor (IDE adapter)

Planned

IDE-specific code/editor operations

Docker Desktop

Planned

Container lifecycle + logs

Legend: Stable = shipping; In progress = active dev; Planned = on the roadmap.

Hybrid Runtime: What It Handles That Screenshots Can't

Scenario

Screenshot Agents

CEL

Browser → Desktop handoff

Lose track when focus leaves the browser

Cortex detects context shift via a11y, continues in native app

Stale state (dynamic content changes between read and act)

Act on where the button was

Freshness model detects staleness, re-reads before acting

Ambiguous targets (8 identical "Delete" buttons)

~12.5% chance of clicking the right one

a11y tree resolves by label, role, and structural context

Unintended side effects (unexpected modal/popup)

Get stuck or blindly click through

Cortex catches the side effect, records it, agent recovers

Impossible actions (auth-blocked, disabled)

Loop forever or timeout

Escalation ceiling: structured → semantic → vision → terminal stop

Run these scenarios yourself: ./scripts/demo.sh — see DEMO.md for the full walkthrough.

What Makes CEL Different

  • Structure-first perception — reads what's actually on screen through OS-level APIs, not what pixels look like. Vision is the fallback, not the foundation.

  • Hybrid runtime with strategy router — per-action routing: structured → semantic → vision → refresh → terminal failure. Escalation ceiling prevents infinite loops.

  • Continuous awareness — Cortex tracks what changed, not just what's there now. Freshness model (fresh / soft-stale / hard-stale) prevents acting on stale state.

  • Works everywhere — browsers, desktop apps, terminals, legacy software. One runtime, not separate products for browser vs. desktop.

  • Model-agnostic — works with any LLM. Sends structured text, not screenshots. A local 7B model works for most workflows.

  • Agent-agnostic — LangGraph, Mastra, Claude Code, Codex, GPT, Gemini, Cursor, n8n, or future runtimes should all be able to use CEL.

  • 200x cheaper — structured context extraction eliminates expensive vision model inference on every step.

The Problem

Agentic computer use — AI that operates software through the UI — is the defining trend in AI. But it does not work reliably yet.

In browsers, agents have the DOM but still produce unstable results because they depend entirely on LLM interpretation. Outside the browser — on desktop apps, terminals, native software — it's far worse. Agents rely on screenshots alone, feeding pixels to vision models and hoping they correctly identify buttons, fields, and values.

Meanwhile, rich structured information already exists on every computer: accessibility trees, native application APIs, network traffic, input events. No tool combines these signals into a standard format that any agent can consume.

MCP solved this problem for tool access. CEL solves it for computer use.

The Solution: CEL

CEL (Context Execution Layer) is both a context extraction and execution layer. It fuses five streams into a single structured JSON output with per-element confidence scoring:

Stream

What it provides

Vision

Screen capture + vision model analysis

Accessibility tree

Platform APIs (AT-SPI2, AXUIElement, UIA)

Native API bridge

App-specific adapters (Excel COM, SAP Scripting, etc.)

Input layer

Mouse/keyboard — injected, intercepted, logged, replayable

Network layer

Traffic monitoring for state change detection

The agent calls getContext() and gets structured JSON with confidence scores — regardless of which source provided the data. Then it executes actions through CEL using the same multi-source approach. Workflows become replayable sequences of structured contexts and actions, not brittle screenshot-to-click chains.

Works on any interface: browser, terminal, Finder, Excel, SAP, Bloomberg — any OS, any application.

Unlike screenshot-only approaches that route every action through expensive LLM inference, CEL uses structured sources (accessibility tree, native APIs) first and escalates to vision models only when needed. Faster, cheaper, more predictable — and capable of running fully offline.

Use CEL with Claude Code (MCP)

CEL ships as an MCP server with 4 tools. Connect it to Claude Code, Cursor, or any MCP client:

# Build everything
pnpm install && pnpm -r build

# Build native module (macOS)
cargo build --release -p cel-napi
cp target/release/libcel_napi.dylib cel/cel-napi/cel-napi.darwin-arm64.node
codesign -fs - cel/cel-napi/cel-napi.darwin-arm64.node

Pick an LLM provider — the fastest path is the interactive setup (writes ~/.cellar/config.toml):

cellar init

Options: paste a Gemini / Anthropic / OpenAI API key, or install Gemma 4 E4B locally via Ollama for fully-private runs. If you'd rather configure via .mcp.json directly (see below), skip init.

Configuration hierarchy

Environment variables override ~/.cellar/config.toml, which overrides compiled defaults.

# ~/.cellar/config.toml
[llm]
provider = "gemini"          # openai | anthropic | gemini | ollama | compatible
api_key  = "your-key"
model    = "gemini-2.0-flash"

[audio]                      # optional — enables audio transcription in the Cortex
whisper_endpoint = "https://api.openai.com/v1/audio/transcriptions"
whisper_api_key  = "sk-..."
whisper_model    = "whisper-1"
# whisper_language = "en"   # ISO 639-1 hint — improves accuracy

Full variable list: docs/api-reference.md.

Add to .mcp.json in your project root:

{
  "mcpServers": {
    "cellar": {
      "command": "node",
      "args": ["/path/to/cellar/mcp-server/dist/index.js"],
      "env": {
        "CEL_LLM_PROVIDER": "gemini",
        "CEL_LLM_API_KEY": "your-api-key",
        "CEL_LLM_MODEL": "gemini-2.0-flash"
      }
    }
  }
}

Restart Claude Code and you'll have four tools:

Tool

What it does

Modes/Actions

cel_see

Read the screen — structured elements with types, labels, bounds, confidence scores

14 modes

cel_act

Click, type, scroll, drag — by coordinates, element ID, or accessibility API

11 actions + CDP eval

cel_think

Plan, remember, track runs, autonomous execution (run_goal)

16 modes

cel_perceive

Always-on perception engine (Cortex) — continuous screen awareness

7 modes

On startup, the Cortex boots automatically (screen model is warm before your first call) and Chrome CDP is auto-detected.

See docs/quickstart.md for the full setup guide and docs/mcp-server.md for the complete tool reference.

Current State

Cellar is in prototype phase on macOS. The bar for exit is defined in docs/PROTOTYPE_EXIT_CRITERIA.md; the curated regression suite that gates it lives in eval/prototype-subset/.

Gated today (macOS local):

  • Local execution on macOS via AX + CDP + screen capture + input injection

  • MCP server with 4 composable tools: cel_see / cel_act / cel_think / cel_perceive

  • Cortex — always-on perception with background event streams

  • Autonomous execution (run_goal) over the prototype scenario suite: browser happy-paths, grounding, ambiguity, recovery, browser-to-desktop handoff

  • CLI entry points: cellar init (setup) and cellar run-goal "<goal>"

  • BYOK providers (OpenAI, Anthropic, Gemini) and local Ollama (Gemma 4 E4B default)

  • Per-role LLM routing — Planner / Observer / Vision / Validator

Built but outside the prototype exit bar:

  • Audio capture + Whisper transcription fused into the Cortex world model

  • Embedded SQLite + FTS5 for memory / semantic search

  • First-party adapters — Excel, SAP GUI, Bloomberg, MetaTrader

  • Recorder, live-view, and the wider benchmarks/ suite (50+ tasks + hybrid scenarios)

  • napi-rs Rust ↔ Node.js bridge

Later phase — explicitly not prototype work (see docs/ROADMAP.md):

  • Linux accessibility (AT-SPI2) and Windows UI Automation bridges

  • Remote worker / Docker image / managed VMs (cellar-worker/ exists in-tree as a preview; not wired into prototype gates)

  • Managed cloud, control plane, billing

  • Production confidence calibration

  • Portable context maps, community workflow registry

Architecture

cellar/
  cel/                  ← Cortex + perception layer (Rust, Apache 2.0)
    cel-accessibility/  ← accessibility bridge (AXUIElement, AT-SPI2)
    cel-context/        ← unified context API + multi-source fusion + references
    cel-display/        ← screen capture (xcap)
    cel-input/          ← input injection (enigo)
    cel-vision/         ← vision model integration (multi-provider)
    cel-network/        ← traffic monitoring + idle detection
    cel-store/          ← embedded SQLite + FTS5 (memory, knowledge)
    cel-llm/            ← LLM provider abstraction
    cel-planner/        ← built-in planner / runner code (useful, but not the repo's main value)
    cel-napi/           ← Node.js native bindings (napi-rs)
  agent/                ← agent integrations and runtime experiments
  mcp-server/           ← generic tool surface for external agents
  adapters/             ← app-specific adapters (browser, Excel, SAP)
  benchmarks/           ← eval harness (50+ tasks + 5 hybrid scenarios)
  live-view/            ← real-time debug surface (screen + runtime decisions)
  cli/                  ← `cellar` CLI

Getting Started

See docs/quickstart.md for the full step-by-step guide. The short version:

# 1. Build
pnpm install && pnpm -r build
cargo build --release -p cel-napi
cp target/release/libcel_napi.dylib cel/cel-napi/cel-napi.darwin-arm64.node
codesign -fs - cel/cel-napi/cel-napi.darwin-arm64.node

# 2. Configure .mcp.json (see quickstart for full config)

# 3. Grant Accessibility permissions in System Settings

# 4. Restart Claude Code — tools are ready

Quickstart — see what the agent sees

No Rust build needed. Just Node.js 20+ and pnpm:

pnpm install && pnpm -r build
npx tsx examples/quickstart.ts https://github.com/login

This launches a browser, extracts DOM elements as structured ContextElements with confidence scores, and shows the kind of context any external agent runtime would receive.

Prerequisites

  • Node.js 20+ and pnpm 9+ (TypeScript packages)

  • Rust 1.75+ (CEL core, accessibility bridge, native bindings)

  • macOS 13+ with Accessibility permissions

  • Chrome (optional, for CDP features)

Build

# Build everything
make build

# Or separately
make build-rust    # cargo build --workspace
make build-ts      # pnpm install && pnpm build

# Run tests
make test

CLI

cellar init                    # Interactive first-run setup (pick LLM provider or install Gemma 4)
cellar setup                   # Configure AX + CDP permissions on this machine
cellar context                 # Show unified context with confidence scores
cellar context --json          # Output raw JSON
cellar context --watch         # Live-update context in terminal
cellar capture                 # Capture screenshot to file
cellar action click 500 300    # Click at coordinates
cellar action type "Hello"     # Type text
cellar action key Enter        # Press a key
cellar action combo Ctrl C     # Key combination
cellar mcp                     # Start MCP server (stdio)
cellar mcp install             # Print Claude Desktop config
cellar run <workflow>          # Execute a saved workflow
cellar train                   # Enter training mode

Benchmarks

Hybrid Runtime Scenarios (CEL advantage)

5 scenarios designed to test where multi-source perception matters. Run them: ./scripts/demo.sh

Scenario

What breaks screenshot agents

CEL metric

Browser → Desktop handoff

Lose context across app boundary

sideEffectWarnings

Stale state (2s shuffle)

Click where button was

staleRecoveries, refreshRoutes

Ambiguous targets (8 similar names)

Can't distinguish identical buttons

semanticRoutes

Side-effect detection (unexpected modal)

Stuck or blindly proceed

sideEffectWarnings

Terminal failure (auth-blocked)

Loop forever

terminalFailures

General Web Tasks

We also benchmark on 50+ general web tasks against other tools:

Tool

Approach

Cellar

Multi-source fusion (DOM + a11y + vision + network), confidence scoring, incremental updates

Anthropic Computer Use

Screenshot-only, pixel-coordinate actions via API

Browser-Use (OSS)

Hybrid screenshot + DOM (Python)

Browserbase + Stagehand

Cloud CDP + AI SDK

Browser-Use Cloud

Managed browser-use + custom model

Measured on Apple M-series (arm64, 12 cores, 18GB RAM), April 2026. Hybrid suite: 5 tasks testing browser-desktop handoff, stale state recovery, ambiguous targets, side-effect detection, terminal failure. All local tools use Gemini 2.5 Flash. Computer Use locked to Claude Sonnet.

Benchmark results (April 2026 — Hybrid Suite, 5 tasks)

Tool

Avg Time

LLM Calls

Cost/Task

Success

CEL

20.8s

1.4

$0.0005

100%

Browser-Use OSS

23.4s

3.0

$0.001

100%

Stagehand v3

35.6s

18.2

$0.005

20%

Computer Use

36.2s

6.2

$0.155

100%

Browser-Use Cloud

46.5s

5.6

$0.003

100%

CEL vs the field:

vs

Speed

Cost

Accuracy

Computer Use (Anthropic)

1.7x faster

310x cheaper

Same (100%)

Browser-Use Cloud

2.2x faster

6x cheaper

Same (100%)

Stagehand v3

1.7x faster

10x cheaper

5x better (100% vs 20%)

Browser-Use OSS

1.1x faster

2x cheaper

Same (100%)

Why CEL wins:

  • 1 LLM call per task — structured context means most tasks extract data in a single pass. Competitors need 3-18 calls.

  • $0.0005/task — Gemini Flash + context distillation. At 1000 tasks: CEL $0.50 vs Computer Use $155.

  • Structured context is free — 500+ elements extracted in 100-400ms via Rust-native DOM fusion, no LLM required.

  • Full Rust execution loop — perceive, plan, execute, verify all in Rust. No FFI in the hot path.

  • For building reliable automation (not one-off tasks), structured context is the foundation

See benchmarks/README.md for full methodology, per-task breakdown, and how to reproduce.

Roadmap

The forward plan lives in docs/ROADMAP.md. Two related references:

Contributing

See CONTRIBUTING.md for how to get started, and DEVELOPMENT.md for build instructions and conventions.

We welcome contributions — especially:

  • Accessibility bridges (macOS AXUIElement, Windows UI Automation)

  • New application adapters — see docs/building-adapters.md

  • MCP tool improvements

  • Test coverage for platform-specific code

  • Documentation and examples

Platform Support

Platform

Status

macOS

Primary platform. AXUIElement bridge, Cortex, MCP server — all fully functional.

Linux

AT-SPI2 accessibility bridge working

Windows

Planned (UI Automation bridge designed, not yet implemented)

License

  • Everything OSS-destined (cel/, agent/, cli/, mcp-server/, cellar-worker/, live-view/, recorder/, registry/, docs/, benchmarks/, examples/, e2e/, tests/, box/): Apache License 2.0.

  • Community adapters (adapters/): MIT.

  • Commercial-only (app/, future control-plane/, cloud/, billing/): proprietary — not covered by this license.

See docs/oss-boundary.md for the full license map and what stays private.

Available Tools

4 tools
cel_actCEL ActA

Execute actions on the screen: mouse clicks, keyboard input, accessibility actions, drag & drop, and direct value setting. Always use cel_see first to understand the screen.

For click/move: provide (x, y) coordinates or a target_ref from cel_see make_reference. For form filling: prefer set_value over type — faster and more reliable. For buttons/checkboxes: prefer ax_action over click — uses native accessibility API.

Coordinate Actions (x,y or target_ref): click, right_click, double_click, mouse_move.

Keyboard: type (text string), key_press (single key: Enter, Tab, Escape, etc.), key_combo (modifier combinations: ['Ctrl','C'], ['Cmd','Shift','S']).

Accessibility API (preferred for reliability): ax_action — native a11y actions on element_id: click, activate, press, increment, decrement, cancel, show_menu, scroll_to_visible, raise, pick, delete. set_value — direct value injection on element_id: text for fields, 'true'/'false' for checkboxes.

Deterministic spreadsheet actions: write_cells (atomic Numbers cell writes with optional readback verification), read_cells (read Numbers cell values from the document model instead of guessing from AX text).

Other: scroll (dx,dy at optional x,y), drag (from_x,from_y to to_x,to_y), cdp_eval (execute JavaScript in browser via CDP — best for cookie banners, iframes, overlays, and elements invisible to the accessibility tree).

Batching: pass array of 1-4 actions for sequential execution (100ms default delay). Re-observe with cel_see after each batch to avoid stale-state cascading failures.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries full burden. It describes each action type, mentions deterministic spreadsheet actions, batching with default delay, and warns about stale-state cascading failures. Side effects (UI mutation) are implied, and no contradictions exist.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is lengthy but well-structured with bullet points and sections. It front-loads purpose and general guidance. Some redundancy exists (e.g., repeating 'prefer'), but overall it is organized and earn its detail for the variety of actions.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema is provided, and the description does not explain what the tool returns. Additionally, the input schema is empty, creating a mismatch with the description that implies parameters. The missing return value and schema inconsistency reduce completeness.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Input schema has zero parameters, so baseline is 4 per instructions. The description adds substantial meaning by detailing all action types and their required coordinates, target_ref, element_id, etc., far beyond the empty schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states that the tool executes actions on the screen including mouse clicks, keyboard input, accessibility actions, drag & drop, and direct value setting. It also distinguishes itself from siblings by advising to use cel_see first, making its purpose distinct and specific.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Extensive guidelines are provided: always use cel_see first, prefer set_value over type for form filling, prefer ax_action over click for buttons/checkboxes, and detailed recommendations for each action type. Batching and re-observing instructions are also given, offering clear when-to-use and when-not-to-use guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

cel_perceiveCEL PerceiveA

Always-on perception engine (Cortex). Maintains a continuously-updated mental model via background event streams with periodic accessibility tree refreshes on significant changes, and vision/screenshots when flagged as needed.

IMPORTANT: Singleton — only one perception session can be active at a time. cel_see 'watch' mode is unavailable during an active session.

Modes:

  • start: Boot the cortex with a goal. Set enable_suggestions=true (default) for LLM-powered next-action recommendations on each read.

  • read: Get the mental model snapshot (instant — model is kept warm by background events).

  • feed: Report an action you took (action, target, expected outcome). Cortex waits for screen to settle, diffs against current model, returns verification.

  • checkpoint: Summarize completed work and reset action history. Use between phases of multi-step tasks.

  • configure: Update goal or enable_suggestions mid-session.

  • status: Cortex health — confidence score, uptime, cycle count, element counts (stable vs volatile), temporal state (loading, errors, focus trail).

  • stop: Shutdown the cortex and get a summary.

The model includes temporal awareness (loading states, error persistence, focus trail) and element stability classification (stable vs volatile targets).

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are present, so the description carries full burden. It discloses that the tool is always-on, maintains a background mental model, uses event streams, accessibility refreshes, and optional screenshots. It explains each mode's behavior and side effects (e.g., feed waits for screen settle, diffs model). Minor ambiguity about whether feedback modifies state, but overall highly transparent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is somewhat lengthy but well-organized with a clear mode list and important constraints upfront. Every sentence adds information, though some details could be tightened. Front-loading the singleton note is effective.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no parameters, no output schema, and no annotations, the description covers all essential information: purpose, modes, constraints, sibling differentiation, and behavioral model. It is fully adequate for an agent to select and invoke the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has zero parameters (100% documented by schema), so baseline is 4. The description adds value by explaining the modes which act as sub-operations, but no parameter details are needed.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool is an 'always-on perception engine' that maintains a mental model, and explicitly lists all modes (start, read, feed, etc.) with specific verbs and resources. It effectively distinguishes from siblings like cel_see by noting that 'watch' mode is unavailable during active session.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides explicit guidance on when to use each mode, including when not to use certain modes (e.g., 'cel_see watch mode is unavailable during an active session'). It also highlights the singleton constraint, aiding the agent in choosing this tool appropriately.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

cel_seeCEL SeeA

Read and observe the current screen state. Returns structured UI elements, window lists, screenshots, CDP page content, accessibility element details, and screen change events. Always use this BEFORE acting to understand what's on screen.

Screen Context: context (elements with filter/compression — use detail 'compact' to save tokens), screenshot (PNG capture), windows (visible window list), monitors (display list).

Element Inspection: focused (high-fidelity detail for one element_id), element_at (hit-test x,y coordinates), is_settable (check if set_value works), make_reference (resilient ref that survives across snapshots), cursor_position.

Browser (CDP): cdp_status (debug targets & connection state), cdp_page (full page content as text).

Observation Recall: observation (load a persisted context snapshot by observation_id).

Waiting & Watching: wait_for_element (poll for element by type/label, default 10s timeout), wait_for_idle (poll until screen stabilizes — requires 2 consecutive stable polls), watch (event-driven — 18 event types: tree_changed, network_idle, focus_changed, value_changed, window_created, menu_opened, menu_closed, sheet_created, layout_changed, title_changed, app_activated, app_deactivated, window_moved, window_resized, window_minimized, window_restored, selection_changed, row_count_changed). Note: watch is unavailable during an active cel_perceive session.

Limits: CDP enrichment caps at 50 text_blocks, 50 interactive_elements, 3000 char body_text.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Despite no annotations, the description provides rich behavioral details: default timeout for wait_for_element (10s), requirement for wait_for_idle (2 consecutive stable polls), 18 event types for watch, CDP limits (50 text_blocks, etc.), and conflict note about cel_perceive. This goes far beyond simple annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well-organized with clear sections (Screen Context, Element Inspection, Browser, Observation Recall, Waiting & Watching, Limits). Each sentence adds value, providing necessary detail without redundancy. Front-loaded with purpose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no output schema, the description lists categories of returned data but does not fully specify output structure. However, it covers key aspects like limits and sub-function behaviors. It feels complete for a read tool, though a more structured output spec would be even better.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There are zero parameters in the input schema, and the description compensates by thoroughly explaining all the tool's sub-functions (Screen Context, Element Inspection, etc.). According to guidelines, 0 params = baseline 4; this description exceeds that with detailed breakdown of capabilities.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description explicitly states 'Read and observe the current screen state' and lists many capabilities. It distinguishes from siblings by saying 'Always use this BEFORE acting', making clear this is the observation tool while cel_act is for actions and cel_perceive for perception.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear usage advice: 'Always use this BEFORE acting'. It also notes a limitation (watch unavailable during cel_perceive session). However, it does not explicitly state when not to use or provide direct comparison with cel_perceive.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

cel_thinkCEL ThinkC

CEL's cognitive layer: delegated autonomy, planning, knowledge, run tracking, and LLM passthrough.

Efficiency rule: if the MCP host already reasons well step-by-step, prefer cel_see + cel_act and keep planning in the host. Use run_goal only when you intentionally want CEL to take over the control loop.

Delegated Autonomous Execution: run_goal — give a natural language goal, CEL runs a full internal see→plan→act loop autonomously. This can be convenient, but it adds an internal planner loop and may be slower or more expensive than host-driven execution. Only goal, max_steps (default 80), and timeout_ms (default 900_000) are tunable — vision, self-healing, decomposition, and notebook are implicit in the canonical loop and no longer per-invocation knobs (see docs/canonical-agent-plan.md).

Planning: plan (LLM-powered step planning with optional history for multi-step context), plan_with_vision (plan with screenshot — use for visual/spatial tasks).

Knowledge Store (persisted to ~/.cellar/cel-store.db): store_knowledge (save facts with source and optional tags), search_knowledge (FTS5 full-text search, default 10 results, scope by workflow).

Working Memory: memory_get, memory_set (per-workflow scratchpad, not persisted across sessions).

Observations: observe (record insight with priority high/medium/low), get_observations (retrieve, default 50).

Run Tracking: run_start, run_finish, run_log_step (per-step with confidence score), run_history, run_steps.

LLM Passthrough: llm_complete (text, 4096 tokens default), llm_complete_with_image (vision, 4096 tokens default).

Maintenance: eviction (TTL cleanup — default 90 days runs, 365 days knowledge).

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description details many capabilities but fails to disclose what happens on invocation without arguments. No annotations are provided to clarify behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is excessively long and poorly structured, lacking front-loading. It lists many sub-functions without clear organization.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity of the described features and the lack of parameters or output schema, the description is incomplete for an agent to know how to effectively use the tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has no parameters, so schema-description coverage is 100%. No parameter semantics are needed, baseline score of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description lists many sub-operations but does not state what the tool does when invoked with no parameters. The purpose is vague and ambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description includes efficiency guidance preferring cel_see+cel_act in some cases, but does not clarify how to invoke any of the listed sub-operations since the tool takes no parameters.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 4 tool updatesv0.1.0
    • First observedcel_act
    • First observedcel_perceive
    • First observedcel_see
    • First observedcel_think

TDQS

A3.7/5.0

Scored across 4 tools

Disambiguation5/5

Each tool has a distinct role: cel_act for executing actions, cel_perceive for continuous perception, cel_see for reading screen state, and cel_think for cognitive planning. Descriptions clarify boundaries despite some perceptual overlap.

Naming Consistency5/5

All tool names follow the consistent pattern 'cel_verb' (act, perceive, see, think), using lowercase with underscores throughout. No deviations or mixed conventions.

Tool Count4/5

With only 4 tools, the set is compact but each encapsulates many sub-operations via parameters and modes. The count is slightly low but appropriate for the server's focused domain of screen automation and perception.

Completeness4/5

The tools cover perception, action, and cognitive planning comprehensively for UI automation. Minor gaps exist (e.g., no explicit system-level operations), but core workflows are well supported and no obvious dead ends.

Maintenance

ActivityInactive
ResponsivenessSlow

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    F
    maintenance
    An open-source MCP server for macOS and Windows that provides native desktop control via Accessibility APIs, OCR, and Chrome CDP. It enables AI agents to interact with applications, manage browser sessions, and automate workflows with high-speed native UI actions.
    91 npm
    15
    AGPL 3.0
  • A
    license
    Not graded
    quality
    A
    maintenance
    Gives AI agents and MCP clients direct control over native desktop apps, Chrome/Electron browsers, and Android devices with screenshots, OCR, accessibility-based element lookup, input simulation, window management, CDP, and ADB in one local server.
    133
    MIT
  • A
    license
    Not graded
    quality
    B
    maintenance
    MCP server that gives AI agents hands and eyes on macOS, enabling them to see and operate native and Electron applications via structured accessibility queries, screenshots, OCR, and application-specific skills.
    4
    MIT
  • F
    license
    A
    quality
    A
    maintenance
    Cross-platform desktop automation MCP server that lets AI agents capture screenshots, run OCR with UI-element classification, control mouse/keyboard, and launch programs on Linux, macOS, and Windows.
    20
    1
    -