Skip to main content
Glama

Agent Decision Kit

Fast, local-first decisions and browser actions for coding agents.

CI License: Apache-2.0 Node.js 20.19+

Quick start · Browser demo · Agent setup · Privacy and cost

Local browser demo screenshot

Agent Decision Kit is an experimental Apache-2.0 MCP server and CLI for the small decisions inside agent loops: which visible browser action to take, which file to inspect, what context to keep, whether a diff deserves review, and whether supplied evidence supports a completion claim.

It uses ordinary code to constrain choices and execute actions. Its default decision backend is a small model that runs locally through Transformers.js. Optional Jev and OpenAI-compatible providers are adapters, not requirements.

What it does

  • Browser loop first. Playwright inspects a bounded set of visible controls, chooses from those controls, and returns a short page delta. Only when a page has no semantic controls, it also detects text targets whose only interaction hint is an explicit CSS pointer cursor. These non-semantic custom targets always require confirmation. Launch an isolated profile or list and explicitly select a local Chrome tab over loopback CDP. Read-only date widgets are marked in the snapshot and opened through a separate click before choosing a visible date; they are never passed to browser_fill.

  • Approval for consequential actions. Payment, sending, publishing, deletion, form submission, and non-semantic CSS pointer targets pause for a distinct confirmation call. Password and file inputs are excluded. Text and date/time fields can be filled, and native select options can be chosen, without submitting the page. Contenteditable drafts are masked from DOM text; private field values and action destinations are hashed locally to invalidate stale approvals, never returned to the agent.

  • Structured decisions. Batch up to eight Choice, Score, and yes/no (noul) questions in one request, with probability estimates, confidence source, and explicit calibration label.

  • Warm before a latency-sensitive run. Call provider_warmup once inside the same MCP session to initialize the local model before ambiguous decisions. Its first call may download weights; it is optional and does not promise a speed target.

  • Fast local browser paths. A unique exact control label or non-semantic CSS pointer target in a simple click/tap/open request (quoted or unquoted), numbered tabs, named checkboxes, and clear expand-then-submit steps can resolve without model inference. Duplicate labels remain ambiguous. These rule-based choices return no probability; custom targets and sensitive actions still wait for approval.

  • Coding workflows. Find relevant files, prune context while keeping requested strings verbatim, suggest a model route, pre-screen a diff, rerank, classify, screen, extract from caller-supplied candidates, and check completion evidence.

  • CLI and MCP. Use the same functions in a shell pipeline or from an MCP-compatible coding agent.

  • Fast visual text fallback. When a page has no semantic controls or clear CSS pointer targets, browser_decide_and_act returns immediately and points to browser_visual_text. It runs local Tesseract OCR, masks editable text fields, and returns bounded lines and word boxes. If the first OCR pass misses the requested phrase, browser_visual_action tries one slower sparse-text pass. A click proposal still requires one exact, unique match and a separate browser_confirm call; stale screenshots are rejected. OCR can miss or misread text; the screenshot stays local while recognized text enters the agent context.

  • Optional visual question answering. browser_visual_inspect can answer a question about the screenshot with SmolVLM2 500M, an Apache-2.0 vision-language model. First use downloads model files; CPU inference can take tens of seconds. The screenshot stays local. Its description is uncalibrated and never triggers an action.

Architecture diagram

Calibration is visible

confidence and probability values are estimates, not a promise that the selected action is correct. Each model answer identifies confidenceSource as provider-reported, maximum-probability, or unavailable, separately from its calibration label (uncalibrated-estimate, posthoc-calibrated, provider-calibrated, or unavailable). Deterministic browser matches report not-applicable-rule and do not invent probabilities. The default MiniLM embedding baseline is uncalibrated and is not a drop-in Jev replacement. No speed, accuracy, or parity claim is made before an independent evaluation.

Related MCP server: agentboost

Quick start

Requirements: Node.js 20.19 or newer. The repository is in experimental alpha; the npm package is not published yet. Run from a local checkout:

npm install
npm run build
npm run browser:install
npm run mcp

Configure your agent to start node /absolute/path/to/agent-decision-kit/dist/cli.js mcp. Agent-specific files and examples are in docs/agents/. When the package is published, the shorter command will be npx -y agent-decision-kit mcp.

For a one-shot CLI decision, pass JSON on stdin:

cat examples/decision.json | node dist/cli.js decide

For local semantic filtering of newline-delimited items:

cat examples/tasks.txt | node dist/cli.js filter --query "is a browser automation task"

The first decision call downloads the configured embedding model unless its files are already cached. It runs locally after that. To use a local Ollama or another OpenAI-compatible server instead, set AGENT_DECISION_PROVIDER=openai-compatible; the default endpoint is http://127.0.0.1:11434/v1.

Browser demo

The demo is a local static page with fictional tasks. Start any static server from the repository root, then launch the MCP server and ask your agent:

Open the local browser demo, inspect the visible tasks, and mark the setup task complete. Show me the page change.

The accessible demo is examples/browser-demo.html. The screenshot above and short recording at website/public/images/browser-demo.gif were captured with Playwright. The second example draws its interface into a canvas, so Playwright has no DOM controls to inspect. It demonstrates fast local OCR and an exact-text click proposal that requires separate approval:

Open the canvas-only demo · npm run ocr:verify checks OCR, cancellation, approval, and stale-screenshot rejection · npm run vision:verify runs local visual question answering (first use downloads model weights).

Canvas-only fake task board used to check local screenshot understanding

Both demos use synthetic data. They do not send, publish, charge, or delete anything outside the page.

The browser tool does not bypass CAPTCHAs, site access controls, or authentication. For remote model providers, browser labels are blocked by default; set AGENT_ALLOW_REMOTE_BROWSER_CONTEXT=true only when you intend to share bounded page labels with that provider.

Connect your agent

Agent

Setup guide

MCP configuration format

Claude Code

Guide

CLI or .mcp.json

Codex CLI

Guide

~/.codex/config.toml

Cursor

Guide

.cursor/mcp.json

Gemini CLI

Guide

settings.json or gemini mcp add

Windsurf Cascade

Guide

mcp_config.json

VS Code / Copilot Chat

Guide

.vscode/mcp.json

GitHub Copilot CLI

Guide

~/.copilot/mcp-config.json

Cline

Guide

CLI MCP wizard or IDE settings

OpenCode

Guide

opencode.json

All integrations use the standard MCP stdio transport. CI starts the actual CLI subprocess, discovers its MCP tools, and calls a local workflow tool; separate protocol tests cover the in-memory transport. In a local Windows check, Claude Code 2.1.218 reported the isolated server as connected; Codex CLI 0.154.0 loaded an isolated entry as enabled, but that check did not run an agent turn or tool call. The other vendor clients have not been runtime-tested here.

Claude Code users can also opt into the prompt-routing hook. It adds a local, unbenchmarked route suggestion to submitted prompts; it never changes the active model.

Privacy and cost

  • No account, hosted inference endpoint, product telemetry, or central server is required.

  • The default provider sends state to the local Transformers.js model after its weights are downloaded. The optional local OpenAI-compatible adapter also defaults to loopback.

  • Remote OpenAI-compatible and Jev providers send the decision request to the configured provider. Browser context is separately blocked for remote providers unless explicitly opted in.

  • Jev is optional and can incur TypeSafe charges. Its API endpoint and request format follow TypeSafe's API docs. Jev outputs are not stored as training data, used to tune the local model, or used to build an imitator; review the TypeSafe agreement before enabling that adapter.

  • Chromium uses a separate persistent profile at ~/.agent-decision-kit/browser-profile; set AGENT_DECISION_BROWSER_DIR to change it. CDP attachment uses a dedicated Chrome profile and explicit tab selection. Cookies, storage, URL credentials/query/hash, local file paths, and editable form values are not returned by DOM tools. A non-reversible digest of form state and action destinations is kept locally only to expire stale approvals; raw field values are not returned. Local OCR masks editable fields and returns bounded visible page text to the calling agent; that text may enter its model context. Treat it as untrusted page content. The screenshot itself stays local. Remote decision-provider browser context still requires explicit opt-in. Closing an attached session disconnects instead of closing the selected Chrome context.

Benchmarks and honest claims

benchmarks/ contains labeled fixtures, an evaluation protocol, raw records, and metric definitions. The eight-task MiniWoB smoke suite completed 7/8 on earlier Windows and Linux runs and 8/8 on multiple Linux CPU repeats of the same seed-7, five-action task set. The latest diagnostic-enabled repeat passed 8/8; end-to-end latency was p50 1,096 ms / p95 1,492 ms, and browser decision/action latency was p50 305 ms / p95 699 ms. A preceding run completed 7/8 when click-test terminated before the decision/action call after 61.8 seconds; an isolated retry passed 1/1 and the later full rerun passed 8/8, so that failure remains recorded as an unreproduced anomaly. These curated runs are integration evidence, not a representative benchmark or general speed claim. A separate 4/5 multi-tool harness exercises form entry, date-picker selection, native selects, paginated search, and read-only visual description. A focused OCR link test failed 0/1 because both OCR passes missed the label; the CSS pointer-target path later completed that target 1/1 after explicit local approval (280 ms loop, one sample). These are curated tool-loop checks, not a full agent or representative quality/speed results. A 30-case project-specific decision fixture scored 7/10 on Choice, 5/10 on yes/no, and 1.04 mean absolute error on Score; it is not held out and its probability estimates are uncalibrated. Across one Windows and one Linux CPU repeat, the 20 Choice/yes-no confidence values had descriptive ECE near 0.061; this small non-held-out result does not establish calibration. Report decision accuracy, score error, calibration, browser task success, action count, and latency separately. Broader WebArena/VisualWebArena evaluation and runtime testing of the remaining vendor clients remain evaluation work; CI does not generate fabricated benchmark charts.

We do not claim the 500 ms p95 target, 7-second browsing demo, Jev equivalence, or any other speedup until a reproducible run is published with hardware, versions, sample counts, and raw results.

Development

npm install
npm run typecheck
npm test
npm run package:verify
npm run mcp:verify
npm run hook:verify
npm run agents:verify
npm --prefix website install
npm run website:build

See contributing, security policy, and the changelog. The source is released under Apache-2.0.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    B
    maintenance
    Enables coding agents to perform file, search, patch, git, process, test, package, network, and system operations through 60 typed MCP tools with structured inputs/outputs, structured errors, and a full event journal, replacing terminal use with a typed machine API.
    MIT
  • A
    license
    Not graded
    quality
    C
    maintenance
    Provides coding agents with structured planning, persistent project memory, automated verification, and safety permission controls through MCP tools, enabling better planning, context retention, self-checking, and guarded execution.
    1 npm
    MIT
  • A
    license
    A
    quality
    B
    maintenance
    Enables coding agents to make offline, zero-cost decisions using schema-safe Choice/Score/Noul primitives, a confidence gatekeeper, planning, adversarial red-teaming, research, and RLVR-based self-improvement via 21 MCP tools.
    24
    2
    MIT