Agent Decision Kit
Provides an optional OpenAI-compatible decision backend that can use a local Ollama server (default http://127.0.0.1:11434/v1) for local model inference.
Supports optional OpenAI-compatible providers as a decision backend, allowing remote OpenAI-compatible endpoints to be used instead of the default local Transformers.js model.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Agent Decision Kitwhich files are relevant to the failing auth test?"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Agent Decision Kit
Fast, local-first decisions and browser actions for coding agents.
Quick start · Browser demo · Agent setup · Privacy and cost

Agent Decision Kit is an experimental Apache-2.0 MCP server and CLI for the small decisions inside agent loops: which visible browser action to take, which file to inspect, what context to keep, whether a diff deserves review, and whether supplied evidence supports a completion claim.
It uses ordinary code to constrain choices and execute actions. Its default decision backend is a small model that runs locally through Transformers.js. Optional Jev and OpenAI-compatible providers are adapters, not requirements.
What it does
Browser loop first. Playwright inspects a bounded set of visible controls, chooses from those controls, and returns a short page delta. Only when a page has no semantic controls, it also detects text targets whose only interaction hint is an explicit CSS pointer cursor. These non-semantic custom targets always require confirmation. Launch an isolated profile or list and explicitly select a local Chrome tab over loopback CDP. Read-only date widgets are marked in the snapshot and opened through a separate click before choosing a visible date; they are never passed to
browser_fill.Approval for consequential actions. Payment, sending, publishing, deletion, form submission, and non-semantic CSS pointer targets pause for a distinct confirmation call. Password and file inputs are excluded. Text and date/time fields can be filled, and native select options can be chosen, without submitting the page. Contenteditable drafts are masked from DOM text; private field values and action destinations are hashed locally to invalidate stale approvals, never returned to the agent.
Structured decisions. Batch up to eight
Choice,Score, and yes/no (noul) questions in one request, with probability estimates, confidence source, and explicit calibration label.Warm before a latency-sensitive run. Call
provider_warmuponce inside the same MCP session to initialize the local model before ambiguous decisions. Its first call may download weights; it is optional and does not promise a speed target.Fast local browser paths. A unique exact control label or non-semantic CSS pointer target in a simple click/tap/open request (quoted or unquoted), numbered tabs, named checkboxes, and clear expand-then-submit steps can resolve without model inference. Duplicate labels remain ambiguous. These rule-based choices return no probability; custom targets and sensitive actions still wait for approval.
Coding workflows. Find relevant files, prune context while keeping requested strings verbatim, suggest a model route, pre-screen a diff, rerank, classify, screen, extract from caller-supplied candidates, and check completion evidence.
CLI and MCP. Use the same functions in a shell pipeline or from an MCP-compatible coding agent.
Fast visual text fallback. When a page has no semantic controls or clear CSS pointer targets,
browser_decide_and_actreturns immediately and points tobrowser_visual_text. It runs local Tesseract OCR, masks editable text fields, and returns bounded lines and word boxes. If the first OCR pass misses the requested phrase,browser_visual_actiontries one slower sparse-text pass. A click proposal still requires one exact, unique match and a separatebrowser_confirmcall; stale screenshots are rejected. OCR can miss or misread text; the screenshot stays local while recognized text enters the agent context.Optional visual question answering.
browser_visual_inspectcan answer a question about the screenshot with SmolVLM2 500M, an Apache-2.0 vision-language model. First use downloads model files; CPU inference can take tens of seconds. The screenshot stays local. Its description is uncalibrated and never triggers an action.
Calibration is visible
confidence and probability values are estimates, not a promise that the selected action is correct. Each model answer identifies confidenceSource as provider-reported, maximum-probability, or unavailable, separately from its calibration label (uncalibrated-estimate, posthoc-calibrated, provider-calibrated, or unavailable). Deterministic browser matches report not-applicable-rule and do not invent probabilities. The default MiniLM embedding baseline is uncalibrated and is not a drop-in Jev replacement. No speed, accuracy, or parity claim is made before an independent evaluation.
Related MCP server: agentboost
Quick start
Requirements: Node.js 20.19 or newer. The repository is in experimental alpha; the npm package is not published yet. Run from a local checkout:
npm install
npm run build
npm run browser:install
npm run mcpConfigure your agent to start node /absolute/path/to/agent-decision-kit/dist/cli.js mcp. Agent-specific files and examples are in docs/agents/. When the package is published, the shorter command will be npx -y agent-decision-kit mcp.
For a one-shot CLI decision, pass JSON on stdin:
cat examples/decision.json | node dist/cli.js decideFor local semantic filtering of newline-delimited items:
cat examples/tasks.txt | node dist/cli.js filter --query "is a browser automation task"The first decision call downloads the configured embedding model unless its files are already cached. It runs locally after that. To use a local Ollama or another OpenAI-compatible server instead, set AGENT_DECISION_PROVIDER=openai-compatible; the default endpoint is http://127.0.0.1:11434/v1.
Browser demo
The demo is a local static page with fictional tasks. Start any static server from the repository root, then launch the MCP server and ask your agent:
Open the local browser demo, inspect the visible tasks, and mark the setup task complete. Show me the page change.
The accessible demo is examples/browser-demo.html. The screenshot above and short recording at website/public/images/browser-demo.gif were captured with Playwright. The second example draws its interface into a canvas, so Playwright has no DOM controls to inspect. It demonstrates fast local OCR and an exact-text click proposal that requires separate approval:
Open the canvas-only demo · npm run ocr:verify checks OCR, cancellation, approval, and stale-screenshot rejection · npm run vision:verify runs local visual question answering (first use downloads model weights).

Both demos use synthetic data. They do not send, publish, charge, or delete anything outside the page.
The browser tool does not bypass CAPTCHAs, site access controls, or authentication. For remote model providers, browser labels are blocked by default; set AGENT_ALLOW_REMOTE_BROWSER_CONTEXT=true only when you intend to share bounded page labels with that provider.
Connect your agent
Agent | Setup guide | MCP configuration format |
Claude Code | CLI or | |
Codex CLI |
| |
Cursor |
| |
Gemini CLI |
| |
Windsurf Cascade |
| |
VS Code / Copilot Chat |
| |
GitHub Copilot CLI |
| |
Cline | CLI MCP wizard or IDE settings | |
OpenCode |
|
All integrations use the standard MCP stdio transport. CI starts the actual CLI subprocess, discovers its MCP tools, and calls a local workflow tool; separate protocol tests cover the in-memory transport. In a local Windows check, Claude Code 2.1.218 reported the isolated server as connected; Codex CLI 0.154.0 loaded an isolated entry as enabled, but that check did not run an agent turn or tool call. The other vendor clients have not been runtime-tested here.
Claude Code users can also opt into the prompt-routing hook. It adds a local, unbenchmarked route suggestion to submitted prompts; it never changes the active model.
Privacy and cost
No account, hosted inference endpoint, product telemetry, or central server is required.
The default provider sends state to the local Transformers.js model after its weights are downloaded. The optional local OpenAI-compatible adapter also defaults to loopback.
Remote OpenAI-compatible and Jev providers send the decision request to the configured provider. Browser context is separately blocked for remote providers unless explicitly opted in.
Jev is optional and can incur TypeSafe charges. Its API endpoint and request format follow TypeSafe's API docs. Jev outputs are not stored as training data, used to tune the local model, or used to build an imitator; review the TypeSafe agreement before enabling that adapter.
Chromium uses a separate persistent profile at
~/.agent-decision-kit/browser-profile; setAGENT_DECISION_BROWSER_DIRto change it. CDP attachment uses a dedicated Chrome profile and explicit tab selection. Cookies, storage, URL credentials/query/hash, local file paths, and editable form values are not returned by DOM tools. A non-reversible digest of form state and action destinations is kept locally only to expire stale approvals; raw field values are not returned. Local OCR masks editable fields and returns bounded visible page text to the calling agent; that text may enter its model context. Treat it as untrusted page content. The screenshot itself stays local. Remote decision-provider browser context still requires explicit opt-in. Closing an attached session disconnects instead of closing the selected Chrome context.
Benchmarks and honest claims
benchmarks/ contains labeled fixtures, an evaluation protocol, raw records, and metric definitions. The eight-task MiniWoB smoke suite completed 7/8 on earlier Windows and Linux runs and 8/8 on multiple Linux CPU repeats of the same seed-7, five-action task set. The latest diagnostic-enabled repeat passed 8/8; end-to-end latency was p50 1,096 ms / p95 1,492 ms, and browser decision/action latency was p50 305 ms / p95 699 ms. A preceding run completed 7/8 when click-test terminated before the decision/action call after 61.8 seconds; an isolated retry passed 1/1 and the later full rerun passed 8/8, so that failure remains recorded as an unreproduced anomaly. These curated runs are integration evidence, not a representative benchmark or general speed claim. A separate 4/5 multi-tool harness exercises form entry, date-picker selection, native selects, paginated search, and read-only visual description. A focused OCR link test failed 0/1 because both OCR passes missed the label; the CSS pointer-target path later completed that target 1/1 after explicit local approval (280 ms loop, one sample). These are curated tool-loop checks, not a full agent or representative quality/speed results. A 30-case project-specific decision fixture scored 7/10 on Choice, 5/10 on yes/no, and 1.04 mean absolute error on Score; it is not held out and its probability estimates are uncalibrated. Across one Windows and one Linux CPU repeat, the 20 Choice/yes-no confidence values had descriptive ECE near 0.061; this small non-held-out result does not establish calibration. Report decision accuracy, score error, calibration, browser task success, action count, and latency separately. Broader WebArena/VisualWebArena evaluation and runtime testing of the remaining vendor clients remain evaluation work; CI does not generate fabricated benchmark charts.
We do not claim the 500 ms p95 target, 7-second browsing demo, Jev equivalence, or any other speedup until a reproducible run is published with hardware, versions, sample counts, and raw results.
Development
npm install
npm run typecheck
npm test
npm run package:verify
npm run mcp:verify
npm run hook:verify
npm run agents:verify
npm --prefix website install
npm run website:buildSee contributing, security policy, and the changelog. The source is released under Apache-2.0.
This server cannot be deployed
Maintenance
Related MCP Connectors
The project brain for AI coding agents — memory, decisions, sprints, knowledge base via MCP.
Shared control plane for AI coding agents — tasks, memory, decisions, file locks. 12 tools.
Governed app access for AI agents: 1,000+ apps & 12,000+ tools via Code Mode MCP.
Remote MCP learning coach for coding agents.
Related MCP Servers
- AlicenseNot gradedqualityBmaintenanceEnables coding agents to perform file, search, patch, git, process, test, package, network, and system operations through 60 typed MCP tools with structured inputs/outputs, structured errors, and a full event journal, replacing terminal use with a typed machine API.MIT
- AlicenseNot gradedqualityCmaintenanceProvides coding agents with structured planning, persistent project memory, automated verification, and safety permission controls through MCP tools, enabling better planning, context retention, self-checking, and guarded execution.1 npmMIT
- AlicenseAqualityBmaintenanceEnables coding agents to compact conversation contexts verbatim, make fast decisions through choice, boolean, and rubric scoring, and enforce command safety guardrails.6MIT
- AlicenseAqualityBmaintenanceEnables coding agents to make offline, zero-cost decisions using schema-safe Choice/Score/Noul primitives, a confidence gatekeeper, planning, adversarial red-teaming, research, and RLVR-based self-improvement via 21 MCP tools.242MIT