jev-ultrafast-mcp
This MCP server lets AI control a headless Chrome/Chromium browser via CDP, supporting whole-task handoff as well as manual, step-by-step browser automation.
Open URLs and get stable element reference tables.
Observe pages in full or delta mode, with optional text/JSON element data.
Execute batched actions: click, type, select, toggle, hover, upload, keys, scroll, nav, back/forward/reload, wait, screenshot, tab, and eval.
Run deterministic page assertions on URL, title, text, elements, values, checked state, and counts.
Record, replay, list, inspect, and delete zero-model macros with parameter placeholders.
Hand a high-level goal to a decision model for autonomous end-to-end execution (requires an API key).
Manage tabs and independent sessions; close sessions or shut down a launched browser.
Diagnose environment, browser connection, keys, and policy envelope.
Controls Google Chrome via Chrome DevTools Protocol to automate browser tasks end-to-end, including navigating pages, clicking elements, filling forms, and extracting page content.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@jev-ultrafast-mcpGo to example.com, order the blue shirt, and leave a review"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
⚡ jev-ultrafast-mcp - One Call Finishes Whole Browser Tasks
Welcome to jev-ultrafast-mcp – a tool that lets you hand an entire browser job to a smart helper and get it done in one single request. Instead of telling the computer click-by-clickhire a pro driver who takes the wheelhandles every stepand comes back with the result.
🤔 What Problem Does This Solve?
Imagine you want a computer to fill out a form or test a webpage. Normallyyou would need to give it dozens of tiny instructions: "click herewait now type thatclick submit check the text..." That process is slow and needs a new instruction for every action.
jev-ultrafast-mcp changes that completely. You describe the whole goal at once (like "go to this siteorder the blue shirt and leave a review") and the software figures out all the little steps by itself. It runs the browser on a server (not your screen) using a special decision-making brain that watches the page and clickscorrectly. The entire task costs just one call (one command) not one call for each click. That is a huge speed boost for anyone doing browser work with AI tools.
🧠 Core Benefits at a Glance
✅ One-call flow – Describe the whole job onceand it executes end-to-end automatically
✅ Ref-based element tables – The system tracks every button and box on the page with stable references (no missed clicks from shifting layouts)
✅ Code-checked assertions – After each actionit verifies the page changed correctly (like a self-checking robot)
✅ Zero-model macro replay – Re-run successful flows instantly without using any AI compute (super cheap)
✅ Runs over Chrome DevTools Protocol – Same professional protocol that powers Chrome's own debugging (reliable and standardized)
Related MCP server: surf-mcp
🚀 Getting Started (Windows)
Step 1: Go to the download page
Visit this link to download the application:
👉 CLICK HERE TO DOWNLOAD jev-ultrafast-mcp
After clickingyou will arrive at the project's homepage on GitHub. Look for a green "Releases" section on the right side (or a "Download" button). Click on the newest release file listed there (the file with.zip at the end). If you only see a "Code" green buttonthat is for developers – you want the Releases link instead.
Step 2: Save the file easy to find
Choose a simple spot on your computer to save the downloaded file – like your Desktop or Downloads folder. Do not worry about settings or options – just click Save when the browser asks where to put it.
Step 3: Open the downloaded file
Once the download finishesgo to your Desktop or Downloads folder and double-click the file you saved (it has a name like jev-ultrafast-mcp-v1.0.zip). A window will open showing a few files inside. This is a compressed folder–we need to take the files out.
Step 4: Extract (unzip) the folder
Inside that zip windowlook for the words "Extract All" in the toolbar at the topthen click it. A small window appears asking where you want the files placed. The default location (usually a new folder next to the zip file) is perfect. Click Extract. Your computer will create a normal folder with all the needed files inside.
Step 5: Look for the launcher file
Go into that new extracted folder. You should see a file named**start.bat** or run_windows.bat (or asimilar .bat file). That is the magic button. If you see a file named setup.exe insteaduse that one – but most often it will be a .bat file. Do not confuse this with the source code files (.py or .js files) – you only need to run the .bat file.
Step 6: Run the application
Double-click that .bat file. Windows may show a blue popup saying "Windows protected your PC" (SmartScreen). That is just Windows being cautious about a new file. Click "More info" and then "Run anyway" – this is safe because you downloaded it from the official project page. A black terminal window will open (that is normal) andin a few seconds your tool will be ready. That black window is the control panel – keep it open while you use the software. Close it when you are done.
Step 7: Connect it to your AI assistant the 30-second way
If you use Claude Desktop or Claude Code (Anthropic's assistant) this software automatically shows up in their list of available "tools"once it's running. You might need to restart theai app once after first setup. the
If you use Cursor or VS Code with an AI extensionopen the settings and add this line to the MCP (Model Context Protocol)servers section:
{ "mcpServers": { "jev-ultrafast-mcp": { "command": "your-full-path\start.bat" } } }(Replace your-full-path with the actual folder path you extracted.")
Step 8: Give your first task (try this simple test)
Once connectedtype this to your AI assistant:
"Use the browser tool to go to example.comand tell me the main heading text."
The AI will call jev-ultrafast-mcp onceand the software will handle opening the browser going to the pageand reading the text – all by itself. You will get the answer in just a few seconds. That is the whole experience: describe the goalget the result.
📋 Features Explained for Non-Programmers
🎯 One-Call Whole-Task Execution
You give one instruction ("order a pizza") not fifty steps. The built-in decision model looks at the starting pagefigures out what to clicktypepressand waitsover and over until the task is done. It's like giving a GPS your destination instead of turn-by-turn directions at every corner.
📊 Ref-Based Element Tables (No More "Element Not Found")
The system builds a live list of every buttonlinktext boxand image on the screenusing special IDs (called references). Those IDs stay stable even if the page changes halfway through. That means fewer failures and smoother execution – especially on webpages with moving parts (ads popping upanimation loading€).
✔️ Code-Checked Assertions (Self-Verifying Steps)
After every click or keypressthe software runs a mini-test to confirm the page actually changed as expected. If a click was supposed to open a dialog but nothing happenedthe system knows immediately and retries with a different approach. You never see a half-finished job with a "stuck" browser – you get a clean finish oran honest error.
🪄 Zero-Model Macro Replay (Replay Successful Flows for Free)
Let's say you just completed a perfect flow (like logging into a dashboard and downloading a report€). You can save that entire sequence as a "macro" – essentially recorded keystrokesand element clicks. Next time you need the same task you just replay it directly through the browser torchrome – with zero AI processing. That means it costs almost nothing in time and compute.The replay is exact because it uses the same element referencesnot fuzzy approximations.
🌐 Runs Over Chrome DevTools Protocol (CDP)
CDP is the official language that Chrome browsers speak to developers. Instead of using third-party hacks or flaky screen-simulationthis tool speaks directly to Chromium's core. That gives you rock-solid stability and full access to browser features (tabsnetwork logsJavaScript console€€without attaching to your visible desktop. The browser runs invisibly ona server (headless mode€ when you don't need to watch – fast and undetectable.
es without attaching to your actual remote computer)
🛠️ System Requirements (What You Need)
OS: Windows 10 or Windows 11 (64-bit recommendedbut 32-bit should work too)
RAM: At least 4 GB free memory (8 GB is better for heavy tasks such as multi-page forms or scraping)
Storage: 500 MB free disk space for the software and temporary browser files
Internet: A standard broadband connection (you need to talk to theai services like Claudeand also the browser needs to load pages€)
Pre-installed – Google Chrome or Microsoft Edge (the software will find your browser automatically; no need to do anything manual)
❓ Frequently Asked Questions (Quick Answers)
Q1: Do I need to know how to code to use this?
No. You only need the ability to run a .bat file (double-click a file)and write simple English sentences to an AI assistant. That's it. No Python knowledgeNo command line expertiseNo web development.
Q2: Is it safe to run the .bat file?
Yes. The file comes directly from this repository (the same page you downloaded). It only starts the tool locally on your machine. Windows might show a warning because it's a new program – just click "More info" → "Run anyway." The project is open-source so you0re free to inspect the contents if security is a top concern.
Q3: Which AI assistants work with it?
It is built on the Model Context Protocol (MCP) – that's the industry standard connection method. That means it works well with Claude (Desktop and Code) Cursor VS Code and many other modern AI code editors. If your AI tool supports "MCP servers" (most do these days) then you are good to go.
Q4: What if the task fails halfway through?
The built-in assertion system catches problems early and retries automatically. If it genuinely cannot finish (like the website blocks automation€) it will stop grand tell you exactly why. You can then adjust your instruction and try again. It doesn't waste your time pretending to work.
Q5: Can I run multiple browser tasks at the same time?
Technically yes – simply start the tool multiple times on different portsor use a tool manager that supports concurrency. For a typical userwe recommend sticking to one at a time to avoid confusion. The speed gain from one-call flow is usually enough.
📚 Realistic Example Use Cases
Form Filling & Submission – You give it a list of fields and values ("fill name as John Doe email as j@example.com" and the tool will click each fieldtype correctlysubmit and confirm the thank-you message. All in one request.
Web Scraping for Personal Projects – Ask it to "collect all product names and prices from this page and return them as a table." It runs the headless browseropens the sitescrolls throughand extracts structured dataandsends you back a neat list.
Multi-Step Workflow Automation( e.g., log into a dashboarddownload a CSV email it to yourself). Describe the whole chain oncethe tool clicks login enters credentials (you provide them in the prompt)loads the pageclicks downloadapplies a filterand then composes an email – all autonomously. It can also replay that exact flowlater for free using macros.
🧑🏫 Troubleshooting (3 Simple Fixes)
Issue: The .bat file opens and closes immediately (black window disappears)
Fix: Right-click the
.batfile → click "Edit" (it opens Notepad). Add the wordpauseat the very end of the fileand save.Closeand run again. This time the black window stays open showing the error. Copy that error and ask the AI assistant what it means – it will tell you exactly what missing (usually a path issue).
Issue: The AI assistant says it cannot find the "jev-ultrafast-mcp" tool
Fix: Firstdouble-check that the
.batfile is still running (you should see a black window open). If it closedtry running it again. Then restart your AI application (fully close andreopen – the list of MCP tools reloads on startup). once the window is activeand the AI should see it within a few seconds.
.
If you use Cursor or VS Codecheck your settings JSON that you typed the server path correctly.
🔍 Technical Snapshot (For the Curious)
Under the hoodthis software is a Python-based server that speaks the Model Context Protocol (MCP)to AI clients and the Chrome DevTools Protocol (CDP)to a Chromium browser instance. It maintains a two-way bridge: AI sends a high-level goalthe server breaks it down into a sequence of DOM operations (clickinputnavigate€)and each operation is verified with an assertion check. Key components include:
A decision model that parses the goal into a start state and a terminal condition usingzero-shot planning. It uses a combination of DOM heuristics and accessibility tree structures to identify the correct interactive elements.
A ref-table module that maintains a stable mapping from XPath+/CSS-paths to integer indices (the "ref-based" strategy for robustness against dynamic class names).
An assertion engine that compares DOM snapshots before/after actions using structural diffing to confirm success. .on
A macro engine that stores sequential (action, ref€» pairs in a JSON file for later deterministic replaying (zero-model cost).
🏁 Ready to Turbocharge Your Browser Tasks?
You have everything you need. It's a simple download extract double-click connect-and-go experience. No codingbackground necessary – just an ability to describe what you want done and let the smart driver handle the rest. Thousands of repetitive browser hours become secondsand single calls. Try it now – the onboarding takes under 5 minutes from download to first successful task. Stop baby-sitting a browser and start directing a whole fleet of digital assistants.
⬇️ Download jev-ultrafast-mcp Now(Opens the release page – download the latest .zip file from there)
Keywords: agent-tools, ai-agents, browser-agent, browser-automation, browser-use, cdp, chrome-devtools-protocol, claude, claude-code, cursor, headless-chrome, llm, mcp, mcp-server, model-context-protocol, playwright-alternative, python, vscode, web-automation, workbuddy
Available Tools
10 toolsbrowser_actA
Execute one or more ops in order, then return a delta observation.
Batch ops into a single call — each call is a round trip.
op fields click ref (ref may be "e12", or "e12" of a combobox to open it) type ref, text, [clear=true], [submit=false] select ref, value (option value or label) toggle ref, [state] (checkbox/radio/switch; no state = flip) hover ref upload ref, path | paths[] keys key ("Enter", "Meta+A", "ArrowDown") | keys[] scroll [dir=down|up|left|right], [amount=600], [ref] nav url back | forward | reload wait [ms=500] wait_for_ref ref, [timeout_ms=8000] wait_for_text text, [timeout_ms=8000] wait_for_load [timeout_ms=20000] screenshot [path], [full=false], [format=jpeg] (path names a file in the shots dir) tab action=list|new|switch|close, [index], [url] eval js (only when JEVMCP_ALLOW_JS=1)
Actions matching the confirmation rules (pay, delete account, …) return
needs_confirmation; re-send that op with "confirm": true to proceed. That
covers every op that clicks, not only the one named click -- toggle
presses the control too -- and the role it gates on is read from the
element the server observed, never from the op.
A bare single character in keys is text, not a key press -- it goes into
whatever has focus -- so it is refused (blocked_by_policy) on a field whose
value must not leave the page rather than confirmed. Use type with that
field's ref, which asks for "confirm": true and records {{secret}}.
| Name | Required | Description | Default |
|---|---|---|---|
| ops | Yes | ||
| dry_run | No | ||
| session | No | default | |
| observe_after | No | ||
| stop_on_error | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description goes far beyond the annotations by explaining confirmation behavior, covering ops beyond `click` that press controls, reading the role from the observed element rather than the op, and the `keys` single-character policy for sensitive fields. It also discloses environment-dependent behavior (`eval` only when JEVMALLOW_JS=1) and screenshot file conventions. This is rich behavioral context with no contradiction against the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but fully structured: a one-sentence purpose, a one-line batching rule, a tabular op reference, and focused prose for confirmation and key-handling edge cases. No sentence is filler, and the most important operational guidance is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity, this description is unusually complete: it covers all op kinds, their parameters, default values, confirmation semantics, sensitive-field behavior, environment-gated features, and output style. The presence of an output schema covers the return shape, and the annotations already declare the safety profile, so nothing essential is missing for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description carries nearly the entire semantic burden, and it does so well for the required `ops` array by documenting every op and its fields, defaults, and special values. It does not explain the optional top-level parameters (`dry_run`, `session`, `observe_after`, `stop_on_error`), although their names and defaults make their meaning reasonably inferable.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a clear verb and resource: 'Execute one or more ops in order, then return a delta observation.' It also lists the supported operation types in a compact table, making the tool's scope obvious. It does not explicitly distinguish itself from siblings like browser_macro, which could plausibly also execute sequences of actions, so it misses the top score.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit practical guidance: 'Batch ops into a single call — each call is a round trip,' which tells agents to consolidate actions to reduce latency. It also explains how to handle confirmation-requiring actions and how to properly enter text into protected fields. It does not state when to prefer an alternative sibling tool such as browser_macro or browser_goal.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_assertARead-onlyIdempotent
Verify the current page against deterministic checks. Returns pass/fail.
checks {"type": "url_matches", "pattern": "/checkout"} {"type": "url_contains", "text": "/orders/"} {"type": "title_matches", "pattern": "Order"} {"type": "text_contains", "text": "Thanks", "regex": false} {"type": "text_absent", "text": "Error"} {"type": "element_exists", "role": "button", "name": "Continue"} {"type": "element_gone", "ref": "e12"} {"type": "value_equals", "ref": "e7", "value": "Zurich"} {"type": "checked", "ref": "e9", "state": true} {"type": "count_at_least", "role": "link", "min": 3} {"type": "js", "expr": "document.title.length > 3"}
| Name | Required | Description | Default |
|---|---|---|---|
| checks | Yes | ||
| session | No | default |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and idempotentHint=true, so the description correctly adds value by emphasizing 'deterministic' checks and 'Returns pass/fail'. The enumerated check types materially expand behavioral transparency because the schema otherwise reveals only an opaque array of additionalProperties-objects. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The two-sentence intro is front-loaded with purpose and outcome, and the example block is compact and information-dense. Each line illustrates a distinct check type with no filler or redundant words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description is complete enough for a tool with a polymorphic checks parameter and an output schema. It covers all major assertion categories and return behavior. Minor gaps remain around exact pattern-matching semantics and the precise pass/fail output structure, but those are not critical for selecting and invoking the tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the checks items are 'additionalProperties: true', giving no semantics. The description fully compensates by presenting 11 concrete check object shapes with field names like pattern, text, ref, role, name, min, and expr. The session parameter is adequately covered by its schema default.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('verify') and names the resource ('current page'), and explicitly states the return contract ('Returns pass/fail'). The extensive list of concrete check types makes it unmistakable what the tool does and differentiates it from siblings like browser_observe or browser_act.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly implies this tool is for deterministic assertions on the current page, and the check examples show common scenarios. However, it never explicitly states when to prefer this over siblings such as browser_observe, browser_act, or browser_goal, nor does it mention cases where it should not be used.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_closeA
Close a session's tab. Set shutdown_browser=True to stop the browser too.
Only a browser this server launched is stopped. In attach mode the browser is yours: shutdown detaches and leaves it, and every other window, running.
| Name | Required | Description | Default |
|---|---|---|---|
| session | No | default | |
| shutdown_browser | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the annotations (readOnlyHint=false, destructiveHint=false), the description reveals a non-obvious behavioral trait: only a server-launched browser is actually stopped, while attach-mode browsers are left running. This is valuable context; it could go further by stating what happens to the session after the tab closes, but the key caveat is present.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is tight and front-loaded: one clear action, one conditional flag, and one caveat. Every sentence earns its place, with no redundant filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given its moderate complexity, an output schema, and annotations, the description covers the main behaviors an agent needs: what is closed, when the browser stops, and how attach mode differs. The main gap is the undocumented session parameter, which keeps this from a 5.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description must carry parameter meaning. It clearly explains shutdown_browser=True and its attach-mode limits, but it leaves the session parameter undefined beyond the schema title 'Session', so an agent must infer what a session is and what values are valid.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening sentence names a specific verb and resource: it closes a session's tab, with an optional flag to stop the browser. This is clear and not a restatement of the tool name, but it does not explicitly differentiate the tool from siblings such as browser_tabs or browser_act.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives conditional guidance for shutdown_browser and explains attach-mode behavior, but it never states when to prefer browser_close over browser_tabs, browser_act, or other siblings. The intended use is implied rather than explicitly contrasted.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_doctorARead-onlyIdempotent
Report environment: browser binary, connection, keys, and policy envelope.
Call this when anything behaves unexpectedly — it separates "no browser" from "blocked by policy" from "no key".
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and idempotentHint=true, so the description doesn't need to restate safety. It adds context by revealing that the tool's behavior is to report environment details and triage failure causes, which is beyond what annotations convey. No contradictions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences with zero fluff. The core purpose is front-loaded in the first sentence, and the usage guidance is in the second. Every word earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a diagnostic tool with no parameters, an existing output schema, and annotations covering safety, the description fully equips an agent to decide when to call it and what it can expect. Nothing essential is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so the description doesn't need to explain any. The baseline for 0-param tools is 4, and there is no missing parameter information to compensate for.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Report') and a specific resource ('environment') plus the scope of that report (browser binary, connection, keys, policy envelope). It clearly distinguishes this diagnostic tool from action-oriented siblings like browser_open, browser_act, and browser_close.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly says when to call the tool: 'when anything behaves unexpectedly.' It also clarifies the diagnostic value by separating failure modes. It does not explicitly name alternatives or when not to use it, but the context is clear enough for an agent to route correctly.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_goalA
Hand a whole browser task over. Needs a decision-model key.
This is the entry point for browser work, not an optimisation on top of the
manual loop. Pass url and the goal and the page is opened and driven to
the end here: one call, one host turn, instead of a turn per click.
Leave url out when the task continues from a page an earlier step left
behind; the goal then runs against whatever the session is already showing.
Each step costs one request (operation + every target head in a single
speculative fan-out). verify runs browser_assert-style checks on the final
page, so the result is a fact rather than a model's claim of success.
The decision model is reachable through two APIs and both are supported here.
JEV_PROVIDER=typesafe uses Jev's own API with TYPESAFE_API_KEY;
JEV_PROVIDER=openrouter routes the same model through OpenRouter with
OPENROUTER_API_KEY and no TypeSafe account. Same model, same contract either
way. Reading a page needs no key at all, so when no key is set the handoff is
unavailable while browser_open, browser_observe and browser_act keep working.
| Name | Required | Description | Default |
|---|---|---|---|
| url | No | ||
| goal | Yes | ||
| verify | No | ||
| session | No | default | |
| verbose | No | ||
| max_steps | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the annotations, it discloses per-step request cost with speculative fan-out, that verify runs browser_assert-style checks producing facts rather than model claims, and the two provider API routes plus no-key behavior. No contradiction with the readOnly/openWorld/destructive hints.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is relatively long but every paragraph earns its place: purpose, usage, cost, verification, and auth. It is front-loaded with the core idea and organized so an agent can quickly extract selection and invocation guidance.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Together with the output schema and annotations, the description gives enough context for selection, invocation, auth requirements, cost behavior, and fallback options. It is only slightly incomplete because a few auxiliary parameters like session and max_steps are left without explicit semantics.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description adds real meaning for url (include or omit), goal (the whole task), and verify (final-page assertions). However, it does not explain session, verbose, or max_steps, so it only partially compensates for the schema's lack of parameter documentation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states explicitly that the tool hands a whole browser task over, drives the page to completion, and is the entry point for browser work rather than an optimization on top of the manual loop. This clearly distinguishes it from siblings like browser_open, browser_observe, and browser_act.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly positions when to use this tool versus the per-click manual loop, names the fallback tools that keep working when no key is set, and gives a clear condition for omitting url when continuing from an existing session. This is actionable usage guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_macroA
Record, replay, list, or delete a macro — a discovered path with no model calls.
action="record_start" begin capturing ops (needs the session to be driving the task)
action="record_stop" finish and save under name
action="run" replay name; params fills {{placeholders}} in text/url
action="list" | "inspect" | "delete"
A field the page marks as a secret is never written to the macro: its text is
stored as the placeholder {{secret}}, so pass params={"secret": "…"} to
replay it. Leaving that out types the placeholder literally, which fails
visibly rather than leaking — and the extension's replay panel asks for the
same field, because it is an ordinary placeholder.
Replay re-resolves each step by role + name against a fresh observation and refuses to act when the best match is weak or ambiguous.
| Name | Required | Description | Default |
|---|---|---|---|
| goal | No | ||
| name | No | ||
| action | Yes | ||
| params | No | ||
| session | No | default | |
| start_url | No | ||
| threshold | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the sparse annotations (readOnlyHint=false, openWorldHint=true, destructiveHint=false), the description discloses that secrets are never written to the macro and are stored as {{secret}} placeholders, that omitting them 'fails visibly rather than leaking', and that replay refuses to act on weak or ambiguous matches. These are meaningful safety and failure-mode disclosures the annotations do not provide.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description front-loads the core purpose and then organizes into action definitions, secret handling, and replay semantics. It runs long, but each sentence adds information, and the secret-handling and weak-match-refusal passages earn their length given their safety importance.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 7-parameter, multi-action tool, the description explains the essential usage contract: actions, placeholder mechanics, and the safety/ambiguity rules. An output schema exists to cover return values, so the main remaining gap is that several parameters (goal, session, start_url, threshold) are left to an input schema with 0% prose coverage.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description carries the burden. It richly documents `action` (five values with their effects), `name` (save/replay target), and `params` (placeholder substitution including the {{secret}} case). It leaves `goal`, `session`, `start_url`, and `threshold` unexplained, though the weak-match refusal partially hints at threshold semantics.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening line 'Record, replay, list, or delete a macro — a discovered path with no model calls' gives specific verbs, a resource, and a defining trait. The 'no model calls' framing distinguishes this from sibling tools like browser_act and browser_observe, which drive live model-mediated interactions.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives per-action context: record_start 'needs the session to be driving the task', run replays a name with params filling placeholders, and list/inspect/delete are enumerated. It doesn't explicitly name sibling alternatives or state when-not-to-use, but 'a discovered path with no model calls' implies the use case that separates this from live-action tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_observeARead-onlyIdempotent
Re-read the page: new element table, or a delta if little changed.
mode: "auto" (delta when possible), "full" (whole table, e.g. after a big
change), "delta" (force). include_json=True appends a machine-readable
copy of the element table when you want to plan over it programmatically.
| Name | Required | Description | Default |
|---|---|---|---|
| mode | No | auto | |
| session | No | default | |
| include_json | No | ||
| include_text | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already mark the tool as read-only and idempotent, and the description adds meaningful behavioral details: the page is re-read, the result may be a full element table or a delta depending on changes, and include_json appends a machine-readable copy. This gives the agent an accurate picture of what to expect without contradicting the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded: the core behavior appears in the first line, followed by concise parameter guidance. Every sentence adds useful information, and there is no filler or repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description, annotations, and output schema together cover the important aspects: read-only behavior, mode selection, and return-style options. The main gaps are the undocumented session and include_text parameters and the lack of an explicit pointer to sibling tools, but these are lower-risk because the core usage is clear.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description carries the burden. It explains mode and include_json well, but it does not cover session or include_text, leaving those two parameters undocumented in both schema and description. It partially compensates for the schema gap but not fully.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with 'Re-read the page' and names the concrete output ('new element table, or a delta if little changed'), making the operation and resource clear. This clearly differentiates it from siblings like browser_act or browser_assert, which are action/assertion tools rather than observation tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly explains when to use each mode: 'auto' for delta when possible, 'full' after a big change, and 'delta' to force. It also tells the agent when to set include_json=True, namely when planning programmatically over the element table.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_openA
Open a URL in a new owned tab and return the element table.
Use hint to restate the goal in one line; it is echoed back so the next
step has the goal in context without re-reading this call.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | ||
| hint | No | ||
| session | No | default |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already signal this is not read-only and may have external effects. The description adds useful context about creating an 'owned tab' and echoing back the hint, but it does not disclose details like behavior when the URL fails, whether the current tab changes, or session implications. Acceptable but not rich.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tightly written sentences. The first front-loads the core behavior, and the second adds targeted guidance on `hint` without wasting words. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple open action with an output schema and annotations, the core behavior is covered. However, the description omits the `session` parameter semantics and does not clarify how this relates to the sibling tools, so an agent may still have unanswered questions about multi-tab or multi-session workflows.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It explains `hint` well but says nothing about `url` or `session`. The `url` is required and only typed as string; the description does not add format, protocol, or usage guidance. Only one of three parameters gains real semantic value from the text.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'Open a URL in a new owned tab and return the element table.' This clearly distinguishes it from sibling tools like browser_tabs or browser_observe by emphasizing the new owned tab and the returned element table.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description makes clear this is the entry action for opening a URL and gives specific guidance for using the `hint` parameter. It does not explicitly name alternatives or exclusions, but the open-and-return-table purpose is unambiguous enough for an agent to select it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_sessionsARead-onlyIdempotent
List open sessions (independent owned tabs).
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint and idempotentHint, so the safety profile is covered. The description adds a bit of conceptual context by explaining that sessions are 'independent owned tabs,' but it does not disclose return format, pagination, or any other behavioral nuance.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One short sentence with no filler or redundant wording. The key action and resource definition are front-loaded, and every word earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter, read-only listing tool with an output schema present, the description is nearly complete. It would be stronger if it explicitly mentioned the relationship to browser_tabs, but nothing needed to make the call itself is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has no parameters, so the description does not need to explain parameter semantics. Baseline 4 applies for a zero-parameter tool.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('List'), a resource ('open sessions'), and defines that resource as 'independent owned tabs,' which distinguishes it from the sibling browser_tabs. An agent can tell what this tool does and roughly what it returns.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is given about when to use this tool versus alternatives such as browser_tabs. The parenthetical helps define the concept but does not explicitly state when this tool should be selected or when another tool is preferred.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_tabsA
List, open, switch to, or close tabs.
Tabs opened by the page show up in observations on their own. To act on one,
prefer target_id (the #handle printed by action="list"): indexes are
positional and get renumbered whenever the tab list changes, so an index
read a call ago can address a different tab. index is a convenience when
listing and acting in the same breath; omit both to mean "the current tab".
url is checked against the domain envelope, like every other way of
choosing a destination.
| Name | Required | Description | Default |
|---|---|---|---|
| url | No | about:blank | |
| index | No | ||
| action | No | list | |
| session | No | default | |
| target_id | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description adds meaningful behavioral context beyond the annotations: tabs opened by the page appear in observations automatically, indexes are positional and get renumbered, `target_id` is a `#handle`, and `url` is validated against the domain envelope. This goes well beyond the readOnly/openWorld/destructive hints. It is not contradictory to the annotations; closing tabs is consistent with destructiveHint=false since no data destruction is implied.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact, front-loaded with the primary purpose, and each subsequent sentence earns its place by clarifying tab handles, index semantics, current-tab behavior, and URL validation. There is no filler or repetition of schema data.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description explains the trickiest aspects of tab management and an output schema exists, so return values are not required. But it still omits the allowed `action` values (beyond inferring them from the first sentence) and leaves `session` undefined, so an agent cannot fully determine correct invocation for all intended uses.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description carries the burden of explaining parameters. It does this well for `target_id`, `index`, and `url`, explaining the handle mechanism, positional instability, and domain-envelope check. However, it never documents `action` values beyond `action="list"` or explains the `session` parameter at all, leaving meaningful gaps for a 5-parameter tool.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a clear verb-resource pair: 'List, open, switch to, or close tabs.' This makes the tool's domain obvious and distinguishes it as tab-focused from siblings like browser_open or browser_observe. It loses the top score because it advertises multiple operations rather than one specific verb and does not explicitly contrast itself with overlapping siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives concrete direction on how to use the tool: prefer `target_id` over `index`, use `index` only when listing and acting in the same breath, and omit both to mean the current tab. It also explains the domain-envelope check for `url`. It stops short of explicit when-not-to-use or alternative-tool routing, so it earns a 4 rather than a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
10 tool updates
v0.1.5- First observed
browser_act - First observed
browser_assert - First observed
browser_close - First observed
browser_doctor - First observed
browser_goal - First observed
browser_macro - First observed
browser_observe - First observed
browser_open - First observed
browser_sessions - First observed
browser_tabs
TDQS
Scored across 10 tools
Each tool has a clear primary purpose, but there is some overlap: browser_act includes tab operations that browser_tabs also handles, and browser_close overlaps with tab-closing in browser_tabs and browser_act. The descriptions mostly disambiguate these contexts well, though an agent might occasionally hesitate between them.
All tools share a consistent browser_ prefix in snake_case, which makes the family obvious and predictable. The second part mixes action verbs (open, observe, act, assert, close) with nouns (macro, tabs, sessions, doctor), so it is not a strict verb_noun pattern but remains readable and mostly consistent.
Ten tools is well-scoped for a browser automation server. Each tool occupies a meaningful role: navigation, observation, interaction, verification, tab/session management, recording, autonomous task handling, and diagnostics. None feel redundant or missing.
The tool surface covers the full browser workflow: open, observe, act, assert, manage tabs and sessions, close, record/replay macros, delegate to a decision model, and diagnose environment issues. There are no obvious dead ends or missing operations for the stated domain.
Maintenance
Related MCP Connectors
AI-powered browser automation — navigate, click, fill forms, and extract data from any website.
AI-powered web automation. Navigate websites using AI agents for one page or a thousand
AI-powered web automation. Navigate websites using AI agents for one page or a thousand
- openhelmOAuthai.openhelm
Autonomous cloud agent tasks: real browser + your tools, structured evidence-backed results.
Related MCP Servers
- FlicenseNot gradedqualityCmaintenanceEnables AI assistants to control a browser through a set of tools, allowing them to perform web automation tasks like navigation, typing, clicking, and taking screenshots.-
- AlicenseBqualityDmaintenanceEnables visual browser automation through natural language descriptions, allowing AI to click, type, and navigate web pages by seeing the page.18MIT
- FlicenseNot gradedqualityCmaintenanceEnables AI assistants to control a browser through a set of tools, allowing them to perform web automation tasks like navigation, typing, clicking, and taking screenshots.-
- FlicenseNot gradedqualityCmaintenanceEnables AI assistants to control a browser through tools for navigation, typing, clicking, and taking screenshots to perform web automation tasks.-