jev-ra
This server gives an agent full control over a shared browser session to open pages, pursue goals, and extract or manipulate page content.
browser_open: open a URL and get a page summarybrowser_run: pursue a complete goal, optionally supplying typed values and a step budgetbrowser_act: take a single decided step toward an instructionbrowser_search: search the web, read top results, and rank them against a goalbrowser_observe: list visible controls and page textbrowser_extract: pull structured data—text, elements, links, tables, or main contentbrowser_click: click an element by its refbrowser_type: type text into a field by its refbrowser_select: choose a dropdown optionbrowser_scroll: scroll the page up or downbrowser_press: press Enter, Escape, or Tabbrowser_wait: wait briefly and re-observe the pagebrowser_screenshot: capture the current viewport as a JPEGbrowser_close: close the server's browser session
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@jev-raOpen the Hacker News front page and list the top 5 stories"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.

Site: brnyxx.github.io/jev-ra replays a real recorded run and explains the pipeline.
jev-ra
A fast browser-use layer for CLI coding agents. Claude Code, Codex, or any MCP client hands jev-ra a goal. TypeSafe Jev, a System One decision model, picks the operation and the target element for every step in one round trip. Your agent plans, supplies the text values, reads what the page says, and takes over when jev-ra escalates. No second LLM runs inside the loop.

task | browser-use 0.13.10 + gemini-3-flash | jev-ra | |
Wikipedia: open the Gödel incompleteness article | 23,058 ms | 2,714 ms | 8.50× |
Google Flights ZRH→LON one-way, results on screen | 66,414 ms | 8,888 ms | 7.47× |
Olive Young category: sort by 신상품순 | 15,071 ms | 3,806 ms | 3.96× |
Measured 2026-09-18 on one machine and one dedicated Chrome, both tools through OpenRouter. jev-ra is the median of 5 runs; browser-use is its single recorded run, which was faster than its own 5-run median on every task. Each jev-ra run was verified against the final page; 25 of 25 passed with no text-model calls. Method, p90, cost and raw rows.
Quick start
Claude Code
export OPENROUTER_API_KEY=sk-or-...
uvx jev-ra install claude
# then, in Claude Code: "open wikipedia.org and find the Gödel incompleteness article"Codex
export OPENROUTER_API_KEY=sk-or-...
uvx jev-ra install codex
# then, in Codex: "use jev-ra to open wikipedia.org and find the Gödel incompleteness article"Shell
export OPENROUTER_API_KEY=sk-or-...
uvx jev-ra doctor
uvx jev-ra run https://en.wikipedia.org/wiki/Main_Page "Open the Godel incompleteness article." \
--value "search_query=Godel incompleteness theorems"No Python setup? npx -y jev-ra install claude does the same thing through the npm launcher. The
npm package is a launcher only: it finds uv, offers to install it, and runs the PyPI package pinned
to its own version.
There is no install step either way: uvx runs jev-ra straight from PyPI and registers uvx jev-ra mcp as the server command. For a permanent copy, uv tool install jev-ra. The key is forwarded
from the variable you already exported and is never printed.
No Chrome on the machine either? The repository's Dockerfile builds an image with a Chromium in
it: docker build -t jev-ra . then docker run --rm -e OPENROUTER_API_KEY jev-ra doctor. Chrome's
own sandbox needs a user namespace a container's seccomp profile usually refuses, so jev-ra
launches it once, reads what it said, and starts it again without the sandbox when that is why it
would not run.
Related MCP server: Playwright MCP Server
How it works
One decision per step. The only text typed into the page is text you supplied.
MCP tools
tool | arguments | what it does |
| url | Open a URL in the shared session and summarise the page. |
| goal, values?, max_steps?, resume? | Pursue a whole goal. Supply values for anything that must be typed. |
| query, goal?, max_pages? | Search, read the best results in parallel tabs, rank them against the goal. |
| instruction, values? | Take one decided step towards an instruction. |
| max_elements? | List the observed controls and the visible text. |
| mode? | Structured page data: |
| ref, page_key? | Click one observed element by its ref. |
| ref, text, page_key? | Type into one observed field. |
| ref, option, page_key? | Select an observed dropdown option. |
| direction? | Scroll one viewport step up or down. |
| key | Press Enter, Escape or Tab. |
| - | Wait a moment and observe again. |
| - | JPEG of the current viewport. |
| - | Close the session held by the server. |
Every response carries elapsed_ms, and decisions plus cost whenever Jev was called.
browser_observe returns a page_key; pass it back to browser_click, browser_type or
browser_select and a ref from a page that has changed since is refused as stale instead.
CLI
command | what it does |
| pursue a goal from a URL until it is done or escalates |
| search the web and read the best results |
| open a URL and keep the session for later commands |
| list the controls and text of the open page |
| pull structured data out of the open page |
| take one decided step on the open page |
| click one observed element |
| type into one observed field |
| select an observed dropdown option |
| scroll the open page |
| press Enter, Escape or Tab |
| wait a moment and observe again |
| save a JPEG of the viewport |
| close the session kept by |
| stop what jev-ra started and empty its profile |
| print where a stored run's time went, step by step |
| render a stored run by its run id |
| run the MCP stdio server |
| run the same tools over HTTP and SSE |
| print the agent guide, for saving as a skill file |
| register jev-ra as an MCP server with a coding agent |
| check the key, the endpoint, Chrome and one live decision |
| time the offline fixtures, and the live tasks with --live |
| run the real-site corpus |
open … close share one browser across invocations through a target id in
$XDG_STATE_HOME/jev-ra/session.json. Add --json to any command for the raw payload.
Every run is stored under its own id in $XDG_STATE_HOME/jev-ra/runs/, newest 200 kept.
jev-ra trace RUN_ID prints its step table; --html writes a single-file page - the run's own
JSON inline, nothing to fetch - to send to whoever asked what the run did.
--profile NAME gives a run a Chrome user-data-dir of its own under
$XDG_STATE_HOME/jev-ra/chrome-profiles/NAME, so a login done once on that profile is still
there on the next run; open records the profile and the stateful commands reattach to it. The
MCP tool takes the same thing as browser_open(url, profile). One process drives one browser, so
a server already on a profile refuses a second one instead of answering from the wrong cookies.
serve is the same tool set over MCP's streamable HTTP transport, at /mcp, answering with
server-sent events; /healthz answers without a key. Every request carries a key from
JEV_RA_SERVE_KEYS as Authorization: Bearer <key>, is charged the decisions it spends against
that key's daily quota, and leaves one JSON line on stderr with the run id it also returns in the
x-jev-ra-run-id header. A key is never logged: the line names it by a digest.
Python
from jev_ra import Agent
with Agent() as agent:
result = agent.run(
"Place the order with express shipping.",
values={"name": "Ada Lovelace", "email": "ada@example.com"},
url="https://example.com/checkout",
)
print(result.status, result.elapsed_ms, [step["target_label"] for step in result.steps])Values
TYPE_TEXT needs a string, and jev-ra will not invent one. Jev picks which of your values belongs in
the field it is about to fill, in the same round trip that picks the field. If nothing fits and no
text helper is configured, the run stops with needs_value and reports the field's label, role and
current value. You supply the value and call again. The default install has no text model.
When it hands control back
Result.status is done, blocked, escalate or budget. When a run stops short, reason is one
of needs_value, stuck_loop, unverified_done, stale, invalid_decision, too_many_controls,
provider_error, blocked, blocked_by_site or needs_human. A run that ends on budget names the budget it hit (steps, decisions
or time) in reason instead; provider_error is the provider refusing to answer at all, so check
the key and the route rather than retrying the goal. An escalation
also carries the top eight operation/target candidates with their probabilities, and up to 3,000
characters of page text — enough to decide what to do without observing again.
A goal that reads as a question — it ends in a question mark, or opens with what, which, how many,
when, who or find the — also gets Result.final_answer: one sentence the configured text helper
takes from the page the run finished on, whatever the run ended as. With no text helper it stays
null and detail says so, and a goal that is an instruction never asks for one.
Verification is deterministic: after every action jev-ra compares url, title, text and field state,
and page_changed comes from a semantic page marker, not from the model.
While the page settles after an action, jev-ra asks Jev the next question already, against the page
as it should read with that input applied and nothing else changed. If the settled page offers
anything the guess did not, the answer is thrown away and the question asked again.
Result.speculations and Result.prefetched count how often that paid off.
A site that answers with an error page (HTTP 5xx or 429, or a short page that says so) is waited out
for two seconds and reloaded once before anything is decided on it; if it still answers with an
error, the run stops with blocked_by_site and detail.wall names the status ("http 502"), so a
site's bad minute is never mistaken for a page to act on.
A site that asks for a person - a CAPTCHA widget, Cloudflare's "Just a moment...", a press-and-hold
check - is handed to one. jev-ra backs off and asks for the page once more, as it does for an error
page; if the check is still there, it brings the Chrome window to the front, raises a desktop
notification naming the site (JEV_RA_NOTIFY=0 turns it off), and waits up to 120 s
(JEV_RA_HUMAN_WAIT_S) for the page to stop being a check. Then the same run carries on, with the
same goal, history and values and the steps it has left. The wait is reported apart as
Result.human_wait_ms and spends no step and no time budget. If nobody clears it in time, the run
stops with needs_human: detail names the site and the check and carries resume, the run id.
Once the person has cleared it, browser_run(resume=...) or jev-ra run --resume ID carries the run
on from that page instead of starting over. A headless browser, or one at a remote CDP address, has
nobody to show the check to, so the run stops with needs_human at once and detail.next_step says
how to rerun where a person can see it. A site that refuses outright - "Access denied", a 403 with no
check, a geo block - is still blocked_by_site, with detail.kind set to refusal. jev-ra never
solves a check, never hides that it is automated, and never borrows cookies from another profile.
The navigations a run starts - its first open and the one reload it may spend - are paced per host:
at least 1 s apart (JEV_RA_PACE_S), and a host that answered 429 or put up a check makes the next
one wait twice as long each time in a row. A run opens its address once, so an ordinary run is never
held, and the machine the run is on is never paced.
Benchmarks
Five tasks, five runs each, every run verified against the page it left behind. Measured 2026-09-18 through OpenRouter, on the same machine and Chrome as the browser-use rows:
task | median | p90 | success | decisions | cost | ratio |
Wikipedia article | 2,714 ms | 3,179 ms | 5/5 | 3 | $0.00075 | 8.50× |
Google Flights search | 8,888 ms | 10,573 ms | 5/5 | 14 | $0.00317 | 7.47× |
Olive Young sort | 3,806 ms | 4,858 ms | 5/5 | 4 | $0.00204 | 3.96× |
Search with a citation | 2,416 ms | 2,571 ms | 5/5 | 4 | $0.00035 | no baseline |
Local checkout form | 2,191 ms | 2,338 ms | 5/5 | 5 | $0.00049 | no baseline |
Ratios are against browser-use 0.13.10 + gemini-3-flash flash_mode, its single recorded run of each
task on the same day, machine and Chrome, also through OpenRouter: 23,058 ms, 66,414 ms and
15,071 ms respectively. Text-model calls across all 25 runs: 0. A same-harness re-run of
browser-use, five runs per task that day, was slower still: 9.07×, 8.31× and 7.26×. Against the
fastest of browser-use's six runs of each task (15,759 ms, 49,914 ms and 15,071 ms), our median is
5.8×, 5.6× and 3.96×; no ratio measured that day is below 3.96×.
jev-ra bench --live --runs 5 reproduces this table and prints PASS/FAIL against the v0.1 bar of
≥ 3× on every task with a baseline. On 0.2.5 through the TypeSafe direct route (2026-09-23) the
first three tasks took 4,681 ms, 11,603 ms and 5,675 ms; browser-use was not re-run that day, so
those times are not a like-for-like ratio. Method, the browser-use rows, and how to reproduce
them.
jev-ra on the left, browser-use flash_mode on the right, same task, same Chrome, real time:

A run that finishes without doing the task counts as a failure, not as a time.
Accuracy on real sites
The corpus is 83 tasks on real sites in ten families (search, e-commerce, booking, forms, docs,
news, portals, auth walls, Japanese and Chinese sites), each with a spec that checks the page the
run left behind. On 0.2.4, three runs each through the TypeSafe direct route on 2026-09-23:
213 / 249 = 85.5 %. On the forty tasks both tools were given, browser-use 0.13.10 flash_mode
passed 29 / 40 = 72 % in one run on 2026-09-22, with a median of 19.4 s on the runs that
passed; jev-ra 0.1 passed 102 / 120 = 85 % in three runs on 2026-09-18, median 3.1 s. Those
two rows differ in day, run count and jev-ra version, so they are not a like-for-like comparison.
A single three-run pass moves by about five tasks on site weather alone, so a change counts only
when a per-task rerun agrees.
Per-task rows and the noise measurement.
What it will not do
limit | what happens |
Canvas drawing, games, anything painted rather than marked up |
|
File upload |
|
CAPTCHA and other checks a person can clear | handed to the person at the window; |
Bot walls that refuse outright, stealth |
|
Auth flows |
|
Pop-up windows, multi-tab workflows | the run stays on its own target |
Cross-origin iframes | reported as one opaque element; open shadow roots and same-origin iframes are traversed |
More than 250 visible controls |
|
Each returns an escalation with the page text and the ranked candidates.
FAQ
OpenRouter or a TypeSafe key? Either. jev-ra resolves JEV_RA_API_KEY, then TYPESAFE_API_KEY,
then OPENROUTER_API_KEY. A key starting sk-or- selects the OpenRouter route
(typesafe/jev-1.13); anything else goes direct (jev-latest). JEV_RA_ENDPOINT and
JEV_RA_MODEL override both. OpenRouter is easier to get. Which route decides faster has not been
measured under the same conditions: the upstream jev-ultrafast
recording has a 178 ms
median decision on the direct route, and jev-ra's recorded Flights run a 296 ms median through
OpenRouter (2026-09-18), on different machines and days.
What does a task cost? Through OpenRouter on 2026-09-18, between $0.00035 (a search, 4 decisions) and $0.00317 (the whole Google Flights flow, 14 decisions). Cost scales with decisions, not with page size, because the state sent is the element table and the visible text, never the HTML.
Does it need its own Chrome? It will find or launch one on its own profile
($XDG_STATE_HOME/jev-ra/chrome-profile) and reuse it; when the Chrome it launched dies, the next
command starts another. Point BU_CDP_URL at a different Chrome to override - if nothing answers
there, jev-ra says so rather than launching one behind your back. Do not point it at a browser
signed into anything you would not let an agent operate.
Why no text model? The host agent already has the context. A second model costs one call per
field (675-938 ms per mercury-2.5 call through OpenRouter, five calls measured on 2026-09-18 for
the design) and writes values nobody
supplied. You can still configure one with JEV_RA_TEXT_MODEL.
Configuration
variable | effect |
| key, in that order of precedence |
| override the route |
| path to the browser binary to launch |
| an existing Chrome to drive instead of launching one |
| e.g. |
| the language the browser asks sites for; default |
| budgets (40 / 80 / 120) |
|
|
| egress for a Chrome jev-ra launches, e.g. |
| least seconds between two navigations a run starts on one host; default |
|
|
| how long a run waits for a person to clear a human check; default |
|
|
| search endpoint template, |
| optional text helper, off by default |
| comma list of API keys |
| decisions per key per day for |
| log level for stderr, e.g. |
|
|
$XDG_CONFIG_HOME/jev-ra/config.json sets the same keys; the environment wins.
Credits
jev_ra/browser/snapshot.js and the NEXT_ACTION / TARGET instruction texts are adapted from
browser-use/jev-ultrafast (MIT), where they were
measured. Chrome is driven through
browser-harness (MIT).
See THIRD_PARTY_NOTICES.md.
MIT licensed. Contributing · Security · Agent guide · Usage reference · 한국어
Available Tools
14 toolsbrowser_actC
Take one decided step towards an instruction on the current page.
| Name | Required | Description | Default |
|---|---|---|---|
| values | No | ||
| instruction | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already convey readOnly=false, openWorld=true, idempotent=false, and destructive=false, so the description need not restate those. It adds some useful behavioral context by saying 'one decided step' and 'current page,' but it still does not disclose what types of actions may occur, whether navigation or form submission may happen, or what the tool returns.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single short sentence with no filler, and the key idea 'one decided step' is front-loaded. It is concise, though the phrasing 'towards an instruction' is slightly awkward and the sentence carries less useful content than it could.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a general browser action tool, this description is not complete enough. It leaves unclear how to phrase instructions, what values means, what observable result the agent can expect, and how browser_act relates to the many specific sibling browser_* tools.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate for missing parameter documentation, but it does not explain the required 'instruction' beyond using the same word. It also says nothing about the optional 'values' object, which is opaque and could represent form values, options, or something else.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description identifies the action as taking one step toward an instruction on the current page, so it communicates a general purpose rather than being a pure tautology. However, it is vague about what the step concretely does and does not distinguish browser_act from sibling tools like browser_click, browser_type, or browser_scroll.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use browser_act versus any of the specific sibling tools. The description only implies that the tool executes a single decided step, but it never says how to break down instructions, when a specific sibling should be preferred, or when browser_act should not be used.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_clickA
Click one observed element by its ref. Pass the page_key from browser_observe to refuse a stale ref.
| Name | Required | Description | Default |
|---|---|---|---|
| ref | Yes | ||
| page_key | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description adds a behavioral nuance about stale refs and how page_key addresses that, which goes beyond the annotations. However, it doesn't disclose side effects (e.g., navigation, state changes) or error conditions. Annotations already indicate readOnlyHint=false and destructiveHint=false, so the safety profile is partially covered, but the description adds limited behavioral context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences with no fluff, front-loading the action and the key parameter. Every word earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple click tool, the description covers the essentials: what to click and how to avoid stale refs. It doesn't mention return values (no output schema) or failure modes, but given the tool's simplicity and the presence of sibling tools like browser_observe, it's reasonably complete. A mention of when not to use it (e.g., if you need to run JavaScript) would improve it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema provides no descriptions (0% coverage), so the description must compensate. It explains 'ref' as referring to an observed element and clarifies the purpose of 'page_key' (to refuse stale refs). This adds meaningful context beyond the schema, though it doesn't specify formats or types.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Click one observed element') and the required resource ('by its ref'), with a specific verb and resource. It distinguishes itself from siblings like browser_act by focusing on clicking a specific observed element and referencing browser_observe as the source of the ref.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives a hint about when to use the page_key parameter ('Pass the page_key from browser_observe to refuse a stale ref'), but it doesn't explicitly state when to use this tool versus alternatives like browser_act, nor does it provide exclusions or conditions for not using it. The guidance is implied rather than explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_closeAIdempotent
Close the browser session held by this server.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already provide idempotentHint=true and destructiveHint=false. The description adds only the 'held by this server' qualifier, which clarifies scope but does not disclose side effects like releasing resources or invalidating subsequent browser tool calls. Given the simplicity of the operation and annotation coverage, a modest score is appropriate.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single sentence with no wasted words. It states the action and scope directly and is immediately understandable. There is no redundancy with the schema or annotations.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a parameterless, side-effect-simple tool with annotations covering idempotency and destructiveness, the description is fully sufficient. No output schema exists, and none is needed for a close operation. The agent has everything required to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters and the schema is trivially complete. Per the baseline rule for 0-parameter tools, this receives a 4. The description adds nothing about parameters, which is correct since there are none.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('close') and a clear resource ('browser session held by this server'). It distinguishes from all siblings by naming the lifecycle termination operation, opposite of browser_open.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by the purpose: close the session when done with it. However, there is no explicit guidance on when to use it versus alternatives, no mention of timing (e.g., after all browser actions), and no exclusions. It is not misleading, but the agent must infer the appropriate moment.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_extractBRead-onlyIdempotent
Pull structured page data from the DOM: text, elements, links, tables or main.
| Name | Required | Description | Default |
|---|---|---|---|
| mode | No | text |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and idempotentHint=true, so the safety profile is covered. The description adds minimal behavioral context beyond listing extractable data types; it does not disclose limitations, preconditions, or response format. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single concise sentence, front-loaded with the core action and immediately enumerating the extractable types. No wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple tool with one parameter and no output schema, the description conveys the general purpose but omits critical details like the fact that it operates on the current page (implied but not stated) and the exact accepted values for 'mode'. An agent might need to infer how to use it correctly, so completeness is adequate but not strong.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% (no description for 'mode'), so the description must compensate. It lists potential values ('text, elements, links, tables or main') which hints at valid modes, but does not explicitly map them to the parameter or clarify defaults. This adds some meaning beyond the schema but not complete clarity.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a clear action ('Pull structured page data from the DOM') and specifies the resource (DOM) and the types of data (text, elements, links, tables, main). This is specific and reasonably distinct from sibling tools like browser_observe, though it does not explicitly name a sibling or differentiate.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives such as browser_observe or browser_search. It does not mention conditions, exclusions, or situations where a sibling would be preferred.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_observeARead-onlyIdempotent
List the observed controls and the visible text of the current page.
| Name | Required | Description | Default |
|---|---|---|---|
| max_elements | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is covered. The description adds the concrete behavior of listing controls and visible text, but does not disclose details like pagination, element limits, or how 'observed controls' are determined. It does not contradict the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single front-loaded sentence with no filler. Every word contributes to the core meaning, and it is appropriately concise for a simple observation tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple tool, the description gives a basic idea of what it returns, but it omits any explanation of the sole parameter max_elements and does not clarify the meaning of 'observed controls'. With no output schema, the description should carry more of the semantic load than it does.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description does not mention max_elements at all, so the description does not compensate for the missing schema documentation. The parameter name and type give an indirect clue about its limiting behavior, but the agent gets no explicit semantics or guidance on when to provide it.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('List') and a clear resource ('observed controls and visible text of the current page'), which distinguishes this read-only inspection tool from siblings like browser_click, browser_type, and browser_screenshot. Even without explicitly naming alternatives, the action and scope are unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies the tool is for inspecting the current page's controls and text, which is a sensible use case, but it gives no explicit guidance about when to prefer it over siblings like browser_extract, browser_search, or browser_screenshot. No exclusions or alternative conditions are stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_openAIdempotent
Open a URL in the shared browser session. A named profile keeps its own cookies.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | ||
| profile | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already capture open-world and idempotency traits. The description adds meaningful behavioral context by noting the shared browser session and that a named profile keeps its own cookies. It does not contradict the annotations, though it does not detail navigation side effects or page-load behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two short sentences with no filler. The core action and session context are front-loaded, and the profile/cookie note earns its place as useful additional context.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple two-parameter navigation tool with annotations covering safety and side-effect class, the description supplies the session and cookie context an agent needs. It omits return-value details, but no output schema exists and the action's result is reasonably self-evident, so this is a minor gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It adds meaning for the 'profile' parameter by explaining cookie isolation, and 'Open a URL' implicitly defines the url parameter. However, it does not provide format, allowed values, or behavior for the required url parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Open') and the resource ('a URL') and adds that it operates in the shared browser session. This distinguishes it from the interaction-focused sibling tools like click, type, and scroll, though it does not name a sibling explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives useful context: use this when you want to open a URL in the shared browser session. However, it does not explicitly contrast it with alternatives such as browser_search or browser_run, nor does it state when not to use it. The usage is implied rather than fully specified.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_pressB
Press Enter, Escape or Tab.
| Name | Required | Description | Default |
|---|---|---|---|
| key | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations are minimal (readOnly false, destructive false, openWorld true) and do not describe side effects. The description adds no behavioral context beyond the action itself—no mention of timing, focus requirements, or what happens after the key press. It does not contradict annotations, but fails to disclose any meaningful behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence with zero wasted words. It states the action and the exact keys, making it highly efficient for the agent to parse.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple one-parameter tool with no output schema, the description conveys the core action. However, it lacks any context about when to use it (e.g., after focusing an element) or potential side effects. The agent must infer the browser context from the name alone, which is a notable gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% (no description on the 'key' parameter). The description compensates partially by listing the allowed keys (Enter, Escape, Tab), giving the agent concrete examples. However, it does not clarify whether other keys are accepted, case sensitivity, or formatting, so coverage is incomplete.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action (press) and the specific keys (Enter, Escape, Tab), distinguishing it from sibling tools like browser_click (mouse) and browser_type (text). It is a specific verb+resource, making the tool's purpose unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives. It does not mention any context or exclusion criteria, leaving the agent to infer when pressing a key is appropriate. Sibling tools are not referenced, and no conditions are stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_runA
Pursue a whole goal on the current page. Supply values for anything that must be typed.
Pass resume with the run id a needs_human escalation returned, once the person has cleared the check, to carry that run on instead of starting the goal again.
| Name | Required | Description | Default |
|---|---|---|---|
| goal | No | ||
| resume | No | ||
| values | No | ||
| max_steps | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, openWorldHint=true, idempotentHint=false, and destructiveHint=false. The description adds context about the resume flow and human escalation, which is behavioral. It does not contradict annotations. However, it doesn't elaborate on side effects, success/failure behavior, or page state changes beyond the resume note, so it adds only marginal transparency beyond what annotations provide.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences and front-loads the core purpose. It avoids unnecessary words. The second sentence is slightly convoluted with the phrase 'a needs_human escalation returned' but remains efficient. No filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool that pursues a whole goal with four optional parameters and no output schema, the description is thin. It doesn't explain what happens on success/failure, how max_steps works, error handling, or how it relates to the other browser_* tools. It relies heavily on the annotations for safety context, but the operational behavior is underspecified for an agent to use it confidently in complex scenarios.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must explain parameters. It references 'goal' implicitly via 'Pursue a whole goal', explains 'values' as 'anything that must be typed', and mentions 'resume' explicitly. It does not mention 'max_steps' at all. The explanations are high-level and don't clarify formats or constraints, so it only partially compensates for the schema's lack of descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a clear verb and resource: 'Pursue a whole goal on the current page.' This distinguishes it from the sibling action tools (browser_click, browser_type, etc.) which perform single actions. It's specific and unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives partial usage guidance: it instructs to supply values for typed input and explains the resume mechanism for needs_human escalation. However, it never explicitly states when to use this tool vs the alternative single-action tools, nor does it mention prerequisites or when not to use it. The guidance is useful but incomplete.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_screenshotARead-onlyIdempotent
Capture the current viewport as a JPEG.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and idempotentHint=true, so the safety profile is covered. The description adds useful behavioral detail (JPEG format, viewport scope) but does not reveal the output form, such as a file path or data URL, which would be helpful for a tool with no output schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single sentence with no filler; every word earns its place. The action, scope, and format are front-loaded and immediately parseable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter, annotation-rich, non-destructive tool, this description is nearly complete. The only notable gap is the unspecified return format, but given the low complexity this is a minor omission.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters and schema coverage is 100%, so there is nothing meaningful for the description to add. The 'current viewport' phrase clarifies scope but does not affect parameter usage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States an unambiguous action ('Capture'), a clear resource ('the current viewport'), and a concrete output format ('JPEG'). This is specific enough to distinguish it from sibling tools like browser_click or browser_extract without extra context.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance about when to use this tool versus browser_observe or browser_extract, nor any mention of when it should or should not be used. The description only states what it does and leaves selection entirely to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_scrollB
Scroll the page one viewport step up or down.
| Name | Required | Description | Default |
|---|---|---|---|
| direction | No | down |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already indicate non-read-only, non-idempotent, and non-destructive behavior. The description adds that it scrolls by a viewport step, which is a useful detail. However, it doesn't disclose any side effects like the need for a loaded page or potential changes to focus, so transparency is moderate.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence that front-loads the action and specifies the scope. Every word contributes to understanding, and there is no unnecessary information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple tool with one parameter and no output schema, the description is incomplete. It doesn't specify the exact values for the 'direction' parameter (e.g., 'up', 'down', or other strings), nor any prerequisites like an open page. The agent would need to guess or rely on external knowledge to use it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has zero description coverage, leaving the 'direction' parameter undefined. The tool description mentions 'up or down', which suggests the valid values, but it doesn't explicitly map the parameter to these values. This adds some meaning beyond the schema but leaves ambiguity about accepted input format.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action: scrolling the page, with a specific scope of 'one viewport step' and direction 'up or down'. This distinguishes it from sibling tools like browser_click or browser_type, which perform different actions.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives. It doesn't mention typical use cases (e.g., navigating long pages) or any exclusions (e.g., not for horizontal scrolling). The agent is left to infer its purpose from the name and action alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_searchAIdempotent
Search the web, read the best results in parallel tabs, and rank them against the goal.
| Name | Required | Description | Default |
|---|---|---|---|
| goal | No | ||
| query | Yes | ||
| max_pages | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already provide safety and side-effect hints (readOnlyHint=false, idempotentHint=true, openWorldHint=true, destructiveHint=false). The description adds meaningful behavioral context beyond these: it reads results in parallel tabs and ranks them against the goal. This clarifies that the tool opens tabs and processes results, which is not covered by annotations. No contradictions with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, concise sentence that front-loads the primary action (search) and includes the main purpose (read and rank). There is zero wasted text, and every word contributes to understanding. This is an exemplary model of conciseness.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has three parameters and no output schema or parameter descriptions. The description covers the high-level behavior (search, read, rank) but omits details such as what 'best results' means, how max_pages influences the process, and what the return format is. Given the tool's relative simplicity, the description is adequate but not fully complete; an agent would need to infer some details or rely on the parameter names.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It explains the 'goal' parameter (used for ranking) but leaves 'query' and 'max_pages' undefined. 'max_pages' is particularly ambiguous—it likely controls how many results to read, but the description doesn't say this. The description only partially clarifies parameter semantics, leaving significant gaps.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: search the web, read best results in parallel tabs, and rank against a goal. This distinguishes it from sibling tools like browser_act or browser_click, which focus on page interaction rather than search and ranking. The verb 'search' and the resource 'web' are specific and unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for information retrieval and ranking, but does not explicitly mention when to avoid it or use alternatives. The context is clear (search the web), but there are no exclusions or named alternatives. Since the sibling set includes many interaction tools, the description's focus on search inherently suggests the use case, but it doesn't spell out when not to use it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_selectAIdempotent
Select an observed dropdown option by its value or label, refusing a stale ref.
| Name | Required | Description | Default |
|---|---|---|---|
| ref | Yes | ||
| option | Yes | ||
| page_key | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already indicate readOnlyHint=false, destructiveHint=false, idempotentHint=true, and openWorldHint=true. The description adds meaningful behavioral context beyond annotations: it refuses stale refs, which is a safety-relevant behavior. It also clarifies that selection is by value or label, which is not in the schema. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence that front-loads the core action and includes a critical safety qualifier ('refusing a stale ref'). Every word earns its place; no filler or repetition of schema details.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with no output schema and sparse parameter documentation, the description covers the core action and one behavioral nuance, but leaves gaps: what a 'ref' is, how the option is matched (exact vs partial), and what happens on failure. Given the tool's moderate complexity and the openWorldHint, a bit more context would help an agent invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It explains that 'option' is a value or label, which adds meaning. However, it does not explain what 'ref' refers to (presumably a reference to an observed dropdown element) or what 'page_key' is for. With 0% schema coverage, the description only partially compensates for the undocumented parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Select') and resource ('dropdown option'), and adds a distinguishing detail: it selects by value or label and refuses stale refs. It is clear enough to differentiate from siblings like browser_click or browser_type, though it doesn't explicitly name a sibling alternative.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage context: use this when you have observed a dropdown option and need to select it, and avoid using it with a stale ref. However, it does not explicitly state when to prefer this over browser_act or browser_click, nor does it describe prerequisites like having observed the dropdown first.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_typeA
Type text into one observed field by its ref, refusing a page that moved since that observation.
| Name | Required | Description | Default |
|---|---|---|---|
| ref | Yes | ||
| text | Yes | ||
| page_key | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the annotations, the description discloses a concrete failure mode: it refuses to act if the page has changed since the observation. The annotation set is sparse, so this stated refusal behavior is meaningful; other effects of typing are not described, but the main behavioral caveat is present.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One sentence, front-loaded with the action and object, with the important refusal behavior placed in a dependent clause. Every word contributes to selection or invocation confidence; there is no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a three-argument browser mutation with no output schema, the description covers the core action and one precondition, but leaves page_key unaddressed and does not state what a successful or refused call returns. It is minimally viable but has clear gaps an agent would have to resolve by trial.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema provides zero descriptions and only the titles 'Ref', 'Text', and 'Page Key', so the description must compensate. The description clarifies 'ref' as the identifier of an observed field and 'text' as what to type, but page_key is left completely unexplained, and no format or limits are given. Coverage is too low for the description to carry the full semantic burden.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly identifies the action ('Type text') and target ('one observed field by its ref'), which conveys the core purpose. It does not explicitly differentiate from sibling tools such as browser_press or browser_run, but the notion of a ref from observation makes the target unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'observed field by its ref' implies this tool is used after a browser_observe-like step and that a fresh ref is needed, but there is no explicit when or when-not guidance. No alternative tools are named, so an agent must infer the appropriate selection context from the wording alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_waitARead-onlyIdempotent
Wait a moment and observe again.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveHint false, so the safety profile is covered. The description adds the behavioral trait that the tool pauses and then re-observes, but it leaves the wait duration vague ('a moment'). No contradiction with annotations is present.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single short sentence with no filler, and the main action appears immediately. It is appropriately sized for a zero-parameter utility tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple zero-parameter, read-only wait-and-observe tool, the description covers the core behavior and implies the result is an observation. It could be more precise about the wait length and whether the output mirrors browser_observe, but these are minor gaps for a tool this simple.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters and no required inputs, so parameter documentation is unnecessary. Schema coverage is vacuously complete, and there is no parameter information the description would need to add.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a clear action chain: wait a moment, then observe again, which tells the agent the tool introduces a delay and produces a fresh page observation. It is distinguishable from browser_open and browser_close, but it does not explicitly differentiate itself from browser_observe, which also produces observations.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no guidance about when to use browser_wait instead of browser_observe or other siblings, nor does it mention waiting for dynamic content to settle before acting. The agent must infer usage from the tool name and workflow context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
5 tool updates
v0.2.8- Changed
browser_click1 field changed- added
Input schema / properties / page_keyAdded value: +{ + "anyOf": [ + { + "items": {}, + "type": "array" + }, + { + "type": "null" + } + ], + "default": null, + "title": "Page Key" +}
- Changed
browser_open1 field changed- added
Input schema / properties / profileAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "title": "Profile" +}
- Changed
browser_run5 fields changed- added
Input schema / properties / goal / anyOfAdded value: +[ + { + "type": "string" + }, + { + "type": "null" + } +] - added
Input schema / properties / goal / defaultAdded value: +null - removed
Input schema / properties / goal / typeRemoved value: -"string" - added
Input schema / properties / resumeAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "title": "Resume" +} - removed
Input schema / requiredRemoved value: -[ - "goal" -]
- Changed
browser_select1 field changed- added
Input schema / properties / page_keyAdded value: +{ + "anyOf": [ + { + "items": {}, + "type": "array" + }, + { + "type": "null" + } + ], + "default": null, + "title": "Page Key" +}
- Changed
browser_type1 field changed- added
Input schema / properties / page_keyAdded value: +{ + "anyOf": [ + { + "items": {}, + "type": "array" + }, + { + "type": "null" + } + ], + "default": null, + "title": "Page Key" +}
14 tool updates
v0.1.0- First observed
browser_act - First observed
browser_click - First observed
browser_close - First observed
browser_extract - First observed
browser_observe - First observed
browser_open - First observed
browser_press - First observed
browser_run - First observed
browser_screenshot - First observed
browser_scroll - First observed
browser_search - First observed
browser_select - First observed
browser_type - First observed
browser_wait
TDQS
Scored across 14 tools
Most tools have clear boundaries: open/search/observe/extract/click/type/select/scroll/press/wait/screenshot/close are distinct actions. The only mild overlap is between browser_observe and browser_extract, but their descriptions separate visible/control inspection from structured data extraction.
Every tool follows the same browser_<verb> pattern with lowercase snake_case throughout. The naming is predictable and makes the action type immediately clear.
14 tools is well within the ideal range for a browser automation server. Each tool covers a meaningful interaction or observation primitive without unnecessary bloat.
The set covers the full core workflow: open, search, observe, interact, extract, screenshot, and close. Minor gaps like explicit back/forward/refresh or tab management exist but are workable through the provided tools.
Maintenance
Related MCP Connectors
Stealth web browser for agents: search, fetch, click, download and type in persistent MCP sessions.
Web search, browser automation, scraping, crawling and CAPTCHA solving for AI agents.
Hyperbrowser MCP — wraps the Hyperbrowser AI-agent browsing API
AI-powered browser automation — navigate, click, fill forms, and extract data from any website.
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceEnables real browser automation as tools in Cursor, Claude Desktop, Windsurf, and any MCP-compatible client, allowing AI agents to interact with web pages through natural language.17 npmMIT
- FlicenseNot gradedqualityDmaintenanceEnables browser automation through the MCP protocol, allowing AI agents to control a real browser using accessibility snapshots and natural language commands.-
- AlicenseNot gradedqualityDmaintenanceEnables browser automation through the Claude Chrome Extension, allowing agents to navigate websites, fill forms, take screenshots, and debug web apps via standard MCP protocols.1MIT
- FlicenseNot gradedqualityCmaintenanceEnables remote browser automation via MCP, allowing models to open pages, read snapshots, click, fill, and select elements using Playwright, with built-in security restrictions against sensitive actions.-