jevnav
Drive a browser for agents/tests with intent-based decisions, gated actions, and rich page facts.
Navigate and manage tabs:
goto,new_page,select_page,close_page,tabs.Inspect pages as facts:
page_state,outline,styles,console,network,network_detail,dialogs,summary.Act:
browseandgoalresolve intents with Jev; plusfill_form,press_key,drag,upload_files,scroll,wait_for,read_js.Gate risky/uncertain actions: verdicts
auto/review/blocked, confidence thresholds, risky patterns; onlyautoexecutes.Test and emulate:
resize,emulate,route,unroute,dialog_policy.Profile and trace:
perf_metrics,heap_snapshot,lighthouse,trace_start,trace_stop,screenshot.Audit sessions:
summaryreports steps, gate counts, cost, latency; decisions are traced for replay/CI.
jevnav
Page truth for browser agents — and decisions that replay, test and audit.
https://github.com/user-attachments/assets/502550ec-77d4-439e-b334-fe7007f946b9
jevnav is a browser layer for agents and tests. It reads a page as facts
(structure, computed styles, the controls on screen), lets
Jev — TypeSafe's model for structured
questions — pick the element for an intent with a calibrated probability,
gates risky or uncertain actions to a human, and records every decision in a
trace that replay re-checks offline in CI.
Page truth, not pixels.
outline,stylesanddiffreturn what the browser resolved —font-size 32px → 28pxbetween a mockup and the running app is something an agent can fix. No screenshots in the decision loop.Evidence, not confidence. Each decision carries its probability, the gate verdict and its cost. A site change that breaks a recorded decision makes
replayexit 1 — no model call, no API key.Works where your agent works. CLI, Python/pytest, an MCP server for Claude Code, Cursor, Codex and friends, and a GitHub Action.
Why it exists — selector tests break when a label changes; LLM browser agents
are confident, unauditable and occasionally wrong: docs/why.md.
Install
uv tool install jevnav # or: pip install jevnav (the MCP server is included)
playwright install chromium # one-time browser downloadRequires Python 3.10+. Deciding (go, run, browse, goal) needs a TypeSafe
API key in TYPESAFE_API_KEY or ~/.config/typesafe/apikey.txt. replay,
diff, outline and styles need no key — that is the point.
Related MCP server: playwright-mcp-server
Quickstart
1. Let Jev drive — state a goal and the outcome that proves it:
jevnav go --goal "sign in with the demo account and open the pricing page" \
--start https://app.example.com/login \
--context email=demo@example.com --context password="${ACME_PASSWORD}" \
--success "#pricing.visible" \
--report goal.mdstatus: done — outcome verified against the page
steps: 5 — auto 4, review 0, blocked 0One Jev request per step; every step is gated and traced. The loop stops when
the goal is met, when nothing on the page can make progress (stuck), when the
gate wants a human (review), when the page stops changing (no_progress), or
at --max-steps. done is a claim — --success turns it into evidence
(verified, or unverified and the run fails). --dry-run decides without
acting.
2. Or script the flow and let Jev resolve each intent:
# flows/acme-login/flow.yaml
id: acme-login
start: https://app.example.com/login
steps:
- intent: "Sign in to the existing account"
action: click
- intent: "Type the password"
action: fill
value: "${ACME_PASSWORD}" # read from the environment, never written to the trace
- intent: "Submit the login form"
action: clickjevnav run flows/acme-login/flow.yaml --report run.md3. Replay it in CI — offline, deterministic, no key:
jevnav replay acme-login.trace.jsonl # re-resolve every recorded decision
jevnav replay acme-login.trace.jsonl --execute # also re-run the actions + check --successsteps 3 verdicts: ok 3Change Sign in to Log in on the site and the same replay fails:
[01] changed Sign in to the existing account
no element now has 'button|sign in' (was 'Sign in' / 'button')Exit code 1, with the reason. That is the regression test.
Page truth for your agent
The facts a coding agent needs about a rendered page, without a screenshot:
outline(selector) for a region's structure (tags, headings, text, boxes),
styles(selector, props) for the computed values, page_state() for the
controls jevnav can act on.
jevnav diff compares a mockup with the running app as facts and exits 1 on
drift:
jevnav diff new-ui.html http://localhost:3000 --report ui-diff.mdelement | property | mockup | app |
h1 [Pricing] | font-size | 32px | 28px |
button#cta [Start free] | border-radius | 8px | 4px |
The report also lists structure differences (missing, new and moved elements;
boxes compared with a 4px --tolerance). The loop for "here is a new UI, update
the codebase": the agent reads both pages with jevnav, edits the code itself,
re-runs diff until it exits 0, then pins the outcome with
goal(..., success="<selector>") so replay --execute keeps checking it.
jevnav reports; it never edits your repository and never compares pixels.
Gates
Every decision gets one of three verdicts:
verdict | meaning |
| confidence at or above the threshold and nothing risky — the action runs |
| a human confirms first: low |
| no decision was possible (the model answered |
review and blocked never execute. Thresholds and risky patterns live in an
optional gates.yaml; defaults ship for nine languages:
# flows/acme-login/gates.yaml
min_confidence: 0.9 # scripted flows: one question per step, well calibrated
loop_min_confidence: 0.5 # goal loop: four questions at once, p runs lower
risky:
- "\\b(delete|remove|purchase|pay)\\b" # matched against intent + element name + role
intents:
"delete the *": { min_confidence: 0.99 }
truncated: review # the page had more than 255 candidatesThe goal loop uses a lower threshold on purpose: measured correct loop decisions
land at p 0.41–0.99 and wrong ones at 0.39–0.47, so its safety comes from
deterministic checks instead — fill on a button is refused, a field with no
context value is blocked, two steps that change nothing stop the run, risky
patterns always go to review, and the outcome is verified against --success.
MCP server
claude mcp add --scope user jevnav -- uvx jevnav mcpOr, for Cursor, Claude Desktop, VS Code and other clients:
{
"mcpServers": {
"jevnav": {
"command": "uvx",
"args": ["jevnav", "mcp"],
"env": { "TYPESAFE_API_KEY": "..." }
}
}
}No URL or flags needed: the agent opens pages with goto, one server serves
every site, and each session writes an auditable jevnav-session.trace.jsonl
(--no-trace opts out). The deciding tools are what no other browser MCP has:
tool | what it does |
| one step: Jev picks the element, the gate decides, only |
| drive the whole way; returns |
| open a page, list what jevnav can act on, session totals |
Plus 28 acting and inspecting tools (forms, keys, uploads, tabs, console,
network, styles, outline, emulation, tracing, Lighthouse), each with MCP
annotations so the host knows which calls change state. A decision costs about
$0.00004 and ~330ms, and the page never enters the LLM's context. Full tool
reference, security flags and when to pick jevnav vs. Playwright or
chrome-devtools-mcp: docs/mcp.md.
CI — GitHub Action
- uses: dtduc-git/jevnav@v0
with:
trace: examples/local-demo/demo.trace.jsonl
execute: "true" # also re-run the recorded actions
report: replay.mdNo model call, no API key, ~30 seconds. Fails when a recorded target changed,
became ambiguous, or a recorded --success selector is no longer visible.
Inputs: trace, report, execute, json, version (default latest from
PyPI, or local for a checkout). @v0 floats; pin a release tag such as
@v0.2.2 for fully reproducible CI.
pytest
The jev fixture ships with the package: an ordinary Playwright test gets Jev
decisions, and every test writes a trace that replays in CI.
def test_sign_in(jev):
jev.goto("https://app.example.com/login")
jev.fill("the email address", "demo@example.com")
jev.fill("the password field", "${DEMO_PASSWORD}")
jev.click("the sign-in button")
jev.expect("#welcome")DEMO_PASSWORD=... pytest --jev-trace-dir=traces
DEMO_PASSWORD=... jevnav replay --execute traces/test_sign_in.trace.jsonl # offline, no keyA review verdict fails the test before the action runs, ${VAR} values are
recorded by name only, and jev.page is the real Playwright page for everything
else. Runnable example with a committed trace:
examples/pytest-interop/.
Your own browser
jevnav go --goal "..." # fresh headless Chromium (default)
jevnav go --goal "..." --user-data-dir ~/.cache/jevnav-profile --headed # persistent profile
jevnav go --goal "..." --cdp http://127.0.0.1:9222 # attach to a running ChromeLog in once with --headed and every later run reuses the profile; --cdp
drives the Chrome you already have open, keeps its own settings (so it refuses
--user-data-dir, --locale, --timezone and --user-agent) and never
closes it. Both work on run,
go, replay and mcp, and so does --browser for firefox or webkit (plus
--locale, --timezone, --user-agent). Profile paths and cookies never reach
a trace.
Evidence
what | result | source |
element picks on real sites (9 sites, | 41/41 scored cases correct; 30/30 at the | |
decision latency and cost | p50 365ms, $0.000153 per decision | same |
goal loop (local fixture, 4 goals × 2 wordings) | 8/8 goals correct, incl. the impossible one ( | |
driving tasks vs. chrome-devtools-mcp (same LLM, n=2) | jevnav 8/8, chrome-devtools-mcp 6/8; chrome-devtools 2.4× faster end to end | |
replay | deterministic: offline, no key, exit 1 on a broken decision | run it on your own traces |
Small samples with a single annotator: read them as direction, not proof.
jevnav's advantage is decision cost and evidence, not wall-clock speed on small
pages. Method, caveats and the tool-level comparison:
docs/evidence.md.
How it works

A shortlist, not the page. Visible interactive elements from every frame and open shadow root, ranked by how likely a human would act on them and capped at 120 (
--max-candidates, hard cap 254). Each carries role, accessible name, type, href, placeholder and a scope, so three "Email" fields stay distinguishable.One question per step. The shortlist plus
nonebecomes a choice question; Jev answers with one element and its probability.Fingerprints, never positions. An element's identity is
role|name. The trace stores every candidate's fingerprint as the model saw it, and actions and replay resolve by fingerprint with a uniqueness check, so a shifted page cannot click the wrong thing.Replay verdicts.
ok,moved,changed,ambiguous,error— the last three fail.--normalize REGEXrelaxes known churn such asCart (3)→Cart (4); strict is the default.
The trace format is a public contract: SPEC.md. Interactive
architecture diagram: docs/architecture.html.
Beyond the web: docs/games.md (Jev playing a game from
structured state, measured against a random control).
Non-goals
Not a planner.
go/goaldrive toward a goal you state; deciding what to do stays with you or your agent — jevnav decides where, and records why.No pixel decisions. No screenshots or canvas vision in the decision loop (
screenshotexists for humans), and no text generation —fillvalues come from your flow, context or environment.No hosted service, no telemetry. Nothing leaves the machine except the question sent to your configured Jev endpoint.
Privacy
Traces contain page URLs, element names and your actions — never screenshots.
The goal loop also sends a short digest of the page's visible text and current
form values (passwords masked); scripted flows send neither. ${ENV} values are
recorded by name only. Add *.trace.jsonl to your .gitignore and audit a
trace before sharing it. See SECURITY.md for the authority
boundary of the MCP server.
Suite
jevnav is the browser piece of a verification stack: mcplint (MCP configs), harnessguard (agent harnesses), jevassert + jev-packs (calibrated decision packs), and jev-table.
License
Apache-2.0 · Releases and changelog
Available Tools
33 toolsbrowseADestructive
Find the element matching an intent and, if the gate allows it, act on it.
Returns the gate verdict (auto / review / blocked), the confidence, the
target element and its Playwright selector. Only ``auto`` decisions are
executed. On large pages even precise intents score 0.8-0.95, so pass
``min_confidence`` to set your own bar; risky patterns and the
deterministic checks still apply and cannot be overridden.
| Name | Required | Description | Default |
|---|---|---|---|
| value | No | ||
| action | No | click | |
| intent | Yes | ||
| min_confidence | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already mark the tool as destructive and non-read-only. The description adds meaningful behavioral detail: gate verdicts (auto/review/blocked), confidence scoring on large pages, and that risky patterns/deterministic checks cannot be overridden. This goes well beyond the annotation flags.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Approximately 90 words with the primary purpose front-loaded. Each sentence contributes distinct behavioral or configuration detail; no obvious redundancy or irrelevant filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The output schema covers return values, and the description covers gate behavior and confidence thresholds. However, the missing semantics for `value` and the valid range of `action` are notable gaps that could lead to incorrect invocation despite the rich behavioral context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema has 0% description coverage, so the description must compensate. It explains min_confidence with concrete score ranges and references intent, but it does not clarify the meaning of the `value` parameter or allowed values for `action`, leaving two of four parameters under-specified.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a clear purpose: find an element matching an intent and act on it if the gate allows. It references a gate and Playwright selector, which helps distinguish it from simpler action tools, though the phrase 'act on it' is somewhat broad.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for intent-based element discovery and gated action, and gives parameter guidance around min_confidence. However, it never explicitly contrasts browse with sibling tools like fill_form or press_key, nor does it state when not to use it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
close_pageADestructive
Close the tab at this index (see tabs). If it was the tab being driven, jevnav drives the last remaining one; closing the only tab, or an out-of-range index, is an error. Returns the tabs that remain.
| Name | Required | Description | Default |
|---|---|---|---|
| index | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already flag destructive, non-idempotent, open-world behavior, and the description adds real edge-case context beyond them: the fallback when the driven tab is closed (drive the last remaining one) and the two explicit error conditions. That is exactly the kind of post-condition detail annotations cannot convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, front-loaded with the core action and then the edge cases; every clause carries information. The awkward line-wrapping/indentation slightly hurts readability but nothing is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive mutation tool it covers the mutation result, the fallback behavior, and error cases; an output schema exists so return-value explanation is not strictly required (though it restates it). The one remaining gap is index-base/range semantics, which neither schema nor description pins down.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% and the single 'index' parameter is undocumented in the schema, so the description must carry it. It clarifies that the index refers to a tab index from the 'tabs' tool, but leaves 0-based vs 1-based indexing and valid range unspecified, so the compensation is partial.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Close') and resource ('the tab at this index'), and points the agent at the sibling 'tabs' tool for resolving the index. An agent can distinguish this from siblings like new_page, select_page, and tabs without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explains the boundary conditions under which the tool fails (closing the only tab, out-of-range index) and what happens to the driven tab, plus a pointer to 'tabs' for the index. It stops short of explicitly contrasting with alternatives, but the driving/selection context is clearly implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
consoleARead-onlyIdempotent
Recent console messages and page errors, newest last (observation only; not part of a trace). Returns {messages:[{type, text}]}, where type is the console method (log, info, warning, error, debug, ...) or pageerror; only_errors keeps warning, error and pageerror. Read it after an action to see what the page complained about.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| only_errors | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already carry the safety profile (readOnlyHint=true, idempotentHint=true, destructiveHint=false), so the bar is lower. The description adds genuine value beyond annotations: 'observation only; not part of a trace' clarifies the data source, and the return-shape disclosure ({messages:[{type, text}]}) plus filter behavior ('only_errors keeps warning, error and pageerror') tells the agent what to expect. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Roughly 60 words, with the core purpose front-loaded first, followed by the observation qualifier, return shape, parameter behavior, and usage timing. Every clause earns its place; there is no filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present and strong annotations, the description covers purpose, ordering, filter semantics, and usage timing. The only meaningful gap is explicit semantics for limit. For a simple two-parameter read-only tool, this is nearly complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It does explain only_errors well — noting it keeps warning and pageerror in addition to error, which is non-obvious from the parameter name alone. However, limit is left to inference from the name and the 'newest last' ordering; the description never states that limit truncates the returned list. Partial compensation only.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific resource ('Recent console messages and page errors') with ordering ('newest last') and a verb phrase that marks it as an observation. The 'observation only; not part of a trace' qualifier distinguishes it from tracing and network siblings, so an agent can pick it apart from network, dialogs, or page_state without opening schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides clear contextual guidance: 'Read it after an action to see what the page complained about.' This tells the agent when in a workflow to invoke it. It does not explicitly name alternative tools or state when NOT to use it, but the 'not part of a trace' note partially covers exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
dialog_policyADestructive
Answer dialogs from now on. match=None sets the session default; with a match, adds a rule for dialogs whose text contains it (an earlier rule with the same text is replaced). Applies to future dialogs only; dialogs already seen stay recorded.
| Name | Required | Description | Default |
|---|---|---|---|
| match | No | ||
| action | No | accept |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations indicate destructiveHint=true, readOnlyHint=false, and the description adds valuable behavioral details: it applies only to future dialogs, existing dialogs stay recorded, and earlier rules with the same text are replaced. This goes beyond annotations and clarifies side effects. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three concise sentences, front-loaded with the core purpose ('Answer dialogs from now on') followed by precise rule semantics. No filler or redundancy; every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the output schema exists (so return values are covered) and annotations provide safety hints, the description covers the key behavioral scoping (future-only, replacement) and explains the match parameter well. It omits a full list of possible actions and how to clear the policy entirely, but for a policy-setting tool this is arguably adequate. The session-scoped nature is mentioned via 'session default'.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate. It explains 'match' thoroughly: null sets the session default, a string adds a rule based on text containment, and replacements occur. However, 'action' is only mentioned via its default 'accept' and the phrase 'Answer dialogs'; possible values or their effects are not explained. This leaves a gap for the agent.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: 'Answer dialogs from now on' – a specific verb and resource. It distinguishes itself from sibling 'dialogs' (which likely lists current dialogs) by focusing on setting future policy. The semantics of match and action are introduced, making the tool's role unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context: it sets a session default or adds a rule, and applies to future dialogs only. It does not explicitly name alternatives or exclusions, but the purpose is distinct enough that an agent can infer when to use it (e.g., to automate dialog handling). No explicit 'when not to use' is given, but the context is sufficient.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
dialogsARead-onlyIdempotent
Every alert/confirm/prompt seen, with the policy that resolved it.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint and idempotentHint, covering the safety profile. The description adds that it includes the policy that resolved each dialog, which is meaningful, but it does not disclose potential limitations (e.g., only shows dialogs from the current page, ordering, or whether it includes previously resolved dialogs). It is not contradictory to annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single concise sentence, front-loaded with the core content. No redundant words or unnecessary detail.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's simplicity (no parameters, strong annotations, and an output schema), the description fully conveys what the tool returns. The agent has enough to understand the tool's purpose and invoke it correctly without ambiguity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, and schema coverage is 100%, so the baseline is 4. The description need not explain parameter formats since there are none.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states the specific resource (dialogs) and the content (every alert/confirm/prompt seen, along with the resolving policy). It is clear and distinguishes from the sibling `dialog_policy` (which likely sets policies), though it lacks an explicit verb like 'returns' or 'lists'.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this tool versus alternatives like `dialog_policy` or `page_state`. No context about typical use cases, prerequisites, or exclusions is provided, so the agent must infer when this applies.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
dragADestructive
Drag the element at source_selector onto target_selector. The drag runs immediately with no confirmation, so a wrong target can change page state.
| Name | Required | Description | Default |
|---|---|---|---|
| source_selector | Yes | ||
| target_selector | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already mark the tool destructiveHint=true and readOnlyHint=false. The description adds useful behavioral context by disclosing that the drag executes immediately, without confirmation, and that a wrong target can change page state. This goes beyond the annotation flags without contradicting them.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences with no filler. The main action is stated first and the immediate no-confirmation behavior follows, so every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-param destructive tool with an output schema and annotations, the description covers the essential invocation detail and the key risk. It is complete enough for an agent to call correctly; omitted details like selector format are not severe enough to lower it further.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description supplies the core relationship: source_selector identifies the element being dragged and target_selector identifies the drop destination. However, it does not specify selector syntax, constraints, or behavior for invalid selectors, leaving the schema titles to carry part of the semantic load.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb and resource: 'Drag the element at source_selector onto target_selector.' This unambiguously defines the operation and separates it from sibling navigation, form, and browsing tools, none of which perform drag-and-drop.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The use case is implied by the action name and description, but the description does not explicitly state when to prefer drag over other interaction tools or mention any prerequisites. The warning about page state is risk guidance, not usage-selection guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
emulateAIdempotent
Emulate media, geolocation ("lat,lon"), connectivity and a slow device. Media overrides (color scheme, reduced motion, forced colors) apply to the current tab; geolocation and offline apply to the whole browser context, so later tabs and loads keep them. Fields you omit are left unchanged. The position persists context-wide, but only an http(s) origin that called emulate can read it (that origin gets the geolocation permission); call emulate again after navigating elsewhere. cpu_throttle slows the current tab's CPU by that factor (4 = four times slower; 1 restores full speed). network_conditions throttles the current tab's network with a preset ("Slow 3G", "Fast 3G", "Slow 4G", "Fast 4G", or "No throttling" to switch it off) or with explicit latency_ms/download_kbps/upload_kbps. A network call replaces the tab's whole network profile (unset latency is 0, unset throughput is unthrottled) rather than editing it. Throttling lasts until changed and needs a chromium server (it goes through CDP); on firefox/webkit it is refused before anything else is applied.
| Name | Required | Description | Default |
|---|---|---|---|
| media | No | ||
| offline | No | ||
| latency_ms | No | ||
| geolocation | No | ||
| upload_kbps | No | ||
| color_scheme | No | ||
| cpu_throttle | No | ||
| download_kbps | No | ||
| forced_colors | No | ||
| reduced_motion | No | ||
| network_conditions | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Goes well beyond the annotations by disclosing real runtime behaviors: geolocation persistence across tabs, the http(s)-origin permission/read restriction, cpu_throttle factor semantics (4 = four times slower, 1 = full speed), and crucially that network_conditions replaces the whole profile rather than editing it. It also flags the chromium-only CDP dependency and firefox/webkit refusal, which no annotation conveys.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Dense but tightly front-loaded, leading with the core purpose before the per-parameter behavior. For an 11-parameter tool with interlocking persistence rules, essentially every sentence carries necessary information; no padding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given an output schema exists, return values need not be explained, and the description covers scope, persistence, permissions, mutual-evaluation order (throttling refused before anything else applies on non-chromium), and the replacement semantics of network profiles. Nothing an agent needs to invoke it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema coverage across 11 params, the description carries the burden and largely succeeds: it gives the geolocation "lat,lon" format, network preset names, explicit latency/download/upload usage, cpu factor, and the omit-means-unchanged rule. It loses a point because the `media` parameter's own values (screen/print) are never defined; color_scheme, reduced_motion and forced_colors are instead listed parenthetically as overrides, which muddies what `media` alone accepts.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb (emulate) and enumerates the exact resources it controls: media, geolocation, connectivity, and a slow device (CPU/network). An agent immediately knows the scope. It stops short of the top score because it never names a sibling (e.g. route, network, resize) to distinguish itself from them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Strong contextual guidance: media applies to the current tab while geolocation/offline apply context-wide, fields omitted are left unchanged, and the agent is told to call emulate again after navigating. It gives clear when-conditions but no explicit when-not or named alternatives, so it misses the top band.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
fill_formADestructive
Fill several fields in one call. fields_json is a JSON list of {selector|intent, value, action?}; action is fill (default), select, check or type. fill and select replace the value, type appends, check ticks when value is true-like ("true", "1", "yes", "on" or omitted) and clears otherwise. An intent is resolved by Jev against the page's text, search, combobox, checkbox and radio inputs — or, when the page has none of those, any interactive element. Returns {filled:[...]}; a field with neither selector nor intent is an error.
| Name | Required | Description | Default |
|---|---|---|---|
| fields_json | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description discloses precise action semantics: fill/select replace, type appends, check treats true-like values as ticks and clears otherwise. It also explains how intents resolve against page elements and what the return value is. This adds significant behavioral detail beyond the annotations' readOnlyHint/destructiveHint flags, without contradicting them.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but each sentence earns its place, moving from high-level purpose to parameter format to action semantics to return/error handling. It is not overly verbose, though it could be slightly better structured with explicit examples or bullet points. Still, it remains readable and compact for the amount of behavior it documents.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers the only parameter's structure, all action behaviors, return shape, and error case. It does not include a concrete JSON example, which would help, but given the complexity is moderate and the output schema reportedly exists, the description is largely complete. The openWorldHint could be expanded but is not critical for invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the schema only states 'fields_json' is a string. The description fully compensates by defining the JSON list format, each field ({selector|intent, value, action?}), the action defaults, and the semantics of each action. This is essential meaning that the schema alone cannot convey.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description explicitly states a specific verb and resource: 'Fill several fields in one call.' It then defines the exact JSON structure and supported actions, making the tool's purpose unambiguous. No sibling tool has a similar batching/fill behavior, so the purpose is well distinguished.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'Fill several fields in one call' implies when to use it (when multiple values need to be set together) but it never names alternatives or exclusion conditions. There is no mention of press_key, upload_files, or other sibling tools that could handle similar tasks, leaving the agent to infer the boundary.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
goalADestructive
Drive the browser towards a goal: Jev decides every step, jevnav acts.
``context_json`` is a JSON object of values the goal may need, e.g.
{"email": "a@b.c", "password": "${PW}"}. ``success`` is a selector that
must be visible when the goal is done: pass it and the outcome comes
back verified or the run is reported as unverified. Returns the outcome
(done, stuck, review, ...), the steps taken, cost and the verification.
Risky steps stop the loop and come back unexecuted.
| Name | Required | Description | Default |
|---|---|---|---|
| goal | Yes | ||
| success | No | ||
| max_steps | No | ||
| context_json | No | {} |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, readOnlyHint=false, openWorldHint=true, idempotentHint=false. The description adds valuable context: risky steps stop the loop and come back unexecuted, success verification is built-in, and the outcome includes done/stuck/review states. This goes beyond the annotations and helps the agent understand the tool's safety and control flow.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded with the core purpose. The parameter explanations are integrated naturally. It's slightly dense with the 'Jev decides, jevnav acts' phrasing, which may be jargon, but it's not bloated. Every sentence adds value.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is complex (autonomous loop, verification, risky-step handling) and has an output schema, so the description doesn't need to detail return values. It covers the main behavioral aspects: goal-driven execution, success verification, risky-step stopping, and context injection. It could mention max_steps behavior or what 'review' means, but overall it's complete enough for an agent to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It explains context_json (JSON object of values the goal may need, with an example) and success (selector that must be visible when done, with verification behavior). It doesn't explain goal, max_steps, or the exact format of context_json beyond the example, but the key parameters are covered. Given the low schema coverage, this is a strong effort.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Drive the browser towards a goal') and resource, and explains the high-level behavior: Jev decides, jevnav acts. It distinguishes itself from siblings like goto, browse, and fill_form by being an autonomous goal-driven loop rather than a single action. However, it doesn't explicitly name a sibling alternative, so it's clear but not fully differentiated.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explains when to use this tool: when you have a goal and want the agent to decide steps. It also implies when not to use it: for direct, single actions like goto or fill_form. It doesn't explicitly state exclusions or alternatives, but the context is clear enough for an agent to select it appropriately.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
gotoAIdempotent
Open a URL in the current tab, replacing its content, and wait until the DOM is ready. Returns {url, title, elements}, where elements is the number of interactive elements found; call page_state for the candidate list.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover safety (readOnlyHint=false, destructiveHint=false) and openness (openWorldHint=true), so the bar is lower. The description adds useful behavioral context beyond the annotations: it replaces the current tab content, waits for DOM readiness, and returns an element count. This gives the agent meaningful expectations about side effects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences with no filler. The primary action and its key behavioral constraint ('replacing its content') are front-loaded, and the return summary is compact. Every sentence adds value.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple one-parameter navigation tool, this is nearly complete: it states what happens, the wait condition, the return shape, and the next recommended action. The output schema handles detailed return structure, so the description does not need to. Minor gaps like URL format or error behavior are not critical given the annotations.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the description does not elaborate on the url parameter beyond its name. There is no mention of expected URL formats, supported schemes, or whether relative URLs are allowed. With only one parameter, the impact is limited, but the description still fails to add meaning beyond the bare 'url' property.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb and resource: 'Open a URL in the current tab, replacing its content, and wait until the DOM is ready.' It clearly distinguishes itself from siblings like new_page by emphasizing 'current tab' and 'replacing its content,' so an agent can choose it accurately.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context for when to use this tool—navigating the current tab rather than opening a new one—and it directs the agent to 'call page_state for the candidate list' afterward. It does not explicitly list exclusions or alternative tools, but the current-tab scoping implicitly differentiates it from new_page and browse.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
heap_snapshotADestructive
Write a Chromium heap snapshot of the current page to a file (default: jevnav-heap-.heapsnapshot under --file-root, which defaults to the working directory and must be set explicitly if the server runs from / or $HOME; the path must stay inside it), for Chrome DevTools > Memory. A snapshot is a full JS heap dump, often hundreds of MB; an existing file at the same path is overwritten. Use it to chase a memory leak, not for routine inspection — it changes nothing on the page, and it is chromium-only.
| Name | Required | Description | Default |
|---|---|---|---|
| path | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already indicate destructiveHint=true and readOnlyHint=false, but the description adds crucial specifics: it overwrites an existing file, produces a large dump (hundreds of MB), and does not modify the page. It also notes Chromium-only behavior. These details go beyond the binary annotations, giving the agent actionable expectations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is slightly dense but well-structured. It leads with the action, then provides default, constraints, size, overwrite behavior, and usage note in a logical flow. Each sentence adds value; although verbose, it remains efficient and front-loaded with the core purpose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with one optional parameter and an output schema, the description is largely complete. It covers the primary purpose, default behavior, constraints, size impact, overwrite warning, and platform limitation. It does not detail the return value, but that is handled by the output schema, so no gap exists there. The only minor omission is explicit instruction on how to pass the path parameter, but constraints imply relative/full path usage.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description carries the burden. It explains the default path (jevnav-heap-<timestamp>.heapsnapshot under --file-root), the requirement that the path stay inside file-root, and the default working-directory behavior. This meaningfully supplements the bare schema type string/null, though it does not describe the format for explicit path input.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's function: 'Write a Chromium heap snapshot of the current page to a file.' It specifies the output format and destination, and the purpose ('for Chrome DevTools > Memory'). It distinguishes itself from other tools by noting it is 'chromium-only' and 'changes nothing on the page', which separates it from page-mutating actions among its siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides explicit guidance: 'Use it to chase a memory leak, not for routine inspection.' This tells the agent when to invoke it and when not to, though it does not name specific sibling alternatives. The 'changes nothing on the page' note also clarifies it is safe for state, adding context for selection.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
lighthouseAIdempotent
Run Lighthouse (through npx) against the current or given URL and return the scores. Needs node/npx on PATH (npx fetches Lighthouse on first use), takes tens of seconds, and does not change the page.
| Name | Required | Description | Default |
|---|---|---|---|
| url | No | ||
| categories | No | performance,accessibility,best-practices,seo |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description transparently discloses non-obvious behaviors beyond the annotations: it requires node/npx on PATH, fetches Lighthouse on first use, takes tens of seconds, and does not change the page. These details add real value and are consistent with the idempotentHint and destructiveHint annotations, with the page-unchanged claim clarifying the readOnlyHint=false nuance (external side effects like npx downloads).
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with zero filler: the first states the core action and output, the second adds prerequisites, duration, and side-effect caveat. Information is front-loaded and every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, return values need not be explained. The description covers purpose, URL source, prerequisites, runtime, and page safety. The only meaningful gap is the semantics of the 'categories' parameter, which prevents a perfect score but does not undermine overall viability.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has no parameter descriptions (0% coverage), so the description must compensate. It clarifies the 'url' parameter via 'current or given URL', but the 'categories' parameter is not mentioned at all, leaving its format (e.g., comma-separated vs array) and exact meaning unstated. This is a notable gap for one of two parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Run'), a resource ('Lighthouse'), and the output ('return the scores'). It also distinguishes itself from browser-level tools like perf_metrics or trace_start by specifying that it invokes Lighthouse via npx, which clearly separates its purpose.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage (for getting Lighthouse audit scores) and provides practical context like needing node/npx and taking tens of seconds. However, it does not explicitly say when to prefer this over sibling tools such as perf_metrics or trace_start, or mention any exclusions/alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
networkARead-onlyIdempotent
Recent network requests, newest last; only_failed keeps 4xx/5xx and transport errors. Each entry carries method, url, status, resource, error (for failed requests) and id — pass that id to network_detail for headers and body.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| only_failed | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover readOnlyHint, idempotentHint, and destructiveHint, so the safety profile is known. The description adds valuable behavioral context: the ordering (newest last), the exact filter semantics (only_failed keeps 4xx/5xx and transport errors), and the per-entry fields (method, url, status, resource, error, id). This goes beyond annotations without contradicting them.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with zero filler. The first sentence front-loads the core purpose and filter behavior; the second describes the output fields and the routing to network_detail. Every word earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple list tool with two parameters and an output schema, the description covers ordering, filtering, output fields, and the relationship to network_detail. The only minor omission is an explicit note about the limit parameter's effect, but given the output schema exists and the tool is straightforward, this is not a critical gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description must explain parameters. It explains only_failed thoroughly ('keeps 4xx/5xx and transport errors'), but limit is left unexplained beyond its name and default. While limit is fairly self-explanatory, the description could have added a phrase like 'controls the number of entries returned.' The partial coverage earns a middle score.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool lists recent network requests, newest last, and distinguishes itself from network_detail by noting the id can be passed there for headers/body. It is specific about the resource (network requests) and the verb (list/return), making it easy for an agent to select correctly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context on when to use this tool (to get a list of network requests, optionally filtered to failures) and points to the sibling network_detail for more detailed headers/body. However, it does not explicitly state when not to use it or mention alternatives beyond network_detail, leaving a small gap in exclusion guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
network_detailARead-onlyIdempotent
Headers and (text) body of one recorded request: pass the id from network's output, or url_contains for the newest matching URL. Reads jevnav's own network buffer; nothing is re-requested.
| Name | Required | Description | Default |
|---|---|---|---|
| id | No | ||
| url_contains | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description adds meaningful behavioral context beyond annotations by stating that it reads jevnav's own network buffer and that nothing is re-requested. It also clarifies the selection behavior for url_contains (newest matching URL), giving the agent insight into how the tool executes.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two concise sentences with no filler. The first sentence names the resource and selection methods, the second adds the important local-buffer note. Every element earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the presence of an output schema and the tool's simple nature, the description is complete. It explains what is returned, how to identify the request, and the key behavioral trait (no re-request). No important calling detail is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description carries the burden of explaining parameters. It explains that 'id' comes from network's output and 'url_contains' selects the newest matching URL, providing useful semantics. It does not specify the matching format or behavior when both are omitted, but the core meanings are clear.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool returns the headers and text body of one recorded request, with a specific selection mechanism (id or url_contains). It distinguishes itself from the sibling 'network' tool by clarifying that this is the detail view for a single recorded request, not the list.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides direct usage instructions: pass the id from network's output, or use url_contains to get the newest matching URL. It does not explicitly discuss when not to use the tool or name alternatives, but the context makes the appropriate use case clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
new_pageA
Open a new tab (optionally at a URL) and drive it from now on.
| Name | Required | Description | Default |
|---|---|---|---|
| url | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare not read-only, not destructive, and not idempotent, but do not mention context switching. The description adds the important behavioral fact that after opening, the tool 'drives it from now on', indicating a persistent change in the agent's active target. This goes beyond annotations and provides valuable context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, tightly worded sentence. The core action ('Open a new tab') is front-loaded, and the secondary clause ('drive it from now on') adds essential context without redundancy. No filler or unnecessary details.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with only one optional parameter and an output schema, the description covers the primary behavior and the crucial context switch. It does not mention edge cases like opening a blank tab (implied by 'optionally at a URL') or how the previous tab is affected, but these are not critical for correct invocation. Overall, it provides sufficient information for an agent to use the tool effectively.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has a single optional nullable url parameter, and description coverage is 0%—the description does not mention the parameter by name. However, it says 'optionally at a URL', which conveys the parameter's optionality, matching the schema. It adds minimal meaning (that a URL can be given) but does not explain format or behavior when omitted. Since coverage is low, the description should compensate more; this is adequate but not thorough.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool opens a new tab and takes control of it. The verb 'open' and resource 'tab' are specific, and the phrase 'drive it from now on' distinguishes it from sibling tools like select_page or close_page, making the purpose unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use it (to create a new tab and make it the active context) but does not explicitly contrast with alternatives like goto (navigate current tab) or select_page (choose existing tab). There is no explicit 'use this when...' or 'instead of...' guidance, leaving some inference to the agent.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
outlineARead-onlyIdempotent
Structural outline of a page or region: the shape an agent can act on, without a screenshot. Returns {selector, count, elements}, one entry per heading, landmark, section, form, label, control, link, image or text block inside selector (default "body"), capped at limit (default 200) in document order; each entry carries tag, level (h1-h6), text, ownText, leaf, name, id, classes and box [x, y, width, height]. A selector that matches nothing falls back to the whole body (the result still echoes the selector you asked for). Use it before editing a region or diffing a mockup against the app — page_state is the clickable list, styles the computed CSS, screenshot for humans. Reads only.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| selector | No | body |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Behavior is disclosed well beyond the read-only/idempotent annotations: it describes document-order capping at limit, per-entry fields, and the fallback to the whole body when a selector matches nothing while still echoing the requested selector. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The core purpose is front-loaded in the first clause, then the description adds required detail in a compact, information-dense sequence without repeating schema defaults or annotation data. Every sentence contributes either return semantics, fallback behavior, or usage routing.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a relatively simple read-only outline tool with two optional parameters, the description covers input semantics, output shape, edge behavior, and when to prefer it over siblings. The output schema exists, so the listed return fields are a bonus rather than a burden, and nothing necessary for a correct call is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description carries the full burden, and it delivers: selector and limit are both explained with defaults and behavior, including the no-match fallback for selector and the 200 cap for limit. An agent can fill both parameters correctly without needing extra documentation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a precise definition: 'Structural outline of a page or region: the shape an agent can act on, without a screenshot.' It names the exact output shape and explicitly contrasts itself with sibling tools (page_state, styles, screenshot), so an agent can tell which tool to use from the description alone.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It says 'Use it before editing a region or diffing a mockup against the app' and then names alternatives: 'page_state is the clickable list, styles the computed CSS, screenshot for humans.' This is explicit when-to-use guidance with sibling differentiation and no ambiguity.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
page_stateARead-onlyIdempotent
Current URL, title and the interactive elements jevnav can see. The elements it can act on carry a jevnav data-jevcid stamp.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotations already declare readOnlyHint and idempotentHint, so the safety profile is covered. The description adds useful behavior beyond annotations by noting that only interactive elements carrying a jevnav data-jevcid stamp can be acted upon, clarifying what 'visible/actionable' means in this tool's context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, no filler. The core output (URL, title) is front-loaded, and the important actionable-element detail is stated second. Every sentence carries meaning.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter, read-only state tool with a full output schema and strong annotations, the description is complete. It explains what the agent receives and the actionable-element stamp behavior without needing to enumerate return fields, which the output schema already covers.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters and the schema has 100% coverage by virtue of being empty, so there is nothing for the description to add about inputs. The baseline of 4 applies here since parameter semantics are not applicable.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly identifies the tool as a read-only state snapshot: current URL, title, and the interactive elements jevnav can see. It lacks an explicit verb like 'Get' and does not explicitly differentiate itself from siblings such as outline or styles, but the scope is specific and understandable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies this is the tool to use when you need the current page URL, title, or visible actionable elements, but it does not state when to use it instead of alternatives like browse, outline, or wait_for. No exclusions or alternative routing are provided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
perf_metricsARead-onlyIdempotent
Chromium performance counters for the current page, read over CDP: DOM nodes, JS heap size, layout and task durations. Cheap telemetry for spotting growth between actions; for real profiling use chrome-devtools. Reads the page, changes nothing, chromium only.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint and idempotentHint, so the description's 'changes nothing' is reinforcing but not the main contribution. It adds valuable context beyond annotations: the data is read over CDP, it is Chromium-only, and it is cheap telemetry. These operational details help an agent understand platform and protocol requirements.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded: the first sentence states the resource and values, the second gives the use case and alternative, and the third summarizes safety and platform constraints. Every clause earns its place with no redundant filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter read-only tool with an output schema, this description is complete. It covers what data is returned, the operational context (CDP, Chromium only), the intended use case, and the safety profile. Nothing needed to select and invoke the tool correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There are zero parameters and the schema is empty, so there is nothing to document; baseline 4 is appropriate. The description compensates by enumerating the kinds of data returned (DOM nodes, heap size, durations), which helps an agent understand what the tool actually provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool reads Chromium performance counters for the current page, listing specific metrics: DOM nodes, JS heap size, layout and task durations. It differentiates itself from real profiling tools by framing itself as cheap telemetry and pointing to chrome-devtools for deeper profiling.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly says when to use this tool: 'for spotting growth between actions', and gives the alternative: 'for real profiling use chrome-devtools.' The 'chromium only' constraint also sets clear boundary conditions for when this tool should not be used.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
press_keyADestructive
Press a key or a combination ("Control+A", "Shift+Enter"). With a selector, presses on that element (first match); without, on the page. Acts immediately — a shortcut or Enter can submit or delete — and returns {key, selector}.
| Name | Required | Description | Default |
|---|---|---|---|
| key | Yes | ||
| selector | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already indicate destructiveHint=true, but the description adds important context: it acts immediately and can submit or delete, which aligns with the destructive hint. It also clarifies the return value {key, selector} and the behavior with/without selector. No contradiction with annotations; the description reinforces and extends the safety profile.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three sentences, each serving a purpose: specifying input format, targeting behavior, and warning of side effects with return value. It is efficient and front-loaded with key usage. Only minor fluff: 'with a selector' phrasing could be tightened, but overall structured effectively.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's simple parameter set (key, selector) and the presence of an output schema (which explains the return), the description covers all essential aspects: input format, targeting, immediate side effects, and return. It could add details like how to escape special characters or whether modifier keys are required, but it is sufficient for correct invocation in most cases.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It explains that 'key' is a key or combination (e.g., 'Control+A') and 'selector' targets an element on first match, which adds meaning beyond the schema's minimal labels. However, it doesn't detail format edge cases or prerequisites (e.g., element must be visible), so it only partially compensates for low coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states the verb (press) and the resource (key or combination on element/page), making the tool's purpose clear. It distinguishes from siblings like scroll, fill_form, and goto by focusing on key press actions, though it doesn't explicitly name a sibling. The mention of selector differentiates it from other input tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage context: press keys for shortcuts like Control+A, Shift+Enter, and with a selector to target a specific element. It notes that it acts immediately and can submit or delete, warning of side effects. However, it doesn't explicitly state when NOT to use this tool versus alternatives like fill_form or click actions, and no siblings are named.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
read_jsADestructive
Evaluate a JS expression in the page and return its value. This is arbitrary JavaScript: an expression can change page state, so treat it as an action and keep it for reading values only (disable with --no-eval).
| Name | Required | Description | Default |
|---|---|---|---|
| expression | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already indicate destructive=true and readOnly=false; the description adds valuable context by explicitly warning that arbitrary JavaScript 'can change page state' and explaining that the tool should be treated as an action. The --no-eval flag is extra operational detail beyond the structured annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the primary purpose, followed by a precise risk warning and operational hint. Every sentence earns its place with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present and a single, well-explained parameter, the description covers the key behavioral and safety aspects an agent needs. It might have mentioned that evaluation is synchronous or returns a JSON-serializable value, but those are minor given the output schema and annotations.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The single parameter 'expression' has 0% schema description coverage, but the description defines it as a 'JS expression' and explains its behavior and risks. The parameter name and type string are self-explanatory, and the description adds the crucial arbitrary-code caveat, so it compensates for the missing schema description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'Evaluate a JS expression in the page and return its value.' It clearly identifies the tool's function and distinguishes it from navigation, clicking, and page-state tools by focusing on expression evaluation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear usage guidance: 'treat it as an action and keep it for reading values only' and mentions the --no-eval disable flag. It implies the tool should be used for reading rather than mutating state, though it does not name specific alternative sibling tools or explicit when-not-to-use conditions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
resizeAIdempotent
Resize the browser viewport to width x height. The size persists for the session and can change what the page renders (responsive layout).
| Name | Required | Description | Default |
|---|---|---|---|
| width | Yes | ||
| height | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description adds genuine behavioral context beyond the annotations: it discloses that the size persists for the session and that it can alter page rendering due to responsive layout. This complements the idempotentHint=true and readOnlyHint=false annotations without contradicting them.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with no waste: the core action is front-loaded, and the behavioral implications (persistence, rendering impact) follow. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple two-parameter tool with an output schema and safety-related annotations, the description covers the essentials well. Minor gaps remain—units are unspecified and the relationship to the overlapping 'emulate' sibling is unaddressed—but nothing critical blocks correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema coverage, the description carries the burden, and the phrase 'to width x height' does connect the two integer parameters to their role as viewport dimensions. However, it adds no units, bounds, or constraints, leaving the semantics only minimally richer than the raw schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb and resource ('Resize the browser viewport to width x height'), making the tool's function immediately clear. It is unambiguous, though it doesn't explicitly differentiate from the sibling 'emulate' tool, which could also affect viewport size.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use, when-not-to-use, or alternative guidance is provided. The persistence note ('The size persists for the session') implies a usage context, but there is no explicit routing relative to overlapping siblings like 'emulate', leaving the agent to infer when this tool is preferred.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
routeA
Stub or block requests matching a URL pattern (testing). The stub persists until unroute; the call is recorded as an action, never part of a replay path.
| Name | Required | Description | Default |
|---|---|---|---|
| body | No | ||
| abort | No | ||
| status | No | ||
| pattern | Yes | ||
| content_type | No | application/json |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description discloses non-obvious stateful behavior beyond the generic annotations: the stub persists until unroute, and the call is recorded as an action but never part of a replay path. This adds meaningful behavioral context that the annotations alone do not provide.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences with the core action front-loaded and no filler. Each clause earns its place by adding either the purpose, the persistence rule, or the recording semantics.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the output schema and annotations, the description covers the most important non-obvious behavior: lifecycle, action recording, and replay exclusion. It could be more complete by explaining how status/body/abort shape the stubbed response, but the parameter names and defaults largely compensate.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description had the burden of explaining parameters, but it only mentions the URL pattern conceptually and says nothing about status, body, abort, or content_type semantics. The parameter names and defaults provide some hints, but the description does not compensate for the missing schema descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names the action ('Stub or block'), the resource ('requests matching a URL pattern'), and the context ('testing'). It also differentiates from the sibling unroute by explicitly noting that the stub persists until unroute.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The '(testing)' tag and the mention of unroute give useful situational context, but the description does not explicitly state when to choose route over alternatives like network or how it relates to replay behavior. Usage guidance is implied rather than fully spelled out.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
screenshotADestructive
Save a PNG of the page (or one element) to a path, for a human to look at.
| Name | Required | Description | Default |
|---|---|---|---|
| path | No | ||
| selector | No | ||
| full_page | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already mark the tool as destructive and non-read-only, so the description does not need to repeat that. It usefully adds that the output is a PNG saved to a path and can target an element. However, it does not mention potential overwriting of existing files or that full_page affects capture dimensions, leaving some behavior implicit.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, front-loaded sentence conveys the core action, target, format, destination, and purpose with no filler. Every phrase earns its place, and the key information is immediately visible.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with three optional parameters and an output schema, the description gives the essential overview but omits important call details like full_page semantics and what happens when path is null. An agent could make a reasonable first attempt, but might mis-handle full_page or default path behavior without additional context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has zero description coverage, so the description must carry the burden. It does clarify path and selector ('to a path', 'page or one element'), but it does not explain the full_page parameter or the meaning/behavior of null defaults. This is partial compensation, not complete.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Save'), identifies the resource ('page or one element'), specifies the output format ('PNG') and destination ('path'), and clarifies the intended audience ('for a human to look at'). This clearly distinguishes screenshot from sibling tools like outline, styles, or page_state, which serve different inspection purposes.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context: use this tool when visual evidence is needed for a human to review. It implies a distinction from machine-readable inspection tools, though it does not explicitly name alternatives or state when not to use it. The lack of explicit exclusions prevents a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
scrollA
Scroll the page down or up by pixels of document height.
| Name | Required | Description | Default |
|---|---|---|---|
| amount | No | ||
| direction | No | down |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already indicate this is not read-only (readOnlyHint=false) and not destructive (destructiveHint=false). The description adds minimal behavioral context beyond the action itself, such as the unit (pixels of document height), but does not disclose effects like whether it scrolls the entire page or only a scrollable element, or if there is any animation. It does not contradict annotations, and the bar is lower because annotations exist, but it adds limited extra transparency.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, focused sentence that states the action, the target, and the primary parameters. It is front-loaded with the verb and resource, and there is no extraneous information. It is efficiently concise.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple scroll tool, the description is mostly complete: it names the action, the target (page), the direction (up/down), and the unit (pixels). The output schema exists and likely covers return values, so the description does not need to explain those. The openWorldHint suggests possible side effects, but the description does not mention any; however, given the simplicity of scrolling, the core usage is clear. The only gap is the lack of explicit parameter value details, but that is partially covered by the description's 'down or up' and 'pixels' hints.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% (no descriptions in the schema), so the description must compensate. It partially does: 'amount' is implied to be pixels (from 'by pixels of document height'), and 'direction' is implied to accept 'down' or 'up'. However, it does not specify the exact valid values for direction (e.g., case sensitivity), the behavior of negative amounts, or how the default values work. This is adequate but not comprehensive.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('scroll') and resource ('page'), and clarifies the two possible directions (down or up) with a unit of measurement (pixels of document height). It clearly differentiates from siblings like goto (navigation), drag (pointer), and press_key (keyboard) by describing a page-level scroll action.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this tool versus alternatives. It does not mention when scrolling is appropriate, whether it should be used for specific scroll contexts (e.g., inside an iframe), or when other tools like drag might be more suitable. The description is purely definitional.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
select_pageAIdempotent
Drive the tab at this index (see tabs). The switch is immediate; the tab keeps its state, and an out-of-range index is an error.
| Name | Required | Description | Default |
|---|---|---|---|
| index | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare idempotentHint=true and destructiveHint=false, so the safety profile is covered. The description adds meaningful behavioral context: the switch is immediate, the tab keeps its state, and out-of-range indices are errors. This goes beyond what annotations provide, though it doesn't mention what happens to the current tab or whether the selected tab becomes active in the UI.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences with no wasted words. The core action is front-loaded, and the behavioral notes (immediate switch, state preservation, error condition) are packed efficiently. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter tool with an output schema and annotations covering idempotency and safety, the description is nearly complete. The main missing piece is the zero-based vs. one-based indexing ambiguity, which is critical for correct invocation. The reference to 'tabs' helps an agent discover the related tool, but the indexing convention should be explicit.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for the single 'index' parameter. The description explains that index refers to a tab position and that out-of-range values are errors, which adds meaning beyond the bare integer type. However, it doesn't specify whether the index is zero-based or one-based, which is a meaningful gap for a single-parameter tool.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Drive the tab') and resource ('at this index'), and references the 'tabs' sibling for context. It distinguishes itself from navigation tools like goto/browse by focusing on tab selection rather than URL navigation, though it doesn't explicitly name a sibling alternative.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context: it operates on a tab index, the switch is immediate, and the tab keeps its state. It doesn't explicitly say when to use this vs. alternatives like goto or new_page, but the 'see tabs' reference and the focus on index-based selection imply the usage context. It also notes an out-of-range index is an error, which is a useful boundary condition.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
stylesARead-onlyIdempotent
Computed styles for the elements matching a selector (the facts behind a visual diff).
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| props | No | ||
| selector | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and non-destructive behavior. The description adds minimal behavioral context (the 'facts behind a visual diff' framing) but does not disclose how limit or props affect results, nor any edge cases like empty selector matches. It does not contradict annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, tightly written sentence that front-loads the core purpose with zero filler. It is concise and structured well for quick comprehension.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
While the tool is simple and an output schema exists, the description fails to explain the limit and props parameters, which are essential for correct invocation. Without schema descriptions, the agent lacks enough context to use the tool properly. The core purpose is clear, but parameter semantics are incomplete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must clarify parameters. It only explains 'selector' as a matching mechanism, but gives no meaning for 'limit' (number of elements) or 'props' (which CSS properties to return). An agent would be left guessing on those two parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool returns computed styles for elements matching a selector, with a specific use case ('the facts behind a visual diff'). It is a distinct action from siblings like page_state or screenshot, and the verb+resource combination is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies a use case (visual diffing) but provides no explicit guidance on when to prefer this tool over alternatives like page_state or read_js. No exclusions or comparative context are given, so an agent must infer when it is appropriate.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
summaryARead-onlyIdempotent
This session so far: steps, auto/review/blocked counts, cost, latency.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover read-only, idempotent, non-destructive behavior; the description adds the concrete content returned (steps, counts, cost, latency), which is useful context beyond the annotations. It does not detail edge cases or how metrics are computed, but that is acceptable given annotation coverage.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is one short, front-loaded sentence that lists what the summary contains. There is no filler or redundant information; every word adds value.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-argument, read-only summary tool with annotations and an output schema, the description covers the key categories of information. It does not over-explain, and the output schema can handle detailed return values. The phrase 'so far' provides the temporal scope.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so the baseline is 4. There are no parameters requiring additional explanation, and the description complies with that.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool provides a session summary with specific data points (steps, counts, cost, latency), and the phrase 'session so far' distinguishes it from page-specific sibling tools. It lacks an explicit verb like 'Returns' or 'Shows', but the intent is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description offers no direct guidance on when to use this tool versus siblings such as page_state or perf_metrics. The only hint is 'session so far,' which implies but does not explicitly state that this is for accumulated session data rather than current page state.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
tabsARead-onlyIdempotent
List the open pages and which one jevnav is driving.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, openWorld, idempotent, and non-destructive. The description adds that it lists pages and highlights the active one, which is useful context beyond annotations. However, it does not disclose edge cases like hidden tabs or ordering.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, focused sentence with no filler. It directly states the action and the result.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has no parameters, strong annotations, and an output schema, the description sufficiently explains what the tool does and what it returns. It is complete for an agent to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes no parameters, so the schema is empty and coverage is 100%. The description adds no parameter info, which is appropriate given there are none. Baseline for zero parameters is 4.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool lists open pages and identifies which one is currently being driven (active). This is a specific verb and resource, and it distinguishes from siblings like select_page or close_page.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives such as page_state or select_page. It does not mention any exclusions or conditions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
trace_startA
Start a Playwright trace (open it later with npx playwright show-trace).
| Name | Required | Description | Default |
|---|---|---|---|
| screenshots | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already indicate a state-changing operation (readOnlyHint false, openWorldHint true). The description adds the useful follow-up command but does not disclose additional side effects such as resource usage, the need to stop the trace, or performance impact.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single concise sentence that front-loads the purpose and includes a practical follow-up command. No unnecessary words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description is adequate for a simple start action with one optional parameter, but it omits any guidance on the screenshots parameter and does not mention pairing with trace_stop, which is important for correct usage.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the description does not mention the 'screenshots' parameter at all. The agent is left to infer its meaning from the name and default, which is insufficient given the low coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb (Start) and resource (Playwright trace), and the follow-up command 'open it later' clarifies the action. It implicitly distinguishes from the sibling trace_stop by focusing on the start action.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies the tool is for starting a trace but does not explicitly state when to use it versus alternatives (like trace_stop) or any conditions/exclusions. It lacks clear guidance on the appropriate context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
trace_stopADestructive
Stop the Playwright trace started by trace_start and write the zip to
path (default: jevnav-trace-.zip under --file-root, which
defaults to the working directory and must be set explicitly if the
server runs from / or $HOME; the path must stay inside it).
The zip holds snapshots and, if enabled, screenshots of everything since
trace_start — open it with npx playwright show-trace <path>. Stopping
without a started trace is an error; passing the same path again
overwrites that zip. Chromium, Firefox and WebKit all support it.
| Name | Required | Description | Default |
|---|---|---|---|
| path | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the annotations (destructiveHint=true, idempotentHint=false), the description discloses concrete behavioral details: passing the same path overwrites the zip, stopping without a trace errors, the default path format, the file-root restriction, and the zip's contents. This gives the agent a realistic model of side effects and failure modes.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but every sentence adds necessary operational detail: default path, file-root rule, zip contents, viewer command, error condition, overwrite behavior, and browser support. The main purpose is front-loaded, and no sentence is filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with one optional parameter and an output schema, the description covers all decision-relevant context: prerequisites, side effects, path constraints, and error cases. An agent has enough information to invoke it correctly and anticipate the outcome.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has one optional path parameter with 0% description coverage, so the description carries the full burden. It explains the default value, the file-root constraint, the overwrite behavior, and how the path is used. This fully compensates for the schema's lack of parameter documentation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Stop the Playwright trace started by trace_start and write the zip to path.' It clearly identifies the tool's role as the counterpart to trace_start, so an agent can distinguish it from the sibling tools without ambiguity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description makes the usage context explicit: it must be called after trace_start, and stopping without a started trace is an error. It also states the path constraint relative to --file-root. It does not name alternative tools, but the lifecycle pairing with trace_start is clear enough that no alternative would apply.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
unrouteAIdempotent
Remove one route stub, or all of them.
| Name | Required | Description | Default |
|---|---|---|---|
| pattern | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover idempotency and the mutating nature of the operation. The description adds the key behavioral distinction that omitting the pattern removes all stubs, and that the tool affects 'stubs' rather than live routes or other resources, going slightly beyond the structured annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence that front-loads the operation and its scope. Every word adds meaning, with no filler or repetition of the tool name.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a low-complexity tool with one optional parameter, an output schema, and annotations covering mutability and idempotency, the description is nearly complete. It conveys the one/all distinction and the removal behavior; the only slight gap is not explicitly stating that omitting pattern is what triggers 'all', though the schema's null default supports that interpretation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must carry parameter meaning. It indicates that a single pattern removes one stub, while the tool can also remove all stubs, which maps directly to the optional 'pattern' parameter and its null default. It does not explicitly name the parameter, but the mapping is clear for a single-parameter tool.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Remove') and resource ('route stub') and immediately conveys the two possible scopes: one stub or all stubs. This makes it clearly distinct from the sibling tool 'route', which installs stubs, and leaves no ambiguity about the operation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context: use unroute when you want to remove either a specific route stub or every route stub. It does not explicitly name alternatives or state when not to use it, but for a direct inverse of the 'route' sibling, the usage context is clear enough.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
upload_filesADestructive
Set files on a file input, chosen by a CSS selector or by an intent Jev resolves. With an intent and several file inputs, the choice goes through the gate first. Every path must exist and sit inside --file-root (default: the working directory; set it explicitly if the server runs from / or $HOME); the call replaces the input's current selection. Returns {files, status, reason, target, executed} — status "review" or "blocked" means nothing was set.
| Name | Required | Description | Default |
|---|---|---|---|
| paths | Yes | ||
| intent | No | ||
| selector | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true and readOnlyHint=false, so the mutation is known. The description adds valuable context: it replaces the input's current selection, requires paths to exist and sit inside --file-root, and explains that status 'review' or 'blocked' means nothing was set. This goes beyond the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded with the core action. It packs a lot of behavioral detail into a few sentences. The return format is listed at the end, which is useful. Slightly dense but every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the output schema exists, the description doesn't need to explain return values in depth, but it does list the return fields. The tool has 3 params, 1 required, and the description covers the key semantics. The gate behavior and file-root constraint are important context that is included. Minor gaps: no mention of what happens when selector matches multiple inputs, or how intent resolution works in detail.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It explains 'paths' (must exist, inside --file-root), 'intent' (resolved by Jev, goes through gate), and 'selector' (CSS selector). It doesn't detail the exact format of paths or how intent resolution works, but it adds meaning beyond the bare schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'Set files on a file input, chosen by a CSS selector or by an intent.' It clearly distinguishes the two selection mechanisms and explains the gate behavior for intents. This is specific enough to differentiate from siblings like fill_form or drag.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explains when to use this tool: to set files on a file input, with a selector or intent. It doesn't explicitly name alternatives or exclusions, but the context of file inputs is clear. The gate behavior for intents is a useful usage detail, though it doesn't say when NOT to use it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
wait_forARead-onlyIdempotent
Wait until text or a selector appears, then report the page like page_state (element stamps included).
| Name | Required | Description | Default |
|---|---|---|---|
| text | No | ||
| selector | No | ||
| timeout_ms | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, covering safety. The description adds behavioral context: it waits for a condition and then reports page state including element stamps. It does not mention timeout failure behavior, but the annotations cover the non-destructive nature. This adds value beyond the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence that is front-loaded with the core action. It is concise, efficient, and contains no filler. Every part adds value.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is an output schema present, so return values need not be described. The description covers the main behavior (waiting and reporting) and references page_state for context. It does not detail what happens on timeout or if no condition is provided, but these are edge cases that an agent could infer from the parameters and general browser automation knowledge. Overall, it is sufficiently complete for typical use.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It explains 'text' and 'selector' as conditions to wait for, but does not mention the timeout_ms parameter at all. The description gives partial parameter meaning but misses the timeout, which is a significant gap given the lack of schema descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb (wait) and resource (page state), and clearly distinguishes itself from page_state by noting it waits first. It also mentions 'element stamps' which adds specificity about the output content. This is not a tautology and provides a clear purpose.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage: use this tool when you need to wait for a text or selector before getting a page snapshot, as opposed to page_state which reports immediately. It references page_state as a comparison, giving a sense of when to use this over the sibling. However, it does not explicitly state exclusions or alternative tools for other conditions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
1 tool update
- Changed
emulate5 fields changed- added
Input schema / properties / cpu_throttleAdded value: +{ + "anyOf": [ + { + "type": "number" + }, + { + "type": "null" + } + ], + "default": null, + "title": "Cpu Throttle" +} - added
Input schema / properties / download_kbpsAdded value: +{ + "anyOf": [ + { + "type": "integer" + }, + { + "type": "null" + } + ], + "default": null, + "title": "Download Kbps" +} - added
Input schema / properties / latency_msAdded value: +{ + "anyOf": [ + { + "type": "integer" + }, + { + "type": "null" + } + ], + "default": null, + "title": "Latency Ms" +} - added
Input schema / properties / network_conditionsAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "title": "Network Conditions" +} - added
Input schema / properties / upload_kbpsAdded value: +{ + "anyOf": [ + { + "type": "integer" + }, + { + "type": "null" + } + ], + "default": null, + "title": "Upload Kbps" +}
1 tool update
v0.2.0- Changed
network_detail2 fields changed- added
Input schema / properties / idAdded value: +{ + "anyOf": [ + { + "type": "integer" + }, + { + "type": "null" + } + ], + "default": null, + "title": "Id" +} - removed
Input schema / properties / indexRemoved value: -{ - "anyOf": [ - { - "type": "integer" - }, - { - "type": "null" - } - ], - "default": null, - "title": "Index" -}
33 tool updates
v0.1.0- First observed
browse - First observed
close_page - First observed
console - First observed
dialog_policy - First observed
dialogs - First observed
drag - First observed
emulate - First observed
fill_form - First observed
goal - First observed
goto - First observed
heap_snapshot - First observed
lighthouse - First observed
network - First observed
network_detail - First observed
new_page - First observed
outline - First observed
page_state - First observed
perf_metrics - First observed
press_key - First observed
read_js - First observed
resize - First observed
route - First observed
screenshot - First observed
scroll - First observed
select_page - First observed
styles - First observed
summary - First observed
tabs - First observed
trace_start - First observed
trace_stop - First observed
unroute - First observed
upload_files - First observed
wait_for
TDQS
Scored across 33 tools
Most tools have distinct purposes and detailed descriptions that clarify boundaries (e.g., page_state vs outline vs styles, browse vs goal). A few adjacent tools like wait_for/page_state and the many diagnostic tools could still be confused, but the set is largely unambiguous.
All names use snake_case, but conventions are mixed: action verbs (goto, scroll, press_key), noun-like observation tools (page_state, console, network), and noun_verb pairs (trace_start/trace_stop). Readable but not a consistent verb_noun pattern.
33 tools is too many for the apparent scope. While browser automation is broad, many advanced diagnostic tools (perf_metrics, heap_snapshot, lighthouse, trace_start/stop, network_detail) could be consolidated or made optional, making the surface heavy.
Core browser automation is well-covered: navigation, element interaction, forms, tabs, network/dialog observation and interception, emulation, and performance/tracing. Minor gaps exist (e.g., explicit hover/double-click, cookie/storage, iframe handling) but agents can often work around them.
Maintenance
Related MCP Connectors
Headless browser primitives for AI agents when sites need real JS rendering.
Stealth web automation for AI agents. Login, signup, navigate, screenshot.
Stealth web automation for AI agents. Login, signup, navigate, screenshot.
Provides cloud browser automation capabilities using Stagehand and Browserbase, enabling LLMs to i…
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceProvides complete browser automation capabilities for AI agents via 44 tools, including navigation, element interaction, state management, and session recording.387 npm1Apache 2.0
- AlicenseAqualityFmaintenanceStateful MCP server wrapping Playwright for browser automation. Provides tools to navigate, interact, and extract data from web pages via a persistent browser session.191,006 npmMIT
- AlicenseAqualityCmaintenanceEnables AI agents to fully control a browser for web automation, including navigation, clicking, typing, scrolling, screenshots, and DOM inspection, with session persistence and anti-bot bypass.4033 npmMIT
- AlicenseAqualityDmaintenanceProvides an LLM with a real Chromium browser to perform web tasks, recording every action into a structured trace for later verification of goal completion.537 npmMIT