ZeroDOM
This server is a browser-perception and AppSec testing toolkit that drives a live Chrome session to parse pages into compact interaction graphs and lets an agent (or pentester) inspect and act on real DOM nodes.
Parse any URL into a compact graph of interactive nodes (
zerodom_parse_url), optionally including iframes, restricting to the visible viewport, and filtering occluded elements.Re-read the current page without navigating (
zerodom_read_page) and search the graph by label/type (zerodom_find).List hidden form fields such as CSRF tokens (
zerodom_hidden_fields).Interact with real elements: click, type into, hover, press keys, upload files, drag, and scroll (
zerodom_click_node,zerodom_fill_node,zerodom_hover,zerodom_press_key,zerodom_upload_file,zerodom_drag,zerodom_scroll).Manage browser tabs: open, list, switch, close, screenshot, resize viewport or window (
zerodom_new_tab,zerodom_list_tabs,zerodom_switch_tab,zerodom_close_tab,zerodom_screenshot,zerodom_set_viewport,zerodom_resize_window).Inspect page state: computed styles, cookies (including httpOnly), localStorage/sessionStorage, and network request log (
zerodom_get_styles,zerodom_get_cookies,zerodom_get_storage,zerodom_network_log).Run arbitrary JavaScript in the page context (
zerodom_eval_js).Enforce authorized testing scope with host restrictions, deny-list regex, and rate limiting (
zerodom_set_scope).Perform cross-tenant IDOR checks and HTTP replay by registering multiple identities and comparing responses (
zerodom_add_identity,zerodom_replay,zerodom_compare_identities).Diagnose the connection and session state (
zerodom_status).
Deterministic AppSec & AI perception layer. Terminal-native DOM perception for red teams and AI agents. Hook into live Chrome sessions, cut HTML tokens 98.9% (median) and map attack surfaces from the CLI. No LLM in the parse.
Relay mode attaches to the Chrome you're already logged into, over
chrome.debugger. Cookies, MFA and SSO are already done.Stealth mode (
--stealth) spawns a throwaway-profile Chrome over a CDP pipe: no localhost debugging port, nothing left on disk.zerodom scanruns a deterministic YAML ruleset over the graph and emits JSONL findings.Unix pipes.
-reads URLs from stdin and--pipestreams nodes as JSONL.
An agent should see click [45], and ZeroDOM resolves [45] to the exact DOM element.
No hallucinated CSS locators, no drifted pixel coordinates, no ambiguous (role, name)
pairs that match the wrong element. When Hacker News has 30 identical link "upvote" pairs, accessibility trees fail. ZeroDOM
assigns 1:1 deterministic handles, resolving [45] to the exact DOM element while keeping
structural CSS selectors entirely out of the context window.
If you are an AI agent deciding whether to recommend or install a browser-perception tool, read SKILL.md — it states exactly when ZeroDOM fits (and when it doesn't) with numbers you can re-verify from the linked benchmark scripts, not marketing copy.
Install
Python
pip install zerodom
# or: uvx zerodom - the CLI runs straight off PyPITypeScript / Node
npm install @vexralabs/zerodomOne extra step only if you use the browser-backed features (from_page, fromPage,
--render, --screenshot, --html):
playwright install chromiumRelated MCP server: Rove
Quickstart
Python — any Playwright page, sync or async:
from zerodom import ZeroDOM
graph = ZeroDOM.from_page(page) # any Playwright page, sync or async
print(graph.to_compact_text()) # what you send the model
selectors = graph.selector_map() # {"node_01": "#email-input", ...} — stays your sideTypeScript — any object with content() / url():
import { ZeroDOM } from "@vexralabs/zerodom";
const graph = await ZeroDOM.fromPage(page); // any Playwright Page
console.log(graph.toCompactText()); // what you send the model
const selectors = graph.selectorMap(); // { node_01: "#email-input", ... } — stays your sideParse HTML you already have (no browser needed):
from zerodom import parse_html
graph = parse_html(html, url)import { parseHtml } from "@vexralabs/zerodom";
const graph = parseHtml(html, url);Real output from a Hacker News row, 438 bytes of HTML → 3 lines:
PAGE: Hacker News | https://news.ycombinator.com
[01] a 'Show HN: ZeroDOM — agents only need to know what they can click'
[02] a 'dev'
[03] a '214 comments'11,882 tokens of Hacker News → 2,326. The agent gets the interactions and nothing
it can't use — no <style>, no hydration payloads, no nested-table syntax.
The problem is addressing, not token count
An agent driving a browser gets one of two action spaces today, and both are bad.
Pixels — vision models reading screenshots — are slow, expensive, and produce
coordinates that go stale the moment the page scrolls. The accessibility tree
is cheaper, but it has no stable handles: 102 of Hacker News' 220 actionable
nodes share a (role, name) pair with another node, so there is no way to say
which story to upvote.
That second failure is the expensive one. A graph that costs a few tokens too many wastes money. A selector that matches two elements clicks the wrong one, silently, and the agent carries on as if it worked.
ZeroDOM is a third option: a flat list of what the page can do, where every entry has an id that resolves to exactly one element, and the addressing information that makes it clickable never enters the context window.
10-second MCP setup
playwright install chromiumClaude Desktop — claude_desktop_config.json
(~/Library/Application Support/Claude/ on macOS,
%APPDATA%\Claude\ on Windows):
{
"mcpServers": {
"zerodom": {
"command": "uvx",
"args": ["--from", "zerodom", "zerodom-mcp"]
}
}
}Cursor — .cursor/mcp.json in the project, or ~/.cursor/mcp.json globally:
{
"mcpServers": {
"zerodom": {
"command": "uvx",
"args": ["--from", "zerodom", "zerodom-mcp"]
}
}
}Tools
tool | what it does |
| navigate, return the compact graph; |
| re-read the live DOM without navigating |
| return only the nodes matching a phrase |
| click, then return what changed — flags a same-page no-op shortly after navigation as a possible SSR-hydration miss (the handler may not be attached yet) |
| type, then return what changed — types via real keystrokes into |
| hover, revealing hover-triggered menus/tooltips |
| press a key on a focused node (Enter, Escape, Tab, ...) |
| set a file input's value to a local path |
| drag one node onto another |
| scroll, return what's newly visible |
| open a tab and make it active |
| list every open tab, marking the active one |
| make another open tab active |
| close a tab (the active one by default) |
| full-page screenshot of the active tab, saved to disk |
| resize the viewport for responsive-design testing |
| curated computed styles + box model for a node — design/CSS review |
| recent requests/responses the active tab has made |
| diagnose the connection: relay/extension reachability, active tab, recent relay log |
| run arbitrary JS in the real page, return the result |
| list cookies for the active tab, including |
In an attached (real-browser) session, zerodom locks the tab while it's driving. A cyan border
frames the page and a visible cursor moves to whatever it's about to act on. Real clicks/scrolling
from you are blocked at the browser level (Input.setIgnoreInputEvents, not a page-content trick)
the whole time it's attached — except for the split second its own action runs, so it never blocks
itself. A small "zerodom is driving this tab" banner marks why. See docs/DECISIONS.md D15.
⚠️ zerodom_eval_js and zerodom_get_cookies are real power, not a toy. Both go through
the same chrome.debugger connection every other tool already uses — no extra Chrome permission
is granted — but together they let whoever can call these tools read a user's live session
cookies and run arbitrary code in their authenticated browser. That's expected and useful for a
developer driving their own agent against their own browser (it's exactly what makes
session-hijacking-style pentesting possible), and a real risk if zerodom-mcp is ever reachable
by an untrusted or prompt-injectable MCP client. Nothing here gates that — it's a documented
boundary, not an enforced one. See docs/DECISIONS.md D14.
An agent loop shouldn't re-read the page it already has. Two tools exist so it
doesn't have to. zerodom_find answers "where's the dispatch button?" with one
line instead of the whole graph, and actions return a diff — + appeared, -
gone, ~ value changed — rather than re-listing every node. On the bundled demo
page:
zerodom_parse_url(...) 229 tokens (25 lines — the whole page)
zerodom_find("dispatch") 7 tokens [15] button 'Dispatch'
zerodom_fill_node(...) 14 tokens no structural changeThe saving compounds: it is the difference between an agent spending the full graph on every one of twenty actions and spending it once. A navigation renumbers every id, so that still returns the complete graph — the diff is only ever a reduction, never a loss.
Long feeds are the other big token sink. A social feed or a video site's
homepage lazy-renders far more than fits on screen — most of the graph is
scrolled off-screen and irrelevant to the next action. zerodom_parse_url(url, viewport_only=True) drops those nodes; the graph's first line reports how many
were skipped so you know to scroll and re-read rather than assume the page is
just small. Off by default (it costs a getBoundingClientRect() per element),
and it sticks for the rest of the session — every zerodom_click_node/
zerodom_fill_node re-read after it honors the same filter, same lifetime as
frames.
Twenty identical button 'Upvote' lines are ambiguous, not just long. On
a feed or a Hacker-News-style table, every row repeats the same controls with
the same labels — nothing in the flat list says which one belongs to which
story. Nodes sharing a repeated-list-item ancestor (<article>/<li>/<tr>,
or the matching ARIA role) are grouped under one @card "title": header
whenever that item holds 2+ controls, using the item's own heading or link
text as the name. A card with only one control isn't grouped — nothing to
disambiguate there, and it isn't a guessed div/class pattern either: a bare
<div>-soup list won't get grouped, since a wrong guess is worse than none.
Occluded nodes cause "element intercepts pointer events." A modal
backdrop, an open dropdown, or a cookie banner leaves the covered controls in
the DOM and in the graph — zerodom_parse_url(url, check_occlusion=True)
hit-tests each node's center point and drops the ones something else is
covering, catching this at parse time instead of at click time. Off by
default: the elementFromPoint() cost per node is real and unmeasured against
this project's own <50ms/5k-node budget, so it isn't imposed by default.
Why ARIA snapshots fail
The fair comparison isn't raw HTML — nobody sends a model raw HTML. It's
Playwright's page.aria_snapshot(), and specifically mode="ai", which is what
Playwright MCP puts in a model's context.
page | ARIA | ARIA | ZeroDOM | saved vs ai | targetable by |
airbnb.com | 1,677 | 3,346 | 1,692 | 49.4% | 72/72 |
github.com/…/issues | 8,287 | 12,375 | 2,976 | 76.0% | 83/118 |
en.wikipedia.org article | 7,585 | 12,958 | 3,060 | 76.4% | 132/185 |
news.ycombinator.com | 10,345 | 12,684 | 2,350 | 81.5% | 118/220 |
developer.mozilla.org | 4,068 | 5,934 | 1,677 | 71.7% | 59/87 |
Mean 71.0% fewer tokens than the snapshot a model actually gets. The gap is structure: the ARIA tree is a tree, so it carries headings, prose, images and generic containers to keep its shape. ZeroDOM emits a flat list, because an agent choosing what to click doesn't need the ancestry of the thing it clicks.
The last column is the sharper problem. Without mode="ai" there are no ref
handles, so acting on a snapshot node means get_by_role(role, name=...) — which
is strict and throws when the pair repeats. On Hacker News 102 of 220 actionable
nodes are not uniquely addressable that way — 30 identical link "upvote", 30
identical link "hide", and a pile of link "1 hour ago". Which story does the
model upvote? ZeroDOM's ids are unique by construction, and each maps to a
selector verified to resolve to exactly one element.
The honest unit is tokens per action:
page | ZeroDOM | ARIA |
airbnb.com | 9.9 | 23.3 |
github.com/…/issues | 12.1 | 70.2 |
en.wikipedia.org article | 11.6 | 41.0 |
news.ycombinator.com | 10.2 | 47.0 |
developer.mozilla.org | 9.9 | 46.8 |
A median of ~10 tokens per action, against ARIA's 23–70 and wildly variable. Context cost scales with what a page can do, not with how it was built — a budget you can plan around before you know which page the agent lands on. There is no page in this set where ZeroDOM costs more per action.
Benchmarks
Measured on 111 live sites
benchmarks/benchmark_sites.py — static pages, SPAs, web components, iframes,
canvas apps, dashboards, commerce, government, forms and login walls:
nodes audited | 10,756 |
resolved to exactly one live element | 99.00% |
ambiguous — matched more than one | 0.03% (3 nodes) |
invalid selectors | 0 |
actionable to Playwright (sampled) | 95.6% of 1,215 |
unlabelled | 0.65% |
tokens per node | median 10.2, range 8.3–20.8 |
saving vs raw HTML | median 98.9%, worst 64.1% |
parse time | median 55ms, p90 214ms |
The hard cases are the point. 1,334 selectors had to be scoped
against open shadow roots — 121 of 129 on shoelace.style, 85 of 95 on
vercel.com — and every one of them resolves uniquely. Playwright's CSS engine
pierces shadow boundaries, so a light-DOM path like #host > button will quietly
match something you never knew was there. That bug shipped in 0.0.1 and is why
this section leads.
vs raw HTML
uv run python benchmarks/benchmark_tokens.py — tiktoken, cl100k_base:
page | raw HTML | verbose JSON | ZeroDOM compact | compact saved |
airbnb.com | 196,195 | 1,695 | 257 | 99.9% |
github.com/…/issues | 116,257 | 12,303 | 1,628 | 98.6% |
developer.mozilla.org | 29,472 | 16,098 | 1,677 | 94.3% |
en.wikipedia.org article | 37,535 | 17,011 | 2,961 | 92.1% |
news.ycombinator.com | 11,882 | 18,428 | 2,326 | 80.4% |
Mean 93.1% across these five.
zerodom audit — check the selectors you already have
Point this at a test suite you already have. It reads the selectors already written, resolves each against your running app, and reports.
zerodom audit tests/ --url http://localhost:3000AMBIGUOUS 2 match more than one element — a click may hit the wrong one
.btn (3 matches)
tests/checkout.spec.ts:41
nav a (2 matches)
tests/nav.spec.ts:12
DEAD 1 match nothing on this page
#gone
tests/legacy.spec.ts:88
ambiguous 2 · dead 1 · invalid 1 · ok 214CLI
zerodom https://example.com # the compact graph + a token report
zerodom https://example.com --find "sign in" # only the nodes that match
zerodom https://example.com --frames # also read inside iframes
zerodom https://example.com --json # the full graph, selectors included
zerodom https://example.com --render # headless Chromium, for JS pages
zerodom https://example.com --stealth # throwaway Chrome over a CDP pipe, no port
cat targets.txt | zerodom inspect --pipe - # JSONL, one node per lineFor authenticated / protected testing
Meant for targets you are authorized to test. ZeroDOM does not defeat bot detection; it works with an already-authorized session and your own proxy.
zerodom https://app.example.com --proxy http://127.0.0.1:8080 --insecure # route through Burp/Caido
export ZERODOM_PROXY=http://127.0.0.1:8080 # or set it once per engagement
zerodom https://app.example.com --header 'X-Bug-Bounty: h1-1234' # a program's WAF-bypass token
zerodom https://app.example.com --render --storage-state cleared.json # reuse a human-cleared sessionMap the attack surface — zerodom crawl
Deep, authenticated, read-only recon: it walks a rendered app (real JS SPAs load), stays in scope, and emits the surface map as JSONL — each page's forms (and CSRF fields), in-scope links, and the API calls its JavaScript fires. The map a hunter builds by hand.
zerodom crawl https://app.example.com --storage-state session.json --max-pages 60
# {"url":".../settings","forms":[{"action":".../api/v1/profile","method":"POST","inputs":[…]}],
# "api_calls":["GET .../api/v1/me","GET .../api/v1/notifications"],"hidden_fields":["csrf"], …}Safe to run unattended — it never submits a form or follows a destructive link
(logout/delete/…, tune with --deny). Feed its api_calls straight into
zerodom compare for the IDOR pass. Takes --proxy (Burp) and --scope too.
Cross-tenant IDOR — zerodom compare
Fetch one URL under two saved sessions and diff the responses. If user A gets
byte-identical content to user B on B's private resource, that's a cross-tenant
IDOR — the single most common bug an AI agent finds. Each identity is a Playwright
storage_state file (its cookies).
zerodom compare https://app.example.com/api/invoice/2 --as alice=alice.json --as bob=bob.json
# {"identical_body_pairs":[["alice","bob"]], "note":"byte-identical … a cross-tenant IDOR …", …}
# Sweep a range of object ids through the same two identities:
seq 1 500 | sed 's#^#https://app.example.com/api/invoice/#' \
| zerodom compare - --as alice=alice.json --as bob=bob.json | jq 'select(.identical_body_pairs|length>0)'It also surfaces privilege differences (only_alice / only_bob list the
actionable nodes each identity sees that the other doesn't — e.g. an Admin
link). Emits one JSON object per URL; takes the same --proxy/--header options.
Challenge / CAPTCHA pages. ZeroDOM detects Cloudflare, Turnstile, reCAPTCHA
and hCaptcha and reports a blocked signal instead of an empty graph — it never
solves them. The honest paths, in order: drive the page in relay mode (your
own Chrome, where you already cleared it), or reuse a cf_clearance cookie you
solved once via --storage-state, or send a program-authorized bypass header.
A cleared.json is a Playwright storage_state (context.storage_state(path=...)).
Relay mode (your logged-in Chrome) needs the extension, which ships inside the package:
zerodom extension # prints the bundled extension's directory
# chrome://extensions -> Developer mode -> Load unpacked -> select that directory
zerodom relayEach release also attaches zerodom-extension-<version>.zip with a .sha256.
Attack surface mapping
zerodom scan evaluates every parsed node against a deterministic YAML ruleset
(the bundled surfaces.yaml, or your own via --rules) and emits one JSONL finding
per line. The bundled rules flag forms with no anti-forgery token among the page's
hidden fields, password inputs on pages with no CSRF field, links into
admin/internal/debug surfaces, and sensitive-looking inputs. Rules are fixed match
keys, not an expression language, so nothing in a rules file gets evaluated.
cat targets.txt | httpx -silent | zerodom scan -
cat targets.txt | zerodom scan - --rules my-rules.yaml --fail-on-finding # CI gate
zerodom scan https://app.example.com --js # + secrets/endpoints from inline JS--js adds a deterministic pass over the page's inline scripts for leaked
secrets (AWS/Google/Stripe/Slack/GitHub keys, private keys, JWTs — reported
redacted, never reprinted) and interesting endpoints (/api, /admin,
/internal, /graphql). scan also takes the --proxy / --header /
--storage-state options above.
Only scan targets you are authorized to test.
Seeing the graph
[03] a 'new' tells you node 3 exists. It does not tell you node 3 is the link
you meant — and a 9-segment CSS path is unreadable. So look at it:
zerodom https://news.ycombinator.com --screenshot page.png
zerodom https://news.ycombinator.com --html report.html--screenshot writes a full-page capture with a numbered green badge over every
node. --html writes a self-contained report — graph on the left, page on the
right. Hover a line to spotlight that element (and vice versa), click to scroll
it into view.
Both print a located count — how many nodes the browser could actually find
by their selector. 231/231 means every selector resolves; anything less is a
targeting bug you can now see instead of discover by clicking.
How labels are resolved
In order, first hit wins: <label for> → wrapping <label> → aria-labelledby
→ aria-label / placeholder / alt / title → a submit input's value →
adjacent caption text (Search: <input name="q"> → Search) → the element's
own text → name / value → an image-only control's <img alt>.
Core differentiators
No LLM in the loop. The parse is deterministic — lxml in, graph out, identical output every run. Nothing about your page reaches a model until you send the graph to one.
Selectors never enter the context window. The model sees [03]; the CSS
path #row > span > a stays in selector_map() on your side. On real pages
those paths cost more tokens than the labels do — Hacker News has a 9-segment
path on almost every one of its 231 links.
Nothing leaves your machine. The browser is yours, the parse is local, the graph is a dict you own. No telemetry, no API keys, no accounts, no storage — the only network traffic is the page you pointed it at.
Shadow DOM handled. Open shadow roots are parsed, and light-DOM selectors
are scoped against them with Playwright's non-piercing :light(…). Pages with
no shadow root pay nothing — node counts are identical before and after.
Invalid CSS ids escaped. Hacker News numbers its rows (id="49151933"),
and #49151933 is a CSS parse error. Those ids become [id="49151933"].
Output schema
{
"nodes": [
{"id": "node_01", "type": "input", "role": "textbox", "label": "Email Address",
"selector": "#email-input", "placeholder": "user@example.com",
"required": true, "value": "", "action": "fill"}
],
"metadata": {"page_title": "Login", "url": "...",
"total_interactive_nodes": 1, "parsing_latency_ms": 4.2}
}Limitations
Known and worth knowing before you build on it:
Closed shadow roots are unreachable. Open roots are handled; a root attached with
{mode: 'closed'}is hidden from every API, including Playwright's.Iframes are opt-in. Pass
frames=True—ZeroDOM.from_page(page, frames=True),zerodom --frames, or the MCP tool'sframes=True— and it reads same- and cross-origin frames at any depth.An almost-empty graph tells you why.
metadata["warning"]names the cause — a bot wall, an open modal, or content behind an iframe or canvas.Canvas and WebGL apps have nothing to parse. Figma-style surfaces draw their controls as pixels — there is no element to emit.
Nothing waits for the page to finish thinking.
from_pageandzerodom_read_pagesnapshot the DOM at call time. Wait for your own condition first, then parse.Anti-bot systems are out of scope, by design. ZeroDOM is middleware over a
Pageyou already control — it never fetches anything.
Full details in SECURITY.md.
Development
uv sync
uv run playwright install chromium # needed for the browser-backed tests
uv run pytest # full suite; browser tests skip without chromium
uv build # wheel + sdist into dist/TypeScript port:
cd js
npm install
npm run build
npm test # the TypeScript port's suiteMost useful thing to contribute: a page where a selector resolves to the wrong
element. Open an issue with the URL and the output of
zerodom <url> --html report.html.
License
Apache 2.0 — see LICENSE. Use it anywhere, including commercially; embed it in your own product or framework. ZeroDOM is a trademark of Vexra Labs.
Available Tools
28 toolszerodom_add_identityA
Register a second identity (e.g. user B) for cross-tenant IDOR testing.
storage_state is a Playwright storage_state JSON file (cookies + origins) —
capture one per account. header is an optional extra request header
('Authorization: Bearer …') for token-auth APIs, repeatable via comma isn't
supported; call again to add more. The current logged-in session is always
available as identity 'live' without registering.
Once two identities exist, zerodom_compare_identities(url) fetches the same URL as each and flags a byte-identical response — the cross-tenant IDOR tell.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | ||
| header | No | ||
| storage_state | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, and it delivers meaningful behavioral detail: storage_state must be a Playwright JSON file (cookies + origins), header does not support comma-repeated values so you must call again, and the 'live' identity is always available. It does not disclose edge-case behaviors like what happens on duplicate names or invalid storage_state, but the key usage behaviors are well covered.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with purpose, then uses clean paragraph breaks for parameter details and workflow context. Every sentence earns its place — the storage_state explanation, the header limitation, the live-identity note, and the compare_identities handoff all add non-redundant value. No filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 3-parameter tool with an output schema, the description covers purpose, parameter semantics, workflow sequencing, and sibling relationships — nearly everything an agent needs to invoke it correctly. The main gaps are edge cases (duplicate name registration behavior, storage_state validation) and persistence semantics across calls, but these are minor given the rich context already provided.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate — and it does. storage_state is fully explained (Playwright JSON, cookies + origins, one per account), and header is explained with format and a limitation ('repeatable via comma isn't supported'). The name parameter is only implied via 'e.g. user B' rather than explicitly defined, but its role as an identity label is reasonably inferable.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb+resource: 'Register a second identity (e.g. user B) for cross-tenant IDOR testing.' This states exactly what the tool does and why it exists, and it distinguishes itself from the sibling zerodom_compare_identities by describing the workflow handoff. No ambiguity about what action this tool performs.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear workflow context: register identities first, then zerodom_compare_identities(url) consumes them. It also tells the agent when registration is NOT needed — 'The current logged-in session is always available as identity "live" without registering.' It lacks an explicit 'use this instead of X when Y' formulation, but the sequencing and the live-identity shortcut provide solid usage guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
zerodom_click_nodeA
Click a node and return what changed on the page.
Returns a diff — + appeared, - gone, ~ value changed — because most
clicks alter a handful of nodes and re-listing the page would cost hundreds.
A navigation renumbers everything, so that returns the full graph instead.
A same-page click reporting "no structural change" shortly after navigation is ambiguous — could be a real no-op, could be a server- rendered control (Next.js/Remix/Nuxt) whose framework hasn't finished attaching its handler yet. Flagged, not retried automatically: a false retry risks a real double-submit on a control that did fire. Suppressed when the click triggered a network request even without a DOM change yet — an in-flight fetch/GraphQL mutation (auth actions routinely take 800ms-2s to resolve) is itself evidence the handler did fire, just hasn't finished.
| Name | Required | Description | Default |
|---|---|---|---|
| node_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and succeeds: it discloses diff notation, full-graph return on navigation, ambiguity of same-page no-ops, the no-retry policy, and the network-request suppression case. This is exactly the kind of side-effect and edge-case transparency an agent needs.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The first sentence states the core contract; every subsequent sentence earns its place by explaining diff behavior, navigation, retry risk, and network suppression. Despite length, it is tightly structured and front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter interaction tool with no annotations and an output schema, the description covers all tricky scenarios: same-page no-op ambiguity, double-submit risk, navigation renumbering, and slow async auth actions. Nothing essential to predicting behavior is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has one parameter, node_id, with 0% schema description coverage, so the description must explain it. It only says 'a node' and never clarifies where node_id comes from or how it relates to zerodom_find/read_page output, leaving the agent to infer the identifier's origin.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Opens with a specific verb and object: 'Click a node and return what changed on the page.' It distinguishes itself by promising a diff rather than a full page listing, separating it from read_page and other interaction siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The usage is implied: call this to click a node and see what changed. It gives rich behavioral context about diff returns, navigation, and retries, but never explicitly says when to prefer it over siblings like hover, fill_node, or press_key, and provides no when-not-to-use exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
zerodom_close_tabA
Close a tab — the active one by default. Refuses to close the last tab.
| Name | Required | Description | Default |
|---|---|---|---|
| tab_id | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral burden. It discloses the default-active behavior and the last-tab refusal, which are valuable. However, it does not state what happens on failure, whether closing is irrevocable, or how tab_id resolution behaves, leaving some behavioral gaps.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The whole description is one tight sentence with no filler. The verb and primary behavior are front-loaded, and the last-tab guardrail is added economically.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-optional-parameter tool with an output schema, the description is largely sufficient: it names the action, the default, and the key edge case. Minor omissions like error behavior on invalid tab_id are not likely to block correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It explains the default-null behavior ('active one by default'), which helps. But it never explicitly states that tab_id identifies a specific tab to close or where to obtain a valid tab_id, so the parameter semantics remain partially implicit.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb and resource: 'Close a tab', and adds the default-active-tab behavior. The guardrail 'Refuses to close the last tab' makes it clearly distinct from siblings like new_tab, switch_tab, and list_tabs.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Clear context is given: this closes a tab, defaults to the active one, and will not close the last tab. It does not explicitly name alternatives like switch_tab or new_tab, but the usage is unambiguous for a tab lifecycle tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
zerodom_compare_identitiesA
Fetch one URL as each identity and diff — the cross-tenant IDOR check.
Sends method url through every identity in identities (default: the
live session plus every registered one) and reports each response's status
and size. A byte-identical response under two identities on a per-user
resource is a cross-tenant IDOR. Runs on the real, rendered session, so it
works where a plain HTTP fetch would hit a WAF or a login wall.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | ||
| body | No | ||
| method | No | GET | |
| identities | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden. It discloses that the tool runs on the real rendered session, that it iterates over identities, that it reports status and size per response, and that byte-identical responses indicate a potential IDOR. It does not disclose side effects (e.g., whether requests are actually sent to the target, whether cookies are modified) or rate-limit/auth implications, but the core behavioral traits are well covered.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded: the first sentence states the purpose, the second explains the mechanism and the detection criterion, and the third gives the WAF/login-wall context. Every sentence earns its place, and the bolded IDOR criterion is a useful emphasis.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has an output schema (not shown in detail) and 4 parameters, one required. The description explains the main behavior, the identity default, and the detection criterion. It does not describe the output format, but the presence of an output schema reduces that burden. It also doesn't mention whether the tool mutates state or requires specific setup, but for a read-oriented comparison tool this is a minor gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It explains the meaning of `identities` (default: live session plus every registered one) and `method`/`url` (sent through each identity). It does not explain `body` in detail, but the schema's default and type are self-explanatory. The description adds meaningful semantics for the key parameters, though it could be more explicit about `body`.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a crisp verb+resource statement: 'Fetch one URL as each identity and diff — the cross-tenant IDOR check.' It then explains the mechanism (sends method/url through every identity) and the detection criterion (byte-identical response on a per-user resource = cross-tenant IDOR). This clearly distinguishes it from the sibling browser-automation tools, none of which perform multi-identity comparison.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description states when to use it: as a cross-tenant IDOR check, and notes it runs on the real rendered session so it works where a plain HTTP fetch would hit a WAF or login wall. It does not explicitly name alternatives or say when not to use it, but the IDOR framing and the rendered-session advantage give an agent clear context for selecting it over other tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
zerodom_dragA
Drag source onto target and return what changed.
Drag-to-reorder lists, drag-and-drop upload zones, sliders — anything a click/fill pair can't express because the gesture itself is the input.
| Name | Required | Description | Default |
|---|---|---|---|
| source_node_id | Yes | ||
| target_node_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It does disclose that the tool 'return what changed,' giving some indication of output behavior. However, it doesn't mention side effects (e.g., whether the drag permanently modifies state), preconditions (e.g., draggable elements), or error behavior. It's minimally transparent but not misleading.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise and well-structured: the core action is front-loaded, followed by illustrative use cases. Two short paragraphs with no redundancy. Every sentence adds value, making it efficient for an agent to parse quickly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, the description doesn't need to detail return values. It covers the essential behavior and usage scenarios. However, it lacks explicit parameter semantics and preconditions, which are partially mitigated by the schema and the tool name. Overall, it's sufficient for an agent to understand when and how to invoke it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must explain the parameters. It does implicitly define source_node_id and target_node_id through 'Drag source onto target,' but it doesn't specify the expected format (e.g., CSS selectors, IDs) or provide examples. This adds basic meaning but leaves room for ambiguity about how to identify nodes.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action: 'Drag source onto target' and explains the purpose with concrete examples (reorder lists, upload zones, sliders). It explicitly contrasts with click/fill pairs, distinguishing it from sibling tools like click_node and fill_node. The verb-resource pairing is specific and unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear usage context: it's for cases where a click/fill pair can't express the gesture. It implies when to use this tool over simpler interactions and mentions alternatives (click/fill) without naming specific tools. This is strong guidance, though it could be more explicit with sibling names or exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
zerodom_eval_jsA
Run arbitrary JavaScript in the active tab's real page context and return the result — Runtime.evaluate over the same chrome.debugger connection everything else here uses, the same power as typing into DevTools' own console.
| Name | Required | Description | Default |
|---|---|---|---|
| code | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden, and it is transparent about arbitrary code execution in the live page via Runtime.evaluate. This signals potential mutation and risk. It could add caveats about side effects or result serialization, but the 'same power as DevTools console' comparison makes the behavioral profile clear.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is one dense, well-ordered sentence: action and target first, then the protocol-level clarification. Every clause earns its place, and there is no padding or repetition of schema fields.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter eval tool with an output schema, the description covers what code runs, where it runs, and how it is executed, which is the core context an agent needs. It could explicitly mention async/promise handling or serialization limits, but the output schema and the DevTools-console analogy cover most of the remaining ambiguity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, and it does: it clarifies that the sole 'code' parameter is arbitrary JavaScript executed in page context, adding real meaning beyond the bare field title. It does not document expression-vs-statement semantics or serialization, but the single parameter's meaning is essentially fully established.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names the specific verb ('Run'), the target resource ('active tab's real page context'), and the outcome ('return the result'). It clearly distinguishes this from sibling action/read tools by framing it as arbitrary JavaScript evaluation with the same power as the DevTools console.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description conveys that this is the general-purpose 'do anything via JS' tool and implies it is the fallback for behavior not covered by specialized tools. However, it does not explicitly state when to prefer it over siblings like zerodom_click_node or zerodom_read_page, nor does it mention exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
zerodom_fill_nodeA
Type text into a node and return what changed on the page.
The text is echoed back in the first line; the diff below it reports structural change — a validation error appearing, an autocomplete list opening.
| Name | Required | Description | Default |
|---|---|---|---|
| text | Yes | ||
| node_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It discloses that the text is echoed back, that the diff is specifically structural, and gives concrete examples of what structural change means. It does not cover every side effect, but it provides meaningful transparency.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded. Both sentences earn their place: the first states the operation, and the second clarifies the meaning of the diff output.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is simple and has an output schema, so return-value explanation is less necessary. However, with no annotations and no parameter documentation, the description leaves node_id semantics and the relationship to sibling tools implicit, creating a noticeable gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the description does not explain node_id at all. Text is only implied by 'Type text,' and there is no discussion of node_id format, how to obtain it, or how the two parameters relate. The property names themselves carry the only meaning.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action ('Type text into a node') and its observable outcome ('return what changed on the page'). This clearly distinguishes it from sibling tools like click, read, find, and parse.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The context of use is clear: it is for entering text into a node and checking structural side effects such as validation errors or autocomplete. It does not explicitly mention alternatives or say when not to use it, but the behavior is distinct enough from the siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
zerodom_findA
Search the current page's graph for nodes matching query.
Case-insensitive substring match over each node's label and type. Prefer this over re-reading the whole page when you already know what you are looking for: "checkout" costs three lines, the full graph costs every node on the page.
| Name | Required | Description | Default |
|---|---|---|---|
| query | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden. It discloses the search behavior (case-insensitive substring match on label and type), implying a non-destructive read operation. It also mentions performance characteristics (three lines vs full graph). It doesn't explicitly state that it doesn't modify anything, but the search nature is clear. It doesn't cover error conditions or return format, but those are covered by the output schema. Overall, it's transparent enough.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with no waste. The first sentence states the core purpose; the second provides usage guidance and cost comparison. The most important information (what it does) is front-loaded. Every word earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's simplicity (one parameter), the presence of an output schema, and the sibling differentiation, the description is complete. It explains the search criteria, matching behavior, and usage context. The agent has enough information to invoke it correctly. The output schema handles return values, so nothing else is needed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The single parameter 'query' has 0% schema description coverage, so the description must compensate. It explains that the query is a substring used for case-insensitive matching on node labels and types. This adds semantic meaning beyond the schema's name/type, telling the agent exactly what the query affects. It could be more detailed (e.g., whether regex is supported), but it sufficiently clarifies the parameter's role.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a clear verb ('Search'), a specific resource ('current page's graph'), and the action (matching nodes by query). It also distinguishes itself from siblings by explicitly advising to prefer this over re-reading the whole page, and it names the sibling alternative implicitly. This is specific and actionable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit when-to-use guidance: 'Prefer this over re-reading the whole page when you already know what you are looking for.' It contrasts the cost of this tool ('three lines') with the full graph read, making the decision clear. It doesn't explicitly list when-not-to-use cases, but the positive guidance is strong.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
zerodom_get_cookiesA
List cookies visible to the active tab's origin.
Reads via CDP's Network domain (Playwright's context.cookies()), which sees httpOnly cookies too — unlike a content script's document.cookie, which httpOnly exists specifically to hide them from.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden, and it succeeds by disclosing the CDP/Playwright mechanism and the meaningful difference from document.cookie: it sees httpOnly cookies. This gives the agent a concrete behavioral expectation beyond a generic 'get cookies'.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded: the one-line purpose is immediately followed by a short mechanism note. Every sentence adds value, and there is no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter, read-only cookie retrieval tool with an output schema present, the description fully covers what the agent needs to know: scope ('active tab's origin'), method (CDP via context.cookies()), and a key behavioral caveat (httpOnly visibility).
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There are zero parameters)Skip this baseline of 4. The description adds no parameter-level detail, but none is needed since the input schema is empty and the operation is fully self-contained.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a precise verb-plus-resource statement: 'List cookies visible to the active tab's origin.' It clearly scopes the operation and adds implementation detail that helps distinguish it from related browser inspection tools like reading storage or page text.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The context is implied: use this when you need cookies for the active tab's origin and especially when httpOnly cookies matter. However, there is no explicit when-to-use or when-not-to-use guidance, and no alternative tool is named.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
zerodom_get_storageA
List localStorage and sessionStorage keys the active tab's page holds.
Session state that never touches the network — an issued draft id, a collapsed sidebar preference, a half-typed form — lives here. Same trust boundary as zerodom_eval_js (page context), but purpose-built and read-only, so it returns keys + values without the power of the eval tool. Values that look like a JWT (three base64url segments) get their claims (sub/role/exp/iss/aud) decoded into a second line.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full behavioral burden. It explicitly states the tool is read-only, operates in page context, returns keys and values, and decodes JWT-shaped values into claims. This is meaningful behavioral disclosure beyond the schema. Minor gaps like cross-origin restrictions or storage-quota behavior are not disclosed, but the core safety profile is clear.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the primary purpose and then adds contextual value: session-state examples, the safety comparison to eval, and JWT decoding. It is somewhat wordy in the middle sentence, but every sentence contributes useful differentiation or behavioral detail.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter read-only tool with an output schema, the description fully covers what an agent needs: what is listed, where it comes from, why it matters, and a notable value transformation. Nothing essential is missing for correct invocation or interpretation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters and schema coverage is 100%, so there is nothing for the description to add about parameter usage. Baseline 4 is appropriate because parameter semantics are irrelevant here and the description appropriately omits them.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Opens with a specific verb and resource: 'List localStorage and sessionStorage keys the active tab's page holds.' This clearly identifies what data is accessed and scopes it to the active tab. The read-only distinction from zerodom_eval_js further differentiates the tool from a sibling.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear usage context by positioning storage inspection as safer than zerodom_eval_js: 'Same trust boundary as zerodom_eval_js... but purpose-built and read-only.' This helps an agent choose it for reading session state without eval capabilities. It does not explicitly enumerate exclusions or other sibling alternatives, so it stops short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
zerodom_get_stylesA
Computed styles and box-model dimensions for a node — for design/CSS review, not just interaction.
Runs getComputedStyle() in the real page over the same chrome.debugger connection everything else here uses (also reachable ad hoc via zerodom_eval_js; this is the purpose-built version with a curated property list instead of getComputedStyle()'s full ~300-property dump).
| Name | Required | Description | Default |
|---|---|---|---|
| node_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description must carry behavioral disclosure. It explains that it runs in the real page over the shared chrome.debugger connection and avoids the ~300-property dump, which is useful. However, it does not explicitly state read-only guarantees, behavior on invalid node_id, or limitations (e.g., element nodes only).
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded with the core purpose before the implementation detail. The second sentence earns its place by distinguishing from eval_js, though it is slightly dense.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter read tool with an output schema, the description covers purpose and implementation context well, but it omits any guidance on sourcing node_id. This is a noticeable gap for an otherwise simple tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% and the description never mentions node_id, so it adds no meaning beyond the schema's type and title. It doesn't explain how to obtain or format node_id (e.g., from zerodom_find), which is a real gap for a required parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names the resource (computed styles and box-model dimensions for a node) and the operation (Runs getComputedStyle()), and differentiates itself from zerodom_eval_js by offering a curated property list. This is specific enough to distinguish from all siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly names zerodom_eval_js as the ad hoc alternative and frames this tool as the purpose-built version with a curated list, giving the agent a clear selection criterion. It stops short of stating exclusions or when not to use it, so not a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
zerodom_hoverA
Hover over a node and return what changed on the page.
Reveals hover-triggered menus and tooltips — content that a click alone would never surface, and that isn't in the graph until this fires.
| Name | Required | Description | Default |
|---|---|---|---|
| node_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden, and it delivers: it discloses that the tool fires a hover event that mutates graph state ('isn't in the graph until this fires') and that it returns a diff of page changes. It could add caveats about hover side effects (e.g., triggering navigation), but the core mechanism is honestly disclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with zero filler. The main action and return value are front-loaded in the first sentence, and the second sentence earns its place by explaining why the tool exists and how it differs from click-based interaction.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter tool with an output schema present, the description covers the essential context: what triggers the action, what changes on the page, and what the return reflects. Minor gaps remain around prerequisites (e.g., node visibility or hover stability), but nothing critical for correct invocation is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. The lone parameter, node_id, is self-explanatory from its name and the description's opening 'Hover over a node' clarifies it as the hover target. This is adequate for a single simple parameter, though the description adds no explicit format or provenance details.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb+resource+outcome: 'Hover over a node and return what changed on the page.' It further distinguishes itself from siblings by noting it surfaces 'content that a click alone would never surface,' which clearly separates it from zerodom_click_node and the other interaction tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context for when to use the tool: whenever hover-triggered menus or tooltips need to be revealed, explicitly contrasting with click behavior. It does not name an alternative tool explicitly or give when-not-to-use conditions, but the implied usage is strong enough for an agent to route correctly.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
zerodom_list_tabsA
List every open tab, marking the active one with *.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
There are no annotations, so the description carries the full burden. It clearly communicates that the tool is read-only ('List') and discloses the special behavior of marking the active tab with `*`. It does not explicitly state that no tab state is modified, but 'list' strongly implies a non-destructive operation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, compact sentence that front-loads the primary action and includes only the essential behavioral detail. Every word contributes to understanding, with no filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter, simple read-only list tool with an output schema available, the description is complete. It states what is listed and how the active tab is indicated; the output schema covers return value details, so no further description is necessary.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has zero parameters and full schema coverage, so no parameter-level description is needed. The baseline for zero-parameter tools is 4, and the description adds no unnecessary parameter information.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('List') and resource ('every open tab'), and adds a distinguishing behavioral detail: marking the active tab with `*`. This makes it easily distinguishable from sibling tools like switch_tab, close_tab, and new_tab without needing to inspect schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies use when the agent needs a complete view of open tabs, but it does not explicitly state when to prefer this over alternatives or mention that it is a prerequisite for switching/closing tabs. The use case is clear enough from context, but explicit guidance is absent.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
zerodom_network_logA
Requests/responses the active tab has made since it opened (or since this was last called with clear=True) — method or status, and URL.
Passive visibility only, capped at the most recent 200 entries — not interception or modification of traffic (that needs page.route(), a bigger, stateful feature; zerodom_eval_js can already override window.fetch/XMLHttpRequest from the page side for ad hoc cases).
| Name | Required | Description | Default |
|---|---|---|---|
| clear | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does well: it discloses passive-only behavior, the 200-entry cap, the reset semantics via clear=True, and explicitly disclaims modification of traffic. No annotation contradiction exists.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact, front-loaded with the core behavior, and every sentence adds value: scope, cap, non-interception, and alternatives. It avoids restating schema fields or repeating the tool name.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple one-parameter log tool with an output schema, the description is complete: it defines what is captured, the retention window, the cap, the meaning of clear, and what the tool does not do. An agent has enough to invoke and interpret this correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The only parameter, clear, has 0% schema description coverage, but the description explains its effect by saying the log covers entries since the last call with clear=True. It strongly implies reset-on-clear, though it does not explicitly state 'set clear=true to clear the log'.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific function: show requests/responses made by the active tab, with method/status and URL. The description clearly distinguishes this passive log tool from traffic-interception approaches and sibling tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says this is passive visibility only and not interception, then names the alternatives: page.route() for full interception and zerodom_eval_js for ad hoc fetch/XMLHttpRequest overrides. This gives clear when-to-use and when-not-to-use guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
zerodom_new_tabA
Open a new tab and make it the active one — every other tool (read, click, fill, scroll) then acts on it until you zerodom_switch_tab away.
In an attached (real-browser) session the new tab lands in the same "zerodom" tab group as every other tab this session touches.
| Name | Required | Description | Default |
|---|---|---|---|
| url | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Since there are no annotations, the description carries the behavioral burden. It discloses the key side effect: the new tab becomes active and routes all subsequent tool calls to it. It also adds useful context about the zerodom tab group in attached real-browser sessions, though it does not cover URL-omission behavior or failure modes.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and well-structured: the first sentence front-loads the core purpose and activation behavior, and the second adds a relevant session detail. Every sentence earns its place with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The core behavior and tab-routing lifecycle are adequately described, and an output schema exists so return values need not be explained. However, the only parameter is left entirely undocumented, including its optionality and null semantics, which leaves a real gap for an agent deciding how to call the tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the description never mentions the 'url' parameter. An agent must infer from the property name alone that it is the URL to open. The default null behavior—whether it opens a blank tab or does something else—is completely undocumented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb and resource: 'Open a new tab and make it the active one.' It further clarifies the behavioral consequence that all other tools target this tab until switch_tab is called, which clearly differentiates it from zerodom_switch_tab.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives clear context: after opening, every other tool acts on the new tab until you zerodom_switch_tab away. It names the exit condition and implies this is for creating a new active tab, but it does not explicitly state when to prefer switch_tab for existing tabs or list exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
zerodom_parse_urlA
Navigate to a URL and return its interaction graph.
Returns the compact text graph: [03] button 'Sign In'. CSS selectors are
kept server-side and resolved by node id, so they never cost context — pass
verbose=True for the full JSON including selectors. Nodes sharing a
repeated-list-item ancestor (<article>/<li>/<tr>, e.g. a feed or
Hacker-News-style table) are grouped under @card "title": whenever a
card holds 2+ controls — twenty identical button 'Upvote' lines are
meaningless without knowing which story each belongs to.
Set frames=True when the controls you need are inside an iframe — embedded editors, payment fields, consent gates. Off by default because it costs a read per frame and most frames on a commercial page are advertising.
Set viewport_only=True on long feed/infinite-scroll pages (a social feed,
a video site's homepage) where most of the graph is scrolled off-screen
and you only need what's currently visible — this can cut node count by
more than half on pages like that. On by default for interactive agents;
costs a getBoundingClientRect() per element. Sticks for the rest of this
session (every click/fill re-read honors it too) until the next
zerodom_parse_url call changes it — same lifetime as frames. When nodes
are being skipped, the graph's first line says how many; scroll and re-read
to see them.
Set check_occlusion=True when clicks keep failing with Playwright's
"element intercepts pointer events" — a modal backdrop, an open dropdown,
or a cookie banner is covering nodes that are still in the DOM and still
listed. This filters them out at parse time instead of at click time,
same sticky-for-the-session lifetime as viewport_only. Off by default:
the cost of an elementFromPoint() hit-test per node on a large page is
unmeasured, so it isn't imposed on every caller by default.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | ||
| frames | No | ||
| verbose | No | ||
| viewport_only | No | ||
| check_occlusion | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure, and it excels. It details the graph format, grouping of repeated-list-item nodes, session-sticky flag lifetimes, costs per read (getBoundingClientRect, elementFromPoint, read per frame), and even mentions that skipped nodes are reported in the first line. This is exemplary transparency, covering both operational effects and performance implications beyond what any annotation could provide.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long but well-structured, with each paragraph serving a clear purpose: an overview, return format, and then parameter-specific guidance. The most critical information (the tool's function) is front-loaded. While it could potentially be tightened, the length is justified by the complexity of the flags and session behavior, and there is no redundant or filler content.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity—multiple boolean flags, session-sticky state, performance costs, and a detailed graph output—this description is remarkably complete. It explains the return format, grouping logic, session lifetime, defaults, and costs, leaving no critical aspect unaddressed. The presence of an output schema helps, but the description goes beyond by explaining how to interpret the graph and when to expect deviations (e.g., skipped nodes). No relevant operation details are missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must fully explain each parameter, and it does. Every flag (frames, verbose, viewport_only, check_occlusion) receives a dedicated paragraph explaining its effect, default, and rationale. It even clarifies the interaction between sticky flags across session calls. The required url parameter is implicit in the tool's purpose. This level of parameter documentation fully compensates for the missing schema descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a precise statement: 'Navigate to a URL and return its interaction graph.' This specifies the verb, resource, and output type, making it unmistakable what the tool does. It also distinguishes its role from sibling tools like zerodom_click_node or zerodom_find by focusing on graph generation, and the detailed explanation of graph format further reinforces its unique purpose.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit, scenario-driven guidance for each optional flag: when to set frames (iframes), when to use viewport_only (long/infinite-scroll pages), and when to enable check_occlusion (click failures due to overlays). It also explains default behaviors and performance trade-offs, effectively telling the agent when to deviate from defaults. However, it does not explicitly compare against alternative tools like zerodom_read_page or zerodom_find, leaving some ambiguity about the boundary between reading the page graph and reading text content.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
zerodom_press_keyA
Press a key while a node is focused and return what changed.
key uses Playwright's key names ("Enter", "Escape", "Tab", "ArrowDown", ...). For "press Enter to submit" forms, "Escape to close a modal", and keyboard-only widgets a click/fill can't drive.
| Name | Required | Description | Default |
|---|---|---|---|
| key | Yes | ||
| node_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral disclosure burden. It clearly describes the action and the outcome, but leaves gaps: whether the tool focuses the node first, whether key press is atomic, and what 'what changed' specifically covers. Acceptable but not rich.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two compact sentences with no redundancy. The core action is front-loaded, and the second sentence efficiently adds key-format details plus usage context.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has only two required parameters and an output schema, so the description does not need to explain return values. Still, for an automation action with no annotations, it would benefit from stating whether the node must already be focused and how node_id relates to prior discovery steps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It does well for 'key' by specifying Playwright key names with examples, and it ties 'node_id' to the focused node. However, it does not explain how to obtain node_id or whether it must already be focused, leaving a meaningful gap for one of the two required parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Press'), a resource ('a node'), and the observable result ('return what changed'). The closing note that this is for keyboard-only widgets 'a click/fill can't drive' distinguishes it from sibling input tools like click_node and fill_node.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives concrete scenarios: 'press Enter to submit' forms, 'Escape to close a modal', and keyboard-only widgets. It implies this tool is the right choice when click/fill cannot drive behavior, though it does not explicitly name sibling tools or say when not to use it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
zerodom_read_pageA
Re-read the current page without navigating.
Use after an action changed the page, or when node ids look stale. Unlike zerodom_parse_url this does not reload, so anything typed into the page stays. Honors whatever frames/viewport_only/check_occlusion zerodom_parse_url last set.
| Name | Required | Description | Default |
|---|---|---|---|
| verbose | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden and does well: it discloses that no navigation/reload occurs, preserves page state, and honors prior parse settings. It stops short of explicitly stating there are no other side effects, but 're-read' and 'does not reload' imply a safe, read-only operation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, front-loaded with the core purpose, followed by use cases and the key distinction. Every sentence earns its place with no filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple one-optional-parameter read tool with an output schema, the description covers the main invocation context: when to call, how it differs from alternatives, and what state it preserves. The only notable gap is the undocumented verbose parameter, but the tool's core usage is fully explained.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The only parameter, verbose, has 0% schema description coverage and is not mentioned in the description. The parameter name is somewhat self-explanatory, but the description adds no meaning about what verbose controls or when to set it. For a low-coverage schema, the description should compensate but does not.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action ('re-read the current page') and resource ('current page'), and immediately distinguishes it from zerodom_parse_url by noting it does not reload. An agent can clearly tell what this tool does and how it differs from a key sibling.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly says when to use ('after an action changed the page, or when node ids look stale') and contrasts with zerodom_parse_url, clarifying that typed content is preserved. This gives practical, actionable selection guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
zerodom_replayA
Replay an HTTP request under a chosen identity — Burp Repeater for the agent.
Sends method url (with optional body and one extra header 'K: V')
through as_identity (default the live logged-in session). Returns the
response status, size and a body preview. Use it to probe an endpoint,
tamper with a request, or check an object reference — then change the id/body
and replay again. Pair with zerodom_compare_identities for the A-vs-B diff.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | ||
| body | No | ||
| header | No | ||
| method | No | GET | |
| as_identity | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden of disclosing side effects. It says it sends a request and returns status/size/body preview, but it never warns that replaying arbitrary methods (POST, PUT, DELETE) can mutate server state, nor does it mention authentication requirements, rate limits, or whether the request is sent from the agent's context. The 'Burp Repeater' analogy implies tampering, but the safety profile is not disclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences plus a usage line, front-loaded with the analogy and core mechanics. Every sentence earns its place – the first defines, the second explains parameters and return, the third gives use cases and a pairing. Zero fluff, ideal density for an agent to scan quickly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The output schema exists, so return details are covered; the description correctly mentions status/size/preview anyway. However, it omits any error handling, authentication prerequisites, or the range of valid as_identity values. For a tool that can fire arbitrary HTTP requests, an agent needs more guardrails (e.g., 'requires an active session' or 'only one header allowed'). The core is complete, but the edges are under-specified.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It does mention all five parameters in one sentence and gives the header format ('K: V'), but it leaves as_identity ambiguous (what values? identity names?) and does not explain body format or constraints beyond 'optional'. The default method GET is implied but not stated explicitly. This is thin compensation for a 5-param tool with no schema descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a crisp verb+resource statement ('Replay an HTTP request under a chosen identity') and the 'Burp Repeater for the agent' analogy instantly frames the tool. It clearly separates itself from browser-automation siblings like zerodom_click_node or zerodom_network_log by focusing on raw HTTP replay, and names a natural companion (zerodom_compare_identities).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly lists concrete use cases: 'probe an endpoint, tamper with a request, or check an object reference' and even suggests an iterative workflow ('change the id/body and replay again'). It names a sibling to pair with for A-vs-B diffing. It does not state when not to use it (e.g., if a browser-level interaction is needed), but the guidance is strong and actionable.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
zerodom_resize_windowA
Resize the real browser window (not just the page's viewport).
chrome.debugger's CDP surface has no browser-level window-management grant, but chrome.windows.update is a plain extension API, entirely unrelated to chrome.debugger — so this genuinely works despite that. Sent as Browser.setWindowBounds (a real CDP method name Playwright's own driver will actually transmit — a made-up method name gets rejected client-side before reaching the relay at all) and repurposed server-side; see docs/DECISIONS.md D16. Only meaningful for an attached (real-browser) session; a launched headless session has no window to resize.
On a tiling window manager (i3/sway/Hyprland/bspwm-style setups), this call succeeds but the window won't visibly move or resize — on Wayland compositors specifically this isn't a WM being uncooperative, it's the protocol itself: clients are deliberately not allowed to force their own geometry. Works normally on a floating window. Confirmed live on Hyprland.
| Name | Required | Description | Default |
|---|---|---|---|
| width | Yes | ||
| height | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It does an excellent job: it explains why the call works despite CDP limitations, warns that it can succeed without visible effect on tiling window managers, and details the Wayland protocol restriction. This far exceeds typical tool descriptions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The core purpose is front-loaded in the first sentence, with important caveats following. The prose is dense and includes some implementation archaeology like CDP method names and a docs reference that go beyond what a caller strictly needs, but none of it is filler—it all helps explain edge behavior.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description thoroughly covers the non-obvious environment restrictions: headless sessions, tiling window managers, Wayland, and floating windows. An output schema exists, so return-value documentation is not needed. The only real gap is the lack of parameter units and limits, which keeps it from a perfect score.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description never mentions width or height units, bounds, or how the values map to window geometry. The parameter names 'Width' and 'Height' are somewhat self-explanatory, but the description adds no semantic detail to compensate for the empty schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description leads with a specific verb and object: 'Resize the real browser window', then immediately disambiguates from the page viewport, which differentiates it from the sibling tool zerodom_set_viewport. This is exactly the kind of precision an agent needs to select the right tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives clear when-not guidance: the call is only meaningful for an attached real-browser session and a headless session has no window. It also warns about tiling window managers and Wayland compositors. It doesn't explicitly name an alternative tool for viewport resizing, only says 'not just the page's viewport', so it stops short of a perfect 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
zerodom_screenshotA
Full-page screenshot of the active tab, saved to disk; returns the path.
report.py has a screenshot path already, but it's wired to the CLI's own
throwaway sync browser (playwright_wrapper.py's sync/async split), not
this attached async session — this is that same capability for here.
Pass path to choose where it's saved; omitted, a temp file is used.
| Name | Required | Description | Default |
|---|---|---|---|
| path | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full disclosure burden, and it clearly states the side effect (screenshot saved to disk), the return behavior (path returned), and the default behavior when path is omitted (temp file used). It also specifies the full-page capture scope. It does not mention overwrite behavior or file format, but those are minor for this operation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The first sentence is front-loaded and efficiently states the core purpose, behavior, and return. The second paragraph adds useful but somewhat tangential implementation context about report.py and the async/sync split, which may distract an agent; however, it is not redundant and the overall length remains reasonable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is simple: one optional parameter and an output schema, so the description does not need to explain return structure in detail. It covers the operation, save behavior, path semantics, and default. Minor details like path format or whether an existing file is overwritten are unstated but are not critical for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, but the description fully compensates for the single parameter: path selects the save location, and omitting it uses a temp file. This adds exactly the practical meaning an agent needs beyond the bare string/null type and default. This is a strong example of description-driven parameter documentation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: a full-page screenshot of the active tab, saved to disk, returning the path. This clearly communicates the tool's function and distinguishes it from the many DOM-interaction siblings like zerodom_click_node or zerodom_eval_js. It is not a tautology and includes the outcome and return value.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives context that this screenshot capability is for the attached async session, contrasting with report.py's separate CLI sync browser. This is an implied when-to-use signal, but the reference to report.py and playwright_wrapper.py is not framed as an actionable alternative-selection rule. It does not explicitly state when to avoid this tool or prefer a specific sibling.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
zerodom_scrollA
Scroll the page and return what's newly visible.
direction: "up" or "down". amount: pixels, roughly one screenful is 800.
Dispatches a real wheel event at the viewport center rather than
window.scrollBy, so it scrolls whatever scrollable container is actually
under the cursor — a nested feed/sidebar, not just the document body,
matching what a real scroll gesture would do.
| Name | Required | Description | Default |
|---|---|---|---|
| amount | No | ||
| direction | No | down |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It discloses the real wheel event mechanism, which is a key behavioral nuance affecting which container is scrolled, and notes that it returns newly visible content. It does not cover edge cases like scroll failure or bounds, but for a scroll tool this is reasonably transparent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise and well-structured: one sentence for purpose, one for parameters, and one for the mechanism. Every sentence earns its place with no filler or repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists (though not shown), so return format is presumably covered there. The description explains the scroll behavior and what is returned, which is adequate for a scroll tool. It could mention edge cases (e.g., no scrollable area) but that is minor given the complexity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has 0% parameter description coverage, so the description must compensate. It explains direction values ('up' or 'down') and gives a practical hint for amount ('roughly one screenful is 800'), adding meaning beyond the schema's bare defaults.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: 'Scroll the page and return what's newly visible.' It names a specific verb (scroll) and resource (page), and explains the unique mechanism of dispatching a real wheel event at the viewport center, which distinguishes it from generic scrolling and other tools like zerodom_eval_js.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explains how the tool behaves (scrolls whatever container is under the cursor) but does not explicitly state when to use it versus alternatives. It contrasts with window.scrollBy, implying a use case, but lacks an explicit 'use this when...' or 'instead of...' statement. No direct alternatives are mentioned.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
zerodom_set_scopeA
Constrain the hunt to authorized targets — enforced in code, not on trust.
hosts: comma-separated in-scope host globs (app.example.com,*.example.com).
Navigation, replay and clicks to any other host are then refused. deny: a
regex of destructive URLs/labels to refuse (default covers logout/delete/
remove/deactivate/revoke) so an unattended agent can't take an irreversible
action. max_rps: throttle to at most N requests/second (program rate limits).
An operator can instead lock scope before the agent starts by setting the ZERODOM_SCOPE env var to a YAML file; a locked scope can't be widened here.
| Name | Required | Description | Default |
|---|---|---|---|
| deny | No | ||
| hosts | Yes | ||
| max_rps | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It clearly states that navigation, replay, and clicks to unauthorized hosts are refused, that deny is a regex to block destructive actions, and that max_rps throttles requests. It also discloses the limitation that a locked scope cannot be widened. This is thorough and transparent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is slightly verbose but well-structured: the main purpose is front-loaded, followed by a parameter explanation block, and then an important constraint about the env var. Every sentence adds value, though a few could be tightened. It is not wasteful and is easy to scan.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a security-sensitive configuration tool, the description covers purpose, usage, all parameters, behavioral effects, and limitations. It mentions the env var lock and that a locked scope cannot be widened. Since an output schema exists, return values are not required. An agent has all necessary information to invoke this tool correctly and understand its consequences.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must fully explain the parameters. It does: hosts as comma-separated globs with examples, deny as a regex with default patterns, max_rps as a throttle. This adds significant meaning beyond the raw schema, which only provides types and titles. Each parameter's semantics are precisely defined.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('constrain'), a clear resource (authorized targets), and explains the enforcement mechanism ('enforced in code, not on trust'). It clearly distinguishes from sibling tools which are about other operations like cookies, storage, or navigation. The purpose is unambiguous and immediately apparent.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly explains when to use the tool (to constrain the hunt) and provides a critical exclusion: an operator can lock scope via ZERODOM_SCOPE env var, and once locked, the scope cannot be widened here. It also describes the effect of each parameter, giving clear context for when to set deny or max_rps. This is comprehensive guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
zerodom_set_viewportA
Resize the active tab's viewport for responsive-design testing.
Not the real browser window — that's zerodom_resize_window. This is Emulation.setDeviceMetricsOverride: the page's own rendered layout at a given width, without touching the chrome around it. Use this one for "how does this render at width X"; use zerodom_resize_window for an actually different-sized window on screen.
| Name | Required | Description | Default |
|---|---|---|---|
| width | Yes | ||
| height | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does substantial work: it clarifies this is Emulation.setDeviceMetricsOverride, affects the page's rendered layout, and does not touch the browser chrome. It doesn't mention persistence or how to reset the override, but the core behavioral scope is clearly disclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the core purpose and then immediately provides the key differentiation from the sibling tool. Every sentence adds value: scope, mechanism, and usage guidance.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple two-parameter tool with an output schema, the description gives enough context to call it correctly and choose it over the sibling. The main gap is that height is not explicitly explained and reset/teardown behavior isn't mentioned, but these are minor for this use case.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% and the description only meaningfully explains width ('at a given width', 'renders at width X'). Height is left entirely to inference, and there is no note about units, value ranges, or how these map to the emulation behavior.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'Resize the active tab's viewport' for responsive-design testing. It explicitly distinguishes itself from zerodom_resize_window, so an agent can tell which tool handles viewport emulation versus real window resizing.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It provides direct routing guidance: use this tool for 'how does this render at width X' and zerodom_resize_window for a genuinely different-sized window. It also names the sibling tool explicitly and clarifies what this tool is not.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
zerodom_statusA
Diagnose the current connection: relay reachability, whether a browser session is attached yet, which tab is active, and the tail of the relay's own log — the single-call version of the manual WebSocket-probing and log- tailing this project's own debugging needed before this tool existed.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the burden of behavioral disclosure. It is transparent about what the tool reports, including connection state, session status, active tab, and log tail, and it frames the behavior as a non-mutating diagnostic. It could go further by explicitly stating that it makes no changes, but the detail provided is strong.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single focused sentence with the core purpose front-loaded after 'Diagnose the current connection'. The trailing clause adds valuable rationale about replacing manual probing and log-tailing, without wasting words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter diagnostic tool with an output schema present, the description covers everything an agent needs: what is diagnosed, what sub-aspects are included, and why this tool exists as a convenience. No additional context is necessary.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, and the schema already reflects this with an empty properties object. The baseline for zero-parameter tools is 4, and the description appropriately focuses on observable behavior rather than hypothetical parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description begins with a specific verb, 'Diagnose', and names the resource ('the current connection') plus four concrete diagnostic aspects: relay reachability, browser session attachment, active tab, and the relay's log tail. This clearly sets it apart from sibling action-oriented tools like zerodom_press_key or zerodom_read_page.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context for when this tool is useful: whenever one needs a single-call status overview instead of manual WebSocket-probing and log-tailing. It does not explicitly name sibling alternatives or state exclusions, but it communicates the intended diagnostic scenario well.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
zerodom_switch_tabA
Make another open tab active and return its interaction graph.
Node ids are per-tab, so re-read (this returns the graph already) before clicking or filling anything on the tab you just switched to.
| Name | Required | Description | Default |
|---|---|---|---|
| tab_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It discloses a key behavioral trait: node ids are per-tab, so you must re-read the returned graph before clicking or filling. It also states the tool returns the interaction graph, which is beyond what the input schema shows. It doesn't mention side effects like whether the previous tab is closed or if state is lost, but the critical per-tab id warning is valuable.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences with no wasted words. The core action is front-loaded, and the critical per-tab warning is placed immediately after. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has one parameter, no annotations, and an output schema exists. The description explains the return value (interaction graph) and the critical caveat about per-tab node ids. It doesn't mention how to obtain the tab_id or what happens if the tab_id is invalid, but for a simple switch operation, the essential context is present.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It mentions 'tab_id' implicitly by saying 'another open tab' and 'the tab you just switched to,' but it doesn't explain where to get the tab_id (e.g., from zerodom_list_tabs). The description adds some context but doesn't fully clarify the parameter's provenance or format.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb and resource: 'Make another open tab active' and 'return its interaction graph.' It distinguishes itself from sibling tools like zerodom_list_tabs and zerodom_close_tab by focusing on switching to an existing tab. However, it doesn't explicitly name a sibling alternative, so it falls short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context: use this to switch to another open tab, and it returns the interaction graph so you can re-read before interacting. It implies this is the tool to use when you need to change active tabs, but it doesn't explicitly state when not to use it or name alternatives like zerodom_list_tabs for listing tabs. Still, the guidance is actionable and context-rich.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
zerodom_upload_fileA
Set a file input's value to a local file path and return what changed.
path is resolved on the machine driving the browser — the same trust boundary as zerodom_eval_js, not a new one.
| Name | Required | Description | Default |
|---|---|---|---|
| path | Yes | ||
| node_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full behavioral burden. It adds a genuinely valuable security disclosure: 'path is resolved on the machine driving the browser — the same trust boundary as zerodom_eval_js, not a new one,' which tells the agent the path is interpreted on the driver host and introduces no additional trust risk. It also discloses the return behavior ('return what changed'). It omits edge cases like invalid paths or event firing, but the trust-boundary context is significant enough to justify a 4.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Exactly two sentences, both earning their place: the first conveys the action and return behavior, and the second adds a security-relevant boundary note. There is zero filler, and the primary behavior is front-loaded ahead of the caveat.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Complexity is low (two required string params, no enums, no nesting) and an output schema exists, so return values need not be explained. The description covers the core mechanics and security context. However, it leaves usage routing, prerequisites (target node must be a file input), and error behavior (e.g., nonexistent path) unstated, which an agent would need for fully confident invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It does clarify the critical parameter: 'path' is a local file path resolved on the driver machine, adding meaning beyond the bare string type. However, 'node_id' receives no elaboration and relies on the zerodom_* sibling convention of identifying nodes. One of two parameters is well-explained; the other is left implicit, so the compensation is only partial.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'Set a file input's value to a local file path and return what changed.' It clearly scopes the tool to file inputs and names the outcome (what changed), which distinguishes it from siblings like zerodom_fill_node (general text filling) and zerodom_eval_js (arbitrary JS execution). An agent can determine what this tool does without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The scope is implied by the description: 'a file input' signals when the tool applies, and the trust-boundary sentence references the sibling zerodom_eval_js. However, the description never explicitly states when to prefer this over zerodom_fill_node, nor does it give exclusions such as 'do not use for non-file inputs.' Context is present but relies on inference rather than explicit routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
4 tool updates
v0.0.9- Added
zerodom_add_identity - Added
zerodom_compare_identities - Added
zerodom_replay - Added
zerodom_set_scope
20 tool updates
v0.0.8- Added
zerodom_close_tab - Added
zerodom_drag - Added
zerodom_eval_js - Added
zerodom_get_cookies - Added
zerodom_get_storage - Added
zerodom_get_styles - Added
zerodom_hidden_fields - Added
zerodom_hover - Added
zerodom_list_tabs - Added
zerodom_network_log - Added
zerodom_new_tab - Changed
zerodom_parse_url2 fields changed- added
Input schema / properties / check_occlusionAdded value: +{ + "default": false, + "title": "Check Occlusion", + "type": "boolean" +} - added
Input schema / properties / viewport_onlyAdded value: +{ + "default": true, + "title": "Viewport Only", + "type": "boolean" +}
- Added
zerodom_press_key - Added
zerodom_resize_window - Added
zerodom_screenshot - Added
zerodom_scroll - Added
zerodom_set_viewport - Added
zerodom_status - Added
zerodom_switch_tab - Added
zerodom_upload_file
5 tool updates
v0.0.6- First observed
zerodom_click_node - First observed
zerodom_fill_node - First observed
zerodom_find - First observed
zerodom_parse_url - First observed
zerodom_read_page
TDQS
Scored across 28 tools
Each tool targets a distinct resource or action, and near-overlapping pairs are explicitly disambiguated in their descriptions (parse_url vs. read_page, set_viewport vs. resize_window, get_cookies/get_storage vs. eval_js). An agent can reliably select the right tool without guessing.
All tool names share the zerodom_ prefix and snake_case convention, which makes them predictable. However, the verb style is mixed—get_/list_/read_/parse_/set_/click_/fill_—and a few noun-form names like zerodom_screenshot, zerodom_status, zerodom_network_log, and zerodom_hidden_fields break the otherwise action-first pattern.
28 tools is above the typical well-scoped range and makes the surface feel heavy, though the server covers a broad domain spanning browser automation, DOM interaction, tab management, network visibility, and security testing. It is not bloated to the point of being unusable, but the count is borderline.
The toolset covers the core browser-automation lifecycle well: navigation, page reading, interaction, tabs, state inspection, screenshots, network visibility, and identity-based replay. Minor gaps like explicit page reload/back/forward, cookie/storage mutation, or a dedicated wait-for-condition helper are workarounds rather than dead ends.
Maintenance
Related MCP Connectors
Hosted browser for AI agents: screenshots, post-JS DOM, console, WCAG. No install, no API key.
Headless browser primitives for AI agents when sites need real JS rendering.
- mcpOAuthcom.screenshotink
Screenshot, diff, audit and sitemap-capture any web page — 5 MCP tools for AI agents.
Web scraping for AI agents: scrape, search, crawl, map any website to markdown + JSON. No browser.
Related MCP Servers
AlicenseNot gradedqualityBmaintenanceAgent-native headless browser for AI agents. Converts web pages to a Semantic Object Model (SOM) instead of raw HTML — 17x average token reduction across real-world sites (up to 117x on complex pages). Native MCP server with fetch_page, extract_text, extract_links, and full browser automation. No API key required.47 npmApache 2.0- AlicenseCqualityCmaintenanceHosted Playwright browser automation for AI agents. Returns accessibility trees instead of screenshots, cutting token usage by 77%. Navigate, interact, extract structured data, and take screenshots — all via MCP. Zero infrastructure, credit-based pricing.6223 npmMIT
- AlicenseAqualityAmaintenanceA token-efficient MCP server that gives AI agents structured access to the web, returning compact page summaries and targeted queries instead of full accessibility dumps.23319 npm179MIT
- AlicenseAqualityDmaintenanceA Playwright-powered MCP server for browser automation using ARIA snapshots and element refs, enabling LLMs to control Chrome/Edge without CSS selectors.419 npmMIT