lyra-browser
OfficialClick on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@lyra-browseropen my Gmail and hand over when it asks for my password"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
lyra-browser
A headful, collaborative browser MCP server for the VEGA agent harness. VEGA drives a single visible Chromium window that the user can watch and take over at any time — co-browsing, not headless scraping.
The same server runs headless under Hermes, where the user is reached through chat rather than at a window — see docs/HERMES_INTEGRATION.md.
Renamed from vega-browser. The old
VEGA_BROWSER_*environment variables and an existing~/.vega-browserdata directory are still honoured; theLYRA_BROWSER_*names win when both are set.
Built on Playwright (headful, persistent profile) and exposed over FastMCP, so VEGA picks it up through its existing MCP integration with near-zero glue.
Why
VEGA is local-first and model-agnostic. This plugin gives it a browser the user shares: the agent reads the page, clicks, and types, while the user can step in for the things an agent must not do alone (passwords, CAPTCHAs, 2FA, payments) and grab full control whenever they want. Every action is written to an append-only audit trail.
Related MCP server: Web Bridge
Architecture
VEGA agent loop ──(MCP: stdio/http)──> lyra-browser server ──> Playwright ──> visible Chromium window
│ ▲
├─ audit.jsonl (every action) │ user watches / takes over
└─ approval + takeover gating ──────────┘Standalone MCP server — registered in VEGA's
mcp.json(no VEGA code change).Python + Playwright — same runtime as VEGA; pins track the harness (
fastmcp>=3.2,playwright>=1.59— one minor above the 1.58.0 VEGA bundles, because element refs need it; see Addressing an element).Dedicated headful window —
launch_persistent_context(headless=False)with a persistent profile, so logins survive across sessions. One server at a time holds that profile (anflockon<data_dir>/profile.lock, taken when the browser opens, not when the server starts). A second server on the same data dir opens a private, empty profile at<data_dir>-instances/<pid>/profileinstead of failing, andopen_browsersays so (profile: "instance", "no saved logins in this instance").No setup tax — reuses the user's installed Chrome/Edge by default (
channel="chrome"→"msedge"→ bundled Chromium), so an end user with only VEGA.app needs no pip and noplaywright install. If no browser is found, tools return abrowser_unavailableenvelope for VEGA's UI to prompt an install.Full preset — human-in-the-loop (highlight, ask-the-user, approval gate), takeover/handoff, audit trail, CI + lint + pre-commit.
Tools
Group | Tools |
Navigation |
|
Tabs |
|
Waiting |
|
Interaction |
|
Downloads |
|
Forms |
|
Reading |
|
Collaboration |
|
Filling a real form
click and type_text cover what every page has, but a form on a real site needs
more. Each of these was measured on KVR's developer forms before it was
written:
read_form— a form'sname,type,labeland<select>options.read_pagereturnsinnerText, so a 188-field product form arrives as a wall of labels with nothing to address. Hidden inputs are counted, not dumped. Every field carriesusable— false means it is in the DOM but hidden or disabled right now, which is how a form keeps a section it has not revealed (measured on KVR:event_countrysits in the Deal/Offer block). Checkusableinstead of finding out through a 30s locator timeout.select_option— a<select>is not a click target. Chooses byvalue,labelorindex, and fireschangethe way a page listens for it. Passsubmits=truewhen the choice posts. Answershiddenordisabledby name when the control cannot be used as it stands.set_editor— CKEditor/TinyMCE keep the editable body in aniframeand leave the<textarea>the form posts hidden and empty, sotype_texton that textarea does nothing.html=truewrites markup and syncs the editor's own data. Leaveselectorempty to use the first editor on the page.upload_file— attaches local files to aninput[type=file]. A styled picker usually keeps it hidden, which is fine. Missing paths are reported before anything is attached.save_draft— sends the form while keeping the item private. It names the control that sends (submit, fromread_form'ssubmitText), first confirms the publish control is still on its private setting and refuses otherwise, and asks forSUBMITbut neverPUBLISH— so it cannot release an item, and a page left armed by anything else will not slip out through it.read_draft— everything a human needs to approve an item, in one answer: the filled fields, each editor's real content (a separate documentread_pagenever shows), the attachment slots, the submit control, and whether the item is a draft or already public. A read, so it needs no approval.publish— the one action that sends an item to an audience. Kept apart from every other tool so that writing a draft can never release it: it asks forPUBLISH, which nothing implies, and it is the only tool that does. A site splits the act across two controls and means neither alone — measured on KVR, the publish radio only arms the form and the item travels on the form's own Submit — so it takes bothselector(the live control) andsubmit(the one that sends). Giving only the first is refused rather than left half-done: arming a form with no way to send it would leave the page primed to publish on someone else's next click.
Drafts and release
Writing an item and releasing it are different acts, and the tool surface keeps
them apart. Every writing tool can be driven to completion with the item still
private; publish is the only way out, and it is PUBLISH-gated and single-use.
The intended flow is: fill the form → read_draft → show the human → only then,
on an explicit yes, publish.
On KVR this matches the page: a new item opens on Draft and stays there until
the publish radio is switched, so a form that is merely filled is not visible to
anyone. A form sent without the publish control set saves as a draft — which also
means every writing tool must stay away from Submit, since KVR has no separate
"save draft" button. That is why publish sends the form itself rather than
leaving the send to click.
UPLOAD and PUBLISH are single-use: a file handed to a page, and content shown
to an audience, cannot be recalled.
Reading and acting on a page
Reading: read_page
read_page has two modes, both read-only (no approval, works during a takeover):
mode="text"(default) — the rendered text (innerText), not HTML.links=trueadds the visible links as{text, href}(absolute, de-duplicated, at most 100, withlinks_truncated).mode="tree"— one line per thing that can be acted on (links, buttons, inputs, selects, checkboxes, tabs, menu items), each with a ref:aria-ref=e12 link "Pricing" -> /pricing aria-ref=e15 button "Save" aria-ref=e16 checkbox "Remember me" [checked]Input values are never shown. The reply carries
elementsandrefs, plusin_viewport(on-screen elements come first) on Playwright 1.60+.
selector limits either mode to one region (first match; not_found at once
when nothing matches). offset and max_chars page through long output:
total_chars is the full length, truncated says more follows and next_offset
is where to continue. Tree pages break on whole lines, so a ref is never cut in
half.
Addressing an element: aria-ref
read_page(mode="tree") → pass the token aria-ref=eN, exactly as written, as
the selector of click or type_text. A ref reaches what a CSS or text
selector cannot name: icon-only links, several links with the same text, elements
inside open shadow roots and iframes. A ref is a handle into the latest tree
read, not a query over the live page: a newer read (a region read included), a
navigation or a removed frame ends it, and so does the element being removed. A
stale ref answers not_found (or hidden for an element that is gone but still
counted) with a hint to read the tree again; refs get no waiting window, because
they cannot appear later.
Requires playwright>=1.59. Refs come from aria_snapshot(mode="ai"), which
1.59 introduced; pyproject.toml still floors at 1.58 because that is what a
VEGA runtime ships. On 1.58 the tree still works but answers refs: false and its
lines carry no ref — address elements with role=link[name="Pricing"] or text=
selectors instead.
When an action cannot be done: click / type_text
Both fail fast and explain, instead of holding the call for Playwright's 30s:
| Meaning |
| Nothing matches the selector, even after up to ~3s for a page that builds its controls late |
| The match is in the page but stayed invisible / disabled for that window (present-but-not-ready controls are re-checked until ready, so a button enabled 0.8s after load still works) |
| The action outran |
| Covered by another element (the hint names it), detached, read-only, not a text field, unstable or outside the viewport |
| The tab died under the call (a popup that closed itself, a closed window) |
Each carries selector, url and a hint with the next move. timeout_ms
(default 10000, at most 30000; 0 is raised to 1ms, never "no limit") bounds the
action itself. The usability check runs before permission is asked, so an
unusable control neither prompts nor leaves a single-use grant behind; the same
driver failures during the action are converted into these envelopes (and
audited) rather than raised. A successful click also reports matches (how
many elements the selector hit — the first is clicked) and clicked
(tag, role, name of what was hit), with a hint when matches > 1.
What an action set off: blocked_by_policy, new_tab
click, type_text, press_key and hover answer when the driver reports the action
done, and two things can still be true of that moment:
The guard refused a navigation the action caused (a link to a site no approval covers, a form sent without
submits=true, a page that redirected itself). The 204 leaves the tab where it was, so the reply used to readokwith the sameurland the agent had no idea why nothing happened — the audit trail was the only account. It is nowblocked_by_policywithurl(where the tab stands), ahintand, for a refused redirect,redirected_to; the fields the reply already had (matches,clicked,dialogs) are kept, and the audit row of the action carries the same status. The hint is the way out:navigateto the destination (that asks the user), or repeat the action withsubmits=true(submit=trueontype_text) when it was meant to send a form. A file the call saved staysokwith itsdownload, and a download the browser started and the call cancelled keepsdownload_blocked. Inobservemode nothing is refused, so the reply staysokand the audit sayswould_deny.The action opened a tab (
target=_blank,window.open). The session follows onto it whileurlis still the page the action was made on, so the reply carriesnew_tab: trueandtab_count, and a hint to usetabs.get_urlandopen_browsercarrytab_counttoo. A popup whose first navigation the guard refused opens no tab at all, and the reply isblocked_by_policyinstead.
Neither costs a click a wait of its own. The driver holds a click until the navigation it
started has been judged, so that verdict is in before the call returns (0.5 to 48 ms
before, measured over playwright and patchright, both guard backends, headless and
headful). What it does not hold for lands 3 to 22 ms after — a popup's tab, a form posted
from a frame, a page's own setTimeout(0) — inside the download_settle_s window every
click and key press already listens for. hover and type_text have no such window, so
they linger up to 40 ms after the driver returns, and only while nothing has shown (their
navigation reaches the guard 1 to 8 ms after they return); a page timer that fires after
that is only on the audit trail. Numbers in scripts/verify_click_refusal_e2e.py
(INFO lines).
Waiting: wait_for
wait_for holds until the page is ready instead of a sleep — or a click used to
pass time. Give exactly one condition: text (in the visible text read_page
shows; state="hidden" waits for it to vanish), selector (reaches state:
visible, hidden, attached, detached), url (a glob over the whole URL,
e.g. **/checkout**) or load_state (load, domcontentloaded,
networkidle). It waits up to timeout_ms (default 10000, capped at 30000) and
returns ok with waited_ms. Running out of time is an answer, not an error:
{"status": "timeout", "waited_ms", "last_seen"} with a short page excerpt (the
URL for a url wait). Bad arguments return error. It only observes: no
approval, and it keeps working while the user holds a takeover.
Tabs: tabs
The session follows onto any tab a page opens (target=_blank, window.open),
so the agent is never left reading the page that opened it. A click or
press_key that opened one says new_tab: true and tab_count, and get_url and
open_browser carry tab_count. tabs covers the rest:
action="list"(default) —tabs: [{index, url, title, active}]plustab_countandactive_index. A read; a tab stuck in a script loop is listed without a title rather than hanging the list.action="switch",index=N— make that tab the one to read and act on, and bring it to the front so the user sees what the agent sees.action="close",index=N— close it. Closing the active tab returns to the tab that opened it, else the newest one; closing the last tab leaves a blank one (closing headful Chrome's last window would quit the browser).
Indexes shift whenever a tab opens or closes, so list again before using one; a
bad index answers not_found with tab_count. Switching and closing are
mutations: they stand down for a takeover, are audited, and ask for INTERACT
on the tab's own site, not the active page's — being approved for one site is
no licence to read or close another's. A tab that closes itself (a sign-in popup)
sends the session back to its opener; if the active page is ever closed, the next
tool moves onto the opener, else the newest tab, else a fresh blank one, instead
of failing on a dead page.
What navigate reports
ok means the navigation happened, not that the page is good. navigate,
go_back and reload_page report http_status (404, 500 …) and content_type
(media type only, lower-cased, e.g. application/pdf) of the document that
answered; both are null when no HTTP response was involved (about:blank,
data:, a page the browser restored from history). wait_until chooses how much
of the load to wait for: domcontentloaded (default), load, commit or
networkidle (which never settles on a page that keeps polling); anything else is
an error before any prompt. A page that draws itself after loading needs
wait_for. A URL that turns out to be a file download does not move the tab: see
Downloads for download=true and download_blocked. Only when the
driver announced a download the browser never reported does it answer
{"status": "download_started", "url", "requested_url", "hint"}; this server then wrote
no file to the download dir.
Hover and scroll
hover— move the pointer over an element, for menus and tooltips that open only while it is on their trigger:hover, thenclickthe item that appears. Nothing is pressed, so no form is sent. Fails fast likeclick(not_found,hidden,disabled,timeout,element_not_actionable,page_closed,timeout_ms) and reportsmatchesandhovered(tag,role,name). NeedsINTERACT.scroll— give exactly one ofto(top/bottom),by_y(pixels, negative up, at most 20000 per call, done with the mouse wheel) orselector(bring it into view). The reply isscroll_y,scroll_heightandat_bottomonce the position has settled. It moves the window only: a feed or panel that scrolls inside its own container needsscroll(selector=...)on an element inside it, and a scroll that moved nothing says so. Lazy lists grow after a scroll, so repeatscroll(to='bottom')untilat_bottomis true.
Downloads
A file a page hands to the browser is kept only when the call declared it. Chromium accepts a download the moment a page offers one, and the navigation guard judges requests, not responses, so this is a separate gate on the file as it arrives:
download=trueonclick,press_keyornavigatedeclares it and buys a single-useDOWNLOADgrant (asked like any other, never implied by a trusted origin) inside that call. The grant pays for exactly one file and is retired when the call stops waiting. The reply carriesdownload:{filename, path, bytes, url_origin}.An undeclared download is cancelled and deleted before anything reaches the download dir, and the call answers
download_blocked(blocked_download:filename,url_origin) — repeat it withdownload=true. A declared call that produced no file answersdownload_not_started; one whose file was not kept answersdownload_failedwith areason(too_large,timeout,error).go_backandreload_pagecannot declare one; their hint points atnavigate. Chrome reports a reload or history step that turns into a download asnet::ERR_ABORTED(not "Download is starting"); an abort that a download follows within 2 s is answered as that download, and any other abort is still an error.With
LYRA_BROWSER_ENFORCEMENT=observe("record, do not block") nothing is cancelled either: an undeclared download is saved like a declared one — same name rules, caps and ledger — and answered asdownloadplusobserved: true, audited aswould_block. Observe never spends a grant (like the navigation guard), so it under-reports a second file in one call that enforce would have cancelled. A download aDOWNLOADgrant covered is auditedsavedand is not flagged.The site picks the file name, so it is treated as hostile: only the last path component survives, control and right-to-left-override characters, characters Windows forbids, leading dots and reserved device names are removed, long names are cut keeping their extension, and a taken name becomes
name (1).ext. Nothing is ever written outsideLYRA_BROWSER_DOWNLOAD_DIR.Size and time are capped (
LYRA_BROWSER_DOWNLOAD_MAX_BYTES,LYRA_BROWSER_DOWNLOAD_TIMEOUT), and the dir is pruned oldest-first toLYRA_BROWSER_DOWNLOAD_KEEPusing a ledger, so files that are not this server's are never deleted.While the user holds a takeover the gate stands aside and their own downloads are saved, recorded as
user_driven, as with navigation.list_downloads— the files this session saved (filename,path,bytes,url_origin), oldest first, plusdownload_dir. A read: no approval, works during a takeover, does not open the browser. Blocked and failed downloads are not listed.
The listener is per tab (popups included), not BrowserContext's download
event: that only exists from Playwright 1.60 while this project supports 1.58,
where a context listener would never fire and every download would go unjudged.
Native dialogs
alert, confirm, prompt and beforeunload stop a page until answered, and a
listener that leaves one open freezes the tab. So one context-level listener
always answers the moment a dialog appears. The default is what a browser with
no listener does: alerts and beforeunload are accepted, confirm is dismissed
(false) and prompt is dismissed (null).
What was raised is reported as
dialogs(type,message, how it was answered) in the reply of theclick,type_textorpress_keythat raised it — only present when there were any. The message is the page's text, not the user's. The form tools (save_draft,publish) use the same handler and fold dialog text into theirerrors.handle_dialog(accept, text)chooses the answer for the dialog the next browser action raises: call it immediately before that action.accept=falseanswers no;textis what apromptreceives (empty accepts the default). Like a one-shot grant it belongs to that one call —click,type_text,press_key,hover,scroll,navigate,go_back,reload_page,select_option,set_editor,upload_file,save_draft,publishor atabsswitch/close — and ends with it, used or not, even when the call was refused (arm again before retrying). A dialog raised after it, by a page's own timer or on a page reached later, gets the default, so a yes armed for one click cannot answer a laterDelete this?. Reading (read_page,screenshot,get_url…) does not use it up, and about 60s is the most it ever waits. It asks nobody (the click that raises the dialog was already asked for), stands down for a takeover, is audited, and the prompt text is never written to the audit log.Nothing here judges the answer: accepting a
confirmthat goes on to post a form is still stopped where the request leaves unless it was declared (submits=true).
Seeing the page
screenshot (the viewport or the full page) and read_image (one element —
an img, a canvas, an svg, a chart) write a PNG to disk and return its
path:
{"status": "ok", "image_path": "…/uploads/browser/page-20260917T123557Z-fd1a0721.png",
"mime_type": "image/png", "width": 1280, "height": 800, "bytes": 115889}The file is the handoff. Measured on VEGA: a tool result is shown to the model
as its text blocks, so a base64 PNG inside the JSON arrives as tens of
thousands of tokens of prose and no picture. A path costs nothing, and the
client attaches the file to the model's next turn as an image. Because VEGA
attaches only files under its own uploads root, captures default to
$VEGA_DATA_DIR/uploads/browser whenever VEGA_DATA_DIR is set — no
configuration for the model to be able to look. inline=true adds the base64
for a client with no access to this filesystem.
Captures are written atomically and pruned oldest-first
(LYRA_BROWSER_CAPTURE_KEEP, default 200), so an agent that looks every turn
does not fill the disk. The audit records the path and size, never the pixels.
read_image fetches nothing — it captures what the element renders, so no
request leaves the browser and nothing is gated.
Permissions
The agent declares what it intends; the browser is judged on what it actually
does. Those are separate layers because a tool cannot know what a page will do
with a click — onclick="form.submit()" submits from a control that looks inert,
an input listener submits while the agent only typed, and an element's type can
change between being checked and being clicked.
Grants are (origin × capability) and short-lived:
Capability | Meaning |
| Load a document from this origin. Implies |
| Click and type while staying on this origin |
| Send a non-idempotent request. Single-use |
| Give a file to a page (a picker, a drop target). Single-use — requested by |
| Make content publicly visible — an item an audience can see. Single-use — requested by |
| Write a file to disk. Single-use — requested by |
Arriving on a site carries permission to use it, so ordinary clicking is not
re-asked. Sending a form and leaving for another site are separate decisions and
are asked — declare them with submits=true (submit=true on type_text).
Enforcement watches outgoing navigations at the request, so it catches a
submission or an escape whatever produced it. A refusal answers with HTTP 204
rather than aborting, which leaves the page standing so the agent can recover.
URLs with no host — file:, data:, javascript: — are never "the same site"
and are asked every time. Where the guard stands — Playwright's route, or a second
DevTools connection that also sees redirect hops — is a choice, see
Guard backends.
Where the line falls, measured rather than asserted:
A same-origin GET navigation is interaction, not a submission. A search box puts what you typed in the query string; gating that would gate every search box and every link with parameters.
SUBMITmeans a body or a mutating method.A submission is authorised by the origin that sends it, not by wherever the form points. Cross-origin form actions are ordinary — payment handlers, SSO, third-party endpoints — and no tool can read a form's
actionbefore the click without the TOCTOU this layer exists to avoid. The destination is recorded in the audit under the labelcross_origin_send— an audit key, not a capability, and never purchased.A page that moves itself to another host is leaving, and is asked. A site bouncing to its own subdomain lands here too; the guard cannot tell that from a hostile page escaping. The refusal is recoverable — the original document is still there and the audit names the destination — so the agent can read where it was being sent and ask for it.
navigate,go_backandreload_pageanswer it asblocked_by_policy, with theurlthe tab still stands on, about a second after the refusal — also when the page did it before it finished loading. The 204 commits nothing, so the driver would otherwise wait out its own 30 s timeout for a navigation that is gone (scripts/verify_scriptredirect_e2e.py; neverssl.com does exactly this).click,type_text,press_keyandhoveranswer it too when the navigation was theirs (see What an action set off); they used to sayok(scripts/verify_click_refusal_e2e.py).A redirect hop is judged like the navigation it leads to. A 301/302/303/307/308 to another origin is classified exactly as a first request there would be: leaving the site is
navigateand is asked, a 307/308 that repeats a POST body issubmit, and a hop that stays on the origin that sent it (or upgrades it fromhttptohttps) is not judged at all. The audit row carriesredirect_from, and an approved site that answers 302 to an unapproved one is refused like the same destination asked for directly:navigate,go_backandreload_pageanswerblocked_by_policyand the tab stays where it was. The agent asked for one URL and was turned away from another, so that answer also carriesredirected_to, the origin it was sent to (never the path or the query), for the agent to ask for. Playwright never shows a hop to a route handler, so the guard hears of it from therequestevent, after the request was sent, and cancels the load there (scripts/verify_redirect_hop_e2e.py; what that leaves open is listed below).
Verified against the shipped defaults: four undeclared POSTs (space key, JS
form.submit() from a type="button", an oninput listener, and an attribute
flipped after the check) are all stopped, while the declared submit goes through.
What this does not cover, stated rather than discovered later:
With the default
routebackend, an HTTP redirect hop is judged after it has left, not before. The request to the first hop that no grant covers is sent — with that site's cookies, and the body of a 307/308 POST — and only the rest of the chain is cut. A server that answers before the cancellation lands commits the page first, and loopback or a LAN, where the targets most worth refusing live, always does: the guard then replaces the tab withabout:blankso the agent does not read it, the audit recordsnot_stopped, and the call still answersblocked_by_policy. Seeing a hop before it leaves is not possible fromcontext.route: Playwright continues every hop itself, also one that follows aroute.fulfill(302), and answering each document from Node (route.fetch) to see the redirect first puts a Node TLS and HTTP/1.1 fingerprint on every page while still leaving the later hops unseen (measured).LYRA_BROWSER_GUARD=cdpjudges every hop before it leaves, so a refused hop's destination sees nothing; see Guard backends.A page can ask the browser to prefetch or prerender another page (
<script type="speculationrules">). Chrome sends those requests, and serves the navigation that later activates one, without an interception point: neither backend is shown the request, nor the click that follows (measured, Chrome 154: a prefetched cross-site link was followed with the guard seeing nothing, and the cross-site prefetch itself reached its server). It is a GET, so it is the same channel as<img src="https://elsewhere/?data">; what it adds is a navigation nobody judged.With
route, a navigation that a service worker answers is never shown to the guard either. Playwright'sservice_workers="block"is an init script that a page defeats by calling the prototype'sregister, and a worker already in the profile is untouched by it (measured: an origin's pre-registered worker served that origin's page to a link click with no judgement). Thecdpbackend turns workers off for every page it guards.A single-page app does its damage over
fetch(), which is not a navigation and is not judged. Deleting mail in a webmail client is an XHR, not a form post.An approved page choosing where to send its form. Judging a submission on its sender is what makes payments and SSO work, and it is also an exfiltration channel: the audit records it (under the key
cross_origin_send, which is only a label on the audit row, not a capability), it does not stop it.A single-use scope is handed back about two seconds after the action that bought it, not instantly, because a click's navigation can leave just after the call returns. A page that submits inside that window rides an approval meant for the click. Two seconds rather than the ten minutes it used to be (
LYRA_BROWSER_RELEASE_GRACE).The approval for a
file:URL names that URL and covers only it, once — everyfile:URL is otherwise the same origin, which would make one yes a yes to the whole disk.The content of a download. It is gated, not inspected (see Downloads): an approved download is bytes from that site, written as-is under a sanitised name.
Enforcement narrows what a compromised turn can reach; it does not make an approved origin safe.
One browser, one session. Reading checks who is driving; only a mutating
tool becomes the driver, so a passive client cannot take the browser by asking
for a screenshot. Reading still counts as using it, so a holder part-way
through a read-only pass is not idle and does not lose the window. Grants are kept per session, but every session drives the
same window, and the route handler has no way back to a request
context — so whichever session called a tool last would decide how the next
request is judged, and one client's single-use approval would be spent by
another's traffic. Rather than offer an isolation it cannot deliver, the server
serves one session at a time and answers the rest with session_conflict.
Ownership lasts as long as the window: close_browser ends a turn, drops the
grants, and lets the next session claim it. A holder that goes quiet for
LYRA_BROWSER_OWNER_IDLE_TIMEOUT (default 900s) also hands it on — closing is
owner-only, so a client that drops with its window open would otherwise lock the
server for the life of the process. The handoff closes the window: passing on
a running browser would leave the previous holder's document loaded and still
able to start navigations, which the guard would then judge, and charge, against
whoever now holds the session.
Who is answering. auto asks the user through MCP elicitation, where the
model cannot forge the answer. Only a client that cannot be asked at all — no
elicitation handler, no live request — falls back to honouring confirm=true,
and the audit records consent_channel=legacy with the reason. A decline, a
cancel and a timeout are answers: the model asserting confirm=true never
overrides one. VEGA does not pass an elicitation handler yet, so it takes the
fallback today; wiring one is what turns the gate from a convention into a
boundary, and needs no change here. Hermes does pass one: its approval prompt
answers, and its value-less accept is read as the yes it is.
Trusted origins. Every new site costs one prompt, and a prompt nobody
answers holds its tool call for LYRA_BROWSER_CONSENT_TIMEOUT (default 300s)
before it becomes a denial. For sites the operator always uses, that answer can
be given in advance: LYRA_BROWSER_TRUSTED_ORIGINS is a comma- or
whitespace-separated list of sites that are pre-approved for NAVIGATE — and the
INTERACT it implies — so they are not asked about again. The operator writes it;
nothing the model or a page says can add to it.
Entry | Trusts |
| exactly that scheme, host and port |
|
|
|
|
Hosts are compared as parsed (scheme, host, port) fields, never as text, so
https://kvraudio.com.evil.io, https://kvraudio.com@evil.io and
https://kvraudio.com:8443 are not kvraudio.com. A site that redirects between
example.com and www.example.com needs both (example.com and
*.example.com). An entry that cannot be read is dropped and the rest still load:
empty and malformed entries, a lone *, single-label wildcards (*.com),
wildcards on an IP address, IPv6 literals, paths, userinfo, a port on a bare host
(write http://localhost:3000), and any scheme but http(s) — file:, data:,
about: and blob: name no site. The list does not know public suffixes:
*.co.uk is two labels and would be accepted, so do not write it.
What it covers is deliberately small. It never answers SUBMIT, UPLOAD,
PUBLISH or DOWNLOAD: those stay single-use and are still
asked on a listed site, every time. (To pre-approve sending from your own product for
end-to-end runs, LYRA_BROWSER_TRUSTED_SEND_ORIGINS is a separate list that answers
SUBMIT and UPLOAD and nothing else.) The list also holds where the browser arrives without
a navigate call -- a redirect hop or a followed link to a listed site -- because the guard
applies it to NAVIGATE too. It does not apply to an opaque origin, while
the user holds a takeover, or when approval is off (that stays
consent_channel=off). A listed site gets a real grant with the same lifetime
and initiator binding an approved one would, so enforcement judges it identically,
and the audit records consent_channel=trusted with the entry that matched.
trusted is a way a decision was reached, not a value of
LYRA_BROWSER_CONSENT_CHANNEL. It applies under auto, elicit and legacy.
Takeover. While the user holds the session, every mutating tool returns
takeover_active and the agent waits — including highlight_element and
ask_user_to_do, which write to the page. Reading tools keep working.
Enforcement steps aside too. Grants record what the agent was allowed to do;
holding a person to them would refuse them their own browser, and takeover exists
for the passwords, CAPTCHAs, 2FA and payments the agent must not do alone — every
one of which is a navigation. Those are recorded as user_driven, so the trail
distinguishes what a person did from what was approved for the agent.
Guard backends
Enforcement has to see every document request before it leaves the browser. There
are two ways to stand there, chosen with LYRA_BROWSER_GUARD (route unless set);
both ask the same judgement (NavigationGuard.decide), so grants, single-use
scopes, takeover, observe mode and the audit trail behave identically. They differ in
what they can see and what they cost.
|
| |
How | Playwright's | A second DevTools connection: |
Redirect hops | Judged after they have left: a | Every hop judged before it leaves: the destination of a refused hop sees nothing |
HTTP cache | Off: Playwright sends | On |
etsy.com (DataDome), headful, fresh profile, N=5 | refused 5/5 | let in 5/5 |
Service workers |
| Bypassed on every guarded page ( |
Open port | none | a debugging port on |
Guard lost | cannot happen | the browser is killed and every call answers |
Launch | Playwright's defaults | adds |
Dependencies | none | none: a stdlib WebSocket client ( |
How cdp guards. Chrome is started with --remote-debugging-port=0 next to the
pipe Playwright uses, and cdp_guard.py opens a second connection to it. It
auto-attaches to every page and out-of-process iframe at browser level with
waitForDebuggerOnStart, so a popup or a cross-site iframe is held until
Fetch.enable is in place on it (a session made through Playwright after the tab
exists cannot promise that: popups' first requests and cross-site iframe requests
escaped in the earlier probe). Each paused request becomes the two objects
classify reads, goes through decide, and is answered Fetch.continueRequest or a
204 Fulfill — never an abort. Nothing is asked of the page while a request is
paused (its renderer does not answer then).
Redirect hops. Each hop is its own Fetch.requestPaused, judged as the
navigation it is: a 307/308 keeps method and body, a 301/302/303 turns a POST into a
body-less GET. A site you approved that answers 302 to an unapproved origin is
refused before the request leaves — the destination sees nothing — and the audit row
names redirect_from, and navigate answers blocked_by_policy with redirected_to. A hop that stays with the origin that issued it (the same
origin, or http→https of the same host on default ports) is not judged again: it is
the request carrying on, and asking again would strand every trailing-slash redirect on
a form.
Uploads. Chrome puts a navigation's whole body into the event (as text and as base64: a 40 MB form is a 98 MB message). The transport reads such a message through without keeping it and the guard judges the event without the body (it only asks whether one exists). A 3, 12 and 40 MB multipart POST were refused when ungranted, delivered when granted, and the guard stayed up.
If the guard is lost. Chrome releases every request it holds, and every target
waiting for the debugger, the moment the connection ends, so a lost connection is a
browser running unguarded. The session therefore kills the browser, records
guard_lost in the audit trail and answers every call guard_lost (also to readers,
and to open_browser) until close_browser acknowledges it; the next browser is a
fresh, guarded one. A socket that ends because the browser quit (a crash, the user
closing the last window) is told apart without waiting: the process is already gone, or it
is running with no open tab and leaves within 300 ms. With a tab open nothing explains the
socket, so it is killed at once. Measured (scripts/verify_cdp_guard_e2e.py): socket
aborted, process gone in 1–2 ms, and a hostile page posting a form every 4 ms got 0 of
them through to its server in 3 of 3 runs (waiting 300 ms first let 72 through). Requests
already in flight at that instant were released by Chrome and are not judged. A sidecar
that cannot start, or a target it cannot guard, is the same event: the window is closed
rather than left open.
Debugging port. --remote-debugging-port is an open, unauthenticated service, and
this is what cdp costs.
What it exposes. Any process on the machine that finds the port can attach to the browser and do what the agent can and more: list tabs, read cookies (HttpOnly ones too), run script in any tab of the logged-in profile. Finding it needs no secret:
GET /json/versionhands out the WebSocket path. Withroutethere is no port; Playwright drives Chrome over a pipe only it holds. Reading the profile directly would also give a same-user process the cookies, so the difference is mostly other users of the host and sandboxed code that has loopback but not the profile directory; on a one-user workstation it is small, on a shared host it is the whole profile.What is done. Random port, chosen by Chrome, bound to
127.0.0.1only (::1refuses; measured). Chrome refuses a handshake that carries anOriginheader (a web page cannot connect, 403) or a foreignHost(no DNS rebinding, 500). The profile directory is made0700before launch;DevToolsActivePort(written0664, and left behind when Chrome exits) is removed before launch and as soon as it is read. The port and the path are never put in an exception, a log line, an audit row or an envelope.What remains. The port exists as long as the browser does; Chrome has no way to close it or to require a token. A second consumer cannot use the pipe (Chrome serves one, and Playwright holds it); putting a proxy in front of a pipe-launched Chrome means launching Chrome ourselves, which this does not do. Chrome 136 and later ignore the flag on their default profile directory because the port is a cookie-theft vector [Chrome's own announcement, not measured here]; this server never uses that directory.
Measured by the
exposuresection ofscripts/verify_cdp_guard_e2e.py, which attaches from a separate process the way a scanner would.
What neither backend sees. fetch()/XHR (not navigations), WebSockets, GETs with the
data in the URL, speculation-rules prefetch/prerender and the navigation that activates
one (above), and data:/about:/blob:/javascript: URLs (not requests). Workers never
issue a document request and are not attached; a page's own back_forward_cache is off
under Playwright's switches, and history navigations are judged with the cache on.
Choosing. cdp for sites behind DataDome, where redirect chains matter, or where
a page's service worker must not answer navigations; route where an open loopback
port is unacceptable. Every scripts/verify_*_e2e.py gate runs on either:
LYRA_BROWSER_GUARD=cdp python scripts/verify_tabs_e2e.py.
How it reaches the user
End users (VEGA.app): they do not run pip or playwright install. VEGA
bundles the playwright Python package in its own runtime, and this server drives
the user's already-installed Chrome or Edge — zero downloads. If they have
neither, a browser_unavailable envelope tells VEGA's UI to prompt a one-click
Chrome install. See docs/VEGA_INTEGRATION.md.
Developers (working on this repo):
pip install -e ".[dev]"
python -m playwright install chromium # only needed if you have no system Chrome
lyra-browser # serve over stdio (what VEGA spawns)
# or: lyra-browser --http --port 8765 # serve over HTTP for devConfiguration (env)
Var | Default | Meaning |
| detected |
|
| — | Base data dir (profile + audit). Overrides everything. |
| — | vega: data goes to |
|
| hermes: data goes to |
| auto | Run without a visible window (see Headless mode). Unset: headless for hermes, and on Linux with no |
|
|
|
| unset | Route the browser through a proxy, e.g. |
|
| Ask before risky actions. |
|
|
|
|
|
|
|
| Seconds a grant stays usable. Single-use ones ignore this. |
| — | Sites pre-approved for |
| — | Sites where sending is also pre-approved — |
|
| Seconds a prompt waits for an answer before it counts as a denial ( |
|
| Seconds before a finished action's unused single-use scope is reclaimed. |
|
| Seconds a session may hold the browser without using it. Handing it on closes the window. |
| — | Where |
|
| Captures kept on disk; oldest are deleted first. |
| — | Where declared downloads are saved. Default |
|
| Saved downloads kept; oldest are deleted first, and only files this server saved (tracked in a ledger in the dir). |
|
| Largest file a download may be (200 MiB); a bigger one is cancelled and answered |
|
| Seconds a declared download may take to arrive once it has started before it is cancelled ( |
| — | Force one channel ( |
|
| Allow bundled-Chromium fallback (only exists after |
|
| Which Playwright to launch with: |
|
| Where the navigation guard stands: |
Headless mode
lyra-browser --headless # or LYRA_BROWSER_HEADLESS=true; the flag wins
lyra-browser --no-headless # force a window even if the env says otherwiseUnset, the mode follows the client and the host: --client hermes is headless
(Hermes reaches the user over chat and gives its servers no display), and so is
any Linux process with neither DISPLAY nor WAYLAND_DISPLAY — launching a
window there would crash, headless is a working session that says so.
Headless does not merely hide the window — it means nobody is watching. The
collaboration tools have no human to reach, so they refuse with an unattended
envelope instead of reporting success for something no one will see:
Tool | Headless behaviour |
|
|
|
|
|
|
Headless Chrome also announces itself: its UA says HeadlessChrome, and
Cloudflare's managed challenge refuses on that token alone (13 of 34 login
pages in the earlier survey). The session therefore sends the UA the same
binary would send with a window — the real version, the real platform, only
that token dropped — so navigator.userAgentData and the Sec-CH-UA-* headers
stay the browser's own. The version is read from the running Chrome on the
first headless launch and remembered in <data_dir>/browser-ua.json; that
first launch, and the one after each Chrome update, relaunch once (about 0.7s).
Navigation, interaction, and reading are unaffected. open_browser reports which
mode you are in via its attended field so the agent never assumes an audience.
Looking like the user's Chrome
Sites that sort agents from people check the launch, not the behaviour. The
session is launched the way Playwright's own MCP server launches — with
--disable-blink-features=AutomationControlled, so navigator.webdriver is
false as in any Chrome a person opens, and without an emulated viewport in
headful mode, so the page sees the real screen instead of a window larger than
the screen it is on. Measured on 42 login pages behind Cloudflare, Akamai,
DataDome, PerimeterX and Kasada: of the 34 plain Chrome passes,
the Playwright defaults failed 6 headful and 19 headless; with these settings
the count is in the survey comment, per site.
What is deliberately not done: the --enable-automation infobar stays (it is
how a person at the shared window is told what is happening), and nothing is
forged — no WebGL or canvas noise, no patched CDP. Route interception and the
service-worker block contribute nothing to the Cloudflare and Akamai refusals; for
DataDome the route does matter, which is why the cdp guard backend exists.
Two vendors still refuse, and the cause of each is measured rather than guessed:
Kasada (hyatt.com) notices one CDP call,
Runtime.enable. Plain Chrome driven over raw CDP passes; the same Chrome withRuntime.enableon is refused; every other call Playwright makes on attach is fine. Playwright enables it on every page and has no switch not to. patchright is a drop-in fork that does not send it and evaluates in an isolated world instead; withpip install "lyra-browser[patchright]"the session launches through it (LYRA_BROWSER_DRIVER=auto), and hyatt passes in both modes. What patchright changes for this server:page.evaluateruns outside the page's own JavaScript world (our evaluate calls only touch the DOM, which is shared) and console messages are not delivered (we read none).DataDome (etsy.com; tripadvisor.com does not discriminate in the measurements) refuses on either of two causes, each sufficient on its own (N=5–6 per cell, headful, one IP):
context.routeswitches the HTTP cache off, so every document request carriesCache-Control/Pragma: no-cache(plain Chrome with onlyNetwork.setCacheDisabledis refused 0/6, while raw CDPFetch.enableon documents or on all requests is fine), and on a GPU-less host--enable-unsafe-swiftshadergives the page a WebGL that plain Chrome lacks. Thecdpguard backend removes both: no route, documents only over a second connection, and the switch dropped. Measured on that backend with the product itself (scripts/measure_compat.py, headful, fresh profile per visit, 12 s apart, N=5): etsy.com let in 5/5 against 0/5 for theroutelaunch; hyatt.com (patchright) 5/5 against 4/5 (the one miss was a 429). Dropping the switch removes WebGL on a host with no GPU, as in plain Chrome there; on a host with a GPU it is inert [INFERENCE: no GPU host here]. Headless is undecidable on the measuring host.
Development
ruff check . # lint
pre-commit install # enable hooksReal-browser gates live in scripts/verify_*_e2e.py, for
example python scripts/verify_browser_e2e.py --headless-only, or headful under
xvfb-run -a.
CDP guard gate —
python scripts/verify_cdp_guard_e2e.py(both modes,--driver) re-runs the browser, script-redirect and tabs gates on thecdpbackend, then checks what only it can do: redirect hops (301/302/303/307/308, with and without a body), popups at N=10, out-of-process iframes, service workers, large uploads, kill-the-socket fail-closed, a sidecar that cannot start, and what another local process can do with the debugging port.scripts/measure_compat.pyis the etsy/hyatt measurement above.Real-sites gate —
python scripts/verify_realsites.py(needs the internet) visits ten live sites and a loopback control, reads each as a tree and as text, clicks its first internal link through thearia-ref=the tree printed, and compares judged and blocked navigations and page sizes withscripts/realsites_baseline.json(±10% plus a small slack; success must match exactly). A site that cannot be reached isSKIP, and the gate fails only when fewer than eight ran. It runs against live sites, so a site editing its header can turn it red with no code change: treat it as a report to read, not a merge gate (the loopbackverify_*_e2e.pygates are the blocking ones).--update-baselinere-records it,--sites a,band--retries Nnarrow and steady a run,--driver patchrightand--headfulpick the driver and the window.
Status
The server, session manager, tools, permission layer, enforcement and audit are
implemented, unit-tested (no browser needed) and exercised end-to-end
against real sites with a real Chrome — ten rounds over eleven sites, 187 judged
navigations. Not yet driven by a live VEGA session; that is the next milestone,
and it is what would let consent_channel=elicit replace the confirm flag.
License
MIT
Available Tools
29 toolsask_user_to_doA
Ask the user to perform a step you must not do yourself.
Use for credentials, CAPTCHAs, 2FA, or payment confirmation. Highlights
the relevant element when selector is given, then returns an
awaiting_user envelope. Relay the instruction, then poll the page
(read_page / get_url) to detect completion.
Requires an attended (headful) session — in headless mode there is no
user to ask, so this returns unattended rather than making you wait.
| Name | Required | Description | Default |
|---|---|---|---|
| selector | No | ||
| instruction | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden and discharges it: it discloses the side effect (highlighting the element when selector is given), the return contract ('awaiting_user' envelope), the environmental prerequisite (attended/headful session), and the degradation behavior ('unattended' rather than waiting). This is exactly the behavior an agent needs before committing to a blocking human interaction.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with purpose, then use cases, then mechanics, then the environment caveat. Every sentence adds a distinct fact and none is redundant.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, yet the description still names the two possible envelopes, which is the key decision-relevant outcome. Combined with the headful prerequisite and the polling instruction, nothing needed to invoke or react to this tool is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate. It meaningfully explains 'selector' (highlights the relevant element when given, optional/default null), but 'instruction' is only implied via 'Relay the instruction' without stating that it is the required text shown to the user. Partial compensation for a low-coverage schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Ask the user to perform a step') and immediately scopes the domain with 'you must not do yourself', which distinguishes it from all the automated siblings like click/type_text. An agent can tell this is the human-in-the-loop escape hatch without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit use cases (credentials, CAPTCHAs, 2FA, payment confirmation) and an explicit when-not: headless mode returns 'unattended' rather than blocking. It also prescribes the follow-up workflow (relay, then poll read_page/get_url), which is actionable operational guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
clickA
Click the first element matching selector (CSS or Playwright text=).
Set submits=true when the control sends a form or completes a
purchase — that asks for the permission it needs. If the page submits
anyway without you declaring it, the request is stopped, so declare it
when you mean it. reason is shown to the user when they are asked.
Fails fast instead of hanging. A selector that matches nothing answers
not_found, and a match that cannot be used hidden or disabled
— each after at most a few seconds' wait for a page that is still
building or enabling its controls, and before anyone is asked. Once
clicking, timeout_ms (default 10000, max 30000) bounds the wait, and
a failure answers timeout (the click may have landed — look at the
page before clicking again), element_not_actionable (covered,
detached) or page_closed.
The reply says how many elements matched (matches; the first is
clicked) and what was hit (clicked: tag, role, name), so a wrong
pick is visible.
If the click makes the page navigate and the browser refuses it — no
approval covers the destination, or a form was sent without
submits=true — the reply is blocked_by_policy (with url, where
the tab still stands, and redirected_to when the site redirected the
browser elsewhere) and the tab has not moved: navigate to the
destination to be asked for it. A click that opens a tab answers
new_tab: true and tab_count; the new tab is the one every later
call reads and acts on (tabs goes back).
Set download=true when the click is meant to save a file: that asks for
permission to write one, and the reply carries download (filename,
path, bytes, url_origin) once it is saved. A click that starts a
download without it is cancelled and answered download_blocked — nothing
is saved, so repeat the click with download=true. The permission pays for
one file and ends with the call.
A native dialog the page raised is answered at once and listed under
dialogs in the reply; handle_dialog chooses the answer beforehand.
| Name | Required | Description | Default |
|---|---|---|---|
| reason | No | ||
| confirm | No | ||
| submits | No | ||
| download | No | ||
| selector | Yes | ||
| timeout_ms | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and delivers: fail-fast behavior, the not_found/hidden/disabled/timeout/element_not_actionable/page_closed failure modes, the policy-stop semantics of submits, one-file download permission that ends with the call, tab-following behavior, and dialog handling. This is unusually rich disclosure of what can go wrong and why.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The core action is front-loaded and most sentences carry distinct behavioral facts (blocked_by_policy, download_blocked, new_tab). It is nonetheless long and dense enough that a few clauses, such as the repeated 'asks for the permission it needs' phrasing, could be tightened.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a six-parameter mutation tool with no annotations and 0% schema coverage, the description is nearly complete, covering failure modes, permissions, and side effects; an output schema exists so return fields need not be detailed, though it still names the key reply fields. The unexplained 'confirm' parameter is the one meaningful hole.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, and it explains reason, submits, download, selector syntax, and timeout_ms (default 10000, max 30000). It never mentions the 'confirm' parameter, so one of six parameters remains undefined anywhere.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening sentence states a specific verb, resource, and selector syntax ('Click the first element matching selector (CSS or Playwright text=)'), which cleanly separates it from siblings like hover, type_text, and select_option. An agent knows exactly what action this performs without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives explicit conditional guidance: set submits=true when the control sends a form or completes a purchase, set download=true when saving a file, and use handle_dialog beforehand for dialogs. However, it never routes the agent to a sibling for adjacent needs (e.g. hover before click, or navigate when blocked), leaving some alternative-selection to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
close_browserC
Close the shared window, discarding every approval it earned.
This is how a session finishes: the window goes away, the grants go with it, and the next session may claim the browser. Without it the refusal another session receives would name a wait that never ends.
| Name | Required | Description | Default |
|---|---|---|---|
| reason | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden, and it does disclose a real consequence: closing discards 'every approval it earned' and 'the grants go with it', plus the cross-session effect on waiting sessions. That is genuine destructive-effect disclosure, but it omits idempotency, behavior when already closed, error conditions, and any auth/permission requirements.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The action is front-loaded in the opening clause, which is good, but the prose is heavily metaphorical and padded ('discarding every approval it earned', 'name a wait that never ends'). Three sentences do more atmospheric work than informational work.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained, and the description does convey the destructive/session-ending nature. But for a terminal, non-annotated tool the safety profile is thin — no note on whether unsaved work is lost, whether double-calling is safe, or what the reason param does — leaving meaningful gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There is one parameter (reason) with 0% schema description coverage, and the description never mentions it at all, so an agent gets no guidance on what the optional reason string is for or how it is used. The description does not compensate for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The name and first verb ('Close') clearly signal terminating the browser session, and the description adds lifecycle framing ('this is how a session finishes'). However, it never plainly says it closes the browser/session — it says 'the shared window' and 'approvals it earned', which is metaphorical and forces inference. It also does little to distinguish itself from siblings like open_browser or the takeover tools beyond the closing action.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied rather than stated: 'This is how a session finishes' gestures at end-of-session use, and the closing sentence implies you must call it so other sessions aren't left waiting. There is no explicit when-not-to-use and no named alternative, so the agent must infer the trigger condition.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_urlA
Return the current URL and page title, and how many tabs are open (tab_count).
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It does disclose the returned fields (URL, title, tab_count), which implies a side-effect-free read, but it says nothing about permissions, freshness, or behavior on non-page contexts (e.g. no tabs open).
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with zero filler. Every clause maps to a distinct returned value; nothing needs trimming.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given zero parameters and an existing output schema, the description need not explain return structure, and the tool is simple enough that little more is required. The only gap is its unstated relationship to the overlapping 'tabs' sibling.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so the schema imposes no documentation burden and the baseline of 4 applies. The mention of the 'tab_count' field is descriptive of output, not a parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Return') and precise resources: current URL, page title, and open tab count. The scope is unambiguous, though it never distinguishes itself from the sibling 'tabs' tool, which plausibly reports similar tab information.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Purely descriptive with no when-to-use guidance, prerequisites, or alternatives. An agent gets no help deciding between get_url, tabs, or read_page for retrieving page context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
go_backB
Go back one entry in history.
Where "back" leads is not known until it happens, so this asks only to
interact with the current site; landing somewhere else is judged then.
Reports http_status and content_type like navigate — both
null when the browser restored the page without asking the server.
| Name | Required | Description | Default |
|---|---|---|---|
| reason | No | ||
| confirm | No | ||
| wait_until | No | domcontentloaded |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It does add genuine value by disclosing outcome uncertainty and the null http_status/content_type semantics when the page is restored from cache, but it says nothing about what 'confirm' triggers, how 'wait_until' affects behavior, or any permission/side-effect profile.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The core action is front-loaded in the first sentence and the text is short overall. The middle clause about 'where back leads' is slightly discursive but does carry real informational content, so it mostly earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so the return-value explanation is somewhat redundant, and the tool's complexity is low. Still, with zero annotations and three fully undocumented parameters, the definition leaves real gaps an agent would need to guess at.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across three parameters (reason, confirm, wait_until), and the description never mentions any of them. The only semantic detail given (http_status/content_type nulls) concerns return values, not inputs, so the description fails to compensate for the schema gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Go back') and resource scope ('one entry in history'), which cleanly distinguishes it from navigate, reload_page, and open_browser. It does not name a sibling explicitly, so it falls just short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by the history semantics and the note that the destination is unknown until it happens, which tells the agent to expect an unverified landing. However, there is no explicit when-to-use / when-not-to-use guidance and no named alternative (e.g. navigate vs reload_page), leaving the routing decision to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
handle_dialogA
Choose how a native dialog is answered — for the NEXT browser action only.
Dialogs never block the page: every alert, confirm, prompt and
beforeunload is answered the moment it appears. Left alone, alerts and
beforeunload are accepted and confirm/prompt are dismissed
(confirm() returns false, prompt() null). What was raised is listed
under dialogs in the reply of the click, type_text or
press_key that follows — its message is the page's text, not the
user's.
Call this immediately BEFORE the one action that raises the dialog.
accept=true (the default) answers yes; accept=false answers no.
text is what a prompt() receives — leave it empty to accept the
prompt's default value; it is ignored when dismissing.
The answer is for the very next browser-acting call (click, type_text, press_key, hover, scroll, navigate, go_back, reload_page, select_option, set_editor, upload_file, save_draft, publish, tabs switch or close) and only for a dialog that call raises while it runs, on any tab. It ends with that call whether or not a dialog appeared — also when the call was refused, so arm again before retrying — and after about a minute at the latest. A dialog raised later, by a page's own timer or on a page you reach afterwards, is answered by default. Reading (read_page, screenshot, get_url, wait_for ...) does not use it up. Arming again replaces it.
| Name | Required | Description | Default |
|---|---|---|---|
| text | No | ||
| accept | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden — and it does: dialogs never block, all types are auto-answered, the default policy per dialog type (accept alerts/beforeunload, dismiss confirm/prompt with false/null), where the result surfaces ('dialogs' in the triggering call's reply), and the ~1-minute expiry. This is unusually complete behavioral disclosure for a mutation-of-state tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core purpose, and most sentences carry needed detail about defaults, scope, and expiry. It is dense and on the long side with some redundancy around expiry/refusal, but no sentence is gratuitous given the tool's subtle semantics.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return-value documentation is not required, yet the description still tells the agent where raised dialogs appear. Combined with full coverage of usage timing, defaults, and parameter behavior, an agent has everything needed to arm a dialog correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate, and it does: accept=true (default) answers yes and false answers no; text is what prompt() receives, empty accepts the prompt default, and it is ignored when dismissing. Both parameters are fully explained beyond the bare schema types.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Choose how a native dialog is answered') and immediately scopes it ('for the NEXT browser action only'). It is clearly distinguishable from every sibling, none of which handle dialogs. An agent knows exactly what this tool controls.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says 'Call this immediately BEFORE the one action that raises the dialog,' then enumerates exactly which calls consume the answer (click, type_text, navigate, tabs switch/close, etc.) and which do not (read_page, screenshot, get_url, wait_for). It also states when the arming expires and that re-arming replaces it, leaving nothing to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
highlight_elementA
Outline an element in the shared window so the user can see what you mean.
Requires an attended (headful) session — in headless mode there is no one
to show it to, so this returns unattended.
| Name | Required | Description | Default |
|---|---|---|---|
| selector | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does well: it discloses the headful-session prerequisite and the exact failure behavior (returns "unattended" in headless mode), which an agent needs before calling. It omits edge cases such as what happens with an invalid or unmatched selector, but output_schema exists so return values need not be restated.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, front-loaded with the purpose then the operative constraint. No filler; every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter visual tool with an output schema already documenting returns, the description covers purpose and the key session constraint adequately. The only real gap is selector semantics, which is left to the schema that does not describe it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the single required parameter "selector" has no schema description, so the description must compensate. It only implies that a selector identifies "an element" without stating the accepted format (CSS, XPath) or matching behavior, leaving the parameter under-specified.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ("Outline an element") plus the reason ("so the user can see what you mean"), which makes the visual-emphasis intent distinct from action siblings like click or hover. It does not explicitly name an alternative tool, so it stops short of full sibling differentiation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a clear usage condition: it requires an attended (headful) session, and in headless mode it is inapplicable. This is an explicit when-not-to-use constraint, though it names no alternative sibling for the headless case.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
hoverA
Move the pointer over the first element matching selector.
For menus and tooltips that only open while the pointer is on their
trigger: hover, then click the item that appeared. Nothing is
pressed, so no form is sent by it.
Fails fast like click: not_found, hidden or disabled
before anyone is asked, then timeout, element_not_actionable
(covered, detached) or page_closed within timeout_ms (default
10000, max 30000). The reply says how many elements matched (matches;
the first is hovered) and what was hovered (hovered: tag, role, name).
A page that navigates where no approval reaches as the pointer arrives is
answered blocked_by_policy (see click).
| Name | Required | Description | Default |
|---|---|---|---|
| reason | No | ||
| confirm | No | ||
| selector | Yes | ||
| timeout_ms | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does so richly: it discloses that nothing is pressed, the fail-fast error sequence (not_found/hidden/disabled before timeout/element_not_actionable/page_closed), the default and max timeout, the reply payload (matches, hovered tag/role/name), and the blocked_by_policy navigation case.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action before the conditional usage tip and error/timeout details. Dense and mostly economical, though the heavy backtick-laden error enumeration is slightly sprawling for a single-sentence action tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, yet the description still summarizes the return fields (matches, hovered) and covers timing, failure modes, and policy behavior. Given zero annotations, this is a complete picture for safely invoking a read-like interaction tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for 4 parameters. It adds real detail for timeout_ms (default 10000, max 30000) and clarifies selector ('first element matching'), but says nothing about reason or confirm, leaving two parameters entirely unexplained in both schema and text.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Move the pointer over the first element matching selector') and explicitly contrasts with the sibling click: 'Nothing is pressed, so no form is sent by it.' An agent can distinguish hover from click and other actions without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a clear scenario ('For menus and tooltips that only open while the pointer is on their trigger') plus the follow-up pattern (hover, then click the item that appeared). It does not state when NOT to use hover or list alternatives beyond the implied click pairing, so it falls short of explicit when/when-not guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_downloadsA
List the files this session has downloaded, oldest first.
Each entry is filename, path, bytes and url_origin (the site
the file came from), the same as the download a call returned when it
saved one. A download that was blocked or failed left no file and is not
listed. This only reads: it asks nobody, works during a takeover and does
not open the browser.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does so thoroughly: it is read-only, "asks nobody" (no prompt/permission interaction), does not open the browser, and is safe during a takeover. It also discloses error-state behavior — blocked or failed downloads leave no file and are excluded from the list.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose and ordering are front-loaded in the first sentence, followed by return shape, exclusion rule, and side-effect profile. Prose is slightly chatty ("it asks nobody") but every sentence contributes; nothing is filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter read tool with an output schema, the description covers everything an agent needs: what it lists, ordering, what is omitted, and that it is safe/non-interactive. Return values are already covered by the output schema, and the description still summarizes the entry fields.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so per the baseline rule a 4 applies. The description adds nothing to explain (nothing to explain), but it does define the returned entry shape, which is extra value.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ("List the files this session has downloaded") plus ordering ("oldest first"). No sibling does anything similar, so it is trivially distinguishable from read_draft, screenshot, tabs, etc.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No explicit when-to-use/when-not statement, but it gives actionable routing context: it "works during a takeover," which tells the agent this tool remains callable in a state where most siblings are blocked. No alternative tools exist to name, so the only gap is the absence of an explicit trigger phrasing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
open_browserA
Launch (or focus) the visible browser window the user shares with you.
If no browser is installed this returns a browser_unavailable envelope
(relay it so the user can install Chrome) rather than failing.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, and it delivers the most important non-obvious behavior: a missing browser yields a 'browser_unavailable' envelope instead of a failure, which the agent must relay to the user. It stops short of covering consent/permission or whether launching an already-open window triggers a takeover.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short, front-loaded sentences with no filler; the primary action comes first and the error-handling caveat second. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return-value detail is unnecessary, and the description adds the key exception case on top. Only minor gaps remain (user-consent or permission implications), acceptable for a zero-parameter launch tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes no parameters, so per the rubric the baseline is 4. There is nothing for the description to clarify beyond what the empty schema already conveys.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'Launch (or focus) the visible browser window the user shares with you.' This is clearly distinct from the obvious sibling close_browser, though the description never names a sibling explicitly to differentiate further.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied rather than stated: launching the browser is obviously a prerequisite for the navigate/click/read_page siblings, but the description never says 'use this first' or names alternatives. An agent can infer the ordering but must supply it itself.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
press_keyA
Press a keyboard key (e.g. 'Enter', 'Escape', 'Control+a') on the page.
No key is treated as special: Enter, Space and anything else all send
whatever the page decides to send, and that is judged when it happens.
Set submits=true if you are deliberately sending a form. A key that
makes the page navigate where no approval reaches is answered
blocked_by_policy and the tab has not moved (see click); one that
opens a tab answers new_tab: true and tab_count.
Set download=true if the key is meant to save a file: that asks for
permission to write one, and the reply carries download once it is
saved. A key that starts a download without it is cancelled and answered
download_blocked; repeat it with download=true.
A native dialog the page raised is answered at once and listed under
dialogs in the reply.
| Name | Required | Description | Default |
|---|---|---|---|
| key | Yes | ||
| reason | No | ||
| confirm | No | ||
| submits | No | ||
| download | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden and does so richly: it discloses blocked_by_policy for navigating keys, new_tab/tab_count for tab-opening keys, the download permission flow with download_blocked, and that native dialogs are answered under 'dialogs'. This is well beyond what any structured field provides.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The core action is front-loaded and each paragraph adds distinct behavioral information rather than restating the schema. Some phrasing is dense and slightly redundant (e.g. 'that is judged when it happens'), keeping it from a 5.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 5-parameter interaction tool with no annotations, the description covers the critical edge-case behaviors thoroughly, and an output schema exists so return values need not be spelled out. The remaining gap is the undocumented reason and confirm parameters.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate. It explains submits and download in detail and gives key format examples ('Control+a'), but leaves two of five parameters (reason, confirm) entirely unexplained, so compensation is partial.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Press a keyboard key ... on the page') with concrete examples like 'Enter', 'Escape', 'Control+a'. It clearly conveys the action, but does not explicitly distinguish itself from siblings like type_text or click, only cross-referencing click for related behavior.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives conditions for the flags ('Set submits=true if you are deliberately sending a form', 'Set download=true if the key is meant to save a file'), which is useful implied usage. However, it never states when to choose press_key over type_text or click, nor any prerequisite context for invoking the tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
publishA
Make a draft public — the one action that sends content to an audience.
Kept apart from every other tool so that writing a draft can never publish
it: this asks for PUBLISH, which is single-use and implied by nothing,
and it is the only tool that buys it.
A site splits this across two controls and means neither of them alone.
Measured on KVR: the publish radio (input[name=is_draft][value='0'] on
a news item, input[name=is_live][value='1'] on a product) only arms the
form — the item travels on the form's own Submit. Switching the radio and
stopping would leave the page primed to publish on someone else's next
click, which is worse than not switching it at all. So give both:
selector— the control that carries the item live, fromread_form'sdraftslist.submit— the control that sends the form. Required:read_formreports the candidates insubmitText, and on KVR they are one button (Submiton a news item,Addon a product). This is the request that actually publicises the item, so nothing else in this tool surface may send a form whose publish control is set.
Call this only when a human has approved the exact item. Show them what
will go public and get their answer first; the approval prompt this raises
is the second gate, not the first. Afterwards read_form reads back the
state the page actually reached.
| Name | Required | Description | Default |
|---|---|---|---|
| reason | No | ||
| submit | No | ||
| confirm | No | ||
| selector | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations exist, so the description carries full burden and does so: it discloses that PUBLISH is single-use and implied by nothing, that the tool is the only buyer of it, that switching a radio alone can leave the page primed for a later publish, and that the approval prompt is a second gate. That is unusually rich behavioral disclosure for a mutating action.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose is front-loaded and the two-bullet parameter layout is efficient, but the middle section carries rhetorical padding ('which is worse than not switching it at all') and restates the single-use PUBLISH idea twice. For a destructive action some verbosity is defensible, yet the prose is longer than the operational content requires.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, return values need no explanation, and the description covers the approval workflow, the two-control model, and post-action readback via read_form. The unexplained confirm/reason parameters and the required-field mismatch with the schema are the remaining gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate. It explains selector (the control carrying the item live, drawn from read_form's drafts) and submit (the request that actually publicises, candidates in submitText) with real operational meaning. However, reason and confirm are never mentioned, and the text labels submit '**Required**' while the schema marks only selector required — a discrepancy worth resolving.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource — 'Make a draft public — the one action that sends content to an audience' — and explicitly carves itself out from siblings: 'Kept apart from every other tool so that writing a draft can never publish it.' An agent can distinguish this from save_draft without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit gating condition ('Call this only when a human has approved the exact item'), describes the alternative behavior it forbids ('nothing else in this tool surface may send a form whose publish control is set'), and names the sibling that supplies the inputs (read_form's drafts/submitText).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
read_draftA
Collect what a form is about to publish, for a human to approve.
Approving an item means reading the whole of it, and on a real form that
means looking in several places at once: the visible fields, the rich-text
body (a separate document read_page never shows), the attachment slots,
and whether the item is currently a draft or already public. This gathers
all of it into one answer so the approval request can carry the actual
content rather than a description of it.
Nothing is changed or sent — this is a read, so it needs no approval.
Fields carry usable; one that is false is hidden or disabled in the
page right now and must be revealed before it can be set.
| Name | Required | Description | Default |
|---|---|---|---|
| max_chars | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the safety burden well, explicitly stating 'Nothing is changed or sent — this is a read, so it needs no approval.' It also discloses the ``usable`` flag semantics — a false value means a field is hidden/disabled and must be revealed first, which is non-obvious behavior beyond the schema. It omits failure modes (e.g., what happens when the form is already public).
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose is front-loaded in the first sentence, and the rest earns its place by explaining why the multi-source gather matters. There is mild redundancy between 'for a human to approve' and 'so the approval request can carry the actual content', but the prose stays readable and purposeful.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return-value documentation is appropriately omitted, and the description still usefully previews what is gathered (fields, rich-text body, attachment slots, draft/public status). The only real hole is the unexplained max_chars parameter, plus no note on error cases when the item is not a draft.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the single parameter max_chars is never mentioned in the description. The name hints at a character cap, but truncation behavior — whether content is cut, elided, or errors — is entirely undocumented, so the description does not compensate for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource (collect what a form is about to publish) and immediately scopes it for human approval. It explicitly distinguishes narrow-scope siblings by noting the rich-text body is 'a separate document ``read_page`` never shows', so an agent can tell it apart from read_page without inspecting either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear context: gather the full item so an approval request carries actual content, not a description. It names read_page as an inadequate alternative for the rich-text body, but does not address other plausible siblings such as read_form or save_draft, so the when-to-use routing is clear but not exhaustive.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
read_formA
List a form's fields with their names, types, labels and options.
read_page returns rendered text, which loses everything needed to
address a field: a large form arrives as labels with no name, and a
<select>'s options vanish. Use this before filling a form you have
not seen. root is a CSS selector for the form (default: the first
form on the page). Hidden inputs are counted, not listed.
Set whole_page=true for controls a page attached outside the form —
a picker or editor dialog is often appended to the body, and a field
found there is still one the form will post.
Each field carries usable: false means it is in the DOM but hidden or
disabled right now, so choosing it will fail. A form hides sections it has
not revealed yet — measured on KVR, a news form keeps event_country
hidden until the Deal/Offer type is chosen. Check usable before
addressing a field rather than discovering it through a timeout.
| Name | Required | Description | Default |
|---|---|---|---|
| root | No | form | |
| max_fields | No | ||
| whole_page | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden and does so well: hidden inputs are counted not listed, each field carries a `usable` flag whose false value predicts failure, and forms can hide unrevealed sections (illustrated with the event_country example). This is exactly the kind of runtime behavior an agent needs and cannot get from the schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the core purpose, then layers the read_page contrast and the usable/hidden-section guidance in a logical order. It is somewhat verbose, with the KVR anecdote and 'rather than discovering it through a timeout' phrasing adding length beyond the minimum needed.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained, and the description still covers field shape (name/type/label/options), the usable flag, and whole_page scope. The only real gap is the undocumented max_fields parameter, which matters for very large forms.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It explains `root` (CSS selector, default first form) and `whole_page` (controls appended outside the form, e.g. a picker dialog) thoroughly, but `max_fields` is never mentioned in the description or the schema, leaving one of three parameters undocumented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('List a form's fields') and enumerates exactly what is returned: names, types, labels and options. It explicitly contrasts itself with the sibling read_page, explaining what read_page loses (name attributes, select options), so an agent can pick the right tool without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit trigger ('Use this before filling a form you have not seen') and names the alternative read_page with the reason to prefer this one. The whole_page guidance also tells the agent when that mode is needed. Nothing about selection is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
read_imageA
Capture one element — an img, canvas, svg, a chart, a
map — as a PNG file and return its path.
This is how you look at a picture on the page: it captures the pixels
the element renders, so it works for anything drawn, not only img.
Nothing is fetched — no new request leaves the browser. Returns
not_found when selector matches nothing.
| Name | Required | Description | Default |
|---|---|---|---|
| inline | No | ||
| selector | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and delivers real behavioral detail: it captures rendered pixels rather than source, asserts no network fetch occurs ('no new request leaves the browser'), and documents the not_found outcome when selector matches nothing. Missing details about re-render timing or how the returned file is scoped keep it from a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action and return value, with the framing sentence about looking at a picture placed second and the no-fetch constraint third. Four short sentences with no filler, though the parenthetical element list and the 'this is how you look' sentence slightly overlap in function.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists and the description still clarifies that the return is a file path, plus the not_found failure case. For a two-parameter read-only capture tool the coverage is strong; only the undocumented 'inline' flag leaves a gap an agent might need.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and neither parameter is documented in the schema, so the description must compensate. It partially does for selector by explaining the no-match/not_found behavior, but never states that selector is a CSS selector, and the 'inline' parameter is completely unaddressed.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (capture) and resource (one element — img, canvas, svg, chart, map) rendered as a PNG file, and the enumerated element types make the scope unambiguous. Combined with 'one element', an agent can distinguish this from the sibling 'screenshot' (full-page capture) without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'This is how you look at a picture on the page' gives clear positive guidance on when to reach for this tool, and 'works for anything drawn, not only img' widens the applicable cases. It does not, however, name the alternative (screenshot) or state when not to use it, so it stops short of the explicit when/when-not/alternative routing of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
read_pageA
Read the page as text, or as a map of what can be acted on.
mode="text" (default) returns the rendered text (innerText), not
HTML. links=true adds the page's visible links as {text, href}
— absolute, de-duplicated, at most 100 (tree lines carry their own href).
mode="tree" returns one line per visible thing you can act on —
links, buttons, inputs, selects, checkboxes, tabs, menu items — plus
headings::
aria-ref=e12 link "Pricing" -> /pricing
aria-ref=e15 button "Save"
aria-ref=e16 checkbox "Remember me" [checked]Pass the aria-ref=... token exactly as written as the selector of
click or type_text. It reaches what a text selector cannot name:
icon-only links, several links with the same text, elements inside open
shadow roots and iframes. A ref is good for the page it was read from and
only while its element stays; a region read (selector) replaces the
refs of the read before it. When a ref is rejected or matches nothing,
read the tree again. Elements on screen come first (in_viewport
counts them). Input values are never shown. refs: false means this
Playwright predates aria refs: lines then carry none, so address elements
with role=link[name="Pricing"] or text= selectors.
selector limits either mode to one region (first match); it returns
not_found at once when nothing matches. offset and max_chars
page through long output: total_chars is its full length,
truncated says more follows, next_offset is where to continue.
Tree pages hold whole lines, so a ref is never cut in half. Every call
reads the page as it is now, so pages read while it changes or scrolls
may not line up.
| Name | Required | Description | Default |
|---|---|---|---|
| mode | No | text | |
| links | No | ||
| offset | No | ||
| selector | No | ||
| max_chars | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does so richly: refs are valid only for the page/read they came from and go stale, a region read replaces prior refs, input values are never shown, in_viewport elements come first, and not_found/truncated/total_chars/next_offset behaviors are all disclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
It is long but front-loaded and well organized, leading with the default mode then proceeding through links, tree refs, selector, and paging. The examples earn their space; a couple of edge-case sentences (refs:false) are dense but still purposeful.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex browser-reading tool with an output schema, the description is complete: it explains return shapes, staleness, ordering, truncation continuation, and failure modes, so an agent has everything needed to invoke and interpret it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate for all 5 undocumented parameters, and it does: it defines mode's two values, what links=true adds, that selector scopes to the first match and returns not_found, and how offset/max_chars page long output.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening sentence states a specific verb+resource ('Read the page') and immediately splits the two output modes (text vs action map). It clearly distinguishes itself from siblings like read_draft, read_form, and screenshot by describing exactly what page content it exposes.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives strong conditional guidance: use mode=text for rendered text, mode=tree when you need to act on elements, pass the aria-ref token to click/type_text, and re-read the tree when a ref is rejected. The only gap is it doesn't explicitly compare against sibling alternatives like screenshot or read_form for the same intent.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
reload_pageC
Reload the current page. Reports http_status and content_type
like navigate.
| Name | Required | Description | Default |
|---|---|---|---|
| reason | No | ||
| confirm | No | ||
| wait_until | No | domcontentloaded |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries full behavioral burden, and it is thin: it discloses nothing about whether reload discards form/editor state, whether it re-fetches from cache or network, or what the 'confirm' flag guards against. Reporting http_status/content_type is return info that the output schema already handles.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, front-loaded with the core action. There is minimal waste, though the second sentence about reported fields is redundant given the output schema exists.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The output schema covers return values, but three undocumented parameters with no annotations leave real gaps. The 'confirm' parameter in particular hints at destructive or disruptive behavior that the description never explains, which is inadequate for a state-changing browser action.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% across three parameters (reason, confirm, wait_until), and the description explains none of them. Notably it never addresses 'confirm' (a boolean whose purpose is opaque) or the semantics/valid values of 'wait_until', leaving the caller with no guidance.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Reload the current page'), which distinguishes it from the sibling 'navigate' (which goes to a new URL). The reference to reporting http_status/content_type 'like navigate' further clarifies its scope. It stops short of an explicit contrast with navigate/go_back, so no strong sibling differentiation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to reload versus calling navigate again, go_back, or wait_for. The only comparative statement is about output format ('like navigate'), not about when to select this tool. Usage must be inferred entirely.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
request_takeoverB
Hand control to the user. Agent mutations are blocked until they hand back.
Requires an attended (headful) session. In headless mode the takeover is refused and no state changes: blocking every mutation while waiting for a user who cannot see the window would deadlock the run.
| Name | Required | Description | Default |
|---|---|---|---|
| reason | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and mostly meets it: it discloses that mutations are blocked, that headful attendance is required, that headless mode refuses with no state change, and even explains the deadlock rationale. It does not say how control is returned (the resume_after_takeover sibling handles that), leaving one behavioral gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight paragraphs, front-loaded with the primary action and effect before the precondition. The headless explanation is a bit verbose but it earns its place by justifying the refusal behavior.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists so return values need not be explained, and the blocking/precondition behavior is well covered. However, the undocumented required parameter and the absence of explicit routing to resume_after_takeover leave real gaps for a state-gating tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There is a single required 'reason' parameter with 0% schema description coverage, and the description never mentions it at all. Since the schema docs are absent, the description had to compensate and does not, so the agent has no guidance on what to supply.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and effect: 'Hand control to the user' with the consequence that 'Agent mutations are blocked until they hand back.' This clearly distinguishes it from read/interaction siblings, though it never names its natural partner resume_after_takeover, so an agent must infer the pairing.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The headful-session requirement is an implied precondition, and the headless refusal reads as a when-not condition, but the description never explicitly says when to prefer this over ask_user_to_do or how it relates to resume_after_takeover. Usage is implied rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
resume_after_takeoverB
Resume agent control after the user has handed the session back.
Releasing a takeover goes through the consent channel: an agent that can
clear its own lock can ignore the lock. On a client without elicitation
this still succeeds — the release is recorded as the model's own
assertion, which is what consent_channel=legacy in the audit means.
| Name | Required | Description | Default |
|---|---|---|---|
| reason | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does substantial work: it explains that release flows through the consent channel, that it still succeeds on clients without elicitation, and that such releases are recorded as the model's own assertion (consent_channel=legacy in the audit). That is meaningful beyond-schema behavioral context. It does not state what happens on failure or whether the lock state is verified.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The purpose is front-loaded in the first sentence, which is good. The second paragraph is dense and jargon-laden (elicitation, consent_channel=legacy, audit), and while it conveys the key caveat, it reads as implementation commentary rather than guidance for invocation.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be described, and the behavioral caveats are covered. The gaps are the undocumented reason parameter and the absence of explicit routing against request_takeover, leaving the agent to infer the full call contract.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There is one parameter (reason) with 0% schema description coverage and the description never mentions it, so an agent gets no guidance on what to pass. With a real parameter present, the 0-param baseline of 4 does not apply, and the coverage gap is not compensated.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific action (resume agent control) and the triggering condition (after the user has handed the session back), which clearly separates it from the sibling request_takeover. It never names that sibling explicitly, but the state-based framing makes the distinction legible.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by the condition 'after the user has handed the session back,' which is a workable trigger. However, it never states when NOT to call it, nor points to request_takeover or ask_user_to_do as alternatives when the session has not been handed back.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
save_draftA
Save the item without showing it to anyone — the safe way to send a form.
A site with no separate "save draft" button means every form POST is the one that could publish, so the difference has to be read off the page rather than chosen by the caller. This sends the form only after confirming the publish control is still on its private setting, and refuses if it is not — so it cannot be used to publish, and a page left armed by anything else will not slip out through it.
Measured on KVR: a new news item opens with is_draft=1 and posting the
form then saves it privately.
submit names the control that sends the form — read_form reports it
in submitText (Submit on a news item, Add on a product). This
asks for SUBMIT but never for PUBLISH: saving is an ordinary
request, and keeping the two apart is what lets an operator write and keep
an item without ever approving a release.
Use publish instead when the item should go out to an audience.
| Name | Required | Description | Default |
|---|---|---|---|
| reason | No | ||
| submit | No | ||
| confirm | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does substantial work: it discloses the pre-flight check on the publish control and the refusal path if the page is not on its private setting, plus the concrete KVR observation (is_draft=1). It omits auth requirements, idempotency, and what a refusal looks like on the wire, so not a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose is front-loaded in the first sentence, but the body is roughly 180 words of literary rationale ('a page left armed by anything else will not slip out through it', the operator/approval framing) that a caller does not need. Several sentences are rhetorical rather than operational.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The output schema covers return values and the safety invariant is well documented, but for a 3-parameter tool with 0% schema description coverage and no annotations, leaving `reason` and `confirm` undefined is a real gap. An agent can call it, but not confidently with non-default arguments.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate for all three params. It explains `submit` well — the control that sends the form, discoverable via `read_form`'s `submitText` — but `reason` is never mentioned and `confirm` is only obliquely implied by the description of the private-setting check rather than tied to the boolean.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a concrete verb and resource (save the item privately) and explicitly contrasts it with the sibling `publish`, making it distinguishable without opening the schema. The guard behavior ('cannot be used to publish') sharpens the scope further.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It names the alternative tool and the exact condition that selects it: 'Use `publish` instead when the item should go out to an audience.' It also explains why the distinction matters, which is the when-not case an agent needs.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
screenshotA
Capture the page as a PNG file and return its path.
Use it to see the page — layout, images, charts, anything
read_page cannot put into words. The file is what you look at;
inline=true adds base64 only for a client without access to this
machine's filesystem.
| Name | Required | Description | Default |
|---|---|---|---|
| inline | No | ||
| full_page | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the burden. It discloses the output artifact (a PNG path) and the conditional base64 behavior of inline=true, which is genuinely useful context. It stops short of stating permissions or that the operation is read-only/non-mutating, leaving a small gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action, then justification, then the inline caveat. Every clause adds information, though the sentence about the file being 'what you look at' is slightly redundant with the preceding line.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be detailed, yet the path is still surfaced helpfully. Usage is well covered; the only real omission is any explanation of the full_page parameter.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% with 2 parameters. The description explains inline=true (base64 for clients without filesystem access) but says nothing about full_page, leaving half the parameters undocumented in both schema and prose.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Capture the page as a PNG file') plus the return form (its path). It explicitly distinguishes itself from the sibling read_page, so an agent can choose between them without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly names when to use it — to *see* layout, images, charts, anything read_page cannot verbalize — and names the alternative tool. The inline clause further scopes usage to clients lacking filesystem access.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
scrollA
Scroll the page. Give exactly one of to, by_y or selector.
to='top'/to='bottom': jump to that end of the document.by_y: scroll by that many pixels, positive down and negative up (at most 20000 either way per call), with the mouse wheel.selector: bring the first matching element into view.
Pages that load more as you reach the end (feeds, lazy lists) grow
after a scroll, so call scroll(to='bottom') again until the reply
says at_bottom. The reply is scroll_y, scroll_height and
at_bottom once the position has settled. A scroll that moved
nothing says so: the content may sit in an inner scrolling panel, which
selector on an element inside it reaches.
selector fails fast like click (not_found, hidden,
disabled, timeout, page_closed).
| Name | Required | Description | Default |
|---|---|---|---|
| to | No | ||
| by_y | No | ||
| reason | No | ||
| confirm | No | ||
| selector | No | ||
| timeout_ms | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does so: it documents the 20000-pixel-per-call clamp and sign convention, the lazy-list growth behavior, the settled-position reply fields (scroll_y, scroll_height, at_bottom), and the fast-fail codes for selector (not_found, hidden, disabled, timeout, page_closed).
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the mutually-exclusive-argument rule, then uses terse bullets for each mode and a short paragraph for edge cases. Every sentence carries operational information; nothing restates the tool name or schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a scroll tool with an output schema, the description covers the core invocation contract, the lazy-load repetition case, the no-movement fallback, and selector failure modes. The remaining gap is the three unexplained auxiliary parameters (reason, confirm, timeout_ms), which leaves the tool slightly incomplete for an unannotated 6-parameter schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across 6 parameters, so the description must compensate, and it only covers three: to, by_y (with sign/limit semantics) and selector. reason, confirm, and timeout_ms are left entirely undocumented in both schema and description, so an agent cannot infer their expected values.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Opens with a specific verb+resource ('Scroll the page') and immediately narrows the invocation contract ('Give exactly one of to, by_y or selector'). This is clearly distinct from siblings like click, navigate, and read_page, and the three modes are individually defined.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly routes usage: to='top'/'bottom' for document ends, by_y for pixel wheel movement, selector to bring an element into view. It also prescribes re-calling scroll(to='bottom') on lazy-loading pages until the reply says at_bottom, and tells the agent what to do when a scroll moved nothing (try selector for inner panels).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
select_optionA
Choose an option in a <select>, by value, label or index.
Give exactly one of the three. A dropdown is not a click target: this
sets the value and fires change, which is what a page listens for.
Set submits=true when the choice sends the form (an onchange
that posts) — that asks for the permission it needs.
| Name | Required | Description | Default |
|---|---|---|---|
| index | No | ||
| label | No | ||
| value | No | ||
| reason | No | ||
| confirm | No | ||
| submits | No | ||
| selector | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden and does disclose real behavior: it sets the value and fires a change event, and submits=true triggers a permission request. It does not cover failure modes, index base, or what confirm/reason do, leaving gaps.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences, front-loaded with the core purpose and the exclusivity rule. Efficient, with no filler, though the parenthetical about onchange is slightly wordy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists so return values needn't be described, and the core behavior is covered. However, for a 7-parameter tool with 0% schema coverage, unexplained parameters (selector, reason, confirm) leave the definition short of complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate. It explains the value/label/index exclusivity and submits, but says nothing about the required selector, reason, or confirm parameters, leaving several of the 7 parameters undocumented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Choose an option') on a specific resource ('<select>'), and explicitly distinguishes itself from the click sibling: 'A dropdown is not a click target.' An agent can route correctly without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a clear selection rule ('Give exactly one of the three' value/label/index) and the condition for submits=true (an onchange that posts). It also implies when not to use click. No explicit exclusions beyond that, so not a full 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
set_editorA
Write into a rich-text editor — the body a form actually posts.
For CKEditor, TinyMCE and similar the visible editor is an iframe
whose body is editable, and the <textarea> the form submits is hidden
and empty: type_text on that textarea does nothing. Leave selector
empty to use the first editor on the page, or pass the editor's own
iframe or [contenteditable] element.
html=true writes markup (links, paragraphs, images) and syncs the
editor's own data when it is a CKEditor instance; otherwise the text is
written verbatim. Set submits=true if the write sends the form.
| Name | Required | Description | Default |
|---|---|---|---|
| html | No | ||
| reason | No | ||
| confirm | No | ||
| content | Yes | ||
| submits | No | ||
| selector | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, and it does disclose real behavioral traits: writes go to the iframe body, html=true writes markup and syncs CKEditor data, and submits=true triggers form submission. However it is silent on what happens on failure, whether existing editor content is overwritten, and the semantics of confirm/reason, leaving meaningful gaps for a mutation tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the core purpose in the first clause, then uses short paragraphs to explain the iframe/textarea mechanic and the two key booleans. No wasted sentences, though the CKEditor/TinyMCE aside could be tightened slightly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained, and the description covers the primary write workflow well. But with zero annotations, 0% schema coverage and two unaddressed parameters (confirm, reason), an agent still lacks guidance on confirmation/approval semantics and overwrite behavior.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, and it explains the most consequential parameters: selector (default to first editor, or pass iframe/[contenteditable]), html (markup vs verbatim, CKEditor sync), and submits (write sends the form). It does not explain reason or confirm, leaving 2 of 6 parameters undocumented in both places.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (write into) and resource (a rich-text editor, described as 'the body a form actually posts'), and explicitly distinguishes itself from the sibling type_text by explaining why type_text fails on hidden textareas. An agent can identify the tool's job and its niche without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Names the alternative (type_text) and the condition that selects this tool instead (rich-text editors where the submitted textarea is hidden). It also gives routing guidance for selector (empty = first editor, or pass the iframe/[contenteditable]). It stops short of explicit when-not-to-use or other alternatives, so not a full 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
tabsA
List the browser's tabs, switch to one, or close one.
action="list"— every open tab with itsindex, URL, title and whether it is theactiveone (the one you read and act on).action="switch"— make tabindexthe active one, and bring it to the front so the user sees what you see.action="close"— close tabindex. Closing the active tab returns you to the tab that opened it, else the newest one; closing the last tab leaves a blank one.
Indexes shift whenever a tab opens or closes, so list again before using
one. Tabs a page opens are followed automatically, and a popup that
closes itself (a sign-in window) returns you to where you were — this is
for when you have to choose. Switching to or closing a tab asks for
permission to use that tab's site, like any other action there;
reason is shown to the user when they are asked.
| Name | Required | Description | Default |
|---|---|---|---|
| index | No | ||
| action | No | list | |
| reason | No | ||
| confirm | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and discloses rich behavior: closing the active tab returns to the opener else the newest, closing the last tab leaves a blank one, indexes shift after open/close, switching brings the tab to the front, and switching/closing asks for site permission with reason shown to the user. This covers side effects and permission flow thoroughly.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Uses bullet points to organize the three actions and behavioral notes, front-loading the main purpose. Every sentence earns its place by covering index shifting, permission requests, or automatic tab handling without redundant filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a multi-action tool with no annotations and an output schema, the description covers actions, index behavior, permission flow, and automatic following well. It neglects the confirm parameter and any error scenarios, but overall provides enough for an agent to call correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must define parameters. It explains action values (list/switch/close), the meaning of index as a tab index, and the purpose of reason (shown to the user during permission prompts), but completely omits the confirm parameter, leaving one of four undocumented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('List the browser's tabs, switch to one, or close one') and enumerates the three actions with enough detail. An agent can tell this apart from siblings like open_browser, close_browser, and navigate without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides clear context: list to see indexes, switch to activate a tab, close to remove it, and notes that automatic following handles most popups so this tool is for when you must choose manually. It lacks explicit alternatives (e.g., when to use navigate vs switch) but the usage context is well-defined.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
type_textA
Fill selector with value. The value is redacted from the audit log.
Set submit=true to press Enter afterwards, which also asks for
permission to send the form. Note that some pages submit on typing alone;
that is stopped unless you declared it, and answered blocked_by_policy
(the tab has not moved; navigate to the destination to be asked for it).
Fails fast like click: not_found, hidden or disabled
before anyone is asked, then timeout, element_not_actionable
(read-only, not a text field, detached) or page_closed within
timeout_ms (default 10000, max 30000) for each step.
A native dialog the page raised is answered at once and listed under
dialogs in the reply.
| Name | Required | Description | Default |
|---|---|---|---|
| value | Yes | ||
| reason | No | ||
| submit | No | ||
| confirm | No | ||
| selector | Yes | ||
| timeout_ms | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden and does so richly: value redaction from the audit log, submit-on-typing suppression and the ``blocked_by_policy`` outcome, a fail-fast error taxonomy (``not_found``, ``hidden``, ``disabled``, ``timeout``, ``element_not_actionable``, ``page_closed``), timeout defaults/limits, and automatic native-dialog handling surfaced under ``dialogs``.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the core action in the first sentence and keeps every subsequent sentence behaviorally relevant rather than padding. The parenthetical clauses are dense and occasionally hard to parse, but there is little wasted text.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return format needn't be described, and the description still covers errors, timeouts, dialog handling, and audit redaction. It is nearly complete except for the unexplained ``reason`` parameter, which an agent would have to infer.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must document all six parameters. It meaningfully explains ``selector``, ``value``, ``submit``, and ``timeout_ms`` (including the 10000 default and 30000 max), and implies ``confirm`` via 'asks for permission to send the form'. The ``reason`` parameter is never mentioned at all, leaving a real gap given zero schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'Fill ``selector`` with ``value``', which immediately identifies it as a form-field text-entry tool. It implicitly differentiates from siblings like ``press_key`` and ``click`` by naming them, but never clarifies vs. ``set_editor`` or ``select_option``, which an agent could plausibly confuse with text input.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives conditional usage: set ``submit=true`` to press Enter, and routes the agent to ``navigate`` when a page's submit-on-typing behavior is blocked by policy. However, it offers no guidance on choosing this tool over ``set_editor`` or ``select_option`` for other editable widgets.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
upload_fileA
Attach one or more local files to a file picker.
This hands a file on this machine to a page, which is why it asks for its
own permission, once per call. selector is the input[type=file];
a styled picker often keeps it hidden, and that is fine — a hidden input
still accepts files. Missing paths are reported before anything is
attached, so a gated call never leaves a picker half-filled.
| Name | Required | Description | Default |
|---|---|---|---|
| paths | Yes | ||
| reason | No | ||
| confirm | No | ||
| selector | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden, and it discloses genuinely useful traits: a per-call permission prompt and an atomic guarantee that missing paths are reported before anything is attached, so a picker is never left half-filled. It omits what a successful result looks like and what happens if permission is declined, but the disclosure is well above average for an unannotated tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four short sentences, front-loaded with the core action, followed by the permission caveat, the selector note, and the atomicity guarantee. Every sentence carries information; no filler or restated title.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be described. Given that, the description covers the action, its permission gate, the selector nuance, and failure timing adequately; only the semantics of the reason and confirm parameters remain unexplained for a 4-parameter tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate and it only partially does: it explains selector (the input[type=file], which may be hidden) and implies paths semantics via the pre-flight missing-path check. The 'reason' and 'confirm' parameters are never addressed, and only the vague mention of a permission request hints at confirm's role.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Attach one or more local files') and immediately differentiates itself from siblings like click, type_text, or set_editor by naming the target as a file picker. An agent can identify this as the file-upload tool without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides the context needed to invoke it: the target is the input[type=file] selector, hidden inputs are acceptable, and it asks for its own permission once per call. It does not name when not to use it or point to an alternative (e.g. ask_user_to_do), so it stops short of full routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
wait_forA
Wait until the page is ready, instead of sleeping or clicking to pass time.
Give exactly one condition:
text: the phrase is in the page's visible text — whatread_pagereturns. Whitespace does not matter, case does.state="hidden"waits for it to be gone instead (a "Loading..." message).selector: an element matching it (CSS or Playwrighttext=) reachesstate:visible(default),hidden,attachedordetached.url: the tab's URL matches. A glob over the whole URL, so use**/checkout**orhttps://site.example/cart*, notcheckout.load_state:load,domcontentloadedornetworkidle.
Waits up to timeout_ms (default 10000, never more than 30000) and
returns {"status": "ok", "waited_ms"}. Running out of time is an
answer, not a failure: {"status": "timeout", "waited_ms", "last_seen"},
where last_seen is a short excerpt of the page text (the URL for a
url wait) so you can see why. Bad arguments return error.
Use it after navigate or an action when the page draws content late: a
page that renders 300ms after load looks empty to read_page until this
returns. It only reads, so it needs no approval and works while the user
has taken over.
| Name | Required | Description | Default |
|---|---|---|---|
| url | No | ||
| text | No | ||
| state | No | visible | |
| selector | No | ||
| load_state | No | ||
| timeout_ms | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and delivers: it documents the success and timeout return shapes including 'last_seen', states that timeout is 'an answer, not a failure', notes bad arguments return 'error', and discloses that it only reads, needs no approval, and works while the user has taken over. These are meaningful behavioral facts beyond any structured field.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core purpose, then organized as scannable bullets with return-value and usage notes at the end — nothing is wasted. It is on the long side for a single tool, but the length is largely justified by six parameters and two distinct return paths.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a six-parameter, zero-coverage, no-annotation tool, this is complete: it defines every parameter, the mutual-exclusion rule, timeout bounds, both return shapes, and the operational context (no approval, takeover-safe). An agent has everything needed to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate and does: it defines the semantics of each of the six parameters — text (whitespace-insensitive, case-sensitive), state (visible/hidden/attached/detached), selector (CSS or text=), url (whole-URL glob with examples), load_state values, and timeout_ms (default 10000, max 30000). It also warns that exactly one condition must be given, which the flat schema does not express.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Wait until the page is ready') and immediately frames its role against the anti-pattern of 'sleeping or clicking to pass time'. It also references the sibling 'read_page' to anchor what it observes, letting an agent distinguish it from navigation and read tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says when to use it ('after navigate or an action when the page draws content late') and why ('a page that renders 300ms after load looks empty to read_page'). It also supplies the alternative it replaces (sleeping/clicking), so the when-to-use is unambiguous.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
29 tool updates
v0.1.0- First observed
ask_user_to_do - First observed
click - First observed
close_browser - First observed
get_url - First observed
go_back - First observed
handle_dialog - First observed
highlight_element - First observed
hover - First observed
list_downloads - First observed
navigate - First observed
open_browser - First observed
press_key - First observed
publish - First observed
read_draft - First observed
read_form - First observed
read_image - First observed
read_page - First observed
reload_page - First observed
request_takeover - First observed
resume_after_takeover - First observed
save_draft - First observed
screenshot - First observed
scroll - First observed
select_option - First observed
set_editor - First observed
tabs - First observed
type_text - First observed
upload_file - First observed
wait_for
TDQS
Scored across 29 tools
Most tools have clearly distinct purposes within a browser-automation domain. Some potential for confusion exists between read_page (text/tree reading) and read_form (form field listing), or between navigate/go_back/reload_page. However, the detailed descriptions strongly clarify when to use each, keeping overlap manageable.
Mix of verb_noun (open_browser, close_browser, read_page, read_form, read_image) and verb-only (click, type_text, navigate, scroll, hover) patterns. The conventions are still readable and the verbs are clear, but there is no single predictable pattern across all 29 tools.
29 tools is at the high end for a browser automation server, especially since many are fine-grained actions (click, type_text, press_key, hover, scroll) that could be combined into a more general 'interact' tool. The count feels heavy for the domain, though each tool does have a specific role.
Covers a wide range of browser automation needs: navigation, reading, interaction, downloads, forms, tabs, dialogs, and session control. A few gaps remain (e.g., no explicit tool for executing JavaScript or managing cookies/local storage, no direct window resizing), but these are minor for most agent workflows.
Maintenance
Related MCP Connectors
Browser MCP for logged-in tasks. Uses your Chrome — credentials stay local. Zero-token replay.
Stealth web browser for agents: search, fetch, click, download and type in persistent MCP sessions.
Hosted real Google Chrome MCP with per-user persistent state. Navigate, click, type, screenshot.
Live browser debugging for AI assistants — DOM, console, network via MCP.
Related MCP Servers
- AlicenseAqualityAmaintenanceMCP server that lets AI agents drive your real Chromium browser with your existing signed-in sessions, providing visible, local, and inspectable automation for tasks like navigation, clicking, typing, and form filling.251Apache 2.0
- AlicenseAqualityBmaintenanceBrowser automation MCP server that uses a real browser to give agents eyes and hands—open pages, click, fill, screenshot, and run scripts via accessibility-tree snapshots.2235 npmMIT
- AlicenseNot gradedqualityAmaintenanceEnables AI agents to see and control the user's real Chrome/Brave/Edge profile over MCP, so they can read pages, click, type, take screenshots, audit layouts, debug CSS, and scrape paginated or infinite-scroll data. Because it drives the normal browser via the DevTools protocol with real input events, it works on modern JavaScript apps and on sites where the user is logged in.19 npm1MIT
- AlicenseBqualityAmaintenanceA desktop browser shared by a human and an AI — same tabs, same live session. It bundles a stdio MCP server with 24 tools (page snapshot, click, type, screenshot, tabs, annotations, evaluate) so any MCP client can drive the browser the human is watching.2462,632 npm3MIT