KitchenSink4Web
OfficialKitchenSink4Web is an MCP server that lets AI assistants read, inspect, and interact with websites through a local browser, read-only by default, with optional clicking/typing and capability packs.
Read pages as structured, budgeted orientations (get_page_view) with completeness reporting, region pricing, and cheap delta reads via
sincetokens.Find elements by text, role, CSS, XPath, or natural-language description and get actable refs (find_elements).
Extract readable prose, tables, lists, links, and article text; export CSV/JSON with the extract pack.
Navigate, scroll, wait for conditions, manage tabs, and handle popups across browser sessions.
Manage sessions: open/close browsers, save/load auth state, use multiple cookie jars/identities, choose browser lanes, and emulate device/viewport/locale/timezone.
Read the action log (get_audit) to see what was done, with redacted arguments and budget counters.
Get workflow recipes and server info (get_workflows, get_server_info).
Watch URLs for changes with monitors (content hash, text appears/gone, selector count).
Enable optional packs: screenshots/PDF, network inspection, saved sessions, file downloads/uploads, approved JavaScript diagnostics, workflow recording/replay, and WCAG accessibility checks.
Safety features include credential redaction, hidden-text stripping and counting, confirmation gates, and honest refusals.
Drives an installed Firefox browser in a separate profile for reading pages, extracting content, and performing browser actions; the Firefox lane is recommended for research-heavy work because it passes bot checks that block automated Chromium.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@KitchenSink4WebNavigate to example.com and extract the article text"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
🚰 KitchenSink4Web Community Edition
Landing page · llms.txt (machine-readable capability manifest for agents and LLM crawlers)
Read websites and extract data with your AI assistant, read-only by default, with clicking and typing when you allow it.
Read websites and pull out the data you need from Claude Code, Codex CLI, Copilot CLI or any other MCP client that runs local tools. KitchenSink4Web drives a browser on your computer, returns a page within a size budget and says what it left unread, and exports tables to CSV or JSON. It starts read-only; clicking, typing and form filling are switched on only when you choose. Websites receive ordinary browsing requests; what reaches your AI provider is decided by your AI app. The Community edition is free under the AGPL. The Business edition adds a Windows installer, a signed update channel, a license your company can approve and support.
Works on: Windows, macOS and Linux for the browser tools. Image-text recognition uses Windows OCR. A browser download and some optional dependencies may be needed.
Install
Pick the route for your AI app. The commands go in PowerShell on Windows or a terminal on macOS and Linux, not into an AI chat. The package routes need Python 3.12 or newer.
Claude Desktop
Install uv, then quit and reopen Claude Desktop. Download the .mcpb file from KitchenSink4Web releases. In Claude Desktop open Settings, then Extensions, then Advanced settings, then Install extension, and choose the file. The bundle fetches the Python package the first time it starts, so the first launch needs a network connection. Restart your session and check that the tools show as connected.
Claude Code or Codex CLI
Install uv, then run the line for your app and restart your session:
claude mcp add web -s user -- uvx kitchensink4webcodex mcp add web -- uvx kitchensink4webAny other local MCP client
Use uvx as the command and kitchensink4web as its argument, or install the package and use kitchensink4web as the server command:
pip install kitchensink4webThen follow your client's guide for adding a local MCP server. Installing the package on its own does not connect it to an AI app.
Business edition
Compare the editions on the pricing page. Already purchased? Your Windows installer and download link are in your license portal.
Related MCP server: antibrowser-mcp
What it can do
53 tools with every pack and browser actions enabled. The default read-only lite launch exposes 12; lite with actions has 20.
Read a page within a chosen size budget and see what was left unread.
Find one section or element without fetching the whole page again.
Extract tables, lists and article text, then export CSV or JSON.
Take screenshots or save a page as PDF.
Read only what changed on a page you already read.
Watch a page for changes while the server runs.
Turn on clicking, typing and form filling when a task needs them.
Record a browser task and replay it with checks, with the workflow pack.
Run automated accessibility checks with the optional dependency.
What is available depends on the packs you enable and the applications installed. The full tool reference is below.
Business edition
Need a license your company can approve and a setup someone supports? The Business edition pairs these tools with a Windows installer, a signed update channel and support under the Business terms. Update checks tell you when a covered release is available; nothing installs on its own. Compare the options on the pricing page. The Community edition stays free under the AGPL, including business use that meets its terms.
Privacy Policy
The tools run on your computer, and KitchenSink4AI receives no documents and no usage data from them. Your AI app may send prompts, file contents and tool results to its own provider under that app's settings and terms. Installing downloads packages, and the Community version check contacts PyPI unless you disable it; those requests carry connection details such as your network address and never your documents. The browser connects to the websites you ask it to use, and those sites see ordinary browsing requests. A missing tokenizer table may be downloaded once; the text being measured is never sent. Cloud folders and backups follow their own settings. The Privacy Policy covers the product, purchases and support records.
Not affiliated with or endorsed by Google, Mozilla, Microsoft, or any website this server visits. Chrome, Chromium, Firefox, and Edge are trademarks of their respective owners, used nominatively to name the browsers this server can drive.
The cheap first read
get_page_view returns the page's shape: what it is, what is on it section by section, what each
section costs to open, and what the sensible next call would be. You set the budget and the read
never exceeds it, so the question is never whether a page fits, only how much detail you bought.
Every read ends with a completeness block naming what was left out and what it would cost to get
it back. Nothing is swept under the rug.
From there the pattern is cheap and gets cheaper. Expand one region instead of raising the budget. Chain delta reads: pass the token a read returns and the next read prices only what changed. Elements come back as short-lived refs your assistant can act on directly, and searches by text and role survive re-renders that break refs.
Open shadow roots are read and their contents are actable like anything else, which is what makes component-built sites (Reddit, MDN, most of the modern web) readable at all. Closed roots cannot be reached by any tool; they are counted and reported rather than silently skipped. Same-origin iframes are read and searched like the page they sit in, each labeled with its own origin; a cross-origin frame is counted and named, never entered.
The packs
The server starts lite: reading, navigation, and session management, 20 tools. Capability packs are chosen at launch and are fixed for the whole session, which means an absent pack is provably absent, not merely switched off:
Pack | What you get |
extract | Tables, lists, links, and article text, with CSV and JSON export. |
capture | Screenshots (passwords blurred), PDF export, and saving pages as files. |
network | See the requests a page makes behind the scenes. Useful for debugging websites; most people leave it off. |
storage | Use with care. Saves a signed-in session so Claude does not have to log in every time. It saves the session, never your password; treat the saved file like one anyway. |
files | Download files from pages into one folder, and upload files into page forms. Every one asks you first. |
diagnostics | Lets Claude run JavaScript on a page, one script at a time, each one shown to you for approval first. This is the most powerful and most dangerous setting on this page. Leave it off unless you know you need it. |
workflows | Record a multi-step task once and replay it later. Every replay checks that the page still matches before anything runs. |
accessibility | Checks a page against the WCAG accessibility rules using axe-core, groups what it finds by rule, and says plainly what automated testing cannot check. Needs an optional extra installed; the tool tells you how if it is missing. |
Browsers and lanes
The default lane drives the server's own bundled Chromium. It never touches your browser or your profile; sessions open in a fresh profile that is deleted on close unless you explicitly save signed-in state to a file.
Lane B drives a browser you already have installed (Chrome, Edge, or Firefox), still with its own separate profile, never yours. Installed Firefox is the lane to reach for on research-heavy work: sites that turn away automated Chromium routinely serve Firefox normally, and both Firefox lanes pass bot checks that block headless Chromium. That is lane steering, not evasion; the server does not disguise what it is, it just lets you use a browser the site treats better.
Signed-in sessions are saved and reloaded with save_auth_state and auth_state= on open. The
state file records when its session cookies expire, and loading a stale file says so plainly
instead of letting the login fail mysteriously.
manage_session(action='status') reports which browsers are installed and which lane it would
recommend for the page you are on. It never switches lanes for you.
The safety model
A browser is where your logins, your money, and your mistakes all live, so the safety layer is the product, not a feature of it.
Read-only unless you say otherwise. Out of the box the tools that click, type, and submit do
not exist. They are not disabled, they are absent from the tool list, and no instruction on any
webpage can talk your assistant into using a tool that is not there. One switch at install
(allow acting in the bundle, KS4WEB_ALLOW_ACTING=true elsewhere) brings them in.
Credential blindness. The server never reads what you type into password fields, and secrets observed in cookies and site storage are vaulted: they cannot leak into your assistant's conversation, because the text that would carry them is redacted before it leaves the server.
Hidden text is evidence, not instructions. Text a page hides from human eyes is the oldest trick for slipping commands to an AI. This server strips it from the reading channel, counts it, reports the count, and wraps everything a page says in labeled provenance markers, so your assistant always knows which words came from a stranger.
Confirmation gates. Form submissions, payments, and page scripts stop and ask you first, through the client's own confirmation prompt. As of 2026-09 that prompt displays in Claude Desktop and Claude Code; the claude.ai web client does not display it yet, so gated actions refuse there instead of proceeding unconfirmed.
Honest refusals. When something cannot be done, from a login wall to a bot check to a page that crashed its own renderer, the answer names what stood in the way and what to try next. A refusal that explains itself is cheaper than a retry loop.
Two different things ask you for permission
Your MCP client's own permission prompt names a tool (something like mcp__web__type_text) and
comes from the client, not from KS4Web; approving a read-only tool group there is safe and stops
most of the asking. A KS4Web gate names an action in plain words ("submitting this form sends a
password") and cannot be pre-approved away for payments, credentials, posts, deletions, or legal
assent. If a prompt names a tool, it is the client; if it names what is about to happen in the
world, it is KS4Web.
Consent scope
Out of the box the consent scope is research: reading and navigation are free, query-shaped
submissions (a search box) proceed, and everything else asks first. The full scope lets ordinary
form submissions proceed without asking; the irreducible set asks under every scope, always. The
install screen's "Submit routine forms without asking each time" checkbox is the whole choice.
The GET rule is the one place page-authored text classifies down rather than up: a form that looks
like a search is allowed to proceed as one, and what a mistaken call in that direction buys is an
unprompted ordinary submission, never a payment, credential, post, or deletion; those always ask.
In the other direction, a submit button that says nothing recognizable in any shipped language
passes ungated under the full scope; if that risk matters to you, stay on research.
Client compatibility
If a feature refuses with CONFIRMATION_REQUIRED on claude.ai web, it works on Desktop and Code.
Client | What works |
Claude Desktop | Everything, with confirmations rendered. |
Claude Code | Everything, with confirmations rendered. |
claude.ai web | All reading, navigation, and search. Gated actions refuse honestly because no confirmation can render; pre-authorized classes work if configured at launch. |
Two numbers that matter
The same page, read two ways. A raw dump of the Treaty of Versailles article on Wikipedia costs 33,073 tokens of your assistant's memory. This server's first read of it costs 4,445. The difference is not compression, it is a different product: a map of the page with a price on every region, instead of the whole page whether you wanted it or not.
Numbers on this page are measured by scripts in tools/, not written by hand. Re-run them
yourself; the README is regenerated from their output.
.venv/Scripts/python.exe -X utf8 tools/measure_readme_numbers.pyContext cost (measured)
Surface | Tokens |
Lite tool surface | 6.9k |
Full surface (all packs) | 17.9k |
First read of the Treaty of Versailles article on Wikipedia | 4,445 |
Delta read after one click | 82 |
The budget ladder has 16 rungs; every read reports which rung it printed at and what a deeper
read would cost. Inside a subagent, start at budget_tokens=2500; tool results are capped more
tightly there.
Update check (opt-out)
The server compares its version against PyPI's when you call
manage_session(action='status'), and never at any other time: nothing runs at startup and nothing
runs on a background thread. At most one request per seven days, with a two-second timeout. It never
installs anything. The request is a plain HTTPS GET to pypi.org for this package's public release
index; it sends no identifier, no usage data, no page content and no session state. A check that
could not complete says so and says how long ago the last successful one was, rather than showing
nothing. KS4WEB_UPDATE_CHECK=off switches it off, and the status report then says it is off; the
older KS4WEB_NO_UPDATE_CHECK=1 spelling is still honored.
Testing
2,066 tests, of which 733 drive a real browser. Beyond the suite, every release passes gate batteries that re-run the adversarial findings of four attack rounds: prompt injection through page content, hidden-text smuggling, credential-theft attempts, gate bypasses, workflow replay tampering, and resource abuse. The gates are not aspirational; each one exists because an adversarial round or a field tester actually broke something, and the fix is pinned by the test that would catch its return.
The current tuning came out of a live field campaign: two sessions, twelve real sites, three browser engines, and 141KB of logged friction, retests, and design notes. The findings, and what shipped in response, are in the repository history rather than a marketing page.
Maturity: what the version number does and does not claim
This is the newest member of the family, and the version number says so honestly. It grew up under the house rules: every claim measured, every release attacked before it ships, every refusal honest about its reason. What it has not had yet is a long life in strangers' browsers. The safety net while it earns one: it starts read-only, it prices every read before taking it, and it tells you what it could not see. Bring it your strangest pages and file what breaks.
Known limits
Cross-origin iframes are never entered; the completeness block names each frame and says why. Same-origin iframes are read, searched, and acted in, each labeled with its own origin.
Closed shadow roots cannot be reached by any tool. They are counted so you can tell a component-heavy page from an empty one.
Confirmation-gated actions (auth loading, page scripts, submits) refuse on the claude.ai web client until it renders MCP confirmation prompts; Claude Desktop and Claude Code work today.
Bot walls and CAPTCHAs are reported, never solved or evaded. The Firefox lanes get through device checks that block headless Chromium; a wall that blocks everything is a wall.
sendBeaconand ping requests bypass network routing rules; that is a Playwright limit, documented rather than hidden.
Automated accessibility checking finds a minority of real-world issues. Deque, who build the axe-core engine this tool runs, publish 57% as their own figure for their own tooling. A clean report here means the automated checks passed, not that the page is accessible. There is no score, because every 0-to-100 accessibility number is somebody's weighting rather than a measurement.
Nothing taps you on the shoulder. No MCP client in use today delivers a server-initiated
message into a conversation, so monitors run only while KS4Web runs, and the report is how you ask
what happened; after a restart, the report names the window that went unchecked. A monitor watches
one URL for one deterministic change and does not understand what changed: content_hash is noisy
on pages with clocks or counters, only the main document is watched, and a change that appeared and
reverted between two checks is invisible. A monitor cannot get past a login or a bot wall, and one
that keeps failing pauses itself rather than hammering the site. Five minutes minimum between
checks; everything stays on this machine.
A session handle transfer moves a live session between conversations on the same machine, inside the same running KS4Web. Nothing is copied: it is the same browser, so a page the other conversation navigated has lost your refs, and the report says which. It does not survive a server restart, and it does not cross between Claude Desktop's chat and Code panels, which run separate servers. The session dies with KS4Web on purpose; that is what keeps orphaned browsers off your machine. A token works once, expires in an hour, and is not a lock: any conversation on this server can already see the session. The budget travels with the session.
Two identities cost two browser processes and two profile directories, which is real memory. The action budget is shared across them on purpose, so opening a second identity does not double what a session may do to the world, and two contexts visiting the same site cost one origin against the origin budget. Auth state saves and loads per identity. Closing one identity leaves the others working; closing the last one tells you to close the session instead. Every context shares the session's lane and emulation; two lanes means two sessions.
The lane database is a list of which websites this computer has visited with which browser: hostnames and dates, no pages, no addresses, nothing typed. It exists so a site that refused one browser can be read with one that works. It is stored on this machine, it never leaves unless you export it, learning can be turned off with one switch, and one call erases it entirely.
Install options and settings
The install routes are at the top of this file. This section covers what you choose when the server starts.
The Claude Desktop bundle puts the whole configuration on its install screen as checkboxes: one to allow clicking and typing (off means read-only browsing, the shipped default), and one per capability pack. The server drives its own bundled Chromium, never your browser or your profile.
Packs and read-only mode are chosen at launch (--packs, KS4WEB_MODE,
KS4WEB_ALLOW_ACTING) and are identical for every connection to the
process.
Tip for Claude Desktop: in Tool permissions, set this server's Read-only tools group to Always Allow. Those tools cannot change anything on any page, and it stops most permission prompts.
License
KitchenSink4Web is dual-licensed:
AGPL-3.0 (open source). Free for anyone (individuals, academics, and businesses) for any use that complies with the AGPL's terms. Those terms include sharing source, including your modifications, when you distribute the software or make it available over a network.
Commercial license. For organizations that want to build KitchenSink4Web into their own products or services without the AGPL's source-sharing obligations. Contact licensing@kitchensink4.ai.
Copyright (c) 2026 Alvut Consulting, LLC. KitchenSink4AI is a product line of Alvut Consulting, LLC.
Not affiliated with or endorsed by Google, Mozilla, Microsoft, or any website this server visits. Chrome, Chromium, Firefox, and Edge are trademarks of their respective owners, used nominatively to name the browsers this server can drive.
Available Tools
12 toolsfind_elementsFind ElementsARead-only
Find elements by text, role plus accessible name, natural-language
description, CSS, or XPath, and get back refs you can act on plus a note
on what was not searched. role='button' narrows any query to one
element role (field finding: 'Comment' alone matched 12; with the role
filter it matches the one button). This is the cheap targeted follow-up
that pairs with get_page_view: the page view tells you what string to
look for, and this retrieves it for a fraction of a full read.
Ambiguous results are listed rather than resolved, and zero results
come back with the nearest misses so a miss is a one-turn recovery.
The search covers the main document and every open shadow root in it,
and the matches it returns from a shadow root are actable like any
Same-origin iframes are searched too and the result says which ones
it entered. Two things stay out and the result counts both:
cross-origin iframes, which no tool here opens, and closed shadow
roots, which no tool can reach. XPath is the one kind that does not
enter a shadow root. location={'region': 'r7'} (or a ref, form, or table
from a read) narrows the search to that subtree, components inside it
included, and the first result line names the scope that was searched.
| Name | Required | Description | Default |
|---|---|---|---|
| kind | No | auto | |
| page | Yes | ||
| role | No | ||
| limit | No | ||
| query | Yes | ||
| location | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With readOnlyHint=true and openWorldHint=true, the annotations cover safety, but the description adds substantial behavioral detail: ambiguous results are listed rather than resolved, zero results return nearest misses, searches cover shadow roots and same-origin iframes, and XPath does not enter shadow roots. This goes far beyond the annotations and helps the agent predict edge-case behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long but dense and front-loaded: the core purpose appears in the first sentence, followed by the role-filter example, pairing guidance, ambiguity behavior, scope limitations, and location narrowing. Each sentence earns its place, though the paragraph-style formatting and minor punctuation issues slightly hurt scannability.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a six-parameter tool with no output schema, the description gives a strong picture of inputs, scope, limitations, and result behavior (refs, notes on what was not searched, nearest misses, scope named in the first result line). Missing details like what 'kind' does and how pagination/limit behaves keep it from being fully complete, but the essential invocation context is present.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 0% description coverage, so the description carries the full burden for parameter explanation. It does explain query, role, and location well (including an example for role and a location example), but it never mentions the required 'page' parameter, nor does it explain 'kind' or 'limit' beyond the default. This is a meaningful gap for a tool with two required parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Find elements by text, role plus accessible name, natural-language description, CSS, or XPath.' It clearly distinguishes this tool from get_page_view by calling it a 'cheap targeted follow-up,' so an agent knows it is a retrieval/search action rather than a comprehensive read.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly positions this tool relative to get_page_view: 'the page view tells you what string to look for, and this retrieves it for a fraction of a full read.' It also specifies scope exclusions (cross-origin iframes, closed shadow roots) and when XPath behaves differently, giving the agent concrete guidance on when this tool is appropriate.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_auditRead the Action LogARead-only
Read the action log. Returns one record per call, paginated, carrying the timestamp, lane, page, URL, resolved target with its human label, redacted arguments, outcome, any rebind, any confirmation decision, and the budget counters at that moment. This is an operational record for the user, so you can always know exactly what was done even where a web action cannot be undone. It is not forensic and not evidence.
| Name | Required | Description | Default |
|---|---|---|---|
| tool | No | ||
| limit | No | ||
| session | No | ||
| start_index | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, but the description adds significant behavioral context: pagination, specific fields returned, and a clear caveat that the log is 'not forensic and not evidence.' It also explains the purpose of the log for undo scenarios, which goes beyond the annotation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-organized: it starts with the purpose, then lists the returned fields, then gives the operational rationale, and ends with a caveat. Every sentence adds value without redundancy, and it is appropriately sized for the complexity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Since there is no output schema, the description does explain the return content in detail, which is helpful. However, it omits any explanation of the input parameters, leaving agents to guess at filtering and pagination mechanics. For a tool with 4 parameters and no schema descriptions, this incompleteness is notable.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description was expected to compensate for parameter meaning, but it does not mention any of the four parameters (tool, limit, session, start_index). It only implies pagination without explaining how the limit/start_index work. This is a critical gap for correct invocation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb and resource: 'Read the action log'. It elaborates on the content of each record, making it distinct from sibling tools like get_page_view or get_workflows. The purpose is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides contextual guidance ('operational record for the user, so you can always know exactly what was done') but does not explicitly compare to alternatives or state when not to use it. The use case is implied rather than explicitly routed against siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_page_viewGet Page ViewARead-only
Read a page as an ORIENTATION, not a transcript, under a token budget
it never exceeds whatever the page size. Returns identity, landmark
regions each priced with the cost to expand it, the interactive surface
with refs you can act on, a digest or app skeleton, form and table
inventories, an account of what was NOT read and why, and the next call
for anything unexpanded. location scopes to one region ref,
budget_tokens=2500 suits a subagent, mode='links' includes in-prose
links at their real cost. since=<read_token> is the cheap repeat
read: only what changed, refs kept, a few hundred tokens instead of a
fresh read, and it falls back to a full read when the page navigated in
between and nothing survives to diff. Open shadow roots are read and
their contents get refs you can act on; closed roots cannot be reached
by any tool and are counted at creation, so the completeness block
reports both numbers rather than one confident zero. Same-origin
iframes are entered and read, and their contents get refs naming the
frame they came from; a cross-origin frame is never entered, because
its document belongs to an origin the page itself cannot read either,
and the completeness block counts every frame it did not open.
| Name | Required | Description | Default |
|---|---|---|---|
| mode | No | auto | |
| page | Yes | ||
| view | No | auto | |
| since | No | ||
| detail | No | standard | |
| location | No | ||
| budget_tokens | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations provide readOnlyHint=true and openWorldHint=true, and the description reinforces both consistently while adding substantial context beyond them. It discloses the token-budget guarantee, what the completeness block reports for closed shadow roots and cross-origin frames (dual counts rather than a single zero), and the exact fallback behavior of the `since` read. None of this is derivable from the annotations, and it correctly explains the open-world behavior of reporting what was NOT read and why.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Despite its length, every sentence earns its place: the core purpose is front-loaded, followed by return contents, then parameter guidance, then edge-case behavior. There is zero filler or tautology. The density is justified by the tool's genuine complexity, and the structure flows logically from what it returns to how to control it to what its limits are.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description correctly explains the return payload (identity, landmark regions with expansion costs, interactive surface refs, digest/app skeleton, form/table inventories, unread account, and next call). It covers the complex edge cases of shadow roots and iframes thoroughly. The only incompleteness is the unexplained `view` and `detail` parameters, which leaves an agent guessing at part of the input surface. This is a small but real gap in an otherwise comprehensive definition.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description carries the full parameter-documentation burden. It explains `location` (scopes to one region ref), `budget_tokens` (with a concrete subagent value), `mode` (links at real cost), and `since` (cheap repeat read). However, three of the seven parameters—the required `page`, plus `view` and `detail`—receive no explanation. While `page` is fairly self-evident from the name, `view` and `detail` are not obvious and their interaction with the orientation read is undocumented. This is a genuine gap given the zero schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a precise, specific purpose: 'Read a page as an ORIENTATION, not a transcript, under a token budget it never exceeds whatever the page size.' The verb-resource pair is unmistakable, and the 'not a transcript' contrast actively differentiates it from sibling tools like get_text. An agent immediately knows this is the high-level structural read, not a raw text extraction.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description embeds strong usage direction: `budget_tokens=2500` suits a subagent, `mode='links'` includes in-prose links, and `since=<read_token>` is the cheap repeat-read strategy with explicit fallback semantics. The 'not a transcript' phrasing plus the open/closed shadow root and same-origin/cross-origin iframe rules tell an agent exactly when this tool is appropriate versus alternatives. It even specifies when the since-read fails over to a full read, which is precisely the kind of decision guidance that matters.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_server_infoGet Server InfoARead-only
Report what this server is and what it is running: the product and PyPI package names, the version a client was handed at initialize, the homepage and repository, the sibling servers in the same suite, the tool surface registered in THIS process (how many tools, how many of them are the lite core, which capability packs are loaded and which exist, and the read-only grade and consent scope in force), the test figures measured for this release, and the host Python and platform. Counts come from the live registry, so they describe the process you are connected to rather than the product in general. Needs no browser, opens no page, and reveals no path or user name, so a bug report can carry the output as it stands.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish readOnlyHint and openWorldHint, but the description adds meaningful behavior beyond that: it requires no browser, opens no page, leaks no path or username, and reads live registry state rather than reporting product-level generalities. These are exactly the kind of side-effect and privacy disclosures an agent needs before invoking the tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three dense sentences, each earning its place: the first enumerates the full report scope, the second clarifies data freshness and provenance, and the third states safety/privacy implications. Although the first sentence is long, it is information-dense and well structured rather than padded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no parameters, no output schema, and only safety annotations, the description carries full responsibility for making the tool callable. It enumerates all return categories, clarifies that counts describe the live connected process, and reassures about side-effect and privacy boundaries. Nothing essential is missing for an agent to invoke and interpret the result.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters and schema coverage is 100%, so there is nothing for the description to add about parameter meaning. The provided baseline of 4 applies, and the description instead spends its effort on return-content and behavioral detail, which is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Report what this server is and what it is running.' It then enumerates a comprehensive, unique scope—product and package names, version, repo, tool surface, test figures, Python, platform—that clearly separates it from browser-action and workflow siblings. Even without naming other tools, the content makes the tool's distinct purpose unmistakable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear usage context: it is safe for bug reports because it 'needs no browser, opens no page, and reveals no path or user name.' It also clarifies that counts are live from the connected process, preventing misinterpretation. However, it does not explicitly name alternative tools or state when not to use it, stopping short of full when/when-not guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_textGet TextARead-only
Extract readable prose from a page or one region of it, paginated by
start_index so a long article is read in bounded pieces rather than
one unbounded dump. Text arrives as labeled data with its origin stated,
and hidden regions are stripped and counted rather than silently dropped
or silently included. Hidden content IS retrievable, deliberately:
include_hidden=true returns it in a separately labeled section with the
hiding technique named per block. There is no silent middle tier,
because display:none is a real injection channel; the labeled route is
the whole design. Prose inside open shadow roots is read, the same as
get_page_view reads it; closed roots are counted and stay unreadable.
Prose inside same-origin iframes is read after the main document, each
frame under a header naming it and its origin, because one page can
now deliver text from several documents and a single origin in the
label would be a claim about only one of them.
| Name | Required | Description | Default |
|---|---|---|---|
| page | Yes | ||
| location | No | ||
| max_chars | No | ||
| start_index | No | ||
| include_hidden | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare readOnlyHint and openWorldHint. The description goes far beyond that: it explains pagination mechanics, how hidden content is handled (stripped, counted, and retrievable via include_hidden with per-block technique naming), the behavior for open vs closed shadow roots, and iframe handling with origin labels. It discloses the absence of a 'silent middle tier' and explains the design rationale. This is exceptionally transparent about side effects and edge cases.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single dense paragraph, but every sentence adds meaningful behavioral information. It leads with the core purpose, then covers pagination, hidden content, shadow roots, and iframes. There is no filler or repetition, though it could be tightened by moving some design rationale to a separate note. It is front-loaded and structured logically.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with 5 parameters and no output schema, the description is remarkably thorough: it explains the return format (labeled data with origin), pagination, hidden content, shadow roots, and iframes. The main gap is the 'location' parameter, which is only vaguely referenced as 'one region of it' without defining its structure. Also, the exact interaction of max_chars with pagination isn't spelled out, but the general behavior is clear. Overall it covers most of what an agent needs.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has zero description coverage for its 5 parameters. The description compensates for start_index (paginated) and include_hidden (returns hidden content in labeled section), and implicitly covers max_chars via 'bounded pieces'. However, the 'location' parameter is never explained beyond 'one region of it', leaving its structure and semantics undocumented. With 0% schema coverage, the description should have covered all parameters; it covers three partially and one not at all.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Extract readable prose from a page or one region of it'. It immediately distinguishes itself from siblings by mentioning pagination via start_index and explicitly comparing to get_page_view for shadow roots. This is far more specific than a generic 'get text' and leaves no doubt about what the tool does.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context for when to use the tool (for bounded, paginated prose extraction, including hidden content) and even references get_page_view to clarify shadow-root behavior. However, it never explicitly states 'use get_text when you need X, use get_page_view when you need Y' or lists exclusions. The guidance is implicit through behavioral description rather than an explicit decision rule.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_workflowsGet WorkflowsARead-only
Get recipes for this server: the cheap-read-then-act pattern, the auth workflow (headed handoff plus saved state), reading strategy, budgeting, troubleshooting a page that will not read, the subagent budget setting, lanes, what each capability pack contains with the exact launch flag that loads it, what a site profile is and where it lives, and how to record and replay a multi-step flow. Packs are chosen at launch rather than at runtime, so this is where you learn which flag you need before restarting. Tool availability reflects the packs this server was started with.
| Name | Required | Description | Default |
|---|---|---|---|
| topic | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The readOnlyHint annotation already covers the safe read-only nature, and the description adds meaningful behavioral context: tool availability reflects the server's launch-time packs alarm, and flags must be discovered here before restart. It does not describe the response format or invalid-topic behavior, but the annotation burden is lower here.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The core purpose is front-loaded and the topic list is compact. The final two sentences about launch-time pack selection are useful but slightly redundant with each other, which costs a point.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only documentation tool, the description explains what it returns, why it matters, and when to call it. It lacks explicit return-format details and precise parameter behavior, but the optional-topic schema and readOnlyHint annotation make those gaps non-critical.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The only parameter, topic, has 0% schema coverage, so the description must compensate. The enumeration of concrete workflow topics gives the agent plausible values, but the description never explicitly maps these to the topic parameter or explains what happens when it is omitted or null. It partially compensates without fully documenting the parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a clear verb-resource pair ('Get recipes for this server') and enumerates the specific workflow topics returned, so an agent can tell this is a documentation/how-to tool. It does not explicitly contrast itself with sibling tools like get_server_info, but the listed content is specific enough to prevent serious confusion.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear when-to-use context: packs are chosen at launch, so this tool is where an agent learns which launch flag is needed before restarting. It does not explicitly name alternative tools or state when not to use it, so it stops short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
manage_sessionManage SessionBDestructive
One session is one browser. open starts one on a chosen lane (the
first navigate of a conversation can open one for you, and says so when
it does); close ends the browser; with auth state saved, the login
survives; the session never does. status reads liveness first, so a
dead browser says so, then lists every session this server holds with its
pages, budgets, and what is shared: sessions belong to the server
process, not to a conversation. capabilities, budget, and
reset_budgets report and manage limits. handoff hands the window to
the human for a login, an MFA prompt, or a bot wall. export_handle and
import_handle move a live session between conversations on this
machine: see get_workflows(task='session-transfer'). lanes is the
local lane database (status, export, import, erase). profiles reloads
site profiles from disk. On open, auth_state loads a saved login,
contexts=2 gives one session two independent cookie jars, and device,
viewport, locale, and timezone set what pages in this session
believe about their environment.
| Name | Required | Description | Default |
|---|---|---|---|
| op | No | ||
| lane | No | ||
| note | No | ||
| path | No | ||
| site | No | ||
| token | No | ||
| action | No | status | |
| device | No | ||
| locale | No | ||
| reason | No | ||
| context | No | ||
| session | No | ||
| contexts | No | ||
| timezone | No | ||
| viewport | No | ||
| auth_state | No | ||
| expires_minutes | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations (destructiveHint=true, readOnlyHint=false) already flag mutation, and the description adds substantial value on top: close kills the session but saved auth survives, status reads liveness first so dead browsers are reported, and sessions belong to the server process, not the conversation. This ownership and persistence disclosure is exactly the kind of behavior an agent needs and that structured fields cannot convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Every sentence is information-dense with no filler, and the core concept is front-loaded in the opening clause. However, the structure is a single unbroken prose paragraph mixing operation semantics, parameter effects, and cross-references; a dispatcher of this complexity would be far more scannable as bullets or an op-to-semantics table.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with 17 parameters, 11 sub-operations, no output schema, and no enums, the description provides substantial orientation: it discloses status output shape, session lifetime, liveness checking, and shared-state semantics. Yet it leaves major gaps — the op/action dispatch relationship, output of most operations, and the meaning of expires_minutes, note, reason, path, site, and token — so it is only partially complete for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description carries the full burden for 17 parameters, but it only explains a subset: auth_state, contexts, device, viewport, locale, timezone, and lane. Critical dispatcher parameters — op, action, session, note, reason, path, site, token, context, expires_minutes — are never mapped to their effects, leaving an agent to guess at the mechanics of calling this tool.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description anchors on a strong core metaphor ('One session is one browser') and enumerates every sub-operation (open, close, status, handoff, lanes, profiles, etc.), which clearly distinguishes it from siblings like get_page_view, navigate, and monitor. However, 'manage session' is a generic dispatcher verb, and the purpose must be excavated from a dense wall of text rather than stated as a crisp mission.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is real usage context: it tells the agent that the first navigate may auto-open a session, that handoff is for login/MFA/bot walls, and explicitly points to get_workflows(task='session-transfer') for session transfer. But it never states when NOT to use this tool versus siblings such as monitor or manage_tabs, and per-operation selection guidance is minimal for a dispatcher with 11 operations.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
manage_tabsManage TabsADestructive
List, open, select, or close tabs, report which is focused, and
capture pages a click opened in a popup. Mints and returns the explicit
page handles every other tool accepts, which is how browser state
survives across calls without relying on protocol sessions. Closing a
page invalidates its refs and its delta read tokens, and the result says
so rather than leaving a later failure to explain it. On a session with
more than one cookie jar, context says which jar a new tab opens in
and every listed page says which jar it belongs to.
| Name | Required | Description | Default |
|---|---|---|---|
| url | No | ||
| page | No | ||
| action | No | list | |
| context | No | ||
| session | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotations already signal mutation and destructiveness (readOnlyHint=false, destructiveHint=true), and the description adds valuable behavioral detail: closing a page 'invalidates its refs and its delta read tokens,' and the result announces this instead of leaving a later failure. It also reveals the cookie-jar context behavior, going beyond the structured annotation hints.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but each sentence earns its place: the action list, the handle mechanism, the invalidation consequence, and the cookie-jar context. It is front-loaded with what the tool does, though the later sentences are somewhat intricate for the available parameter detail.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
This is a multi-operation, destructive tool with five parameters, no output schema, and no enum constraints, yet the description leaves the parameter model largely implicit. It does not explain the allowed `action` values or how to request 'report which is focused' or 'capture pages a click opened in a popup,' so an agent cannot reliably construct a correct call from the definition alone.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, but it only meaningfully explains `context`. The description implies action-related behaviors and page/session concepts, but it never maps parameters like `action`, `url`, `page`, or `session` to concrete values or usage, leaving the agent to guess how to invoke each operation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb list and resource: 'List, open, select, or close tabs, report which is focused, and capture pages a click opened in a popup.' It also states the tool's unique role by explaining it mints 'explicit page handles every other tool accepts,' which distinguishes it from the sibling tools that consume those handles.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It clearly conveys when to use the tool: to manage tabs and, importantly, to obtain the page handles that other tools depend on for state across calls. It does not explicitly name alternatives or give when-not-to-use guidance, so it stops short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
monitorWatch a URL for ChangeADestructive
Watches one URL for one deterministic change while this server runs. Four conditions: content_hash, text_appears, text_gone, selector_count; the same page state always answers the same way. Nothing is pushed anywhere: report is how you ask what happened, and check_now forces a check. A monitor that could not check reports stale with the failure, never unchanged. Checks run in a dedicated headless session, never the caller's. There is a floor on the interval and caps on monitors and daily checks; a 429 is honored rather than retried; a monitor that keeps failing pauses itself.
| Name | Required | Description | Default |
|---|---|---|---|
| url | No | ||
| label | No | ||
| since | No | ||
| value | No | ||
| action | No | report | |
| monitor | No | ||
| selector | No | ||
| condition | No | ||
| check_interval_minutes | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the annotations, it discloses deterministic evaluation, no push delivery, stale-on-failure reporting, dedicated headless session, interval floors and quotas, honoring 429s without retry, and auto-pause on repeated failure. This is exemplary behavioral disclosure and does not contradict readOnlyHint/destructiveHint.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The text is dense but every sentence adds a distinct constraint or behavior, and the most important information about scope is front-loaded. No filler or redundant restatement of the schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The behavioral context is unusually complete for a 9-parameter tool with no output schema, and it even hints at report semantics. Yet an agent still cannot fully infer how to create/select a monitor, what `since` and `label` mean, or what the exact response shape is, leaving operational gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema coverage, the description must carry parameter meaning, and it does name the four condition values and report/check_now actions while indirectly suggesting url, selector, value, and interval. However, label, since, and monitor remain unexplained, and no parameter requirements per condition are spelled out, so it only partially compensates.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening line 'Watches one URL for one deterministic change while this server runs' states a specific action, resource, and scope, and the four condition names make the function concrete. It does not name a sibling for contrast, so it misses the explicit differentiation needed for a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description makes the intended use implicit: a long-running watcher for one deterministic page change, with 'report' and 'check_now' as the query/force actions. It never says when to prefer this over siblings like wait_for or get_page_view, so the agent is left to infer tool selection.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
scrollScrollA
Scroll by an amount, to a named element, to the end, or inside a specific container, plus a next-chunk mode that remembers position across calls so a long page is walked without re-reading it. Reports how much content is now reachable and how much remains below, and names virtualized containers where the DOM holds far fewer rows than the page claims, rather than presenting a partial list as complete.
| Name | Required | Description | Default |
|---|---|---|---|
| page | Yes | ||
| action | No | by | |
| amount | No | ||
| location | No | ||
| timeout_ms | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the annotations, the description reveals meaningful behavior: it remembers scroll position across calls in next-chunk mode, reports reachable vs. remaining content, and explicitly identifies virtualized containers instead of presenting a partial view as complete. This is valuable transparency that annotations alone do not provide.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two dense sentences with no filler. It front-loads the primary action modes and then adds behavioral reporting details, making it easy to scan while still packing relevant information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
It covers the main behaviors and high-level return information, but because there is no output schema, it lacks specifics about the exact response format. It also does not define the action enum values or the semantics of timeout_ms, leaving some gaps for an agent trying to invoke the tool precisely.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It alludes to action modes and 'location' as a named element/container, but it never explains allowed action values, the exact meaning of amount, how page is used, or timeout behavior. The description gives only loose mapping to the five parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('scroll') and enumerates distinct modes: by amount, to a named element, to the end, within a container, and a next-chunk mode. This makes the tool's purpose concrete and clearly distinguishes it from general navigation siblings like navigate or wait_for.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context for when the next-chunk mode is useful ('a long page is walked without re-reading it') and for virtualized-container handling. However, it does not explicitly state when to choose this tool over siblings such as get_page_view, navigate, or find_elements, nor does it provide exclusions or alternative routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
wait_forWait ForCDestructive
Wait for text to appear or disappear, an element to reach a state, a
URL to match, or a JS predicate to hold. EVERY condition is checked
against the current state first and returns immediately when it already
holds, so a wait issued after the thing already happened costs nothing
instead of timing out (the field's URL wait expired on a navigation
that had finished before the call). A url value without wildcards
matches as a substring; use * and ? for globbing. Real timeouts, and a
failure that says what was awaited and what was observed instead.
| Name | Required | Description | Default |
|---|---|---|---|
| page | Yes | ||
| value | No | ||
| location | No | ||
| condition | Yes | ||
| timeout_ms | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description portrays a passive polling/observational operation that checks state and returns, but the annotations mark destructiveHint=true and readOnlyHint=false, implying possible mutation or destruction. The description discloses no destructive behavior, so it contradicts the annotations. Per the rubric, this is an annotation contradiction.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-structured and front-loaded. The parenthetical example adds concrete value rather than filler, and the URL globbing note is useful. It is slightly dense but earns its length.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and 0% schema coverage, the description should fully explain how to invoke the tool. It covers common wait behaviors but omits the meaning of `page` and `location`, exact condition syntax, and success return behavior. The annotation contradiction also makes expected side effects unclear.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description carries the full burden. It explains `condition` categoriescars, clarifies `value` substring/glob behavior, and mentions real timeouts, but leaves `page` and `location` undefined and does not enumerate accepted `condition` string values. This is insufficient for a 5-parameter tool.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool waits for text appearance/disappearance, element states, URL matches, or JS predicates. It is specific about the resource and the kinds of conditions, though it does not explicitly distinguish itself from the sibling tool 'monitor'.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description directly states when to use the tool: whenever a wait for a condition is needed. It also adds useful context about immediate return when the condition already holds訪. It does not mention alternatives or exclusions, so it stops short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
1 tool update
v1.0.1- Added
get_server_info
11 tool updates
v1.0.0- First observed
find_elements - First observed
get_audit - First observed
get_page_view - First observed
get_text - First observed
get_workflows - First observed
manage_session - First observed
manage_tabs - First observed
monitor - First observed
navigate - First observed
scroll - First observed
wait_for
TDQS
Scored across 12 tools
Most tools have clearly distinct jobs, especially session/tab management versus navigation versus monitoring. The only mild overlap is among get_page_view, get_text, and find_elements, but their descriptions specify different output contracts: orientation, prose, and targeted refs.
Names are readable and semantically grouped, but conventions are mixed: get_* for reads, manage_* for resources, bare verbs like navigate/scroll/wait_for for actions, and find_elements breaks the order. There is no single consistent verb_noun pattern across the set.
Twelve tools is well within the ideal range for a web automation/browser server. Each tool covers a distinct aspect of the lifecycle—session, tabs, navigation, reading, searching, waiting, monitoring, and meta-inspection—without obvious redundancy.
The set is strong on navigation, reading, monitoring, and session management, but there is no click, type, submit, or select tool. get_page_view and find_elements repeatedly return 'refs you can act on,' yet no tool exists to act on them, creating a significant dead end for real web interaction.
Maintenance
Related MCP Connectors
Reliable web access for AI agents: smart HTTP, rotating proxies, and full-browser rendering.
Give agents eyes on any web page: structured context, and changes explained in plain language.
A real browser for your agent: render any page, or 25 pages of a site, to clean text.
E2LLM gives your AI eyes and hands in a real browser: structured perception (SiFR) plus action.
Related MCP Servers
- AlicenseBqualityBmaintenanceEnables LLMs to interact with web pages, take screenshots, and execute JavaScript in a real browser environment10204 npm299MIT
- FlicenseAqualityDmaintenanceEnables LLM agents to browse and interact with web pages using a stealth Chromium browser, with an accessibility-tree interface for low token usage.12-
- AlicenseNot gradedqualityCmaintenanceEnables language models to control a real visible browser with a full set of 24 tools for navigation, clicking, form filling, screenshots, and reading page structure, while preserving login states and supporting configurable browser policies.171 npmMIT
- AlicenseBqualityBmaintenanceEnables LLM agents to automate web browsers with token-efficient perception, stealth anti-detection, captcha solving, network interception, and isolated persistent sessions.51MIT