Skip to main content
Glama
KitchenSink4AI

KitchenSink4Web

Official

🚰 KitchenSink4Web Community Edition

Tests PyPI License: AGPL-3.0

Landing page · llms.txt (machine-readable capability manifest for agents and LLM crawlers)

Read websites and extract data with your AI assistant, read-only by default, with clicking and typing when you allow it.

Read websites and pull out the data you need from Claude Code, Codex CLI, Copilot CLI or any other MCP client that runs local tools. KitchenSink4Web drives a browser on your computer, returns a page within a size budget and says what it left unread, and exports tables to CSV or JSON. It starts read-only; clicking, typing and form filling are switched on only when you choose. Websites receive ordinary browsing requests; what reaches your AI provider is decided by your AI app. The Community edition is free under the AGPL. The Business edition adds a Windows installer, a signed update channel, a license your company can approve and support.

Works on: Windows, macOS and Linux for the browser tools. Image-text recognition uses Windows OCR. A browser download and some optional dependencies may be needed.

Install

Pick the route for your AI app. The commands go in PowerShell on Windows or a terminal on macOS and Linux, not into an AI chat. The package routes need Python 3.12 or newer.

Claude Desktop

Install uv, then quit and reopen Claude Desktop. Download the .mcpb file from KitchenSink4Web releases. In Claude Desktop open Settings, then Extensions, then Advanced settings, then Install extension, and choose the file. The bundle fetches the Python package the first time it starts, so the first launch needs a network connection. Restart your session and check that the tools show as connected.

Claude Code or Codex CLI

Install uv, then run the line for your app and restart your session:

claude mcp add web -s user -- uvx kitchensink4web
codex mcp add web -- uvx kitchensink4web

Any other local MCP client

Use uvx as the command and kitchensink4web as its argument, or install the package and use kitchensink4web as the server command:

pip install kitchensink4web

Then follow your client's guide for adding a local MCP server. Installing the package on its own does not connect it to an AI app.

Business edition

Compare the editions on the pricing page. Already purchased? Your Windows installer and download link are in your license portal.

Related MCP server: antibrowser-mcp

What it can do

53 tools with every pack and browser actions enabled. The default read-only lite launch exposes 12; lite with actions has 20.

  • Read a page within a chosen size budget and see what was left unread.

  • Find one section or element without fetching the whole page again.

  • Extract tables, lists and article text, then export CSV or JSON.

  • Take screenshots or save a page as PDF.

  • Read only what changed on a page you already read.

  • Watch a page for changes while the server runs.

  • Turn on clicking, typing and form filling when a task needs them.

  • Record a browser task and replay it with checks, with the workflow pack.

  • Run automated accessibility checks with the optional dependency.

What is available depends on the packs you enable and the applications installed. The full tool reference is below.

Business edition

Need a license your company can approve and a setup someone supports? The Business edition pairs these tools with a Windows installer, a signed update channel and support under the Business terms. Update checks tell you when a covered release is available; nothing installs on its own. Compare the options on the pricing page. The Community edition stays free under the AGPL, including business use that meets its terms.

Privacy Policy

The tools run on your computer, and KitchenSink4AI receives no documents and no usage data from them. Your AI app may send prompts, file contents and tool results to its own provider under that app's settings and terms. Installing downloads packages, and the Community version check contacts PyPI unless you disable it; those requests carry connection details such as your network address and never your documents. The browser connects to the websites you ask it to use, and those sites see ordinary browsing requests. A missing tokenizer table may be downloaded once; the text being measured is never sent. Cloud folders and backups follow their own settings. The Privacy Policy covers the product, purchases and support records.

Not affiliated with or endorsed by Google, Mozilla, Microsoft, or any website this server visits. Chrome, Chromium, Firefox, and Edge are trademarks of their respective owners, used nominatively to name the browsers this server can drive.

The cheap first read

get_page_view returns the page's shape: what it is, what is on it section by section, what each section costs to open, and what the sensible next call would be. You set the budget and the read never exceeds it, so the question is never whether a page fits, only how much detail you bought. Every read ends with a completeness block naming what was left out and what it would cost to get it back. Nothing is swept under the rug.

From there the pattern is cheap and gets cheaper. Expand one region instead of raising the budget. Chain delta reads: pass the token a read returns and the next read prices only what changed. Elements come back as short-lived refs your assistant can act on directly, and searches by text and role survive re-renders that break refs.

Open shadow roots are read and their contents are actable like anything else, which is what makes component-built sites (Reddit, MDN, most of the modern web) readable at all. Closed roots cannot be reached by any tool; they are counted and reported rather than silently skipped. Same-origin iframes are read and searched like the page they sit in, each labeled with its own origin; a cross-origin frame is counted and named, never entered.

The packs

The server starts lite: reading, navigation, and session management, 20 tools. Capability packs are chosen at launch and are fixed for the whole session, which means an absent pack is provably absent, not merely switched off:

Pack

What you get

extract

Tables, lists, links, and article text, with CSV and JSON export.

capture

Screenshots (passwords blurred), PDF export, and saving pages as files.

network

See the requests a page makes behind the scenes. Useful for debugging websites; most people leave it off.

storage

Use with care. Saves a signed-in session so Claude does not have to log in every time. It saves the session, never your password; treat the saved file like one anyway.

files

Download files from pages into one folder, and upload files into page forms. Every one asks you first.

diagnostics

Lets Claude run JavaScript on a page, one script at a time, each one shown to you for approval first. This is the most powerful and most dangerous setting on this page. Leave it off unless you know you need it.

workflows

Record a multi-step task once and replay it later. Every replay checks that the page still matches before anything runs.

accessibility

Checks a page against the WCAG accessibility rules using axe-core, groups what it finds by rule, and says plainly what automated testing cannot check. Needs an optional extra installed; the tool tells you how if it is missing.

Browsers and lanes

The default lane drives the server's own bundled Chromium. It never touches your browser or your profile; sessions open in a fresh profile that is deleted on close unless you explicitly save signed-in state to a file.

Lane B drives a browser you already have installed (Chrome, Edge, or Firefox), still with its own separate profile, never yours. Installed Firefox is the lane to reach for on research-heavy work: sites that turn away automated Chromium routinely serve Firefox normally, and both Firefox lanes pass bot checks that block headless Chromium. That is lane steering, not evasion; the server does not disguise what it is, it just lets you use a browser the site treats better.

Signed-in sessions are saved and reloaded with save_auth_state and auth_state= on open. The state file records when its session cookies expire, and loading a stale file says so plainly instead of letting the login fail mysteriously.

manage_session(action='status') reports which browsers are installed and which lane it would recommend for the page you are on. It never switches lanes for you.

The safety model

A browser is where your logins, your money, and your mistakes all live, so the safety layer is the product, not a feature of it.

Read-only unless you say otherwise. Out of the box the tools that click, type, and submit do not exist. They are not disabled, they are absent from the tool list, and no instruction on any webpage can talk your assistant into using a tool that is not there. One switch at install (allow acting in the bundle, KS4WEB_ALLOW_ACTING=true elsewhere) brings them in.

Credential blindness. The server never reads what you type into password fields, and secrets observed in cookies and site storage are vaulted: they cannot leak into your assistant's conversation, because the text that would carry them is redacted before it leaves the server.

Hidden text is evidence, not instructions. Text a page hides from human eyes is the oldest trick for slipping commands to an AI. This server strips it from the reading channel, counts it, reports the count, and wraps everything a page says in labeled provenance markers, so your assistant always knows which words came from a stranger.

Confirmation gates. Form submissions, payments, and page scripts stop and ask you first, through the client's own confirmation prompt. As of 2026-09 that prompt displays in Claude Desktop and Claude Code; the claude.ai web client does not display it yet, so gated actions refuse there instead of proceeding unconfirmed.

Honest refusals. When something cannot be done, from a login wall to a bot check to a page that crashed its own renderer, the answer names what stood in the way and what to try next. A refusal that explains itself is cheaper than a retry loop.

Two different things ask you for permission

Your MCP client's own permission prompt names a tool (something like mcp__web__type_text) and comes from the client, not from KS4Web; approving a read-only tool group there is safe and stops most of the asking. A KS4Web gate names an action in plain words ("submitting this form sends a password") and cannot be pre-approved away for payments, credentials, posts, deletions, or legal assent. If a prompt names a tool, it is the client; if it names what is about to happen in the world, it is KS4Web.

Out of the box the consent scope is research: reading and navigation are free, query-shaped submissions (a search box) proceed, and everything else asks first. The full scope lets ordinary form submissions proceed without asking; the irreducible set asks under every scope, always. The install screen's "Submit routine forms without asking each time" checkbox is the whole choice.

The GET rule is the one place page-authored text classifies down rather than up: a form that looks like a search is allowed to proceed as one, and what a mistaken call in that direction buys is an unprompted ordinary submission, never a payment, credential, post, or deletion; those always ask. In the other direction, a submit button that says nothing recognizable in any shipped language passes ungated under the full scope; if that risk matters to you, stay on research.

Client compatibility

If a feature refuses with CONFIRMATION_REQUIRED on claude.ai web, it works on Desktop and Code.

Client

What works

Claude Desktop

Everything, with confirmations rendered.

Claude Code

Everything, with confirmations rendered.

claude.ai web

All reading, navigation, and search. Gated actions refuse honestly because no confirmation can render; pre-authorized classes work if configured at launch.

Two numbers that matter

The same page, read two ways. A raw dump of the Treaty of Versailles article on Wikipedia costs 33,073 tokens of your assistant's memory. This server's first read of it costs 4,445. The difference is not compression, it is a different product: a map of the page with a price on every region, instead of the whole page whether you wanted it or not.

Numbers on this page are measured by scripts in tools/, not written by hand. Re-run them yourself; the README is regenerated from their output.

.venv/Scripts/python.exe -X utf8 tools/measure_readme_numbers.py

Context cost (measured)

Surface

Tokens

Lite tool surface

6.9k

Full surface (all packs)

17.9k

First read of the Treaty of Versailles article on Wikipedia

4,445

Delta read after one click

82

The budget ladder has 16 rungs; every read reports which rung it printed at and what a deeper read would cost. Inside a subagent, start at budget_tokens=2500; tool results are capped more tightly there.

Update check (opt-out)

The server compares its version against PyPI's when you call manage_session(action='status'), and never at any other time: nothing runs at startup and nothing runs on a background thread. At most one request per seven days, with a two-second timeout. It never installs anything. The request is a plain HTTPS GET to pypi.org for this package's public release index; it sends no identifier, no usage data, no page content and no session state. A check that could not complete says so and says how long ago the last successful one was, rather than showing nothing. KS4WEB_UPDATE_CHECK=off switches it off, and the status report then says it is off; the older KS4WEB_NO_UPDATE_CHECK=1 spelling is still honored.

Testing

2,066 tests, of which 733 drive a real browser. Beyond the suite, every release passes gate batteries that re-run the adversarial findings of four attack rounds: prompt injection through page content, hidden-text smuggling, credential-theft attempts, gate bypasses, workflow replay tampering, and resource abuse. The gates are not aspirational; each one exists because an adversarial round or a field tester actually broke something, and the fix is pinned by the test that would catch its return.

The current tuning came out of a live field campaign: two sessions, twelve real sites, three browser engines, and 141KB of logged friction, retests, and design notes. The findings, and what shipped in response, are in the repository history rather than a marketing page.

Maturity: what the version number does and does not claim

This is the newest member of the family, and the version number says so honestly. It grew up under the house rules: every claim measured, every release attacked before it ships, every refusal honest about its reason. What it has not had yet is a long life in strangers' browsers. The safety net while it earns one: it starts read-only, it prices every read before taking it, and it tells you what it could not see. Bring it your strangest pages and file what breaks.

Known limits

  • Cross-origin iframes are never entered; the completeness block names each frame and says why. Same-origin iframes are read, searched, and acted in, each labeled with its own origin.

  • Closed shadow roots cannot be reached by any tool. They are counted so you can tell a component-heavy page from an empty one.

  • Confirmation-gated actions (auth loading, page scripts, submits) refuse on the claude.ai web client until it renders MCP confirmation prompts; Claude Desktop and Claude Code work today.

  • Bot walls and CAPTCHAs are reported, never solved or evaded. The Firefox lanes get through device checks that block headless Chromium; a wall that blocks everything is a wall.

  • sendBeacon and ping requests bypass network routing rules; that is a Playwright limit, documented rather than hidden.

Automated accessibility checking finds a minority of real-world issues. Deque, who build the axe-core engine this tool runs, publish 57% as their own figure for their own tooling. A clean report here means the automated checks passed, not that the page is accessible. There is no score, because every 0-to-100 accessibility number is somebody's weighting rather than a measurement.

Nothing taps you on the shoulder. No MCP client in use today delivers a server-initiated message into a conversation, so monitors run only while KS4Web runs, and the report is how you ask what happened; after a restart, the report names the window that went unchecked. A monitor watches one URL for one deterministic change and does not understand what changed: content_hash is noisy on pages with clocks or counters, only the main document is watched, and a change that appeared and reverted between two checks is invisible. A monitor cannot get past a login or a bot wall, and one that keeps failing pauses itself rather than hammering the site. Five minutes minimum between checks; everything stays on this machine.

A session handle transfer moves a live session between conversations on the same machine, inside the same running KS4Web. Nothing is copied: it is the same browser, so a page the other conversation navigated has lost your refs, and the report says which. It does not survive a server restart, and it does not cross between Claude Desktop's chat and Code panels, which run separate servers. The session dies with KS4Web on purpose; that is what keeps orphaned browsers off your machine. A token works once, expires in an hour, and is not a lock: any conversation on this server can already see the session. The budget travels with the session.

Two identities cost two browser processes and two profile directories, which is real memory. The action budget is shared across them on purpose, so opening a second identity does not double what a session may do to the world, and two contexts visiting the same site cost one origin against the origin budget. Auth state saves and loads per identity. Closing one identity leaves the others working; closing the last one tells you to close the session instead. Every context shares the session's lane and emulation; two lanes means two sessions.

The lane database is a list of which websites this computer has visited with which browser: hostnames and dates, no pages, no addresses, nothing typed. It exists so a site that refused one browser can be read with one that works. It is stored on this machine, it never leaves unless you export it, learning can be turned off with one switch, and one call erases it entirely.

Install options and settings

The install routes are at the top of this file. This section covers what you choose when the server starts.

The Claude Desktop bundle puts the whole configuration on its install screen as checkboxes: one to allow clicking and typing (off means read-only browsing, the shipped default), and one per capability pack. The server drives its own bundled Chromium, never your browser or your profile.

Packs and read-only mode are chosen at launch (--packs, KS4WEB_MODE, KS4WEB_ALLOW_ACTING) and are identical for every connection to the process.

Tip for Claude Desktop: in Tool permissions, set this server's Read-only tools group to Always Allow. Those tools cannot change anything on any page, and it stops most permission prompts.

License

KitchenSink4Web is dual-licensed:

AGPL-3.0 (open source). Free for anyone (individuals, academics, and businesses) for any use that complies with the AGPL's terms. Those terms include sharing source, including your modifications, when you distribute the software or make it available over a network.

Commercial license. For organizations that want to build KitchenSink4Web into their own products or services without the AGPL's source-sharing obligations. Contact licensing@kitchensink4.ai.

Copyright (c) 2026 Alvut Consulting, LLC. KitchenSink4AI is a product line of Alvut Consulting, LLC.


Not affiliated with or endorsed by Google, Mozilla, Microsoft, or any website this server visits. Chrome, Chromium, Firefox, and Edge are trademarks of their respective owners, used nominatively to name the browsers this server can drive.

Available Tools

12 tools
find_elementsFind ElementsA
Read-only

Find elements by text, role plus accessible name, natural-language description, CSS, or XPath, and get back refs you can act on plus a note on what was not searched. role='button' narrows any query to one element role (field finding: 'Comment' alone matched 12; with the role filter it matches the one button). This is the cheap targeted follow-up that pairs with get_page_view: the page view tells you what string to look for, and this retrieves it for a fraction of a full read. Ambiguous results are listed rather than resolved, and zero results come back with the nearest misses so a miss is a one-turn recovery. The search covers the main document and every open shadow root in it, and the matches it returns from a shadow root are actable like any Same-origin iframes are searched too and the result says which ones it entered. Two things stay out and the result counts both: cross-origin iframes, which no tool here opens, and closed shadow roots, which no tool can reach. XPath is the one kind that does not enter a shadow root. location={'region': 'r7'} (or a ref, form, or table from a read) narrows the search to that subtree, components inside it included, and the first result line names the scope that was searched.

ParametersJSON Schema
NameRequiredDescriptionDefault
kindNoauto
pageYes
roleNo
limitNo
queryYes
locationNo

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With readOnlyHint=true and openWorldHint=true, the annotations cover safety, but the description adds substantial behavioral detail: ambiguous results are listed rather than resolved, zero results return nearest misses, searches cover shadow roots and same-origin iframes, and XPath does not enter shadow roots. This goes far beyond the annotations and helps the agent predict edge-case behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long but dense and front-loaded: the core purpose appears in the first sentence, followed by the role-filter example, pairing guidance, ambiguity behavior, scope limitations, and location narrowing. Each sentence earns its place, though the paragraph-style formatting and minor punctuation issues slightly hurt scannability.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a six-parameter tool with no output schema, the description gives a strong picture of inputs, scope, limitations, and result behavior (refs, notes on what was not searched, nearest misses, scope named in the first result line). Missing details like what 'kind' does and how pagination/limit behaves keep it from being fully complete, but the essential invocation context is present.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 0% description coverage, so the description carries the full burden for parameter explanation. It does explain query, role, and location well (including an example for role and a location example), but it never mentions the required 'page' parameter, nor does it explain 'kind' or 'limit' beyond the default. This is a meaningful gap for a tool with two required parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Find elements by text, role plus accessible name, natural-language description, CSS, or XPath.' It clearly distinguishes this tool from get_page_view by calling it a 'cheap targeted follow-up,' so an agent knows it is a retrieval/search action rather than a comprehensive read.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly positions this tool relative to get_page_view: 'the page view tells you what string to look for, and this retrieves it for a fraction of a full read.' It also specifies scope exclusions (cross-origin iframes, closed shadow roots) and when XPath behaves differently, giving the agent concrete guidance on when this tool is appropriate.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_auditRead the Action LogA
Read-only

Read the action log. Returns one record per call, paginated, carrying the timestamp, lane, page, URL, resolved target with its human label, redacted arguments, outcome, any rebind, any confirmation decision, and the budget counters at that moment. This is an operational record for the user, so you can always know exactly what was done even where a web action cannot be undone. It is not forensic and not evidence.

ParametersJSON Schema
NameRequiredDescriptionDefault
toolNo
limitNo
sessionNo
start_indexNo

TDQS

A3.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, but the description adds significant behavioral context: pagination, specific fields returned, and a clear caveat that the log is 'not forensic and not evidence.' It also explains the purpose of the log for undo scenarios, which goes beyond the annotation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well-organized: it starts with the purpose, then lists the returned fields, then gives the operational rationale, and ends with a caveat. Every sentence adds value without redundancy, and it is appropriately sized for the complexity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Since there is no output schema, the description does explain the return content in detail, which is helpful. However, it omits any explanation of the input parameters, leaving agents to guess at filtering and pagination mechanics. For a tool with 4 parameters and no schema descriptions, this incompleteness is notable.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 0% schema description coverage, the description was expected to compensate for parameter meaning, but it does not mention any of the four parameters (tool, limit, session, start_index). It only implies pagination without explaining how the limit/start_index work. This is a critical gap for correct invocation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb and resource: 'Read the action log'. It elaborates on the content of each record, making it distinct from sibling tools like get_page_view or get_workflows. The purpose is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides contextual guidance ('operational record for the user, so you can always know exactly what was done') but does not explicitly compare to alternatives or state when not to use it. The use case is implied rather than explicitly routed against siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_page_viewGet Page ViewA
Read-only

Read a page as an ORIENTATION, not a transcript, under a token budget it never exceeds whatever the page size. Returns identity, landmark regions each priced with the cost to expand it, the interactive surface with refs you can act on, a digest or app skeleton, form and table inventories, an account of what was NOT read and why, and the next call for anything unexpanded. location scopes to one region ref, budget_tokens=2500 suits a subagent, mode='links' includes in-prose links at their real cost. since=<read_token> is the cheap repeat read: only what changed, refs kept, a few hundred tokens instead of a fresh read, and it falls back to a full read when the page navigated in between and nothing survives to diff. Open shadow roots are read and their contents get refs you can act on; closed roots cannot be reached by any tool and are counted at creation, so the completeness block reports both numbers rather than one confident zero. Same-origin iframes are entered and read, and their contents get refs naming the frame they came from; a cross-origin frame is never entered, because its document belongs to an origin the page itself cannot read either, and the completeness block counts every frame it did not open.

ParametersJSON Schema
NameRequiredDescriptionDefault
modeNoauto
pageYes
viewNoauto
sinceNo
detailNostandard
locationNo
budget_tokensNo

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations provide readOnlyHint=true and openWorldHint=true, and the description reinforces both consistently while adding substantial context beyond them. It discloses the token-budget guarantee, what the completeness block reports for closed shadow roots and cross-origin frames (dual counts rather than a single zero), and the exact fallback behavior of the `since` read. None of this is derivable from the annotations, and it correctly explains the open-world behavior of reporting what was NOT read and why.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Despite its length, every sentence earns its place: the core purpose is front-loaded, followed by return contents, then parameter guidance, then edge-case behavior. There is zero filler or tautology. The density is justified by the tool's genuine complexity, and the structure flows logically from what it returns to how to control it to what its limits are.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description correctly explains the return payload (identity, landmark regions with expansion costs, interactive surface refs, digest/app skeleton, form/table inventories, unread account, and next call). It covers the complex edge cases of shadow roots and iframes thoroughly. The only incompleteness is the unexplained `view` and `detail` parameters, which leaves an agent guessing at part of the input surface. This is a small but real gap in an otherwise comprehensive definition.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description carries the full parameter-documentation burden. It explains `location` (scopes to one region ref), `budget_tokens` (with a concrete subagent value), `mode` (links at real cost), and `since` (cheap repeat read). However, three of the seven parameters—the required `page`, plus `view` and `detail`—receive no explanation. While `page` is fairly self-evident from the name, `view` and `detail` are not obvious and their interaction with the orientation read is undocumented. This is a genuine gap given the zero schema coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a precise, specific purpose: 'Read a page as an ORIENTATION, not a transcript, under a token budget it never exceeds whatever the page size.' The verb-resource pair is unmistakable, and the 'not a transcript' contrast actively differentiates it from sibling tools like get_text. An agent immediately knows this is the high-level structural read, not a raw text extraction.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description embeds strong usage direction: `budget_tokens=2500` suits a subagent, `mode='links'` includes in-prose links, and `since=<read_token>` is the cheap repeat-read strategy with explicit fallback semantics. The 'not a transcript' phrasing plus the open/closed shadow root and same-origin/cross-origin iframe rules tell an agent exactly when this tool is appropriate versus alternatives. It even specifies when the since-read fails over to a full read, which is precisely the kind of decision guidance that matters.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_server_infoGet Server InfoA
Read-only

Report what this server is and what it is running: the product and PyPI package names, the version a client was handed at initialize, the homepage and repository, the sibling servers in the same suite, the tool surface registered in THIS process (how many tools, how many of them are the lite core, which capability packs are loaded and which exist, and the read-only grade and consent scope in force), the test figures measured for this release, and the host Python and platform. Counts come from the live registry, so they describe the process you are connected to rather than the product in general. Needs no browser, opens no page, and reveals no path or user name, so a bug report can carry the output as it stands.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already establish readOnlyHint and openWorldHint, but the description adds meaningful behavior beyond that: it requires no browser, opens no page, leaks no path or username, and reads live registry state rather than reporting product-level generalities. These are exactly the kind of side-effect and privacy disclosures an agent needs before invoking the tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three dense sentences, each earning its place: the first enumerates the full report scope, the second clarifies data freshness and provenance, and the third states safety/privacy implications. Although the first sentence is long, it is information-dense and well structured rather than padded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no parameters, no output schema, and only safety annotations, the description carries full responsibility for making the tool callable. It enumerates all return categories, clarifies that counts describe the live connected process, and reassures about side-effect and privacy boundaries. Nothing essential is missing for an agent to invoke and interpret the result.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters and schema coverage is 100%, so there is nothing for the description to add about parameter meaning. The provided baseline of 4 applies, and the description instead spends its effort on return-content and behavioral detail, which is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Report what this server is and what it is running.' It then enumerates a comprehensive, unique scope—product and package names, version, repo, tool surface, test figures, Python, platform—that clearly separates it from browser-action and workflow siblings. Even without naming other tools, the content makes the tool's distinct purpose unmistakable.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear usage context: it is safe for bug reports because it 'needs no browser, opens no page, and reveals no path or user name.' It also clarifies that counts are live from the connected process, preventing misinterpretation. However, it does not explicitly name alternative tools or state when not to use it, stopping short of full when/when-not guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_textGet TextA
Read-only

Extract readable prose from a page or one region of it, paginated by start_index so a long article is read in bounded pieces rather than one unbounded dump. Text arrives as labeled data with its origin stated, and hidden regions are stripped and counted rather than silently dropped or silently included. Hidden content IS retrievable, deliberately: include_hidden=true returns it in a separately labeled section with the hiding technique named per block. There is no silent middle tier, because display:none is a real injection channel; the labeled route is the whole design. Prose inside open shadow roots is read, the same as get_page_view reads it; closed roots are counted and stay unreadable. Prose inside same-origin iframes is read after the main document, each frame under a header naming it and its origin, because one page can now deliver text from several documents and a single origin in the label would be a claim about only one of them.

ParametersJSON Schema
NameRequiredDescriptionDefault
pageYes
locationNo
max_charsNo
start_indexNo
include_hiddenNo

TDQS

A4.3/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare readOnlyHint and openWorldHint. The description goes far beyond that: it explains pagination mechanics, how hidden content is handled (stripped, counted, and retrievable via include_hidden with per-block technique naming), the behavior for open vs closed shadow roots, and iframe handling with origin labels. It discloses the absence of a 'silent middle tier' and explains the design rationale. This is exceptionally transparent about side effects and edge cases.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single dense paragraph, but every sentence adds meaningful behavioral information. It leads with the core purpose, then covers pagination, hidden content, shadow roots, and iframes. There is no filler or repetition, though it could be tightened by moving some design rationale to a separate note. It is front-loaded and structured logically.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with 5 parameters and no output schema, the description is remarkably thorough: it explains the return format (labeled data with origin), pagination, hidden content, shadow roots, and iframes. The main gap is the 'location' parameter, which is only vaguely referenced as 'one region of it' without defining its structure. Also, the exact interaction of max_chars with pagination isn't spelled out, but the general behavior is clear. Overall it covers most of what an agent needs.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema has zero description coverage for its 5 parameters. The description compensates for start_index (paginated) and include_hidden (returns hidden content in labeled section), and implicitly covers max_chars via 'bounded pieces'. However, the 'location' parameter is never explained beyond 'one region of it', leaving its structure and semantics undocumented. With 0% schema coverage, the description should have covered all parameters; it covers three partially and one not at all.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Extract readable prose from a page or one region of it'. It immediately distinguishes itself from siblings by mentioning pagination via start_index and explicitly comparing to get_page_view for shadow roots. This is far more specific than a generic 'get text' and leaves no doubt about what the tool does.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear context for when to use the tool (for bounded, paginated prose extraction, including hidden content) and even references get_page_view to clarify shadow-root behavior. However, it never explicitly states 'use get_text when you need X, use get_page_view when you need Y' or lists exclusions. The guidance is implicit through behavioral description rather than an explicit decision rule.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_workflowsGet WorkflowsA
Read-only

Get recipes for this server: the cheap-read-then-act pattern, the auth workflow (headed handoff plus saved state), reading strategy, budgeting, troubleshooting a page that will not read, the subagent budget setting, lanes, what each capability pack contains with the exact launch flag that loads it, what a site profile is and where it lives, and how to record and replay a multi-step flow. Packs are chosen at launch rather than at runtime, so this is where you learn which flag you need before restarting. Tool availability reflects the packs this server was started with.

ParametersJSON Schema
NameRequiredDescriptionDefault
topicNo

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The readOnlyHint annotation already covers the safe read-only nature, and the description adds meaningful behavioral context: tool availability reflects the server's launch-time packs alarm, and flags must be discovered here before restart. It does not describe the response format or invalid-topic behavior, but the annotation burden is lower here.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The core purpose is front-loaded and the topic list is compact. The final two sentences about launch-time pack selection are useful but slightly redundant with each other, which costs a point.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only documentation tool, the description explains what it returns, why it matters, and when to call it. It lacks explicit return-format details and precise parameter behavior, but the optional-topic schema and readOnlyHint annotation make those gaps non-critical.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The only parameter, topic, has 0% schema coverage, so the description must compensate. The enumeration of concrete workflow topics gives the agent plausible values, but the description never explicitly maps these to the topic parameter or explains what happens when it is omitted or null. It partially compensates without fully documenting the parameter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a clear verb-resource pair ('Get recipes for this server') and enumerates the specific workflow topics returned, so an agent can tell this is a documentation/how-to tool. It does not explicitly contrast itself with sibling tools like get_server_info, but the listed content is specific enough to prevent serious confusion.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear when-to-use context: packs are chosen at launch, so this tool is where an agent learns which launch flag is needed before restarting. It does not explicitly name alternative tools or state when not to use it, so it stops short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

manage_sessionManage SessionB
Destructive

One session is one browser. open starts one on a chosen lane (the first navigate of a conversation can open one for you, and says so when it does); close ends the browser; with auth state saved, the login survives; the session never does. status reads liveness first, so a dead browser says so, then lists every session this server holds with its pages, budgets, and what is shared: sessions belong to the server process, not to a conversation. capabilities, budget, and reset_budgets report and manage limits. handoff hands the window to the human for a login, an MFA prompt, or a bot wall. export_handle and import_handle move a live session between conversations on this machine: see get_workflows(task='session-transfer'). lanes is the local lane database (status, export, import, erase). profiles reloads site profiles from disk. On open, auth_state loads a saved login, contexts=2 gives one session two independent cookie jars, and device, viewport, locale, and timezone set what pages in this session believe about their environment.

ParametersJSON Schema
NameRequiredDescriptionDefault
opNo
laneNo
noteNo
pathNo
siteNo
tokenNo
actionNostatus
deviceNo
localeNo
reasonNo
contextNo
sessionNo
contextsNo
timezoneNo
viewportNo
auth_stateNo
expires_minutesNo

TDQS

B3.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations (destructiveHint=true, readOnlyHint=false) already flag mutation, and the description adds substantial value on top: close kills the session but saved auth survives, status reads liveness first so dead browsers are reported, and sessions belong to the server process, not the conversation. This ownership and persistence disclosure is exactly the kind of behavior an agent needs and that structured fields cannot convey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Every sentence is information-dense with no filler, and the core concept is front-loaded in the opening clause. However, the structure is a single unbroken prose paragraph mixing operation semantics, parameter effects, and cross-references; a dispatcher of this complexity would be far more scannable as bullets or an op-to-semantics table.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with 17 parameters, 11 sub-operations, no output schema, and no enums, the description provides substantial orientation: it discloses status output shape, session lifetime, liveness checking, and shared-state semantics. Yet it leaves major gaps — the op/action dispatch relationship, output of most operations, and the meaning of expires_minutes, note, reason, path, site, and token — so it is only partially complete for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description carries the full burden for 17 parameters, but it only explains a subset: auth_state, contexts, device, viewport, locale, timezone, and lane. Critical dispatcher parameters — op, action, session, note, reason, path, site, token, context, expires_minutes — are never mapped to their effects, leaving an agent to guess at the mechanics of calling this tool.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description anchors on a strong core metaphor ('One session is one browser') and enumerates every sub-operation (open, close, status, handoff, lanes, profiles, etc.), which clearly distinguishes it from siblings like get_page_view, navigate, and monitor. However, 'manage session' is a generic dispatcher verb, and the purpose must be excavated from a dense wall of text rather than stated as a crisp mission.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is real usage context: it tells the agent that the first navigate may auto-open a session, that handoff is for login/MFA/bot walls, and explicitly points to get_workflows(task='session-transfer') for session transfer. But it never states when NOT to use this tool versus siblings such as monitor or manage_tabs, and per-operation selection guidance is minimal for a dispatcher with 11 operations.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

manage_tabsManage TabsA
Destructive

List, open, select, or close tabs, report which is focused, and capture pages a click opened in a popup. Mints and returns the explicit page handles every other tool accepts, which is how browser state survives across calls without relying on protocol sessions. Closing a page invalidates its refs and its delta read tokens, and the result says so rather than leaving a later failure to explain it. On a session with more than one cookie jar, context says which jar a new tab opens in and every listed page says which jar it belongs to.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlNo
pageNo
actionNolist
contextNo
sessionYes

TDQS

A4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The annotations already signal mutation and destructiveness (readOnlyHint=false, destructiveHint=true), and the description adds valuable behavioral detail: closing a page 'invalidates its refs and its delta read tokens,' and the result announces this instead of leaving a later failure. It also reveals the cookie-jar context behavior, going beyond the structured annotation hints.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but each sentence earns its place: the action list, the handle mechanism, the invalidation consequence, and the cookie-jar context. It is front-loaded with what the tool does, though the later sentences are somewhat intricate for the available parameter detail.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

This is a multi-operation, destructive tool with five parameters, no output schema, and no enum constraints, yet the description leaves the parameter model largely implicit. It does not explain the allowed `action` values or how to request 'report which is focused' or 'capture pages a click opened in a popup,' so an agent cannot reliably construct a correct call from the definition alone.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, but it only meaningfully explains `context`. The description implies action-related behaviors and page/session concepts, but it never maps parameters like `action`, `url`, `page`, or `session` to concrete values or usage, leaving the agent to guess how to invoke each operation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb list and resource: 'List, open, select, or close tabs, report which is focused, and capture pages a click opened in a popup.' It also states the tool's unique role by explaining it mints 'explicit page handles every other tool accepts,' which distinguishes it from the sibling tools that consume those handles.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It clearly conveys when to use the tool: to manage tabs and, importantly, to obtain the page handles that other tools depend on for state across calls. It does not explicitly name alternatives or give when-not-to-use guidance, so it stops short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

monitorWatch a URL for ChangeA
Destructive

Watches one URL for one deterministic change while this server runs. Four conditions: content_hash, text_appears, text_gone, selector_count; the same page state always answers the same way. Nothing is pushed anywhere: report is how you ask what happened, and check_now forces a check. A monitor that could not check reports stale with the failure, never unchanged. Checks run in a dedicated headless session, never the caller's. There is a floor on the interval and caps on monitors and daily checks; a 429 is honored rather than retried; a monitor that keeps failing pauses itself.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlNo
labelNo
sinceNo
valueNo
actionNoreport
monitorNo
selectorNo
conditionNo
check_interval_minutesNo

TDQS

A3.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the annotations, it discloses deterministic evaluation, no push delivery, stale-on-failure reporting, dedicated headless session, interval floors and quotas, honoring 429s without retry, and auto-pause on repeated failure. This is exemplary behavioral disclosure and does not contradict readOnlyHint/destructiveHint.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The text is dense but every sentence adds a distinct constraint or behavior, and the most important information about scope is front-loaded. No filler or redundant restatement of the schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The behavioral context is unusually complete for a 9-parameter tool with no output schema, and it even hints at report semantics. Yet an agent still cannot fully infer how to create/select a monitor, what `since` and `label` mean, or what the exact response shape is, leaving operational gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 0% schema coverage, the description must carry parameter meaning, and it does name the four condition values and report/check_now actions while indirectly suggesting url, selector, value, and interval. However, label, since, and monitor remain unexplained, and no parameter requirements per condition are spelled out, so it only partially compensates.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The opening line 'Watches one URL for one deterministic change while this server runs' states a specific action, resource, and scope, and the four condition names make the function concrete. It does not name a sibling for contrast, so it misses the explicit differentiation needed for a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description makes the intended use implicit: a long-running watcher for one deterministic page change, with 'report' and 'check_now' as the query/force actions. It never says when to prefer this over siblings like wait_for or get_page_view, so the agent is left to infer tool selection.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

scrollScrollA

Scroll by an amount, to a named element, to the end, or inside a specific container, plus a next-chunk mode that remembers position across calls so a long page is walked without re-reading it. Reports how much content is now reachable and how much remains below, and names virtualized containers where the DOM holds far fewer rows than the page claims, rather than presenting a partial list as complete.

ParametersJSON Schema
NameRequiredDescriptionDefault
pageYes
actionNoby
amountNo
locationNo
timeout_msNo

TDQS

A4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the annotations, the description reveals meaningful behavior: it remembers scroll position across calls in next-chunk mode, reports reachable vs. remaining content, and explicitly identifies virtualized containers instead of presenting a partial view as complete. This is valuable transparency that annotations alone do not provide.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two dense sentences with no filler. It front-loads the primary action modes and then adds behavioral reporting details, making it easy to scan while still packing relevant information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

It covers the main behaviors and high-level return information, but because there is no output schema, it lacks specifics about the exact response format. It also does not define the action enum values or the semantics of timeout_ms, leaving some gaps for an agent trying to invoke the tool precisely.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It alludes to action modes and 'location' as a named element/container, but it never explains allowed action values, the exact meaning of amount, how page is used, or timeout behavior. The description gives only loose mapping to the five parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('scroll') and enumerates distinct modes: by amount, to a named element, to the end, within a container, and a next-chunk mode. This makes the tool's purpose concrete and clearly distinguishes it from general navigation siblings like navigate or wait_for.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear context for when the next-chunk mode is useful ('a long page is walked without re-reading it') and for virtualized-container handling. However, it does not explicitly state when to choose this tool over siblings such as get_page_view, navigate, or find_elements, nor does it provide exclusions or alternative routing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

wait_forWait ForC
Destructive

Wait for text to appear or disappear, an element to reach a state, a URL to match, or a JS predicate to hold. EVERY condition is checked against the current state first and returns immediately when it already holds, so a wait issued after the thing already happened costs nothing instead of timing out (the field's URL wait expired on a navigation that had finished before the call). A url value without wildcards matches as a substring; use * and ? for globbing. Real timeouts, and a failure that says what was awaited and what was observed instead.

ParametersJSON Schema
NameRequiredDescriptionDefault
pageYes
valueNo
locationNo
conditionYes
timeout_msNo

TDQS

C2.9/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description portrays a passive polling/observational operation that checks state and returns, but the annotations mark destructiveHint=true and readOnlyHint=false, implying possible mutation or destruction. The description discloses no destructive behavior, so it contradicts the annotations. Per the rubric, this is an annotation contradiction.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well-structured and front-loaded. The parenthetical example adds concrete value rather than filler, and the URL globbing note is useful. It is slightly dense but earns its length.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema and 0% schema coverage, the description should fully explain how to invoke the tool. It covers common wait behaviors but omits the meaning of `page` and `location`, exact condition syntax, and success return behavior. The annotation contradiction also makes expected side effects unclear.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description carries the full burden. It explains `condition` categoriescars, clarifies `value` substring/glob behavior, and mentions real timeouts, but leaves `page` and `location` undefined and does not enumerate accepted `condition` string values. This is insufficient for a 5-parameter tool.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool waits for text appearance/disappearance, element states, URL matches, or JS predicates. It is specific about the resource and the kinds of conditions, though it does not explicitly distinguish itself from the sibling tool 'monitor'.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description directly states when to use the tool: whenever a wait for a condition is needed. It also adds useful context about immediate return when the condition already holds訪. It does not mention alternatives or exclusions, so it stops short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 1 tool updatev1.0.1
    • Addedget_server_info
  2. 11 tool updatesv1.0.0
    • First observedfind_elements
    • First observedget_audit
    • First observedget_page_view
    • First observedget_text
    • First observedget_workflows
    • First observedmanage_session
    • First observedmanage_tabs
    • First observedmonitor
    • First observednavigate
    • First observedscroll
    • First observedwait_for

TDQS

A3.6/5.0

Scored across 12 tools

Disambiguation4/5

Most tools have clearly distinct jobs, especially session/tab management versus navigation versus monitoring. The only mild overlap is among get_page_view, get_text, and find_elements, but their descriptions specify different output contracts: orientation, prose, and targeted refs.

Naming Consistency3/5

Names are readable and semantically grouped, but conventions are mixed: get_* for reads, manage_* for resources, bare verbs like navigate/scroll/wait_for for actions, and find_elements breaks the order. There is no single consistent verb_noun pattern across the set.

Tool Count5/5

Twelve tools is well within the ideal range for a web automation/browser server. Each tool covers a distinct aspect of the lifecycle—session, tabs, navigation, reading, searching, waiting, monitoring, and meta-inspection—without obvious redundancy.

Completeness2/5

The set is strong on navigation, reading, monitoring, and session management, but there is no click, type, submit, or select tool. get_page_view and find_elements repeatedly return 'refs you can act on,' yet no tool exists to act on them, creating a significant dead end for real web interaction.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers