Skip to main content
Glama

Server Details

Stealth web browser for agents: search, fetch, click, download and type in persistent MCP sessions.

Glama couldn't complete the latest health check. If this server requires authentication, missing or expired test credentials may be the cause. A test profile lets Glama authenticate for health checks and discover tools; it is separate from your personal connections.

If you are the author, claim ownership, then add or update a test profile under Admin → Test Profile.

Status
Unhealthy
Uptime
99.2% over 42 days
Last Tested
Transport
Streamable HTTP · MCP 2025-11-25
URL
Repository
yinnho/aginxbrowser
GitHub Stars
35
Server Listing
AginxBrowser

TDQS

A3.9/5.0

Scored across 42 tools

Disambiguation4/5

Most tools target clearly distinct resources and actions (account_login vs account_verify vs account_list; session_click vs session_click_xy vs click; fetch vs search vs cache), and the descriptions are unusually detailed. Some boundaries still blur: session_challenges and session_verdict both report anti-bot/risk-control findings, and the many session_* state readers (state, console, network, cookies, storage, screenshot) require careful reading to select correctly.

Naming Consistency4/5

Names follow a mostly consistent snake_case verb_noun pattern with clear domain prefixes (account_, session_, flow_, render_). Minor deviations: a handful of base tools are single verbs (cache, click, download, eval, fetch, search), and stateless counterparts (click/eval) share stems with session_ prefixed tools, but the system remains predictable.

Tool Count2/5

42 tools is heavy for a single MCP server; several session_* diagnostics (challenges/verdict, console/network, cookies/storage) could be consolidated, and the count pushes beyond the 25+ threshold where agent selection load increases. While the broad browser-automation scope justifies many capabilities, the surface feels over-partitioned rather than tightly scoped.

Completeness4/5

The surface covers the full browser-automation lifecycle: fetch/search/cache, session creation/control/state export, account identities, flows, and rendering to markdown/PDF/video. Minor gaps exist (no dedicated session reload/back-forward or explicit cookie-set tool), but they are workable via session_eval/session_navigate, so no major dead ends.

Available Tools

42 tools
account_deleteAInspect

Delete a named login identity: stored record AND live jar. Cookie values are credentials — delete means gone. Sessions currently running as the account keep their in-process jar handle, but nothing writes back. Returns {deleted: name}, or an error naming the account if it does not exist.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYesThe account to delete: stored record AND live jar.

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With only a title annotation, the description carries the full behavioral burden and meets it excellently. It discloses that both the stored record and live jar are removed, that cookies are credentials and deletion is permanent, that running sessions retain an in-process handle but nothing writes back, and the exact return/error shape.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, each earning its place: the first defines scope, the second conveys permanence and gravity, the third covers runtime behavior and return values. Information is front-loaded and there is no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter destructive tool with no output schema, the description is remarkably complete. It covers what is deleted, the irreversible nature, session edge-case behavior, and the expected response and error format. An agent has everything needed to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and the parameter description already states 'The account to delete: stored record AND live jar.' The tool description adds the context that this is a 'named login identity' but does not provide additional format, constraints, or syntax beyond the schema. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Delete') and a specific resource ('a named login identity'), and clarifies scope by naming both the stored record and the live jar. This clearly distinguishes it from siblings like account_list and account_verify.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The destructive intent is unmistakable: 'delete means gone' and 'Cookie values are credentials' make clear this is the permanent removal tool, not list or verify. It lacks an explicit 'use this when...' or 'instead of...' statement, but the context is strong enough that an agent would not confuse it with alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

account_listA
Read-only
Inspect

List named login identities (the multi-account layer) with metadata only: name, cookie domains, cookie count, updated_at, the last account_verify verdict, and the identity's persona User-Agent (each account is one stable device: its own UA and hardware fingerprint, drawn once and reused). Cookie values are credentials and never leave the server. Use to see which identities exist before session_create {account} picks one.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, but the description adds meaningful behavior beyond that: metadata-only results, cookie values are credentials that never leave the server, and each account is a stable device with reused UA and fingerprint. These details help an agent understand data sensitivity and account identity semantics.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a compact, front-loaded paragraph that first states the action, then the returned metadata, then a security note, then usage. Every sentence contributes a distinct, useful piece of information without repetition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, but the description lists the returned fields explicitly: name, cookie domains, cookie count, updated_at, last account_verify verdict, and persona User-Agent. Combined with the readOnly annotation and zero-parameter schema, this gives the agent enough context to invoke the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so there are no parameter semantics to explain. The description does not need to compensate for missing input schema documentation, and the baseline for a no-parameter tool is appropriate here.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Starts with a specific verb and resource: 'List named login identities (the multi-account layer)'. It then enumerates the exact metadata scope, including fields like cookie domains, updated_at, and account_verify verdict, making it clearly distinguishable from related account tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides explicit usage context: 'Use to see which identities exist before session_create {account} picks one.' This tells the agent when to call it, though it does not state when not to use it or name alternative listing tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

account_loginAInspect

Open a site's login page AS a named account and close the login loop. Creates the session as the account (private jar, device persona), navigates to url, and reports which generic login gates the page shows — needs: password | sms | qr | slider (QR scan like xiaohongshu, SMS code, password form, slider/captcha) — plus a session_verdict classification and, when a human step is needed, the session_id and a /live handoff so a person can finish it in the live view (the engine detects and describes; it never fills credentials or solves challenges). With predicate and no human gates detected, waits for the automatic bounce, then teaches + stamps the account's verify spec — later account_verify calls re-check it bare. Cookies write back after every action, so a login finished in /live is already persisted. status is "logged_in" (verify stamped), "waiting" (drive session_wait/session_input on session_id, or re-call account_login after the human finishes — the account's shared jar already holds it), or "opened" (no predicate given).

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYesThe login page URL to open as this account.
nameYesThe account to log in as (created implicitly on first use; 1-64 chars of [a-zA-Z0-9_-]).
predicateNoA JS expression truthy on the page the site lands on AFTER login, e.g. `!!document.querySelector('.user-nick')`. With it and no human step detected, the call waits for the automatic login bounce and stamps the account's verify spec on success. Without it, the call just opens the page and reports what kind of login it sees.
use_proxyNoRoute through the engine proxy. Seeds a fresh account; an account with an existing record reuses its recorded egress.
timeout_msNoWait budget in ms for the automatic-login bounce (default 60000, clamped 1000..120000). Never spent while a human step (QR/SMS/password/slider) is outstanding — those return immediately with a session handoff.

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations provide only a title, so the description carries the full burden of behavioral disclosure. It thoroughly explains that the tool creates a session, never fills credentials or solves challenges, writes cookies back after every action, and reports statuses ('logged_in', 'waiting', 'opened'). It also discloses the handoff mechanism and verify spec stamping. No contradictions with annotations exist.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long but densely packed with necessary information. It front-loads the core purpose, then covers predicate behavior, cookie persistence, and status outcomes. While it could be better structured with bullet points, every sentence adds value and avoids fluff. It is appropriately sized for the tool's complexity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity, the description covers the essential behaviors: login gate detection, session handoff, verify spec stamping, and status values. It mentions the session_verdict classification and the /live handoff but does not detail every edge case, such as error handling or exact return structure. Still, it provides enough for an agent to call the tool correctly in most scenarios.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. The description adds meaningful context beyond the schema: it explains how 'predicate' drives the automatic-login wait and stamps the verify spec, how 'use_proxy' seeds a fresh account, and how 'timeout_ms' is never spent during human steps. These details help an agent choose and configure parameters correctly, exceeding the baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Open a site's login page AS a named account and close the login loop.' It clearly states the tool's primary function and distinguishes it from siblings like account_verify (which re-checks a verify spec) and session_* tools. The purpose is unambiguous and actionable.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear context for when to use the tool, including the predicate parameter for waiting on automatic login bounce, and how to handle human gates by using session_wait/session_input or re-calling account_login. It references account_verify as a follow-up step, but does not explicitly state when to prefer alternatives like session_navigate or session_create. This is strong guidance but not fully explicit about exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

account_verifyAInspect

Check whether a named account is still logged in. Teach-once: the first call passes url + predicate (a JS expression truthy on a logged-in page, e.g. !!document.querySelector('.user-nick')); the spec is remembered and later calls can be bare. Runs in a scratch session AS the account (private jar), so the probe doubles as a cookie refresh. Returns {name, logged_in, url, checked_at}.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlNoTeach-once: the page that shows login state (its login wall if the account is logged out). Remembered after the first call.
nameYesThe account to check.
predicateNoTeach-once: a JS expression that is truthy when logged in, e.g. `!!document.querySelector('.user-nick')`. Remembered after the first call — later calls can pass neither and rerun the spec.

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no readOnly or destructive annotations, the description carries the full burden and excels: it discloses the private-jar scratch session, the fact that the probe doubles as a cookie refresh, and the remembered-spec statefulness. These are exactly the behavioral traits an agent needs beyond the schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three dense sentences, each earning its place: primary purpose, teach-once behavior, and side-effect/return information. The main verb and resource are front-loaded, with no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Although there is no output schema, the description explicitly gives the return shape {name, logged_in, url, checked_at}. Combined with the side-effect, statefulness, and session context, nothing essential is missing for an agent to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the schema already documents url and predicate thoroughly, including the teach-once behavior and example predicate. The description reinforces this but adds no meaning beyond the schema, so it stays at the baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Check whether a named account is still logged in.' This clearly differentiates it from siblings like account_list and account_delete, and the teach-once detail further scopes what the tool does.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives clear operational context: first call passes url and predicate, later calls can be bare, and the check runs in a scratch session as the account. It does not explicitly name alternatives or exclusion conditions, but the tool's unique purpose makes the intended usage unambiguous.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

cacheAInspect

Query the LOCAL CACHE of every page this server has fetched and every search it has run. Check here BEFORE re-fetching or re-searching — a hit is instant and free while a fresh fetch costs 5-60s. Use query for full-text search (works for Chinese substrings and English words), get to pull a page's full cached content, stats for counts, clear to delete rows.

ParametersJSON Schema
NameRequiredDescriptionDefault
allNoWith clear: delete everything cached for this caller
getNoReturn the FULL cached content of this exact URL instead of listing hits
urlNoOnly rows whose URL contains this substring
kindNoWhich rows to search: "auto" (default, pages + searches), "pages", or "searches"
clearNoDelete matching rows instead of returning them (requires url, since_hours, or all)
limitNoMaximum rows returned (default: 10, max 100)
queryNoFull-text search over cached page contents, titles, URLs and past search queries. Omit to list the latest rows.
statsNoReturn row counts and database size instead of rows
since_hoursNoOnly rows stored within the last N hours

TDQS

A4.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only provide readOnlyHint=false, so the description carries most of the behavioral burden. The description clearly discloses the destructive 'clear to delete rows' path, scopes the tool to cached pages and searches, and explains the operational modes (query, get, stats, clear). It does not cover irreversibility or eviction/persistence details, but the destructive behavior is explicit.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is three sentences, front-loaded with the tool's purpose, and every clause earns its place. It communicates scope, usage timing, cost rationale, and mode selection without redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 9-parameter tool with no output schema, the description covers the core invocation model, the main modes, the destructive path, and the decision to check the cache before more expensive operations. Field-level details such as defaults, limit, since_hours, and clear requirements are fully covered by the schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description adds extra semantic value beyond the schema by grouping modes ('Use query for full-text search, get to pull full content, stats for counts, clear to delete rows') and by adding operational details such as Chinese-substring and English-word search behavior.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The opening sentence names a specific verb and resource: 'Query the LOCAL CACHE of every page this server has fetched and every search it has run.' It also differentiates from the sibling fetch/search tools by telling the agent to check here before re-fetching or re-searching.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit usage guidance: 'Check here BEFORE re-fetching or re-searching — a hit is instant and free while a fresh fetch costs 5-60s.' This tells the agent when the cache is the right choice, and why, which is strong routing guidance relative to alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

clickAInspect

Click an element on a one-off page: loads url in a fresh browser context (stateless — no cookies unless passed, no shared state with other calls), waits wait_secs after load before clicking, then fires a DOM click on the first CSS-selector match. The click may trigger navigation (link, form submit) — the response url and text_after are read after that navigation lands. Returns clicked:false when the selector matches nothing. For multi-step interaction on a shared page use session_click instead.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYesThe URL to load
selectorYesCSS selector of element to click
wait_secsNoSeconds to wait for the page to settle after load, before clicking

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Discloses key runtime behavior beyond the schema: a fresh stateless browser context, waiting after load, DOM click implementation, possible navigation, and response read after navigation. Also documents the clicked:false miss case. Since there are no behavioral annotations, this description carries the burden and does so thoroughly.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four dense sentences with no filler; the main action comes first, then side effects, error case, and alternative are sequenced for efficient parsing. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For three simple parameters, the description covers statelessness, clicking behavior, navigation, and the no-match return value. It is nearly complete, but without an output schema it leaves the exact success shape implicit beyond url/text_after and clicked:false.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. The description adds semantics on top: the selector matches the first element, wait_secs governs time after page load before clicking, and url is loaded in a fresh context. This extra context justifies a 4 rather than baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific operation ('Click an element on a one-off page') with a clear resource (the page URL) and mechanism (CSS-selector match). It explicitly contrasts itself with session_click, so an agent can distinguish it from sibling click tools without inferring.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives the exact usage context: one-off, stateless page interaction. It closes with an explicit routing condition: 'For multi-step interaction on a shared page use session_click instead,' which tells the agent when to prefer the sibling.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

downloadAInspect

Download a file over HTTP(S) with streaming to disk (no memory buffering), SHA-256 integrity hash, and optional resume of interrupted transfers. Filename resolution: explicit param → Content-Disposition → URL tail. Use for binaries, archives, datasets, documents - anything where the agent wants the FILE saved, not its text content read.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYesURL of the file to download (http/https)
resumeNoResume an interrupted download when a local partial file exists
cookiesNoCookies to send with the request: `"name=value"` strings or CDP-style objects `{"name","value","domain",...}` for gated downloads
filenameNoExplicit output filename. When omitted: Content-Disposition → URL tail → "download"
use_proxyNoRoute through proxy (default: false; auto-enabled for known blocked domains)

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With only a title annotation, the description carries the behavioral disclosure burden. It does well by revealing no-memory-buffering streaming, SHA-256 hashing, resume behavior, and the filename resolution precedence. It falls short of 5 because it does not disclose the return value, where the file lands on disk, or how failures are surfaced.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences with no filler. The core behavioral constraint is front-loaded, filename resolution is compressed into a clear chain, and the usage guidance closes the description efficiently. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a download tool, the description is largely complete: purpose, major behavioral traits, filename handling, and use cases are all present. However, without an output schema or supporting annotations, the absence of return-value details, target path, and failure-mode expectations leaves a moderate gap for an agent that must verify the tool's result.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3; each parameter is already described in the schema. The description adds meaningful context for 'filename' via the resolution precedence chain (explicit → Content-Disposition → URL tail), but it does not deepen semantics for url, resume, cookies, or use_proxy beyond what the schema provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource ('Download a file over HTTP(S)') and immediately adds distinguishing mechanics: streaming to disk, SHA-256 integrity hash, resume, and filename resolution. It also explicitly contrasts with reading text content, which differentiates it from sibling tools like fetch.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The final sentence gives a clear when-to-use rule: binaries, archives, datasets, and documents where the agent wants the FILE saved. It also states the when-not condition ('not its text content read'), effectively routing the agent toward a fetch-style alternative for text retrieval.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

evalAInspect

Execute JavaScript on a one-off page: loads url in a fresh browser context, optionally waits wait_secs for the page to settle, evaluates script (async/Promise supported) and returns {url, result}. Script-driven navigation (location.href, form submit) is drained and reflected in the returned url. Stateless — no cookies or page state shared with other calls; when the script needs prior page state or a login, use session_eval.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYesThe URL to load
scriptYesJavaScript code to execute (supports async/Promise)
wait_secsNoSeconds to wait before executing

TDQS

A4.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations provide only a title, so the description carries the full behavioral burden. It discloses fresh browser context, optional wait, async/Promise support, statelessness, no cookie/page-state sharing, and navigation being drained into the returned `url`. It does not mention error handling, timeouts, or broader side effects of executing arbitrary JavaScript, but the core execution model is clearly explained.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three efficient sentences with the main action and return value front-loaded, followed by a navigation edge case and a statelessness caveat with the sibling alternative. No filler or redundant restating of the schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no output schema and minimal annotations, the description covers the essential invocation details: return shape, stateless boundary, navigation handling, and when to choose `session_eval`. Minor gaps remain around error/timeout behavior and exact script execution scope, but the description is sufficient for correct use.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, but the description adds meaning beyond the schema: `url` is loaded in a fresh browser context, `wait_secs` is for the page to settle, and `script` supports async/Promise. This contextualizes the parameters better than the terse JSON Schema descriptions alone.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a precise action ('Execute JavaScript'), a specific resource ('a one-off page' loaded from `url`), and a clear return shape (`{url, result}`). It also distinguishes itself from sibling `session_eval` by explicitly noting that this tool is stateless.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly names `session_eval` as the alternative when prior page state or login is needed, giving a concrete when-not-to-use condition. It also implies the appropriate use case: one-off, stateless JavaScript evaluation on a fresh page.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

fetchA
Read-only
Inspect

Fetch a webpage and return clean markdown/html/text. Use whenever the agent needs to READ any web page - blogs, docs, articles, JS-rendered SPAs, Cloudflare-protected sites. Static pages are served over plain HTTP (~100ms tier:"http"); pages that need JS get the full browser (tier:"browser"). render_tier selects auto (default) / http (pure HTTP, refuses the upgrade) / browser (always the JS browser).

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYesThe URL to fetch
formatNoOutput format: "markdown", "html", or "text" (default: markdown)markdown
sanitizeNoStrip prompt-injection payloads from the text output (default true): zero-width/steganographic characters, instruction-shaped lines ("ignore previous instructions", chat markup tokens, CJK variants), and text hidden via opacity:0 / tiny fonts. A `sanitize_report` field counts what was removed — stripping is observable, never silent. Set false for raw output.
selectorNoCSS selector to extract specific content
max_charsNoMaximum characters to return (default: 50000)
use_proxyNoRoute through proxy (for blocked foreign sites)
wait_secsNoSeconds to wait for JS rendering
js_extractNoJS expression to extract from the page after rendering
capture_xhrNoCapture script-initiated API responses: a list of URL substrings (e.g. ["/api/"]) whose matching fetch/XHR bodies come back in an `xhr` array; an empty list captures every XHR/Fetch. Forces browser rendering (script-initiated requests only exist after JS runs).
render_tierNoRendering strategy: "auto" (default), "http", or "browser"auto
tls_fingerprintNoTLS fingerprint override (stealth mode only): "chrome145", "firefox133", etc.
auto_bypass_challengeNoAuto-detect and bypass Cloudflare Turnstile challenges (default: true)

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the readOnlyHint annotation, the description discloses meaningful behavior: static pages use plain HTTP, JS pages use a full browser, and render_tier controls auto/http/browser with the nuance that http 'refuses the upgrade.' This gives the agent a real sense of how the tool behaves at runtime, though some details like sanitization are left to the schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is three sentences, front-loaded with the core purpose, and each sentence adds useful information. The rendering-tier explanation is a bit dense and contains awkward formatting, but there is no wasted prose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with 12 parameters and no output schema, the description covers the essential invocation context: what it returns, when to use it, and how rendering tiers work. It does not enumerate every parameter, but the schema provides that detail, so the description is sufficiently complete for an agent to select and call the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description adds a little extra color around render_tier (e.g., 'refuses the upgrade') but mostly restates what the schema already documents. It does not materially improve understanding of the other 11 parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Fetch a webpage and return clean markdown/html/text.' It clearly positions the tool as the read-only web-fetching option among siblings like download, search, and session_navigate, and the mention of JS-rendered SPAs and Cloudflare-protected sites further distinguishes its scope.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly says 'Use whenever the agent needs to READ any web page' and gives concrete examples (blogs, docs, articles, JS-rendered SPAs, Cloudflare-protected sites). It does not explicitly name alternative tools or exclusion cases, but the context is clear enough for an agent to select this tool over session-based or download siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

flow_installAInspect

Install a third-party flow from DupHub into the server's workflow directory, making it runnable by name via flow_run. Pulls the template's manifest, verifies every file's sha256 against it (a package that fails anywhere is not installed), requires a valid flow.json with a steps array, and lands atomically under workflow// — replacing any previous install, overriding a same-named built-in. The receipt lists files, total bytes, step count, and how many steps carry scripts (eval_steps) — third-party scripts run with this engine's privileges in session pages, so review before running. Source remote is env-configured (AGINXBROWSER_DUPHUB_URL), never a parameter.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYesWorkflow name — the DupHub template to pull and the directory it lands in (workflow/<name>/). Lowercase/digits/dashes.

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With only a title annotation, the description carries the full burden and does so: atomicity, replacement of prior installs, built-in override, sha256 verification with all-or-nothing failure, required flow.json/steps shape, and a privilege-escalation warning about third-party eval_steps running with engine privileges in session pages. This is exceptional disclosure for a mutation tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the action and destination, then layers verification, atomicity, and the security caveat in order of importance. It is dense and long-sentenced but nearly every clause carries operational information; only the built-in override detail is arguably extraneous.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, and the description compensates by enumerating the receipt contents (files, total bytes, step count, eval_steps). Combined with failure semantics and the env-configured source, an agent has everything needed to invoke and interpret this tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and the single `name` param is already well documented there, so the baseline is 3. The description adds genuinely non-schema meaning: the source remote is env-configured via AGINXBROWSER_DUPHUB_URL and is never a parameter, which prevents an agent from trying to pass a URL.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (install) and resource (third-party flow from DupHub) plus the concrete landing location and the effect of making it runnable by name via flow_run. The distinction from the sibling flow_run is made explicit, so an agent can tell the two apart without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Positively routes to flow_run for execution and implicitly frames this as the prerequisite setup step, and the security note ('review before running') gives a when-to-be-careful cue. It stops short of stating explicit exclusions or prerequisites (auth, network access) for the install itself.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

flow_runAInspect

Run a flow — a recorded, editable JSON browser-session script — deterministically, with zero model tokens. Steps are {op, args, expect?, save?}: ops cover navigate/click/click_xy/input/scroll/eval/wait/screenshot/state/cookies; {{var}} placeholders in args are filled from vars; expect asserts (url_contains | selector | text_contains | eval_truthy) abort with evidence on failure; save collects a step's output into the receipt. Source the flow inline via "flow", or by "name" from the server's workflow//flow.json (unknown name → error lists installed workflows). Pass session_id to reuse a live session (e.g. from import_curl) so login state and flows compose. The receipt carries status ok/failed, saved outputs, the session_id (kept alive), and on failure the failing step, reason and a diagnostic screenshot — fix the flow or take the session over from there.

ParametersJSON Schema
NameRequiredDescriptionDefault
flowNoInline flow document: {create?, vars?, steps:[{op, args, expect?, save?}]}
nameNoOr run a server-side workflow/<name>/flow.json asset. An unknown name errors back with the list of installed workflows — that error is the discovery call.
varsNoValues for {{placeholders}} in step args; wins over the flow's own vars defaults.
max_stepsNoOverride the run's step-execution budget (branch loops re-run steps, so every revisit counts). Default 1000, clamped 1..=100000. A flow document may also declare its own max_steps; this wins.
session_idNoReuse a live session (e.g. from import_curl) instead of creating a fresh one — that's how login state and flows compose.

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations provide only a title, so the description carries the full behavioral burden. It discloses that execution is deterministic and uses zero model tokens, that expect assertions abort with evidence on failure, that save collects step output into the receipt, and that on failure a diagnostic screenshot is included. It also mentions session_id is kept alive after the run, implying side effects on session state. This is rich, honest disclosure of execution behavior and failure modes, though it doesn't detail permission needs or reversibility, which are less relevant for a sandboxed browser script.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single dense paragraph but front-loads the core concept, then expands into operational details. Every sentence serves a purpose: defining the script shape, ops, placeholders, expect/save semantics, sourcing options, session reuse, and receipt behavior. It avoids filler, though it could be slightly better structured with bullet points for ops or expectations. As is, it is efficient and scannable.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool of this complexity (5 optional params, no output schema), the description covers the critical usage aspects: how to provide a flow, how to name it, how placeholders and vars work, how assertions and saving behave, and what the receipt contains on both success and failure. It also explains the session_id reuse mechanism. The only minor missing piece is an explicit list of supported 'op' values, though the description enumerates them inline (navigate/click/click_xy/etc.), so it is adequately complete for an agent to understand what it can do.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so a baseline 3 applies, but the description adds substantial meaning beyond the schema. It explains how 'flow' and 'name' are alternatives, what happens on unknown names, how 'vars' overrides defaults, and how 'max_steps' interacts with branch loops and document-level declarations. The schema descriptors are terse; the description resolves their operational intent (e.g., 'max_steps' becomes a budget with branch re-counting). This goes well beyond the raw JSON schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with 'Run a flow — a recorded, editable JSON browser-session script — deterministically, with zero model tokens.' This gives a precise verb (run), a specific resource (flow) and its characterization. It also distinguishes itself from sibling session_* tools by framing flows as composed scripts, not single operations. The explanation of ops, placeholders, and expect/save further clarifies what the tool does.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explains how to source flows (inline via 'flow' or by 'name' from server assets) and explicitly recommends passing session_id to reuse a live session (e.g., from import_curl) so login state and flows compose. It also states that an unknown name errors with a list of installed workflows — turning an error into a discovery call. While it doesn't explicitly say 'when not to use this versus session_* tools', the composition and determinism cues imply it is for scripted multi-step runs, which is enough context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

import_curlAInspect

Import login state from a real browser in one paste. The human logs into a site in their own Chrome (solving the CAPTCHA/SMS once), opens DevTools → Network, right-clicks any authenticated request → "Copy as cURL", and passes the command here. Returns a live session_id already carrying that site's cookies and sitting on the copied request's URL — the agent continues from where the human left off, no password or second login needed. Works with bash, PowerShell and cmd copy flavors.

ParametersJSON Schema
NameRequiredDescriptionDefault
curlYesA "Copy as cURL" command pasted from Chrome DevTools (Network panel → right-click any authenticated request). bash, PowerShell and cmd flavors all parse; the cookie set is injected and the session navigates to the copied request's URL.
accountNoAttach the session to a named account: the imported login lands in the account's private jar and is written back under its name after every action — one import per identity, no clobbering.
use_proxyNoRoute the session's traffic through the engine proxy.

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only provide a title, so the description carries the behavioral burden. It discloses that the tool imports cookies, navigates to the copied request's URL, returns a session_id, and supports multiple shell flavors. It does not mention side effects like session replacement or security warnings, but the core mutation behavior is transparent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, moderately long paragraph but every sentence earns its place. It front-loads the core purpose and then details the workflow and output. Slightly run-on but not wasteful; could be split into bullet points for clarity, but it is effective.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 3-param tool with no output schema, the description covers the input format, the expected usage flow, the return value (session_id), and account attachment behavior. It lacks edge-case handling (invalid cURL, overwrite rules) but is complete enough for an agent to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. The description adds value beyond the schema by explaining the curl param's multi-flavor parsing and the account param's 'private jar' and 'no clobbering' behavior, which clarifies intent beyond the raw property descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb (import) and resource (login state from a browser via cURL), and distinguishes it from siblings like session_create by framing it as 'Import login state from a real browser in one paste.' It precisely explains what the tool does and how it fits into the session workflow.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives a concrete scenario (human solves CAPTCHA/SMS in Chrome, copies cURL, agent continues) and implicitly positions it as the alternative to manual authentication. However, it does not explicitly name when not to use it or mention alternatives like session_create for fresh sessions, so it stops short of full exclusion guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

render_markdownAInspect

Render a markdown document into a deterministic, self-contained HTML artifact - the document layer, so the agent never writes HTML by hand. Prose rides a plain offline shell (no fonts, no scripts); archify fenced code blocks carry typed zero-coordinate diagram JSON (sequence, workflow, architecture, dataflow, lifecycle families) and render to inline SVG via the layout engine. Same input, same bytes: the receipt carries the sha256 so determinism is verifiable. theme picks light (default) or dark; preset picks the palette family — classic (default), signal-flow, blueprint, editorial — orthogonal to theme; colors bake at generation time (presentation attributes, not CSS variables), and the receipt records both preset and theme. quality picks the composition audit profile — standard (default) or showcase, the delivery gate: the receipt's diagrams[].composition grades route crossings, ambiguous corridors, label clearance (2px standard / 4px showcase), route rhythm, and node text projected to the 930px reader width; the audit never changes the artifact bytes. Mermaid sources are the agent's job to translate, not the engine's: flowchart/graph → workflow (lanes + columns), sequenceDiagram → sequence, stateDiagram-v2 → lifecycle (bands), erDiagram/class → architecture (grid + boundaries) — read the topology and emit the matching zero-coordinate archify JSON; the engine accepts only archify JSON. A broken diagram degrades to a visible code block and lands in receipt.diagnostics; an authored route preset that cannot be honored is self-repaired to a verified semantic substitute and disclosed in receipt diagrams[].repairs - the document still renders. A fence may also carry views: [{id,label,nodes,note?}] (node ids of the active family), emitted as guided-view tabs above the diagram plus an inlined viewer script - clicking a tab lights the member nodes and the routes between them (subgraph), clicking a node lights it with its direct neighbors (ego graph), everything else dims; a view's optional note shows as a caption while it is active (the story layer). window.agxViewer in a session drives and reads the same state programmatically: {focus,view,state} as before, plus route(i,from,to) which returns and lights the shortest authored directed path between two nodes (null when unreachable, state untouched), and reach(i,id,down|up) which returns and lights the authored downstream/upstream closure ({nodes,links}); both dim the rest of the diagram. diagrams[].views in the receipt lists the tabs. motion: true bakes an entrance choreography into the artifact: pure-declarative CSS animation with zero scripts - headings split into per-glyph (CJK) / per-word (latin) spans that rise in with expo easing, prose blocks stagger up an nth-child delay ladder, diagram figures grow in with a back ease (GSAP's easing math as public cubic-bezier equivalents, nothing embedded); the diagrams themselves play a flow story on the same clock - nodes land beat by beat, solid edges draw in (dash-offset), dashed returns fade, sequence messages arrive as sent - with a timed caption strip under each figure as the subtitles, which becomes a static transcript under prefers-reduced-motion; the file itself animates in any browser and the receipt records motion plus diagrams[].story (beat times and captions - the hook for muxing voice later). With session_id the artifact is also loaded into that session (local, free) and the reply carries viewport acceptance: scroll extents measured in the live session and graded fits/tall/wide/oversized, telling the agent how to read the page back. Diagram vocabulary adapted from archify (MIT).

ParametersJSON Schema
NameRequiredDescriptionDefault
themeNoColor theme: "light" (default) or "dark" — the shell background/ foreground and every SVG palette slot swap together; the receipt records which theme produced the bytes
motionNoBake the entrance choreography into the artifact (default false): pure-declarative CSS animation — headings split into per-glyph/per- word spans that rise in with expo easing, prose blocks stagger up a nth-child delay ladder, and diagram figures grow in with a back ease (GSAP's easing math as public cubic-bezier equivalents). The diagrams animate too, on one story clock: nodes pop in one beat at a time, solid edges draw themselves (dash-offset drain), dashed returns fade, sequence messages land as they are "sent", and a timed caption strip under each figure subtitles the beats — under prefers-reduced-motion the strip becomes a static transcript. Zero scripts: the file itself animates in any browser, subtitles and all; the receipt records motion (plus diagrams[].story with the beat times, the hook for muxing voice later) so a cached artifact is never mistaken for the static one
presetNoVisual preset: "classic" (default), "signal-flow", "blueprint", or "editorial" — a palette family orthogonal to theme (each preset exists in both light and dark). The receipt records preset and theme separately
qualityNoQuality profile for the composition audit: "standard" (default) or "showcase" — the delivery gate. The audit grades route crossings, corridors, label clearance, rhythm, and projected text size in the receipt (diagrams[].composition); it never changes the artifact bytes, only how findings are severity-rated
markdownYesFull markdown document. Prose rides a plain offline shell; archify fenced code blocks carry typed zero-coordinate diagram JSON and render to inline SVG.
session_idNoOptional session ID: also load the rendered HTML into that live session (local and free) so session_screenshot / session_state can verify the artifact

TDQS

A3.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The annotations carry only a title, so the description bears the full disclosure burden, and it delivers extensively: deterministic bytes with a sha256 receipt for verification, a fully offline shell (no fonts/scripts), graceful degradation of broken diagrams to visible code blocks with diagnostics, self-repair of unhonorable route presets disclosed in repairs, reduced-motion fallback to a static transcript, and session loading described as local/free. This far exceeds the transparency expectations for an unannotated tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The first sentence is well front-loaded, but the description runs roughly 700 words and buries actionable content under implementation minutiae an agent does not need to invoke the tool: GSAP easing math as cubic-bezier equivalents, per-glyph CJK span splitting, an nth-child delay ladder, and full window.agxViewer API signatures (route(i,from,to), reach(i,id,down|up)). Much of the motion prose is duplicated nearly word-for-word in the schema's motion description, so those sentences do not earn their place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's high complexity (6 parameters, a demanding archify JSON input contract, no output schema) and a title-only annotation, the description is exceptionally complete: it specifies defaults for every option, input format requirements, failure and repair semantics, determinism verification via the receipt, post-call viewport grading (fits/tall/wide/oversized), and even licensing provenance. An agent has everything needed to call and verify this tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, and the description adds meaning above it: theme/preset are orthogonal with colors baked at generation time as presentation attributes rather than CSS variables, quality thresholds are quantified (2px standard / 4px showcase, 930px reader width), and failure/repair behavior is tied to each option. Some marginal value is lost because the schema's motion parameter description already repeats nearly the full animation behavior verbatim.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The opening sentence is specific: 'Render a markdown document into a deterministic, self-contained HTML artifact - the document layer, so the agent never writes HTML by hand.' The verb (render), resource (markdown document), and output (self-contained HTML) are unambiguous. However, unlike the strongest examples, it never names its siblings render_pdf/render_video or explains the boundary between them, so the differentiation is implicit rather than explicit.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives strong in-tool input guidance ('Mermaid sources are the agent's job to translate, not the engine's... read the topology and emit the matching zero-coordinate archify JSON; the engine accepts only archify JSON'), which tells the agent what to feed it and what not to feed it. But there is no explicit when-to-use-this-vs-alternatives guidance for tool selection against render_pdf/render_video; the only usage framing is the implied 'use this instead of writing HTML by hand.'

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

render_pdfAInspect

Cut a rendered page into pages and package as PDF, PNGs, PPTX or DOCX. Print mode (no selector) paginates the document into fixed-height pages (default 794x1123, A4 @96dpi), breaking at top-level block boundaries — no half-cut text where a break can land on a block edge. Slides mode (selector set) makes one page per match, sized to that element — generate an HTML deck with one .slide per page and each becomes a deck page. format "pdf" (default) returns base64 image-based PDF; "png" returns one base64 PNG per page in pages_base64; "pptx" returns a base64 PPTX (one slide per page, deck-sized to the largest page); "docx" returns a base64 DOCX (one page-sized section per page, each section keeps its own height). Returns page count and packaging.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYesPage URL to cut into pages.
widthNoPage width in CSS pixels. Default 794 (A4 @96dpi).
formatNoOutput format: "pdf" (default), "png" (one base64 PNG per page), "pptx" (one slide per page, image-based), "pptx-native" (editable: element-level DrawingML — real text runs, gradient shapes, image parts; requires `selector`), or "docx" (one page-sized section per page).pdf
heightNoPage height in CSS pixels — print pagination only. Default 1123.
selectorNoCSS selector; present → slides mode (one page per match, sized to the element). Absent → print mode (fixed-height pages at block boundaries).
max_pagesNoSafety cap on emitted pages. Default 50.
use_proxyNoRoute through proxy (for blocked foreign sites)
jpeg_qualityNoJPEG quality for PDF page embedding (1-100). Default 90.
tls_fingerprintNoTLS fingerprint override (stealth mode only)

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With only a title annotation and no read-only/destructive hints, the description carries the behavioral burden. It discloses pagination boundaries, slide-per-match behavior, image-based PDFs, per-format packaging details, and the return of page count. It stops short of describing side effects like network fetching or permission requirements, but these are largely non-mutating.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but well-structured: purpose first, then mode behavior, then format-specific output details, then return summary. Every clause adds information, though the long middle sentence packs many behaviors together and could be slightly hard to scan.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with 9 parameters, two modes, and five output formats, the description covers the essential decisions and return characteristics. With no output schema, it supplies high-level return information ('page count and packaging') and names pages_base64 for PNGs, but does not fully specify the response envelope.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description adds real meaning beyond the schema by explaining mode-dependent behavior (block-boundary breaks vs element-sized pages), the deck-size behavior of PPTX, and the section-height behavior of DOCX. This is more than a restatement of parameter names.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('cut') and resource ('rendered page') and enumerates the exact output formats (PDF, PNGs, PPTX, DOCX). It clearly distinguishes itself from sibling rendering tools like render_video and render_markdown by describing page-based packaging.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear when-to-use guidance for the two modes: print mode when no selector is set, and slides mode when a selector is provided. It does not explicitly contrast the tool with sibling alternatives, but the mode-level guidance is strong and actionable.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

render_videoAInspect

Render a page's animation timelines to an MP4 video. The page's scripts must expose window.__timelines — objects with duration() and pause(t) (a paused gsap.timeline registered there works as-is). Each frame seeks every timeline to t=i/fps and paints the viewport, so the output is deterministic — no wall clock in the pixel values. Audio: narration[] places TTS/voice clips at start times (mixed into one AAC track), audio adds looped background music, and subtitles_srt muxes an SRT as a soft mov_text track and (by default, burn_subtitles: false to opt out) burns the same cues into the frame pixels — QuickTime, WeChat and most social embeds ignore the soft track. Requires ffmpeg on the server. Returns base64 MP4 (H.264, yuv420p) plus frame count and durations.

ParametersJSON Schema
NameRequiredDescriptionDefault
fpsNoFrames per second. Default 24.
urlYesPage URL whose scripts register timelines in `window.__timelines` (GSAP-style objects with `duration()` + `pause(t)`).
audioNoBackground music: looped to cover the video, volume-scaled, faded out at the tail.
widthNoViewport width in CSS pixels (floored to even — yuv420p). Default 1280.
heightNoViewport height in CSS pixels. Default 720.
narrationNoVoiceover clips, each starting at its own time (any TTS output; mixed into one AAC track).
use_proxyNoRoute through proxy (for blocked foreign sites)
subtitles_srtNoInline SRT subtitles muxed as a soft (toggleable) mov_text track.
burn_subtitlesNoBurn the cues into the frame pixels too (hardsub) — on by default when `subtitles_srt` is present; QuickTime, WeChat and most social embeds ignore the soft mov_text track. `false` keeps the soft track only.
hold_tail_secsNoFreeze the final timeline state for this many extra seconds. Default 0.5.
tls_fingerprintNoTLS fingerprint override (stealth mode only)
max_duration_secsNoSafety cap on timeline + hold tail, seconds. Default 120.
wait_timelines_msNoHow long to wait for `window.__timelines` to appear, ms. Default 10000.
subtitles_languageNoISO language tag for the subtitle track, e.g. "eng" / "zh".

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With only a title annotation, the description carries the full burden and does an excellent job: it discloses the deterministic frame-seeking behavior, the audio mixing rules, the subtitle muxing and default burn behavior, the ffmpeg dependency, and the exact return value (base64 MP4 plus frame count and durations).

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is multi-sentence but every clause earns its place. It is logically organized: core mechanism first, then audio, then subtitles, then requirements and output. No filler or repetition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given 14 parameters, no output schema, and the tool's complexity, the description covers all essential aspects an agent needs to invoke it correctly: what it does, how it works, the audio and subtitle behaviors, the server dependency, and the return format. Nothing critical is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Although the schema already covers all 14 parameters (100% coverage), the description adds critical meaning beyond it: how timelines are sought, how audio clips are mixed, why burn_subtitles defaults to true (social embeds ignore soft track), and the hold-tail freeze behavior. This is far more than the schema provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb-resource pairing: 'Render a page's animation timelines to an MP4 video.' It then details the exact mechanism (seeking timelines per frame) and the output format, which fully distinguishes it from sibling tools like render_pdf and render_markdown.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description clearly implies when to use it — when a page exposes window.__timelines and you need a deterministic MP4 of the animation. It also states the ffmpeg requirement. However, it does not explicitly name alternative rendering tools or state when not to use it, so it falls just short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

session_challengesA
Read-only
Inspect

One-call risk-control report: did this session hit an anti-bot wall? Taobao/tmall's x5 risk control answers 200 like a normal response — either a redirect onto a punish page (tmd/punish, punish.taobao.com) or an MTop API body carrying FAIL_SYS_USER_VALIDATE / RGV587 / x5secdata. Returns {total, events:[{url,method,status,kind,via}]} where via says whether the wall was navigated into ("url") or swallowed by an API response ("body"). When there are hits, the response also carries the account name (which identity got walled) and a handoff instruction: the engine detects and surfaces but does not auto-bypass — a human opens the live view (/live?session= on the engine's HTTP port), solves the challenge in this session, and the retry rides the cookie that solving sets. Detection only; no automated solving or bypass.

ParametersJSON Schema
NameRequiredDescriptionDefault
session_idYesSession ID

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the readOnlyHint annotation, the description thoroughly discloses behavior: the detection signals, the two possible wall locations via 'url' or 'body', the output structure, the account-name/handoff addition, and the explicit statement that the engine only detects and surfaces, never solves or bypasses. This is rich, accurate behavioral context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long but information-dense: every component, from purpose to failure indicators to output shape to human handoff, earns its place. It is also front-loaded with the primary question the tool answers.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema present, the description compensates fully by specifying the return shape and hit-specific fields. It also explains the follow-up human workflow and the tool's non-bypass boundary, making the definition complete enough for correct invocation and interpretation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema already documents the single session_id parameter with 100% coverage. The description only loosely ties it to 'this session' and does not add format, constraints, or lifecycle context. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a clear, specific purpose: a 'one-call risk-control report' that tells whether a session hit an anti-bot wall. It names concrete failure indicators and the expected output shape, and it distinguishes itself from generic session tools by being detection-only.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear context for when to use it: whenever you need to check whether a session was walled, and it explicitly notes that the tool does not auto-bypass. However, it does not name sibling alternatives or state when to prefer a different session inspection tool, leaving some routing to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

session_clickAInspect

Click an interactive element by its index (from session_state output) inside a live browser session: scrolls it into view and fires a DOM click on the session's current page. Before clicking it re-verifies the element in the same frame — if the page changed since session_state (element detached, disabled, hidden, or covered by an overlay), it returns clicked:false with a reason ("detached"/"disabled"/"not_visible"/"covered_by") and, when covered, a covered_by description of the element that would eat the click — never a silent no-op. A submit click may navigate the session — the returned url/text_after reflect the page after the action, and session state (cookies, localStorage, globals) persists for follow-up calls. Indexes come from the most recent session_state; re-list after navigation.

ParametersJSON Schema
NameRequiredDescriptionDefault
indexYesElement index (from /state output)
session_idYesSession ID

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only provide a title, so the description carries the full burden. It thoroughly discloses behavior: scrolls into view, fires DOM click, re-verifies element, returns clicked:false with reasons on failure, never silent no-op, may navigate, and session state persists. This is exceptionally transparent and goes beyond basic expectations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single dense paragraph but well-structured: it opens with the core action and source, then details pre-conditions, failure modes, side effects, and persistence. It is not overly verbose for the amount of critical information it conveys, though breaking it into bullets could improve readability. It earns its length by covering essential behavioral details.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Despite lacking an output schema, the description explains return values (clicked:false, reason, covered_by, url, text_after), failure modes, navigation side effects, and session persistence. For a two-parameter tool, this is remarkably complete—an agent has everything needed to invoke it correctly and interpret outcomes.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% with both parameters described (index from /state output, session_id). The description adds value by clarifying that index must come from the most recent session_state and re-list after navigation, plus explaining how the index is used (re-verification). This enriches the schema meaning beyond the basic descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('Click an interactive element'), the resource (element by index in a live browser session), and the source of the index (session_state output). It also distinguishes itself from sibling session_click_xy by specifying index-based clicking, making it easy to select the right tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear context: indexes come from the most recent session_state, and it advises re-listing after navigation. It implicitly differentiates from session_click_xy (coordinate-based) by emphasizing index-based selection, though it doesn't explicitly name the alternative. It covers when the tool may not work (page changed) and how to handle it, giving sufficient usage guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

session_click_xyAInspect

Click at viewport coordinates (CSS pixels) via real mouse events — pointerdown/mousedown, pointerup/mouseup, then click on whatever element is hit there. For canvas/map surfaces with no DOM element to index. click_count 2 adds dblclick.

ParametersJSON Schema
NameRequiredDescriptionDefault
xYesViewport X coordinate in CSS pixels
yYesViewport Y coordinate in CSS pixels
buttonNoMouse button: "left" (default), "right", "middle"
session_idYesSession ID
click_countNoClick count: 1 single (default), 2 adds dblclick, 3+ sets detail

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only contain a title, so the description carries the full burden. It discloses the event sequence (pointerdown/mousedown, pointerup/mouseup, then click) and the click_count behavior, which adds value beyond the schema. It does not mention side effects like navigation or page state, but the core behavior is well covered.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences with zero filler. The core action and event sequence are front-loaded, the use case is stated clearly, and click_count behavior is appended without unnecessary detail.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a coordinate-based click tool with no output schema and moderate parameter count, the description covers purpose, event behavior, and primary use case. It doesn't address error conditions or prerequisites like page loaded state, but these are not essential for a simple click operation. The distinction from element-based clicking is implied, not fully elaborated.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so baseline is 3. The description adds semantics for click_count ('2 adds dblclick') and reinforces x/y as viewport coordinates, but it does not add meaning for button or session_id beyond what the schema already provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Click at viewport coordinates'), explains the real mouse event sequence, and explicitly scopes to canvas/map surfaces with no DOM element to index. This clearly differentiates it from sibling click/session_click tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides clear context for when to use: 'For canvas/map surfaces with no DOM element to index.' It implies the alternative (element-based clicking) but does not name specific sibling tools or explicitly say 'use session_click instead', so it falls short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

session_cloneAInspect

Derive a new browser session from a live one, carrying the full login state: cookies, localStorage/sessionStorage, viewport pin, dialog policy, proxy and keepalive flags. The source session stays untouched. Use to snapshot a logged-in state before risky actions, or to run the same login in parallel tabs. Returns {session_id (new), cloned_from, url, viewport}.

ParametersJSON Schema
NameRequiredDescriptionDefault
session_idYesSession ID to derive from (stays alive and untouched)

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With only a title annotation, the description carries the behavioral burden. It explicitly guarantees 'The source session stays untouched,' lists the state that is cloned, and gives the return shape. This gives an agent meaningful expectations without contradicting annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three dense sentences with no filler: the operation, the carried state, the source-safety guarantee, the use cases, and the return object are all front-loaded. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-parameter tool with no output schema, the description is complete: it covers purpose, exact behavioral effect, when to use it, and the return fields. An agent can select and invoke it correctly without needing further context.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and the single parameter is already well documented as 'Session ID to derive from (stays alive and untouched).' The description adds no new parameter-level detail beyond contextualizing that the new session is derived from this one, so the schema carries the weight.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Derive') with a clear resource ('a new browser session') and names exactly what is carried over, including cookies and storage. It is distinct from sibling tools like session_create or session_export, so an agent can tell when session_clone is the right call.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly states when to use the tool: to snapshot a logged-in state before risky actions or run the same login in parallel tabs. It does not mention when not to use it or name alternatives, so it stops short of the full 5, but the use context is clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

session_closeAInspect

Close a browser session and free its resources. For a persistent session this also drops the on-disk login snapshot - idle expiry keeps it, an explicit close does not.

ParametersJSON Schema
NameRequiredDescriptionDefault
session_idYesSession ID

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no readOnly or destructive hints in the annotations, the description carries the behavioral disclosure burden. It does so well by revealing that an explicit close drops the on-disk login snapshot for persistent sessions, while idle expiry keeps it. This is a meaningful side effect beyond what the schema or annotations convey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences with no filler. The core action is front-loaded, and the second sentence adds only the critical distinction about persistent session snapshots. Every word earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a low-complexity, one-parameter close tool with no output schema, the description covers the action, resource impact, and a key side effect. It does not mention error behavior or idempotency, but these are not essential for an agent to invoke the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema already provides 100% coverage for the single session_id parameter with the description 'Session ID', so the baseline is 3. The description does not add parameter-level detail, but none is needed given the simple schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Close') with a clear resource ('browser session') and a concrete consequence ('free its resources'). It clearly differentiates this from sibling tools like session_create or session_list, which serve different lifecycle stages.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The intended usage is implied clearly: call this when a browser session should be ended and its resources released. The description adds useful context by contrasting explicit close with idle expiry, but it does not explicitly name alternatives or state when not to use the tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

session_consoleA
Read-only
Inspect

Read the session's recent page console output (log/info/warn/error) as {url, total, matched, messages:[{ts_ms, level, text, url}]}, newest last. Ring buffer of 500 entries; captures output from page scripts, clicks, evals and navigation alike. Optional filters: level (exact, e.g. "error"), since_ts (epoch ms), url_contains (page URL substring), limit (most recent N matches). The fastest way to see WHY a page misbehaves: click the button, call this, read the error.

ParametersJSON Schema
NameRequiredDescriptionDefault
levelNoOnly entries at this level: "log" | "info" | "warn" | "error"
limitNoKeep only the most recent N matching entries
since_tsNoOnly entries logged at or after this Unix epoch millisecond timestamp
session_idYesSession ID
url_containsNoOnly entries whose page URL contains this substring

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, and the description adds meaningful behavioral context: it is a ring buffer of 500 entries, newest last, and captures output from multiple sources. This goes beyond the annotation and helps the agent understand the tool's memory and ordering behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact, front-loaded with the return shape, and every sentence adds value. The final sentence is a memorable usage heuristic, and there is no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only tool with a fully documented schema, the description covers the return shape, ordering, buffer size, capture scope, and filters. It doesn't explain pagination or what happens when the buffer overflows, but those are minor gaps given the tool's simplicity and the readOnlyHint annotation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all five parameters. The description adds a little extra context by grouping filters and giving an example value for level, but it doesn't substantially extend the schema's meaning.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Read'), a clear resource ('the session's recent page console output'), and the exact return shape. It also distinguishes itself from siblings by focusing on console output, not network, state, or storage.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly says this is 'The fastest way to see WHY a page misbehaves' and gives a concrete workflow: 'click the button, call this, read the error.' It also clarifies that it captures output from page scripts, clicks, evals, and navigation, which helps an agent know when to use it over alternatives like session_network or session_state.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

session_cookiesA
Read-only
Inspect

Export the session's current cookies. Default (meta absent/false) is the full Set-Cookie form ("name=value; Domain=…; Path=/", flags included) — the login-state reuse face that round-trips with session_create's cookies field. meta:true switches to the metadata-only view (#102): [{name, domain, path, secure, httpOnly, sameSite, expires (unix secs, null = session cookie), hostOnly}] with NO values — cookie values are credentials and never leave the server, so "which auth state exists" (is the session cookie Secure, when does it expire, which domains landed) is answerable without holding a single one.

ParametersJSON Schema
NameRequiredDescriptionDefault
metaNoMetadata-only view (#102): {name, domain, path, secure, httpOnly, sameSite, expires, hostOnly} per cookie — no values, ever. For checking what auth state exists; leave false for the value-bearing export that round-trips login state.
session_idYesSession ID

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes well beyond the readOnlyHint annotation: it discloses that cookie values are credentials and never leave the server in meta mode, specifies expires as unix secs with null meaning a session cookie, and enumerates the metadata fields returned. This is exactly the security and output-shape context an agent needs and cannot infer from structured fields.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core action, then the two modes, then the rationale. It is dense and information-rich but somewhat long, with nested parentheticals and examples that slightly impede scanning; nearly every clause earns its place, though.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, and the description compensates by fully specifying both return shapes, including the exact meta field list and the caveat that values are omitted. Nothing needed to call or interpret this tool is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the schema already documents both parameters; the description still adds value by clarifying that the default applies when meta is absent OR false (the schema default is null, which is ambiguous), and by explaining the purpose of each mode for the meta flag.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Export the session's current cookies') and immediately splits the behavior into two clearly named modes (default Set-Cookie form vs meta:true metadata-only view), which is enough to distinguish it from siblings like session_storage, session_state, and session_export.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear selection logic between its own two modes: default is the 'login-state reuse face that round-trips with session_create's cookies field', while meta:true answers 'which auth state exists'. It does not explicitly contrast the tool against sibling tools (e.g. when to prefer this over session_storage/session_export), so guidance is strong but not complete.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

session_createAInspect

Create a persistent interactive browser session for multi-step interaction - clicking, typing, scrolling, reading state across page transitions. Use when the agent must INTERACT with a page (login flows, forms, pagination, click-through) rather than read it once. Returns session_id; persists 8 min idle. With persistent:true the login state survives idle eviction and server restarts - the same session_id revives logged-in.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlNoInitial URL to navigate to (optional). `start_url` is honored as an alias (#115) — callers guessing that name must not land on about:blank.
widthNoInitial viewport width in CSS pixels. Pinned for the session's life (survives navigation) so element rects and media queries anchor to the same layout across every page of the visit.
heightNoInitial viewport height in CSS pixels.
mobileNoMobile device emulation (coarse pointer, no hover) for the initial viewport.
accountNoRun as a named login identity (the multi-account layer): a private cookie jar seeded from the account record, write-back to the account store after every action. Concurrent logins (`taobao-scraper` vs `taobao-publisher`) never clobber each other. The account record survives the session — a later create with the same name picks up the warm jar, plus the record's captured localStorage/sessionStorage (replayed onto the captured origin; an explicit `storage` param wins). 1-64 chars of [a-zA-Z0-9_-].
cookiesNoCookies to inject before navigation: `"name=value"` strings or CDP-style objects `{"name","value","domain",...}`. Lets the session start already logged-in. Round-trips with session_cookies.
storageNoWeb Storage to inject after the initial navigation lands: {"local_storage": {"k":"v"}, "session_storage": {"k":"v"}}. For login states that live in localStorage rather than the cookie jar. An optional "origin": "https://site" key holds the injection until a navigation lands on that origin (a session created without a start_url sits on about:blank first). Round-trips with session_storage.
ttl_secsNoIdle time-to-live in seconds before the session is evicted (default: 480, clamped 60..3600). Raise it for long workflows.
keepaliveNoExempt the session from the idle reaper: it lives until session_close or server exit, so a workflow interrupted by long non-browser steps keeps its login state.
use_proxyNoRoute through proxy (default: false)
persistentNoPersist the login state (cookies + localStorage/sessionStorage + viewport + dialog policy) to the server's local store after every action. If the session idles out — or the whole server restarts — the next call with the same session_id revives it logged-in (storageState-style recovery, no re-login). Explicit session_close drops the snapshot.

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No behavioral annotations exist, so the description carries the whole burden, and it does disclose real lifecycle traits: it returns a session_id, idles out after 8 minutes, and with persistent:true survives idle eviction and server restarts (same session_id revives logged-in). It omits operational concerns such as resource/concurrency cost, limits on concurrent sessions, or that explicit session_close is needed to drop the snapshot.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three front-loaded sentences that go purpose -> when to use -> key lifecycle fact, with no padding. The 'persists 8 min idle' and persistent:true sentences are the highest-value facts and are placed last for emphasis without bloat.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For an 11-parameter, all-optional creator with no output schema, the description covers creation purpose, usage context, the returned session_id, and idle/persistence lifecycle. It would be complete with a pointer to the cleanup/companion tools (session_close, session_list, session_clone) an agent will need next.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all 11 parameters (url/start_url alias, account, cookies, storage, ttl_secs, keepalive, persistent, etc.). The description only reiterates the persistent and idle-time concepts and adds no syntax, alias, or format detail beyond the schema, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Specific verb+resource ('Create a persistent interactive browser session') with the interaction scope spelled out (clicking, typing, scrolling, reading state across page transitions). It also draws a sharp line against the read-only siblings ('rather than read it once'), so an agent can distinguish it from fetch/search/session_clone without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit when-to-use with concrete triggers (login flows, forms, pagination, click-through) and an implicit when-not-to (single read). It does not name the alternative tool (fetch, render_markdown, search) that the agent should reach for in the read-once case, so it stops just short of full routing guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

session_dialogAInspect

Inspect or flip the session's dialog policy for window.alert/confirm/prompt. Dialogs never block the page: each is auto-answered (default dismiss) and logged into session_console at level "dialog". action "list" reports {policy, prompt_text, dialogs}; "accept" makes subsequent confirm() true and prompt() return prompt_text (or the call's default argument); "dismiss" restores the default.

ParametersJSON Schema
NameRequiredDescriptionDefault
actionYes"list" reports the policy and dialog history; "accept"/"dismiss" set the answer applied to subsequent window.confirm/prompt calls (alert is always logged, never blocking).
session_idYesSession ID
prompt_textNoWith action "accept": text window.prompt returns once accepted (omitted keeps the current text).

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With minimal annotations (only title), the description carries the full burden. It discloses that dialogs never block the page, are auto-answered with default dismiss, and are logged to session_console at level 'dialog'. It also details the exact behavioral changes for 'accept' and 'dismiss' actions, including how prompt() returns prompt_text.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two well-structured sentences pack all necessary information with zero fluff. The purpose is front-loaded, and behavioral details follow logically. Every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers the tool's purpose, all actions, the non-blocking behavior, logging, and the return structure for 'list'. While it doesn't detail error handling or the exact contents of 'dialogs', the provided information is sufficient for an agent to call the tool correctly given the lack of output schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description adds meaningful context by explaining how 'action' values behave and how 'prompt_text' interacts with the 'accept' action. This goes beyond the schema's terse descriptions, justifying a slightly higher score.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states a specific verb (inspect or flip) and resource (the session's dialog policy for window.alert/confirm/prompt). It distinguishes itself from siblings by focusing solely on dialogs, which no other sibling handles.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explains the three actions and their effects, making it clear when to use this tool (to inspect or change dialog policy). It doesn't explicitly mention alternatives because none exist, but it does refer to session_console for logging, which provides context without contradicting usage.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

session_dragAInspect

Drag the mouse from one viewport position to another: press at from, steps mousemove events, release at to. The trajectory is humanized by default (eased velocity, wobble, jittered timing, overshoot) — the shapes anti-bot checks score for; pass humanize:false for exact linear interpolation. Moves AMarker-style drag targets, canvas selections and captcha sliders that only track while the pointer travels.

ParametersJSON Schema
NameRequiredDescriptionDefault
toYesWhere to release it
fromYesWhere to press the mouse button down
stepsNoInterpolated mousemove events between from and to. Default: 24 with humanize on, 10 without
delay_msNoMean delay between moves in ms (default 18 humanized / 30 linear) — per-step timing is jittered around this when humanizing
humanizeNoHumanize the trajectory: minimum-jerk easing, perpendicular wobble, timing jitter, grip/settle pauses, occasional hesitation and overshoot-and-correct. Set false when a test/tool needs exact linear interpolation. Default: true
session_idYesSession ID

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations are minimal (only a title), so the description carries the full burden. It discloses the humanized trajectory behavior (eased velocity, wobble, jittered timing, overshoot) and explains that it exists to pass anti-bot checks. It also clarifies the behavior of the `humanize` flag. This is substantial behavioral context beyond what the schema provides.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences with zero filler. The first sentence front-loads the core action and mechanics; the second adds crucial behavioral nuance (humanization) and use-case examples. Every word earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex action with six parameters, the description covers the purpose, the humanization behavior, and typical use cases. It does not explicitly mention the session context, but the `session_id` parameter and tool name make that clear. The lack of an output schema is acceptable since the action is a gesture with no meaningful return value. The description is sufficiently complete for an agent to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all six parameters are already documented in the schema. The description adds some contextual meaning (e.g., that `steps` and `humanize` relate to trajectory smoothing), but it does not significantly enhance parameter understanding beyond what the schema provides. Baseline 3 is appropriate given full coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Drag'), a specific resource ('the mouse'), and precise mechanics (press at `from`, mousemove events, release at `to`). It also names concrete target types (AMarker-style drag targets, canvas selections, captcha sliders), which clearly distinguishes it from sibling click tools like session_click and session_click_xy.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear context for when to use the tool—when dragging is needed—and even gives examples of drag targets that only track while the pointer travels. However, it does not explicitly mention when NOT to use it or point to alternative tools (e.g., using session_click for simple clicks). Since the sibling list includes click tools, this guidance is implied but not explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

session_evalAInspect

Execute arbitrary JavaScript in a live browser session and return the result. Runs in the session's current page, so DOM mutations, globals and storage persist across calls — unlike the stateless eval tool, which loads its own throwaway page each call. Script-driven navigation moves the session's URL. JS exceptions are reported with name, line/column and stack.

ParametersJSON Schema
NameRequiredDescriptionDefault
scriptYesJavaScript code to execute
session_idYesSession ID
timeout_msNoAwait budget for the script's promise in ms (default 5000, clamped 100..120000). Pass a larger budget for slow page-side work such as uploads through the page's own fetch; on expiry the tool errors with EVAL_TIMEOUT (the script may still be running) instead of returning a null result.

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

There are no relevant annotations beyond the title, so the description carries the full burden. It discloses persistence of DOM mutations, globals, and storage across calls, the navigation side effect, and exception reporting details. It does not explicitly warn that arbitrary JavaScript can be destructive to the session, but the persistence statement effectively conveys the risk.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is four concise sentences, each adding distinct value: core action, persistence behavior, contrast with eval, and error reporting. There is no filler or redundancy, and the most important information is front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity, the description covers execution context, side effects, and error behavior, which is substantial for a 3-parameter tool with no output schema. It does not detail the exact return serialization format, but for an eval tool that returns a JavaScript result, this is a minor gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents script, session_id, and timeout_ms including defaults and clamping behavior. The description adds execution-environment context but not new parameter-level semantics, so a baseline score of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: it executes arbitrary JavaScript in a live browser session and returns the result. It also explicitly distinguishes itself from the sibling eval tool by contrasting persistent session state with a throwaway page, making the tool's purpose unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description clearly tells the agent when to prefer this tool over the alternative eval: use session_eval when running in the current session page where mutations, globals, and storage persist, and use eval when a stateless isolated page is desired. It also notes that script-driven navigation changes the session URL, which is critical for choosing the tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

session_exportA
Read-only
Inspect

Export a browser session's recorded action log. Format "bash" (default) returns a runnable curl script that replays every recorded action (navigate/click/input/scroll/eval) against a fresh session on this server — hand it to a shell or cron, zero model tokens. Format "jsonl" returns the raw action log, one JSON object per line. Format "json" returns a flow.json document — the same recording as editable ops ({op, args}) with cookies/storage stripped — that flow_run replays server-side.

ParametersJSON Schema
NameRequiredDescriptionDefault
formatNoOutput format: "bash" (default) renders a runnable curl script that replays every recorded action against a fresh session; "jsonl" returns the raw action log, one JSON object per line; "json" returns a flow.json document (editable ops, cookies/storage stripped) for replay via flow_run
session_idYesSession ID

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint=true and no destructive consequences, so the description already starts with a safety baseline. It goes further by disclosing format-specific behaviors: bash produces a runnable curl script that replays against a fresh session, and json strips cookies/storage into editable ops — adding meaningful behavioral context beyond the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The scope statement is front-loaded and each subsequent sentence covers a distinct format with its purpose and output shape — no filler. The content is dense but appropriately sized for a tool that needs to explain three formats.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description carries the burden of explaining return values and does so for all three formats, including the replay mechanics and stripped data. It doesn't cover error cases (e.g., missing session), but for a read-only export tool the essential usage guidance is present.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the schema already documents session_id and format in detail. The description reinforces the meaning of each format value (bash default and runnable, jsonl raw log, json editable for flow_run), adding semantic color that helps the agent select the right value without needing to read the full schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States the verb 'Export' with the specific resource 'a browser session's recorded action log,' which clearly distinguishes the tool from the many session_* siblings. It also contrasts with flow_run by explaining that json output is replayed server-side, so an agent can tell export apart from the sibling that consumes it.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides clear context on when each format is appropriate: bash for shell/cron with zero model tokens, json for replay via flow_run, and jsonl for the raw action log. It names flow_run as the alternative tool for replay, though it doesn't exhaustively cover when not to use each format.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

session_inputAInspect

Type text into an input/textarea element by its index (from session_state output), focusing it and dispatching input/change events. events:"full" is the complete human typing gesture: per-character keydown/keypress/input/keyup cycles, trailing change, then blur — the tail blur commits on forms that save in onBlur (React capture listeners, #100). In full mode a newline in the text types as the Enter key (key:'Enter', keyCode 13 — "type + Enter submit"): a textarea keeps the newline in its value, a single-line input stays single-line, the page's Enter listeners fire before the blur, and if the key wasn't canceled the input's form submits implicitly (Chrome's default action; preventDefault on the key or submit events vetoes it, so chat inputs that swallow Enter stay unsubmitted). Any characters — newlines, quotes, backslashes — ride through as-is. A refused fill answers filled:false with a reason (readonly/disabled/detached/no-element/wrong-tag/script-error) instead of a silent write. Hidden inputs are legitimate targets and are filled normally.

ParametersJSON Schema
NameRequiredDescriptionDefault
textYesText to type into the input field
indexYesElement index (from /state output)
eventsNoEvent fidelity: "full" types one character at a time with a keydown/keypress/input/keyup cycle per character, for pages whose listeners key on keyboard events (e.g. keypress-Enter login forms). Default fires a single input+change pair after the value is set.
session_idYesSession ID

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations carry only a title, so the description bears the full burden and delivers: the exact event sequence in full mode, the trailing blur/onBlur commit behavior, Enter/newline semantics including implicit form submission and veto conditions, passthrough of arbitrary characters, and an explicit refusal contract (filled:false with readonly/disabled/detached/no-element/wrong-tag/script-error reasons). This is exceptionally rich behavioral disclosure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose is front-loaded and virtually every clause carries substantive information (blur commit, Enter handling, refusal reasons). It is a dense single paragraph with heavy parenthetical nesting, which slightly impairs scanability, but there is little wasted text.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema and no annotations, the description compensates fully: it documents side effects (focus, dispatched events, blur), edge cases (hidden inputs, newlines/quotes/backslashes), Enter-submit behavior, and the failure response shape. An agent has everything needed to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description goes beyond the schema by elaborating what 'events:"full"' actually does (per-character keydown/keypress/input/keyup cycles, trailing change, blur) and by clarifying the index origin, adding real meaning over the schema's terse enum-like description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource — 'Type text into an input/textarea element by its index' — and clarifies the index source (session_state output), which cleanly separates it from siblings like session_click or session_eval. An agent can identify what it does without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explains what the 'events:"full"' mode is for ('type + Enter submit', keypress-Enter login forms) and where the index comes from, which is implicit usage guidance. However, it never explicitly states when to choose this tool over session_eval, session_click, or other input paths, and offers no when-not-to-use conditions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

session_listA
Read-only
Inspect

List live browser sessions with idle age and the time left before auto-eviction. Use to discover a session to reuse instead of creating a new one; sessions expire after 8 min idle.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The annotations already declare readOnlyHint=true, so the read-only nature is covered. The description adds meaningful context about live sessions, idle age, auto-eviction, and the 8-minute idle limit, which helps the agent understand the lifecycle without contradicting annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two purposeful sentences. The first states the core function and output fields, and the second gives usage guidance and expiration policy. No filler or redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter, read-only list tool with no output schema, the description is complete. It tells the agent what the tool returns (idle age, eviction time), when to use it, and a key behavioral constraint (8-minute idle expiry). Nothing essential is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, so the schema is trivially complete at 100% coverage. Baseline for 0 params is 4; the description adds no parameter details because none are needed, and that is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('List') and resource ('live browser sessions') and clearly states what information is returned ('idle age and the time left before auto-eviction'). It distinguishes itself from session_create by framing the tool as a discovery mechanism for reuse.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly states when to use the tool: 'discover a session to reuse instead of creating a new one'. It also provides a critical operational condition ('sessions expire after 8 min idle'), giving the agent a clear decision rule.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

session_navigateAInspect

Navigate a browser session to a new URL.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYesURL to navigate to
session_idYesSession ID

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description discloses the core behavior—navigation to a URL—which is the main observable effect. However, annotations provide only a title, so the description carries the full burden of behavioral disclosure and does not mention whether the tool waits for page load, how it handles invalid URLs, or whether it returns after navigation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, compact sentence that front-loads the verb and target resource. Every word is necessary, and there is no filler or redundant context.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple two-parameter action, the schema covers the parameters completely, and the description gives the essential operation. Still, without an output schema and without behavioral annotations, an agent lacks return-value and postcondition context that would make the tool fully self-sufficient.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema already fully documents both parameters with descriptions: 'URL to navigate to' and 'Session ID'. The description repeats the URL concept without adding any additional meaning, so it stays at the baseline for 100% schema coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific action—'navigate'—against a specific resource, a browser session, and clarifies the target is a URL. This is clear enough to distinguish session_navigate from sibling tools like session_click, session_eval, or session_wait.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool instead of alternatives. It neither mentions prerequisites, such as an active session, nor contrasts itself with closely related siblings like fetch, click, or session_wait that could also change a browser's state.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

session_networkA
Read-only
Inspect

Read the session's network request log. filter="media" extracts playback/stream URLs (m3u8/HLS, mp4, dash, flv...) actually requested by the page's player at runtime - the reliable way to get a real video link, since links embedded in page HTML are often decoys. Media elements and player iframes the engine never fetches (video/audio/source/iframe src) are merged in as candidates: via="network" entries are confirmed requests, via="dom" entries are candidates carrying their tag (iframes = kind "iframe", navigate into them to sniff). Default returns every request as compact rows (method/url/status/type/size/nav); include_headers=true adds each request's outbound header set to its row (#97). Each row's nav is the navigation generation that issued it and the payload's top-level nav is the current one - a lower row nav belongs to an earlier (e.g. timed-out) navigation attempt (#101). Navigate to the video page first, let it load, then call this.

ParametersJSON Schema
NameRequiredDescriptionDefault
filterNo"media" extracts playback/stream links (m3u8/HLS, mp4, dash, ...) from the requests the page actually issued - the reliable way to get a real video link, since URLs embedded in page HTML are often decoys. Media elements and player iframes the engine never fetches (video/audio/source src, iframe src) are merged in as candidates: entries carry via="network" (confirmed requests) or via="dom" (candidates, with their tag; iframes surface as kind "iframe" - player pages to navigate or sniff inside, not playable URLs). Omit to list every request as compact rows.
session_idYesSession ID
url_containsNoNarrow the `xhr` array to URLs containing this substring.
body_max_charsNoPer-body character cap for the `xhr` array (default 4000).
include_bodiesNoAdd an `xhr` array of background API responses (the page's own fetch/XHR traffic with retained bodies) alongside the request rows — the page's API face is often the cleanest structured read of its data.
include_headersNoAdd each request's outbound header set to its row (`headers` on the compact rows, `request_headers` on the `xhr` rows) — what the page's JS actually sent, signed customs like x-s/x-s-common included. Off by default: headers can carry tokens.

TDQS

A4.3/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare readOnlyHint=true, so the description carries most of the behavioral load and does so richly: default row shape (method/url/status/type/size/nav), via="network" vs via="dom" confirmation semantics, iframe kind handling, the nav generation meaning (plus the #97/#101 issue context), and the token warning around include_headers. This is substantial context beyond the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the purpose well, but the body is a dense run-on with many parentheticals and issue references, and the filter parameter is described twice (description and schema) almost verbatim. Most content is useful given the tool's complexity, but it is longer than it needs to be.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description compensates by describing the returned row structure, the merge of network/DOM candidates, and the headers addition. Combined with 100% parameter coverage and readOnlyHint, an agent has everything needed to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all six parameters. The description largely restates the filter and include_headers semantics already present in the schema; its only genuinely additive parameter context (nav generation meaning) concerns the output, not an input. Baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Opens with a specific verb+resource ('Read the session's network request log') and immediately carves out distinct scope via the filter="media" behavior and the network/DOM merge semantics. An agent can distinguish it from siblings like session_console, session_storage, and fetch without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives a clear workflow prerequisite ('Navigate to the video page first, let it load, then call this') and explains when to prefer the media filter over the default listing. It does not name explicit alternative tools (e.g. fetch, session_console) or state when not to use it, so it falls just short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

session_preloadAInspect

Replace the session's document-start preload group (empty array clears). Sources run before each new document's own scripts — including inline tags — which is the only hook that beats pages whose signing layer captures window.fetch/XHR natives at parse time (xhs's inline jsvmp). Set before the first navigate; applies to every navigation from then on.

ParametersJSON Schema
NameRequiredDescriptionDefault
recipeNoBuiltin recipe name appended after `scripts` (e.g. "xhs-sign"): the maintained document-start wrapper for that site — same source the engine auto-mounts on its navigations, without pasting JS. Unknown names are an error, not a silent no-op.
scriptsNoFull JS sources, in order. Sources run before each new document's own scripts (including inline ones) — the only hook that beats pages whose signing layer captures window.fetch/XHR natives at parse time. `[]` clears the group. Set before the session's first navigate and it applies to every navigation from then on.
session_idYesSession ID

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations beyond a title, so the description carries the full load — and it does disclose key traits: replace semantics, that an empty array clears, that sources run before the document's inline scripts, and that the group persists across all navigations. It omits what happens if a source throws and whether the current document is affected.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the action ("Replace...") and the clearing rule, then the timing rationale. It is denser than needed — em-dashes and stacked subclauses — but every clause is informative and there is no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a stateful session-mutation tool with no output schema and only a title annotation, the description supplies the timing, persistence, and hook-mechanism context an agent needs to call it correctly. Error handling and return behavior are the only notable gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the schema already documents session_id, recipe, and scripts in detail. The description reinforces the ordering/clearing semantics of scripts but adds no syntax or format beyond what the schema states, matching the baseline for high-coverage schemas.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: "Replace the session's document-start preload group." The mechanism description (runs before the document's own scripts, including inline ones) clearly signals how this differs from an eval/post-load hook. It does not name a sibling explicitly, so it falls short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives concrete timing guidance — "Set before the first navigate; applies to every navigation from then on" — and explains why a document-start hook is needed (beating parse-time native capture). This is clear usage context, but it never names alternatives (e.g. session_eval) or states when this is the wrong choice.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

session_screenshotAInspect

Screenshot the session's CURRENT DOM state (mutations from clicks/evals included) as a base64 PNG via the built-in renderer. Width/height default to the session's viewport, so session_viewport + session_screenshot shows the responsive layout. Returns {url, width, height, image_base64, format}.

ParametersJSON Schema
NameRequiredDescriptionDefault
dprNoDevice pixel ratio (#185): rasterize at device resolution (Retina/HiDPI sharp) — the bitmap comes back width·dpr × height·dpr with CSS geometry intact. Clamped to 1.0..=3.0 (default: 1.0)
widthNoRender width in CSS pixels; defaults to the session's current viewport
heightNoRender height in CSS pixels; defaults to the session's current viewport
selectorNoCSS selector: capture only that element's box
full_pageNoCapture the full scrollable page instead of the viewport (default: false)
session_idYesSession ID
selector_allNoWith selector, capture every match (default: first match only)

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations carry only a title, so the description does most of the work: it discloses that it uses the built-in renderer, that pending DOM mutations are included, that sizing falls back to the viewport, and it enumerates the response fields. Missing only operational notes like rate limits or failure behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences, front-loaded with the core action and return shape; no filler. Slightly dense but every clause adds information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, yet the description names the exact return object keys {url, width, height, image_base64, format}, and the 7 parameters are fully specified in the schema, so an agent has everything needed to call and consume it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so dpr clamping, selector, selector_all and full_page are already well documented in the schema. The description restates the width/height viewport default rather than adding new meaning, so baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Specific verb + resource: capture the session's CURRENT DOM state as a base64 PNG. The 'mutations from clicks/evals included' clause makes it clear this reflects live state, not a static render, which differentiates it from sibling render_* tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Offers one concrete usage pattern (pair with session_viewport to inspect responsive layout), which is implied guidance, but never states when NOT to use it or names an alternative among the many session_* siblings (e.g. session_state, session_export).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

session_scrollAInspect

Scroll the page up or down by a number of viewport-heights.

ParametersJSON Schema
NameRequiredDescriptionDefault
amountNoScroll amount in viewport-heights (default: 3)
directionNoScroll direction: "up" or "down" (default: down)down
session_idYesSession ID

TDQS

A3.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With only a title annotation and no readOnly/destructive hints, the description carries the burden of behavioral disclosure. It does state the core behavior (scrolling up/down by viewport-heights), but it omits side effects or limitations such as behavior at the top/bottom of the page, whether scrolling is relative to the current position, or whether the action waits for rendering to settle. This is a reasonable but not thorough disclosure for a simple scroll operation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single, front-loaded sentence that conveys the essential action and unit of measurement with zero filler. Every word earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple tool with fully documented parameters, the description covers the primary action but leaves gaps: no mention of return behavior or output (no output schema exists), no mention of session applicability beyond the parameter, and no guidance on repeated or large scroll amounts. It is adequate but not complete for an agent with no other context.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all parameters (amount, direction, session_id) are already documented with defaults and formats. The description adds no new parameter-level meaning beyond restating 'up or down' and 'viewport-heights,' which the schema already captures. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Scroll'), a clear resource ('the page'), and a precise unit of measure ('viewport-heights'), which fully captures the tool's function. It is readily distinguishable from sibling tools like session_navigate, session_click, and session_viewport, which do not involve scrolling by viewport heights.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage: use this tool when you need to scroll the currently active page vertically. However, it provides no explicit guidance on when to prefer this tool over alternatives, no mention of prerequisites (e.g., an active session), and no exclusions or edge-case conditions. Usage is self-evident from the action but not explicitly framed.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

session_set_filesAInspect

Select files on a file input programmatically (Playwright setInputFiles semantics): builds File objects from base64 content, assigns them to input.files, then dispatches input+change so framework onChange handlers fire. Selector-addressed because file inputs are often hidden and absent from the session_state index.

ParametersJSON Schema
NameRequiredDescriptionDefault
filesYesFiles to select
selectorYesCSS selector for the file input, e.g. "input[type=file]". File inputs are often hidden, so this is selector-addressed rather than using the /state index.
session_idYesSession ID

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations provide only a title, so the description carries the behavioral burden. It transparently discloses the full mechanism: builds File objects from base64, assigns them to input.files, and dispatches input+change events so framework onChange handlers fire. It does not mention failure modes or return behavior, but the core side effects are clear.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two compact sentences front-load the purpose, then efficiently explain the mechanism and the selector rationale. Every clause adds value, with no redundant or filler content.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a three-parameter tool with complete schema coverage and no output schema, this description gives enough behavioral context for correct invocation: the selector approach, event dispatch, and framework-handler implications. It does not cover return values or failure conditions, but that is a minor gap given the detailed mechanism already provided.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all parameters and nested fields. The description adds a useful rationale for the selector parameter and confirms the base64 construction path, but it does not materially extend the schema's parameter documentation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific action—'Select files on a file input programmatically'—and anchors it to Playwright setInputFiles semantics. It also distinguishes itself from state-indexed sibling tools by explaining that file inputs are often hidden and absent from the session_state index, so selector addressing is required.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Clear context is provided: this tool is for attaching files to file inputs, and it explains why selector addressing is used instead of the session_state index. It does not explicitly name an alternative or give when-not-to-use conditions, but the intended scenario is readily inferable.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

session_stateA
Read-only
Inspect

Get the current page state as an indexed list of interactive elements. Returns compact text with [N] indexes for use with click/input tools.

ParametersJSON Schema
NameRequiredDescriptionDefault
session_idYesSession ID

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, so the safety profile is covered. The description adds useful behavioral context by specifying the return format: 'compact text with [N] indexes.' It does not describe edge cases or limitations, but for a simple read-only tool with annotation support, this is adequate.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two concise sentences with no filler. The core behavior ('Get the current page state') is front-loaded, and the return-format detail is provided efficiently. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter read-only tool with no output schema, the description provides enough information to call it correctly: what it returns, the format ([N] indexes), and its intended downstream use. It does not describe the exact structure of an element entry, but this is a minor gap given the simplicity of the operation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The only parameter, session_id, is already documented in the schema with a description, giving 100% schema description coverage. The tool description does not add additional meaning about the parameter. Baseline 3 applies because the schema carries the parameter documentation burden.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'Get the current page state as an indexed list of interactive elements.' It clearly differentiates this from the action-oriented sibling tools (session_click, session_input) by positioning itself as a read-only state retrieval tool. The mention of 'for use with click/input tools' strengthens its identity as the preparatory inspection tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description conveys when to use the tool: before invoking click/input tools, to obtain the [N] indexes needed for those actions. It does not explicitly name alternative tools or state exclusions, but the context is clear and sufficient for an agent to select it over mutation-focused siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

session_storageA
Read-only
Inspect

Snapshot the session's localStorage/sessionStorage for the current origin: {url, local_storage, session_storage}. Feed it back via session_create's storage field to restore a logged-in state in a new session — the half of login state that cookies can't carry (many sites keep the session token in localStorage). Call before the session idles out.

ParametersJSON Schema
NameRequiredDescriptionDefault
session_idYesSession ID

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint: true, and the description reinforces this with 'Snapshot', which implies no mutation. The description adds valuable behavioral context: the exact data returned (url, local_storage, session_storage), its purpose in restoring login state, and a timing constraint (before idle). No contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences that are information-dense and free of redundancy. It front-loads the main action and return shape, then provides the use case and a timing note. Each clause earns its place, though it could be slightly shorter without losing critical context.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given there is no output schema, the description compensates by specifying the return fields. It also explains the intended workflow (session_create) and why the tool matters. While it doesn't cover edge cases like empty storage or error conditions, for a simple read-only snapshot tool this is sufficient for an agent to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has a single parameter 'session_id' with 100% coverage (description: 'Session ID'). The tool description does not add any additional meaning or constraints for this parameter beyond what the schema provides. Per the baseline rule for high schema coverage, a score of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb 'Snapshot' and clearly identifies the resource (localStorage/sessionStorage for the current origin). It also specifies the return shape {url, local_storage, session_storage}, which makes the tool's function unambiguous and distinct from generic session tools like session_state or session_export.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides concrete usage guidance: feed the result back via session_create's 'storage' field to restore a logged-in state, and call it before the session idles out. It also explains why it's needed (cookies can't carry this half of login state). It doesn't explicitly mention alternatives or exclusions, but the use case is well-defined.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

session_verdictA
Read-only
Inspect

One call answers "where did this session land": verdict is one of challenge (risk control engaged — punish page or a 200-status API body that swallowed the wall; the response carries a handoff instruction for a human to solve it in the live view), captcha (explicit CAPTCHA interstitial), login (bounced to a login form — auth expired), empty, landed (normal 2xx content page), or unknown (couldn't classify — read facts; a non-2xx main document lands here with facts.doc_status carrying the number, so a zhihu-style burst 403 is branchable). The facts sheet also carries challenge_events, requests, console_errors and the fired signals. Pure code over signals the engine already holds (current URL, risk-control rows, main document status/size, console errors) — no screenshots, no page evals, single-digit milliseconds. Verdict observes, it never bypasses.

ParametersJSON Schema
NameRequiredDescriptionDefault
session_idYesSession ID

TDQS

A4.1/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The readOnlyHint annotation is reinforced and extended with valuable behavioral detail: it is pure code over existing signals, takes no screenshots, runs no page evals, executes in single-digit milliseconds, and never bypasses protections. The description also discloses that challenge responses carry a handoff instruction for human resolution, adding context beyond the annotation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with the core purpose ('One call answers where did this session land') and then delivers a dense but relevant enumeration of verdicts, facts, and behavior. It is long but every clause earns its place; a list format would improve scannability, but the content is not wasteful.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description carries the burden of explaining return semantics, and it does so thoroughly: all verdict categories, the facts sheet contents, handling of non-2xx documents, and the safety/performance profile. For a one-parameter classification tool, an agent has enough context to call it correctly and interpret the result.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage for the single parameter is 100%, so the schema already documents session_id. The description does not add any additional meaning about the parameter itself, which matches the baseline of 3 when structured schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's job: classify where a session landed, and enumerates the possible verdicts (challenge, captcha, login, empty, landed, unknown) with concrete meanings. It distinguishes itself from action tools by saying it 'observes, never bypasses' and uses 'no screenshots, no page evals,' but it does not explicitly differentiate itself from similar observation tools like session_state or session_network.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear context for when to call it: after a session action, to answer 'where did this session land' and to branch on the resulting verdict/facts, including handling non-2xx documents via facts.doc_status. It does not explicitly name alternative tools or state when not to use it, but the intended use case is strongly implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

session_viewportAInspect

Set the session's viewport (device emulation): scripts see innerWidth/innerHeight move, media queries like (max-width: 600px) re-evaluate, element rects re-anchor, and mobile=true flips pointer/hover matchMedia answers to coarse/none. Omitted width/height keeps the current value.

ParametersJSON Schema
NameRequiredDescriptionDefault
widthNoViewport width in CSS pixels; omit to keep the current width
heightNoViewport height in CSS pixels; omit to keep the current height
mobileNoMobile emulation: matchMedia answers pointer:coarse / hover:none and navigator.maxTouchPoints reports 5 (default: false)
session_idYesSession ID

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With annotations only providing a title, the description carries the behavioral burden and does so well: it reveals that scripts observe innerWidth/innerHeight changes, media queries re-evaluate, element rects re-anchor, and mobile flips matchMedia pointer/hover answers. It does not mention session prerequisites or persistence, but the core state-changing behavior is transparent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two dense sentences deliver high-signal information with no filler. The main action is front-loaded, followed by the most important behavioral effects and the omitted-parameter rule, making it easy for an agent to parse quickly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a state-mutating tool with a complete input schema and no output schema, the description explains the behavioral consequences well enough to invoke it correctly. It lacks explicit return-value or error/edge-case information, but those are not essential for a setter of this kind.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and each parameter already has a clear description. The tool description reinforces the semantics of omitted width/height preserving current values and describes the mobile flag's effect, but it does not add meaning substantially beyond what the schema already provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Set') and resource ('the session's viewport'), and clearly distinguishes device emulation from a simple resize by enumerating the observable effects on scripts, media queries, and element rects. It also clarifies the mobile mode behavior, leaving no ambiguity about what the tool does.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear context for when to use this tool: when the session needs device emulation or viewport resizing, including media-query re-evaluation and pointer/hover emulation. It does not explicitly name alternatives or state when not to use it, but no direct sibling appears to overlap with this capability.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

session_waitA
Read-only
Inspect

Wait until a CSS selector matches or a JS predicate turns truthy, with a timeout. The page's event loop keeps running while waiting (fetches, timers, promise chains progress), so this replaces blind sleeps for async content: navigate, session_wait for '.price-card', then click/read. Returns {matched, elapsed_ms, detail:{tag,text} or the predicate value}; errors with timeout ... naming the selector/predicate on expiry. Exactly one of selector/predicate.

ParametersJSON Schema
NameRequiredDescriptionDefault
selectorNoCSS selector to wait for (e.g. ".price-card")
predicateNoJS expression polled until truthy (e.g. "document.querySelectorAll('.card').length >= 3")
session_idYesSession ID
timeout_msNoGive up after this many milliseconds (default: 10000, max: 120000)

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the readOnlyHint annotation, the description reveals that the page's event loop continues running (fetches, timers, promise chains progress), which is important behavioral context. It also discloses the return shape, error text format, and the exactly-one-of constraint. This is substantial added value beyond the annotation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences, tightly packed with purpose, behavior, usage example, return format, error behavior, and the exclusivity constraint. No filler words; key information is front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with no output schema, the description fully covers what an agent needs: what it waits for, the async behavior, a usage example, the return shape, error format, and the parameter constraint. The only minor omission is polling interval, but that is not essential for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. The description adds a critical non-schema constraint: 'Exactly one of selector/predicate.' It also explains the meaning of selector and predicate in the wait context and what the return detail field contains depending on which is used, going beyond the schema's basic definitions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states a specific action: wait until a CSS selector matches or a JS predicate is truthy, with a timeout. It gives a concrete example workflow (navigate, session_wait for '.price-card'), which distinguishes it from blind sleeps, but it does not explicitly compare itself to sibling tools like session_eval or session_console.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly positions the tool as a replacement for blind sleeps in async scenarios: 'so this replaces blind sleeps for async content'. It provides a typical usage pattern, which gives clear context for when to use it. It stops short of stating exclusions (e.g., when not to use) and does not name alternative tools, so it loses the fifth point.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 1 tool update
    • Changedsession_preload1 field changed
      • addedInput schema / properties / recipe
        Added value: +{
        +  "default": null,
        +  "description": "Builtin recipe name appended after `scripts` (e.g. \"xhs-sign\"): the\nmaintained document-start wrapper for that site — same source the\nengine auto-mounts on its navigations, without pasting JS. Unknown\nnames are an error, not a silent no-op.",
        +  "type": [
        +    "string",
        +    "null"
        +  ]
        +}
  2. 1 tool update
    • Changedsession_screenshot1 field changed
      • addedInput schema / properties / dpr
        Added value: +{
        +  "description": "Device pixel ratio (#185): rasterize at device resolution (Retina/HiDPI\nsharp) — the bitmap comes back width·dpr × height·dpr with CSS geometry\nintact. Clamped to 1.0..=3.0 (default: 1.0)",
        +  "format": "float",
        +  "type": [
        +    "number",
        +    "null"
        +  ]
        +}
  3. 1 tool update
    • Addedflow_search
  4. 4 tool updates
    • Addedflow_install
    • Changedsession_cookies1 field changed
      • addedInput schema / properties / meta
        Added value: +{
        +  "default": null,
        +  "description": "Metadata-only view (#102): {name, domain, path, secure, httpOnly,\nsameSite, expires, hostOnly} per cookie — no values, ever. For\nchecking what auth state exists; leave false for the value-bearing\nexport that round-trips login state.",
        +  "type": [
        +    "boolean",
        +    "null"
        +  ]
        +}
    • Changedsession_create2 fields changed
      • changedInput schema / properties / account / description
        Previous value: -"Run as a named login identity (the multi-account layer): a private\ncookie jar seeded from the account record, write-back to the account\nstore after every action. Concurrent logins (`taobao-scraper` vs\n`taobao-publisher`) never clobber each other. The account record\nsurvives the session — a later create with the same name picks up\nthe warm jar. 1-64 chars of [a-zA-Z0-9_-]."New value: +"Run as a named login identity (the multi-account layer): a private\ncookie jar seeded from the account record, write-back to the account\nstore after every action. Concurrent logins (`taobao-scraper` vs\n`taobao-publisher`) never clobber each other. The account record\nsurvives the session — a later create with the same name picks up\nthe warm jar, plus the record's captured localStorage/sessionStorage\n(replayed onto the captured origin; an explicit `storage` param\nwins). 1-64 chars of [a-zA-Z0-9_-]."
      • changedInput schema / properties / storage / description
        Previous value: -"Web Storage to inject after the initial navigation lands:\n{\"local_storage\": {\"k\":\"v\"}, \"session_storage\": {\"k\":\"v\"}}. For login\nstates that live in localStorage rather than the cookie jar.\nRound-trips with session_storage."New value: +"Web Storage to inject after the initial navigation lands:\n{\"local_storage\": {\"k\":\"v\"}, \"session_storage\": {\"k\":\"v\"}}. For login\nstates that live in localStorage rather than the cookie jar. An\noptional \"origin\": \"https://site\" key holds the injection until a\nnavigation lands on that origin (a session created without a\nstart_url sits on about:blank first). Round-trips with\nsession_storage."
    • Changedsession_network1 field changed
      • addedInput schema / properties / include_headers
        Added value: +{
        +  "default": null,
        +  "description": "Add each request's outbound header set to its row (`headers` on the\ncompact rows, `request_headers` on the `xhr` rows) — what the page's\nJS actually sent, signed customs like x-s/x-s-common included. Off by\ndefault: headers can carry tokens.",
        +  "type": [
        +    "boolean",
        +    "null"
        +  ]
        +}
  5. 1 tool update
    • Changedfetch2 fields changed
      • changedInput schema / $defs / RenderTier / oneOf
        Previous value: -[
        -  {
        -    "const": "auto",
        -    "description": "HTTP-direct first, fall back to diting browser. (default)",
        -    "type": "string"
        -  },
        -  {
        -    "const": "http",
        -    "description": "Pure HTTP, no V8/JS. Fastest; misses JS-rendered content.",
        -    "type": "string"
        -  },
        -  {
        -    "const": "obscura",
        -    "description": "Always use the diting browser (current behaviour pre-tiering).\n\"browser\" is accepted as an alias — agents guess it before \"obscura\".",
        -    "type": "string"
        -  }
        -]New value: +[
        +  {
        +    "const": "auto",
        +    "description": "HTTP-direct first, fall back to diting browser. (default)",
        +    "type": "string"
        +  },
        +  {
        +    "const": "http",
        +    "description": "Pure HTTP, no V8/JS. Fastest; misses JS-rendered content.",
        +    "type": "string"
        +  },
        +  {
        +    "const": "browser",
        +    "description": "Always use the diting browser (current behaviour pre-tiering).\nWire name is \"browser\". \"obscura\" is still accepted and not advertised.",
        +    "type": "string"
        +  }
        +]
      • changedInput schema / properties / render_tier / description
        Previous value: -"Rendering strategy: \"auto\" (default), \"http\", or \"obscura\""New value: +"Rendering strategy: \"auto\" (default), \"http\", or \"browser\""
  6. 1 tool update
    • Changedsession_create1 field changed
      • changedInput schema / properties / url / description
        Previous value: -"Initial URL to navigate to (optional)"New value: +"Initial URL to navigate to (optional). `start_url` is honored as an\nalias (#115) — callers guessing that name must not land on about:blank."
  7. 1 tool update
    • Addedsession_preload

Related MCP Connectors

Related MCP Servers

  • A
    license
    A
    quality
    B
    maintenance
    Provides an MCP-native agent browser that enables autonomous agents to perceive and interact with web pages through stealth browsing, identity borrowing, and WAAP detection.
    15
    1
    MIT
  • A
    license
    Not graded
    quality
    C
    maintenance
    Provides a headless Chromium browser through MCP, enabling AI agents to browse JavaScript-rendered pages, search the web, capture screenshots, extract tables and data, and run stateful multi-step interactions like clicking, typing, and form submission.
    2,021,532 npm
    1
    MIT
  • F
    license
    Not graded
    quality
    C
    maintenance
    Headless browser automation for LLM agents via REST API or MCP tools. Enables navigating pages, reading structured content, clicking elements, filling forms, and executing JavaScript.
    -
  • A
    license
    Not graded
    quality
    C
    maintenance
    Enables MCP clients to perform keyless web searches, extract fully rendered pages into clean Markdown, and automate a shared Chrome browser through page navigation, clicking, typing, reading, screenshots, and backtracking—all without API keys.
    MIT
Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.