ghostfox
Ghostfox MCP server enables self-hosted stealth browser automation, inspection, debugging, and captcha solving through MCP tools.
Session & identity: Create sessions with coherent spoofed identities (
session_create), detect logged-in user (session_me), list tabs/popups (session_pages), generate/audit identities (identity_generate,identity_audit).Navigation & page control: Open URLs (
page_open), wait for elements (page_wait_for), dismiss modals (page_dismiss_modal), take screenshots (page_screenshot).Acting on pages: Click by ref or CSS (
page_click_ref,page_click), type/fill (page_type_ref,page_type,page_fill), press keys (page_press), drag/hover (page_drag,page_move_to), upload files (page_upload_file), run JS or init scripts (page_eval,page_init_script).Semantic observation: Get interactive elements with refs via a11y snapshot (
page_a11y), read full values (page_read_ref), extract visible text (page_snapshot), submit comments (page_comment).Vision & pixel analysis: Render luminance grids (
page_pixels), local contrast (page_contrast), template match (page_match_image), OCR (page_ocr,page_captcha_ocr), text bounding boxes (page_vision).Debug cortex: Read console logs (
page_console), uncaught JS errors/stack traces (page_errors), capture and inspect network responses and bodies (page_network_start/read/body).Captcha solving: Solve GeeTest slide/click (
page_geetest_slide,page_geetest_click), hCaptcha (page_hcaptcha), rotate (page_captcha_rotate), OCR captchas (page_captcha_ocr), or delegate to a provider (captcha_solve).Evidence & safety: Retrieve append-only audit trail/snapshots/identity (
session_evidence), gate dangerous actions (confirm_action).
Enables automated browsing on Cloudflare-protected sites by using engine-level anti-detect fingerprinting and solving Cloudflare Turnstile challenges through a CAPTCHA provider.
Ghostfox
The agent-native stealth browser you can own.
Self-hosted · Open source · MCP-first · Engine-level anti-detect
AI agents get blocked. Headless Chrome triggers Cloudflare 403s on ~20% of the web, and hosted "stealth browsers" route your agent's cookies, identities and sessions through someone else's cloud.
Ghostfox is the alternative: a complete browser stack you run yourself — a fingerprint-coherent stealth engine plus a Rust MCP runtime, in one repo.
Firefox (MPL-2.0)
└─ Camoufox (anti-detect patches, by daijro)
└─ Ghostfox engine engine/ — spoofing at the C++ level
└─ Ghostfox runtime runtime/ — Rust: sessions, identities, MCPGhostfox | Hosted stealth (Browserbase etc.) | playwright-mcp | Anti-detect suites (Multilogin etc.) | |
Self-hosted | ✓ | ✗ | ✓ | partially |
Open source | ✓ | ✗ | ✓ | ✗ |
MCP-native | ✓ | ✓ | ✓ | ✗ |
Engine-level anti-detect | ✓ (C++/Firefox) | vendor partnerships | ✗ | ✓ (closed) |
Coherent identities + auditor | ✓ | ✗ | ✗ | partial |
Runtime language | Rust | — | Node | — |
The moat — why this isn't just another wrapper.
We own the engine. The anti-detect lives in C++ patches inside our own Firefox fork — not in injected JS that detectors can read. Upstream Camoufox has signaled partially-closed patches ahead; wrappers inherit that risk, a fork that owns its engine doesn't.
A living proof corpus. Every claim here has a receipt: bilibili icon-click ×6, hCaptcha on production signups, TikTok OAuth+OTP live, 500/500 identity audits, Docker E2E. Features get copied in a week — verified history can't be.
Agent-native ergonomics. 43 coherent MCP tools, fire-then-verify receipts, evidence recording, and a playbook (
AGENTS.md) distilled from real runs. Agents (and their prompts) build habits on this surface — switching costs are real.Canonical distribution. PyPI, npm, GHCR and the official MCP Registry under one name, with the docs, benchmarks and changelogs to back it. Forks will exist; the verified trunk is here.
Captcha suite — 8 families, solved on-device (v0.6.7+). The runtime ships native MCP solvers with local models — no paid captcha farms, no cloud, no browser rent:
Family | Native tool | How |
GeeTest slide / v4 radar |
| bg-vs-fullbg diff + largest-blob gap detection |
GeeTest icon-click (文字点选) |
| custom-trained YOLOv8s + siamese similarity (rten, CPU) |
Rotate |
| 24-angle sweep + programmatic verdict |
Normal OCR |
| ported ddddocr (CRNN+LSTM, onnxruntime) |
hCaptcha |
| layout router + 553-model QIN2DIM zoo + optional vision-model ensemble |
Cloudflare Turnstile / TikTok |
| proven playbooks in |
E2E-verified against production sites (not vendor demos): bilibili icon-click ×6 "Verification Succeeded", hCaptcha on real signups (dashboard.hcaptcha.com, dosya.co), TikTok OAuth+OTP live session.
Debug cortex — page tools that tell you WHY (v0.7). Agents stop guessing when a page misbehaves:
page_console— everyconsole.log/warn/errorsince loadpage_errors— uncaught JS exceptions with stack tracespage_network_start/read/body— request/response capture with body fetch
That's the DevTools trio, exposed over MCP.
Multi-model vision (optional). page_vision + page_ocr +
page_match_image + page_pixels + page_contrast — wire any
vision-capable model (Cloudflare Workers AI, GLM, Qwen, ...) as
cross-checks for grid puzzles and layout questions. Keys are optional;
the native solvers above run fully local.
Native-sourced observation — page_a11y {"native": true} reads the
engine's OWN accessibility tree (Gecko's DocAccessible walk — the
ariaSnapshot plumbing), so page scripts cannot tamper with what you see:
shadow DOM, iframes and ARIA semantics handled by Gecko itself, richer
states (focused/required/checked/expanded/disabled/level). Native refs
are observation handles; acting goes through role+name anchors or the
JS-walk source.
Eyes for agents — page_a11y. One call returns every visible interactive
element with a stable ref, semantic role, accessible name, live value —
piercing shadow DOM and same-origin iframes, so web-component UIs
(Reddit, modern frameworks) are fully visible. The snapshot also reports
login_state (logged-in / logged-out / unknown), page URL and title —
agents check session health before acting, not after failing.
Agents act by ref: page_click_ref e38, page_type_ref e21 "text" — no CSS
selectors needed. Rich editors (Lexical, Draft, ProseMirror) are handled via
editor-native input paths with fire-then-verify receipts. page_wait_for
replaces manual sleeps. page_read_ref gives full untruncated values.
page_upload_file bypasses native file pickers.
Android personas too — session_create {"platform": "android"} gives
portrait screens, Adreno/Mali GPUs, Android font stacks and Firefox-on-Android
UAs, all audited like desktop identities (500/500 coherent, see
runtime/docs).
One identity, no contradictions. Identities are generated from coherent device presets (platform, screen, GPU, fonts that actually ship together), injected at the engine level, and audited before use — a spoofed browser's worst enemy is itself saying "4 cores on a MacBook".
Quickstart
Pick a distribution:
# Python (Linux x86_64)
pip install ghostfox
python -c "import ghostfox; ghostfox.install_engine(); ghostfox.install_runtime()"
# npm / any MCP host
npm install -g ghostfox # or: npx ghostfox install
{
"mcpServers": {
"ghostfox": { "command": "npx", "args": ["-y", "ghostfox", "mcp"] }
}
}
# Docker (engine + MCP runtime, ubuntu:24.04 base)
docker run -i --rm ghcr.io/autokeren/ghostfox:v0.7.2
# or one-line install of the binary stack:
curl -fsSL https://raw.githubusercontent.com/autokeren/ghostfox/main/install.sh | bashAlso listed on the official
MCP Registry
(io.github.autokeren/ghostfox) — one-click add in registry-aware clients.
From source:
# 1) Get the engine (prebuilt) and unpack it somewhere, e.g. /opt
unzip ghostfox-<ver>-lin.x86_64.zip -d /opt/ghostfox
# 2) Build the runtime
git clone https://github.com/autokeren/ghostfox.git
cd ghostfox/runtime
cargo build --release{
"mcpServers": {
"ghostfox": {
"command": "/path/to/ghostfox/runtime/target/release/ghostfox-mcp",
"env": { "GHOSTFOX_HOME": "/opt/ghostfox" }
}
}
}Then the agent can: session_create → page_open → page_a11y → act by ref.
Portable sessions. session_create accepts profile_dir for persistent
profiles: cookies, storage and the identity TOML live together in one place,
so a login survives restarts. Migrating a session from another
Camoufox-lineage browser? Copy the cookies and match the identity to the
origin device (platform, timezone, locale, screen) — a session that
suddenly changes identity looks like an impossible login and anti-fraud
systems revoke it. Proven flow, see AGENTS.md §8.
Full tool surface (52 tools):
Category | Tools |
Session |
|
See |
|
Wait |
|
Act |
|
Captcha |
|
Debug |
|
Vision |
|
Inspect |
|
Identity |
|
Recipes |
|
Evidence & safety |
|
Every mutation returns a receipt — page_fill reports landed_chars, while
type_ref fire-then-verifies async editors, so a silent page swap can't eat an
edit unnoticed. Sessions can also run headful ({"headful": true}) when
humans want to watch the agent work.
Every run records evidence. Each session writes an append-only event log
(events.jsonl), full page snapshots and the identity it used under
~/.ghostfox/recordings/ — fetch it any time with session_evidence.
Related MCP server: WeaveTab-MCP
E2E without emulators: Flutter web + any web app
Android E2E normally means emulators, Appium and a farm of devices.
Ghostfox takes the web route: build the Flutter app for the web
(flutter build web — same Dart codebase) and drive it with the
same senses used against hostile sites:
Semantics tree, read natively — Flutter's a11y nodes (roles, labels, bounds) land in
page_a11y(native=true)once semantics are enabled (one line in the app:SemanticsBinding.instance.ensureSemantics()).Native clicks —
page_click_nativefires the Flutter buttons at their trusted a11y coordinates (verified: submit triggers, status renders).Receipts +
page_diff+page_mutations+page_wait_stable— the same Act→Observe→Compare evidence loop, no polling guesswork.Visual QA —
page_ui_auditjudges the rendered UI.
Honest scope: logic/flow/UI = fully covered; final visual parity with real devices (Impeller rendering) and platform channels (sensors, camera) still need a device or emulator pass.
Sponsor
Ghostfox is independent and self-funded. If it saves your team from captcha walls or vibe-code UI bugs, help keep the house open:
GitHub Sponsors — monthly support, any amount
Buy Me a Coffee — one-time
Companies: sponsorship + early access to the hosted Visual QA service
is open — sponsor the repo
or open an issue with the title sponsorship.
Or install in one command (Linux x86_64):
curl -fsSL https://raw.githubusercontent.com/autokeren/ghostfox/main/install.sh | bashFrom source end-to-end (build the engine yourself):
see engine/README.md — make dir && make build.
Repository layout
runtime/ Rust: ghostfox-{core,fingerprint,mcp,eval} (MIT OR Apache-2.0)
engine/ Browser fork: patches, branding, build system (MPL-2.0)
AGENTS.md The agent playbook — how AI agents drive Ghostfox like a human
SKILL.md The compact agent skill — the golden loop + tool map at a glance
docs/JOURNEY.md The dev-log, todo list and the road ahead (the five senses)Two directories, two licenses, one product. The runtime speaks Juggler natively — no Node, no Python at runtime.
Using Ghostfox with an AI agent (opencode, Codex, Cursor, Claude Code, ...)? Read
AGENTS.mdfirst — it's the distilled playbook from real agent runs: the READ → REASON → DECIDE → ACT loop, self-health (rate limits, drafts, notifications), rich-editor typing, and every known wall with its proven solution.
Why own the engine?
Anti-detect that survives inspection. Spoofing happens inside the engine (navigator, screen, WebGL, fonts, WebRTC, timezone, audio) — not in injected JS that detectors can read.
No cloud dependency. Your agent's identities and cookies never touch a third-party host.
Upstream insurance.
engine/tracks daijro/camoufox asupstream; Ghostfox applies its own branding and can rebase whenever it wants — including if upstream patches go closed-source.
Status
v0.7 — alpha. Verified: identity coherence (500/500), full MCP round-trip
E2E (create → open → fill → submit), 8 captcha families E2E on production
sites (bilibili, hCaptcha-protected signups, TikTok), debug cortex
(console/errors/network), Docker image E2E (session → open → snapshot
inside a container), portable sessions across Camoufox-lineage browsers.
Distributed via PyPI, npm, Docker (GHCR) and the official MCP Registry.
Known limits are tracked in the changelogs under runtime/ and engine/.
Do not use against targets you don't have permission to test. This is a testing / research tool.
Credits
Ghostfox stands on the shoulders of giants — Camoufox (daijro) for the anti-detect patch stack, Mozilla Firefox for the engine, LibreWolf for the patch tooling lineage, and Playwright for the Juggler protocol.
License
engine/— MPL-2.0 (inherited from Firefox / Camoufox). See engine/LICENSE.runtime/— MIT OR Apache-2.0. See runtime/LICENSE-MIT.
Available Tools
42 toolscaptcha_solveA
Solve a CAPTCHA through the configured provider (env GHOSTFOX_CAPTCHA_PROVIDER=2captcha + GHOSTFOX_CAPTCHA_KEY). Turnstile/hcaptcha: pass sitekey + pageurl; image captchas: pass image_base64. Stealth-first: prefer not being challenged at all.
| Name | Required | Description | Default |
|---|---|---|---|
| pageurl | No | Turnstile/hcaptcha-style: the page URL the challenge lives on. | |
| sitekey | No | Turnstile/hcaptcha-style: the site's sitekey. | |
| session_id | Yes | ||
| image_base64 | No | Image captcha: the challenge image as base64 PNG. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the burden. It discloses the provider configuration via environment variables and a stealth-first preference, which is useful context. However, it does not describe the output (e.g., captcha token), failure modes, network/cost implications, or what happens when stealth succeeds.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and well-structured: it opens with the primary action, states the provider configuration, then lays out the two parameter modes, and ends with a strategic note. Every sentence adds value.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and no annotations, the description should explain what the tool returns and how the required session_id is used. It does neither. It also omits practical details like timeouts, errors, or whether the result is a token to be used elsewhere.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description adds meaning beyond the schema by explaining which parameters should be used together: sitekey + pageurl for Turnstile/hcaptcha, and image_base64 for image captchas. However, it does not explain the required session_id parameter, whose schema description is absent.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's job: solve a CAPTCHA through a configured provider. It also differentiates between Turnstile/hcaptcha and image captchas, which are distinct modes of the same action. The tool is clearly distinct from the sibling page/identity/session tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides mode-specific guidance: use sitekey + pageurl for Turnstile/hcaptcha and image_base64 for image captchas. However, it does not explicitly state when to use this tool versus an alternative or when not to use it; the 'stealth-first' note is more philosophical than operational.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
confirm_actionA
SAFETY GATE: confirm a dangerous action before executing it. Call this BEFORE clicking refs on pages flagged with danger_zone (financial/medical/legal/auth).
| Name | Required | Description | Default |
|---|---|---|---|
| ref | Yes | The ref of the element that will be clicked/activated. | |
| action | Yes | What the agent is about to do (for the evidence log). | |
| page_id | Yes | ||
| session_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of disclosing behavior. It only says 'SAFETY GATE' and 'confirm' – it does not explain what happens when called (e.g., whether it blocks, prompts a user, logs evidence, or returns a result). It does not describe side effects, error conditions, or the outcome of a denial. This is a significant gap for a safety-critical tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, compact sentence with the key phrase 'SAFETY GATE' front-loaded. Every word earns its place: it identifies the tool's role, its purpose, and the exact calling context. No fluff or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has no output schema and no annotations, so the description must explain return values, error behavior, and what the agent should do after calling it. It says nothing about what confirm_action returns (e.g., success/failure), whether it requires user interaction, or how the agent should proceed if confirmation is denied. For a mandatory safety gate with 4 required parameters, this is critically incomplete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With schema_description_coverage at 50%, the description must compensate for undefined parameters. It adds some context: 'refs' and 'pages flagged with danger_zone' give meaning to `ref` and `page_id`, and 'dangerous action' ties to `action`. However, it does not explain `session_id` at all, and the schema already covers `ref` and `action`. Partial compensation but not complete.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a clear verb-resource pair: 'confirm a dangerous action before executing it.' It also specifies the exact trigger condition ('before clicking refs on pages flagged with danger_zone') and lists the relevant categories (financial/medical/legal/auth). This clearly differentiates it from the sibling page_* and session_* tools, which are about navigation, inspection, or input, not confirmation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives an explicit when-to-use condition: 'Call this BEFORE clicking refs on pages flagged with danger_zone (financial/medical/legal/auth).' This is clear and actionable. However, it does not name alternatives or explicitly state when not to use it (e.g., 'do not call for normal pages'), so it falls short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
identity_auditA
Audit an identity TOML for coherence violations (contradictory signals a detector would flag).
| Name | Required | Description | Default |
|---|---|---|---|
| identity_toml | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral burden. The verb 'Audit' implies a read-only analysis and the parenthetical defines what is being detected. Still, it does not disclose whether the tool returns a list of violations, a report, or errors, and it does not specify side effects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is one tight sentence that front-loads the action and object and adds a useful clarifying parenthetical. There is no filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Because there is no output schema and no annotations, the description should at least state what the audit produces so an agent can interpret the result. It does not, leaving the tool's return contract underspecified.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema provides 0% description coverage, and the description only repeats 'identity TOML' rather than explaining the expected format. An agent is left to infer whether identity_toml should be raw TOML content, a path, or something else.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific action ('Audit'), a specific resource ('an identity TOML'), and a concrete objective (find coherence violations/contradictory signals). This clearly separates identity_audit from sibling identity_generate, which creates identities rather than validating them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The usage context is implied: call this when an identity TOML needs validation. However, the description does not explicitly state when to prefer this tool over identity_generate or other siblings, nor does it mention exclusions or preconditions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
identity_generateA
Generate a new coherent browser identity, returned as TOML.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It discloses that the tool generates an identity and returns it as TOML, giving basic transparency, but it does not mention side effects, persistence, randomness, or whether the identity is applied to the current browser session.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, front-loaded sentence with no filler. It states the verb, resource, and output format economically while remaining readable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter, no-output-schema tool, the description provides the essential invocation information. It lacks details about what the TOML identity contains or how it relates to the current session, but the tool is simple enough that this is a minor gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, and the schema fully documents the empty object, so parameter semantics are not a concern. The description adds no parameter details, but none are needed; the baseline for a zero-parameter tool is strong.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description states a specific verb ('Generate') and resource ('new coherent browser identity'), and explicitly notes the output format ('returned as TOML'). This clearly distinguishes it from sibling identity_audit, which implies inspection rather than creation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The word 'new' implies this tool is for creating an identity rather than auditing an existing one, but the description does not explicitly state when to prefer identity_generate over related tools like session_create or identity_audit. Usage context is only implied, not explained.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
page_a11yA
Semantic snapshot of the page: every visible interactive element with a stable ref, role (button/link/textbox/...), accessible name and CURRENT value — pierces shadow DOM, so web-component UIs (Reddit, modern frameworks) are fully visible. Use this instead of guessing CSS selectors.
| Name | Required | Description | Default |
|---|---|---|---|
| page_id | Yes | ||
| session_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the behavioral transparency burden. It discloses meaningful behavior: it pierces shadow DOM, only includes visible interactive elements, and returns current values. It could add an explicit read-only guarantee or limitations around iframes, but the snapshot language strongly implies a non-mutating operation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two dense, front-loaded sentences with no filler. The first sentence states what is returned and key traits, and the second gives a direct usage recommendation.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only snapshot tool with simple parameters and no output schema, the description adequately explains the returned content: refs, roles, accessible names, and current values. It could be more complete by explicitly linking the stable refs to sibling ref-based tools like page_click_ref, page_type_ref, and page_read_ref, but it is sufficient for an agent to invoke the tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There are two parameters with 0% schema description coverage, and the description does not explain session_id or page_id. The names are inferable from the broader tool family, but the description adds no parameter-level meaning and does not compensate for the missing schema descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: a semantic snapshot of visible interactive elements with refs, roles, accessible names, and current values. It differentiates itself from guessing CSS selectors, but it does not explicitly distinguish itself from the sibling tool page_snapshot, which could be confused with a generic snapshot.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives a clear usage directive: use this instead of guessing CSS selectors, especially for shadow-DOM-heavy web-component UIs. It provides useful context for when to invoke it, though it does not name alternatives or state when not to use it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
page_captcha_ocrA
CAPTCHA OCR: classify the text of a normal image captcha (the distorted-text family) with the ddddocr model (CRNN+LSTM trained specifically on captcha text, 8210-char charset). 100% local — runs on onnxruntime inside the runtime. Takes the ref of the captcha element (or the whole viewport) and returns the recognized text. Feed it into the answer field and submit. Far stronger than generic OCR on captcha fonts.
| Name | Required | Description | Default |
|---|---|---|---|
| ref | No | Ref of the captcha <img> element. Omit for the whole viewport. | |
| page_id | Yes | ||
| session_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the behavioral burden. It discloses meaningful traits: 100% local execution via onnxruntime, model architecture, charset size, and how input is scoped (ref or entire viewport). It does not detail error or failure behavior, but the core operational transparency is strong.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded, starting with the tool's purpose and then adding model detail, execution model, input behavior, and downstream usage. Every sentence adds value; the efficiency and 'stronger than generic OCR' note reinforces selection without padding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given there is no output schema, the description still communicates the return value ('returns the recognized text') and the practical next step ('feed it into the answer field and submit'). It provides enough for an agent to invoke correctly, though it could briefly mention limitations or failure cases for unusual captcha types.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 33%, so the description must compensate. It explains the most meaningful parameter, ref, including the omit-for-viewport behavior. However, it adds no meaning for the required session_id and page_id parameters, which remain undocumented in both schema and description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific operation (OCR/classify text) and resource (normal image captcha, distorted-text family), and distinguishes the tool by naming the specialized ddddocr model and contrasting with generic OCR. This clearly separates it from siblings like page_ocr and captcha_solve.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives concrete usage context: use for normal distorted-text image captchas, feed the img ref or whole viewport, and pass the result into the answer field. It does not explicitly name sibling tools or exclusion cases, but the 'normal image captcha (distorted-text family)' phrasing provides a clear selection criterion.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
page_captcha_rotateA
ROTATE CAPTCHA SOLVER: brute-force sweep + human replay. Rotate challenges ask to turn an image upright; standard demos rotate 15deg per button click and answer with a check button. The tool first sweeps every angle programmatically (instant JS clicks, ~30s), reads the visible feedback after each check (success keywords: pass/通过/正确/成功/verif/succeeded), then REPLAYS the winning rotation with real human mouse drags and the final check. Returns JSON {clicks, angle, status}. For fixed-image demos the sweep converges every time.
| Name | Required | Description | Default |
|---|---|---|---|
| page_id | Yes | ||
| check_ref | Yes | Ref of the check/verify button. | |
| reset_ref | Yes | Ref of the reset button. | |
| session_id | Yes | ||
| rot_left_ref | Yes | Ref of the rotate-left button. | |
| rot_right_ref | Yes | Ref of the rotate-right button (15deg per click in the standard demos). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, and it delivers: it discloses the brute-force JS-click sweep, the feedback keyword check, the human-replay phase, the approximate duration (~30s), the JSON return shape, and the convergence caveat. Nothing in the text contradicts the absent annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three dense sentences front-load the tool's purpose ('brute-force sweep + human replay') and each subsequent clause adds new information (keywords, timing, return shape, caveat) without repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Covers the execution flow, return value, and a limitation, which is strong for a captcha solver. It stops short of telling an agent what status values can be, what happens on sweep failure, or how session_id/page_id/reset_ref factor in, so it is not fully complete for a 6-parameter tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema already describes check_ref, reset_ref, rot_left_ref, and rot_right_ref; the description adds only the 15deg standard but does not clarify session_id, page_id, or how refs relate to the stated replay flow. At 67% schema coverage, the description does not fully compensate for the undocumented parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Opens with 'ROTATE CAPTCHA SOLVER' and immediately defines the challenge type ('turn an image upright'), which differentiates it from OCR, slide, and hCaptcha siblings. The sequencing detail (sweep then human replay) makes the verb/resource unambiguous beyond the name alone.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a clear trigger condition: rotate challenges that respond to 15deg-per-click buttons and a check button. It does not name explicit exclusions or alternative captcha tools, so it misses the top bar of the rubric, but the context is enough to route an agent correctly.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
page_clickA
Click an element by CSS selector. For form controls (buttons, inputs), uses a JS click; for links and other elements, dispatches real mouse events at coordinates. Prefer page_click_ref when you have a page_a11y ref — it scrolls into view first and is more reliable on web-component UIs. Returns 'ok' on success.
| Name | Required | Description | Default |
|---|---|---|---|
| page_id | Yes | ||
| selector | Yes | CSS selector. | |
| session_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral burden. It discloses the click mechanism for form controls vs links, that real mouse events are dispatched at coordinates, and the success return value. This is solid, though it omits error/not-found behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, each with a distinct job: core action, behavior detail, and sibling guidance. No wasted words, and the most actionable information is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple three-parameter click tool, the description covers purpose, behavior, return value, and the main alternative. It does not document error cases or prerequisite page state, but these are minor gaps given the tool's simplicity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 33%, yet the description adds no parameter-level meaning beyond what the schema already says. It repeats that the selector is a CSS selector but does not explain session_id or page_id, leaving the agent to infer their roles.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: click an element by CSS selector. It also distinguishes itself from the sibling page_click_ref by noting when that tool is preferable, so an agent can tell them apart.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly routes the agent to page_click_ref when a page_a11y ref is available, explaining that it scrolls into view and is more reliable on web-component UIs. This gives clear when-to-use guidance relative to the closest alternative.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
page_click_refA
Click an element by its ref from page_a11y. Scrolls it into view first. No selectors needed.
| Name | Required | Description | Default |
|---|---|---|---|
| ref | Yes | Element ref from page_a11y (e.g. "e12"). | |
| page_id | Yes | ||
| session_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden of disclosing behavior. It explicitly states that the element is scrolled into view before clicking, which is a meaningful behavioral trait beyond the basic action. It does not mention potential side effects of the click or failure modes, but for a simple click-by-ref tool the disclosed behavior is reasonable.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three short sentences with no filler. The primary action is stated first, followed by the key behavioral detail and the distinguishing constraint. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The core usage is clear and the ref-based semantics are well explained, making the tool minimally viable. However, there is no output schema, no annotations, and two of three parameters remain undocumented, so an agent may be unclear about expected return values or how session_id/page_id are used.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 33%; only 'ref' is described. The description reinforces the ref semantics but adds little beyond the schema, and it does not clarify 'session_id' or 'page_id' at all. With low schema coverage, the description should compensate, and it fails to do so for the required context parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the specific action ('Click an element') and the resource/identifier used ('by its ref from page_a11y'). It also distinguishes itself from sibling tools by noting 'No selectors needed,' making it obvious this is the ref-based counterpart to page_click.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies clear usage context: use this tool when you have an element ref from page_a11y and do not need selectors. It does not explicitly name alternatives or provide exclusion criteria, but the mention of ref-based access and 'No selectors needed' gives sufficient guidance for an agent to select between this and selector-based siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
page_commentA
Submit a comment on the current page. Automatically: 1) detects your username via session_me, 2) checks if you already commented (skips if duplicate), 3) finds the comment editor (Reply or Join conversation), 4) opens it, 5) types your text, 6) clicks submit, 7) verifies. Returns result JSON.
| Name | Required | Description | Default |
|---|---|---|---|
| text | Yes | The comment text to submit. | |
| page_id | Yes | ||
| session_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. It meticulously details the entire sequence: detecting username via session_me, checking for duplicates (skips if duplicate), locating the editor, opening it, typing, submitting, and verifying. It also notes the return of result JSON. This is highly transparent about what the tool does, including potential side effects like avoiding duplicate submissions. It exceeds what annotations would typically provide.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise, using a numbered list to break down the workflow. The first sentence states the purpose, followed by a clear sequence. It is efficiently written with no unnecessary words. The structure front-loads the main action and then details the steps, making it easy to scan.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's moderate complexity (multi-step automation) and the lack of an output schema, the description provides a thorough overview of the process. It mentions dependencies (session_me), duplicate handling, and verification. However, it does not specify the exact structure of the returned JSON or any prerequisites like needing to be on a comment-capable page. These gaps are minor given the description's overall completeness, but they prevent a perfect score.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has only 33% description coverage (only 'text' is described). The description does not explicitly explain 'page_id' or 'session_id'. It mentions session_me, which hints at session_id's role, but does not clarify what page_id represents despite the description saying 'current page' (which contradicts the need for an explicit page_id parameter). The description fails to compensate for the low schema coverage, leaving two of three parameters semantically unclear.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Submit a comment') and the specific resource ('on the current page'). It distinguishes itself from sibling page_* tools by describing a high-level workflow (detecting user, checking duplicates, interacting with the editor) rather than a generic action like typing or clicking. The purpose is unambiguous and specific.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use this tool: to post a comment on the current page. It does not explicitly mention alternatives (like manual page_type/page_click sequences) or exclusions, but the context makes its purpose clear. It lacks explicit 'when not to use' guidance, so it falls short of a 5, but the context is clear enough for an agent to infer the appropriate use.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
page_consoleA
DEBUG CORTEX: read the page's console output (log/warning/error) captured at the PROTOCOL level — the page cannot hide or patch it. Invisible debugger for AI agents: reproduce the bug, read what the page logged. Optionally filter noise by clearing after read. Returns JSON [{kind, text, url, line, ts}]. Auto-starts capture on first call.
| Name | Required | Description | Default |
|---|---|---|---|
| clear | No | Drain the buffer after reading (default: keep entries). | |
| page_id | Yes | ||
| session_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It discloses that capture auto-starts on first call, that the page cannot hide or patch the output, that clearing is optional, and that the return is JSON with a specific shape. It does not mention whether the buffer is bounded or whether repeated calls drain or accumulate, but it covers the most important behavioral traits.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded: it opens with the tool's purpose and key differentiator, then adds the workflow hint, return shape, and auto-start behavior. Every sentence earns its place, and the structure makes the most important information immediately visible.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only console inspection tool with no output schema, the description is largely complete: it states the return shape, the auto-start behavior, and the optional clear workflow. It does not specify buffer limits or whether entries are deduplicated, but those are minor gaps for the tool's apparent debugging use case.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 33% (only 'clear' has a description). The description adds meaning for 'clear' by explaining it as 'filter noise by clearing after read', which aligns with the schema's 'Drain the buffer after reading'. However, it does not add meaning for session_id or page_id beyond what their names imply, and the description does not fully compensate for the two undocumented required parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('read'), a specific resource ('the page's console output'), and a distinctive scope ('captured at the PROTOCOL level — the page cannot hide or patch it'). It clearly distinguishes this from sibling tools like page_errors and page_network_read by emphasizing the console/log capture angle and the invisible-debugger framing.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context for when to use the tool: when debugging by reproducing a bug and reading what the page logged. It also mentions an optional workflow step ('Optionally filter noise by clearing after read'). It does not explicitly name sibling alternatives or state when not to use it, but the protocol-level capture framing implies it is the go-to for console output that the page cannot hide.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
page_contrastA
HIGH-PASS VISION: local-contrast grid of an element's image (|gray - gaussian_blur|). Makes ANYTHING blended into a background VISIBLE — captcha characters on photos, watermarks, hidden strokes. 0 = flat area, 9 = strong edge. This is the native tool born from solving GeeTest icon-click with pure math. Works on canvas / img / background-image (one fetch per challenge).
| Name | Required | Description | Default |
|---|---|---|---|
| ref | Yes | Ref of the element to analyze (canvas / img / background-image). | |
| grid_h | No | Grid height in cells (default 44). | |
| grid_w | No | Grid width in cells (default 100). | |
| page_id | Yes | ||
| session_id | Yes | ||
| blur_radius | No | Gaussian blur radius in px (default 4). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It discloses the algorithm, output value range (0 = flat, 9 = strong edge), supported element types, and the network cost ('one fetch per challenge'). It does not state whether the tool is read-only, but the algorithmic description strongly implies it is a pure analysis operation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four sentences deliver the algorithm, use cases, output scale, accepted element types, and cost, with the key purpose front-loaded. The phrasing is somewhat dramatic ('HIGH-PASS VISION', 'native tool born from solving GeeTest'), but no sentence is pure filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers purpose, algorithm, output scale, and supported elements, but no output schema exists, so the return format is under-specified—it is unclear whether the result is a JSON matrix, per-cell values, or a visual grid. Additional details about how grid dimensions map to output and error behavior would make it more complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 67%; the described parameters (ref, grid_h, grid_w, blur_radius) have schema descriptions with defaults. The tool description adds context about the grid values and what ref refers to, but it does not go beyond the schema for parameter semantics. The two undocumented parameters are standard session/page identifiers, so the gap is minor.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific operation—computing a local-contrast grid via |gray - gaussian_blur|—and a clear resource: an element's image from canvas, img, or background-image. It differentiates itself through the explicit high-pass vision concept and concrete use cases (captcha, watermarks, hidden strokes), though it does not explicitly name sibling tools for comparison.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives clear usage context: use this tool when content blends into a background and needs to be made visible, with examples like captcha characters, watermarks, and hidden strokes. It also scopes applicable element types (canvas / img / background-image), implying when it should be used. However, it does not explicitly contrast with alternatives like page_pixels or page_vision.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
page_dismiss_modalA
Detect and dismiss common modals: cookie banners, consent dialogs, popups, overlays. Returns what was dismissed.
| Name | Required | Description | Default |
|---|---|---|---|
| page_id | Yes | ||
| session_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It discloses that the tool detects and dismisses modals and returns what was dismissed, which is useful. However, it does not mention side effects, failure behavior, or the limitation of what counts as a 'common' modal.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences with the scope front-loaded and no filler. The list of modal examples is compact, and the return-value note earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with no annotations and no output schema, the description communicates the core behavior and the return value. Missing context includes what happens when no modal is detected, prerequisites, and side-effect caveats, but the essential invocation intent is clear.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the description does not explain session_id or page_id. The parameter names are relatively self-explanatory in context, but the description adds no meaning beyond the schema's bare field names.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a clear verb-resource pair: dismiss common modals, and enumerates specific examples (cookie banners, consent dialogs, popups, overlays). It also notes the return value, which distinguishes it from other page interaction tools like page_click or page_press.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The modal types listed imply when to use it, and it is the only dismiss-focused sibling, but there is no explicit when-not-to-use guidance or mention of alternatives. No prerequisites or caveats are stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
page_dragA
Drag with a HUMAN-LIKE movement profile: approaches the source, presses, drags along a bezier arc with ease-in-out velocity and micro-pauses, settles, releases. Pass from_ref + to_ref to drag element onto element, or from_ref + offset_x/offset_y to drag by pixels (slider captchas, resize handles). All from page_a11y refs.
| Name | Required | Description | Default |
|---|---|---|---|
| to_ref | No | Ref of the drop target. Omit to drag by offset instead (sliders). | |
| page_id | Yes | ||
| from_ref | Yes | Ref of the element to grab (from page_a11y). | |
| offset_x | No | Horizontal pixels to drag when to_ref is omitted (positive = right). | |
| offset_y | No | Vertical pixels to drag when to_ref is omitted (positive = down). | |
| session_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden and does so thoroughly. It reveals the movement sequence—approach, press, bezier drag with ease-in-out velocity, micro-pauses, settle, release—and clarifies that refs come from page_a11y. This adds significant behavioral context beyond the raw schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and well-structured: movement profile first, then the two invocation modes, then the ref source. Every sentence earns its place with no redundant restatement of schema fields.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 6-parameter tool with no output schema and no annotations, this covers the essential invocation paths, parameter semantics, and behavioral expectations. It lacks an explicit note about return values or what happens if both to_ref and offset are provided, but the 'or' phrasing makes exclusivity reasonably inferable.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description enriches schema parameters by showing how from_ref, to_ref, offset_x, and offset_y combine into mutually exclusive modes. It also adds real-world context for offsets (slider captchas, resize handles) and states the ref source. However, session_id and page_id remain undocumented in both schema and description, keeping this from a perfect score.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb and resource ('Drag') and details a distinct human-like movement profile. It also clearly separates the two main use cases: element-to-element drag and pixel-based drag, making it distinguishable from sibling tools like page_move_to and page_press.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit mode guidance: use from_ref + to_ref for element-to-element dragging, or from_ref + offset_x/offset_y for pixel-based dragging such as slider captchas and resize handles. It does not explicitly name sibling alternatives, but the intended conditions are clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
page_errorsA
DEBUG CORTEX: read the page's UNCAUGHT JS EXCEPTIONS with stack traces, captured at the protocol level. The error console a developer opens in DevTools — as a tool. Returns JSON [{text, url, line, stack: [frames], ts}]. Auto-starts capture on first call.
| Name | Required | Description | Default |
|---|---|---|---|
| clear | No | Drain the buffer after reading (default: keep entries). | |
| page_id | Yes | ||
| session_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full disclosure burden. It discloses non-obvious behaviors: auto-starting capture on first call, protocol-level capture, and the exact JSON return shape with stack trace frames. It does not discuss buffer retention or permission requirements, but the most important behavioral traits are covered.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, with the core purpose front-loaded in the first sentence. The 'DEBUG CORTEX:' prefix is minor noise, but each subsequent sentence adds distinct useful information: output format and capture timing. Efficient and well-structured.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a diagnostic read tool with no output schema, the description fully explains the return value shape, the scope (uncaught JS exceptions), and the critical timing nuance (auto-start capture). The schema covers the 'clear' parameter's behavior. Nothing essential is missing for an agent to use it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 33%, and the tool description adds no parameter-specific meaning. session_id and page_id are left to inference from their names, and the 'clear' behavior is documented in the schema but not in the description. The description fails to compensate for the low schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('read'), a specific resource ('the page's UNCAUGHT JS EXCEPTIONS'), and the capture mechanism (protocol level, stack traces). The DevTools error-console analogy clearly distinguishes it from sibling tools like page_console and page_network. No ambiguity about what the tool does.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The 'DEBUG CORTEX' framing and the explicit reference to the DevTools error console imply this is the tool for JavaScript exceptions rather than console logs or network activity. It provides clear context for when to use it, though it stops short of explicitly naming alternatives or stating exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
page_evalA
Evaluate a JavaScript expression in the page's main frame and return its JSON value. Read-only introspection is safest; treat results of mutations with care.
| Name | Required | Description | Default |
|---|---|---|---|
| page_id | Yes | ||
| expression | Yes | JavaScript expression; the JSON-ified result is returned. | |
| session_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden. It meaningfully discloses that mutations are possible and should be treated with care, and it scopes evaluation to the page's main frame. It does not cover error behavior or serialization limits, but the explicit safety warning is valuable.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences with no filler. The core action is front-loaded, and the safety guidance is placed immediately after it, making the tool easy to parse quickly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a code-execution tool with no annotations and no output schema, the description covers the evaluation scope and return format but omits details about exceptions, serialization behavior, and the role of session_id/page_id. It is adequate but leaves meaningful gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 33%: only 'expression' has a description, while 'session_id' and 'page_id' are undocumented. The tool description adds little beyond restating the expression parameter and does not compensate for the missing parameter semantics.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Evaluate'), a specific resource ('JavaScript expression in the page's main frame'), and the expected output ('return its JSON value'). This clearly distinguishes page_eval from all page_* siblings, since only this tool evaluates arbitrary expressions.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear usage context: read-only introspection is the recommended use, and mutations are explicitly flagged for caution. It does not name alternative tools or state explicit when-not-to-use conditions, so it stops short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
page_fillA
Set an input's value directly (form fill). Works where key-event typing hits engine bugs; fires input/change events like real edits.
| Name | Required | Description | Default |
|---|---|---|---|
| text | Yes | ||
| page_id | Yes | ||
| selector | Yes | CSS selector. | |
| session_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the behavioral burden. It discloses that the value is set directly rather than via key events, and that it 'fires input/change events like real edits,' which is important for agent reasoning about side effects. This is meaningful added context, though it could mention more about limitations or error behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, focused sentence that front-loads the core action and then adds the most important behavioral detail. There is no redundancy, filler, or unnecessary technical jargon.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's moderate complexity, the description covers the core behavior, the reason to use it over an alternative, and relevant event behavior. The main gap is parameter-level detail, but the required parameters are standard across siblings and mostly inferable from names.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 25% (only selector has a description). The description does not explain session_id, page_id, or text semantics beyond the implied 'value' of an input. While parameter names are somewhat self-explanatory, the description does not compensate for the low schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's function: 'Set an input's value directly (form fill)'. It uses a specific verb and resource, and differentiates itself from key-event typing, which maps to sibling page_type. An agent can tell this tool apart from related siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives explicit context for when to use this tool: 'Works where key-event typing hits engine bugs.' This implies page_type as the alternative and identifies a clear condition for choosing page_fill. It doesn't state an explicit when-not-to-use list, but the guidance is clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
page_geetest_clickA
GEETEST ICON-CLICK SOLVER: solve the GeeTest v3 word-click captcha (characters on a photo, click in the strip's order) from the live page. Fetches the challenge image off the element's background, runs the trained ONNX pair (YOLOv8s char detection + siamese order matching, 100% local CPU, ~0.5s), and returns click targets in page coordinates plus the raw boxes. Call AFTER the challenge popup is open. Then click each 'click' point via page_drag and press the geetest confirm button. Returns JSON {clicks: [{x, y}, ...] in page px, boxes_raw: [[x, y], ...] in image px, rect: {...}, img: {w, h}}.
| Name | Required | Description | Default |
|---|---|---|---|
| ref | Yes | Ref of the .geetest_item_wrap element (carries the challenge background image). | |
| page_id | Yes | ||
| session_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full behavioral burden. It discloses that the tool runs a local ONNX model ('100% local CPU, ~0.5s'), fetches the image from the element background, and only returns coordinates rather than clicking itself ('Then click each... via page_drag'). It does not explicitly state that it never submits or confirms the captcha, which costs it full marks.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but mostly earns its length: it front-loads the purpose, then adds meaningful operation details, prerequisites, follow-up actions, and the return JSON structure. Slight redundancy exists between the all-caps title and the first clause, and the model-architecture detail could be trimmed without losing operational value.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with no annotations and no output schema, the description is unusually complete: it states the prerequisite, the processing model, the page-flow steps, and the exact JSON response shape including coordinate spaces. It still leaves session_id/page_id semantics unexplained and does not cover failure cases, but nothing essential to invoking it is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 33%, and the description does nothing to compensate: it never mentions session_id or page_id, which are both required. Only 'ref' is documented in the schema itself, so an agent must guess what the other two identifiers mean or how they relate to the live page.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with 'GEETEST ICON-CLICK SOLVER' and then states exactly what it does: 'solve the GeeTest v3 word-click captcha... from the live page.' It names a specific captcha variant, a specific element, and a specific output (click targets), making it clearly distinguishable from siblings like page_geetest_slide or page_captcha_rotate.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear sequencing: 'Call AFTER the challenge popup is open' and then 'click each point via page_drag and press the geetest confirm button.' This tells the agent when to invoke it and what to do next, though it does not explicitly contrast it with alternative captcha tools or state when not to use it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
page_geetest_slideA
GEETEST SLIDE SOLVER: solve the GeeTest v3-style slide puzzle from the live page, then perform the human drag itself. Extracts the three canvases (bg, puzzle slice, full reference bg) via toDataURL, finds the hole with |bg-fullbg| diff + largest-blob + morphological closing (the JPEG-noise trap), measures the piece's solid-alpha left edge, and drags the slider by hole_x0 - piece_x0 with the engine's humanized mouse. Returns JSON {drag_x, hole: [x0, x1], piece_x0}. Call AFTER the challenge popup is open. No vision model, pure pixel math.
| Name | Required | Description | Default |
|---|---|---|---|
| page_id | Yes | ||
| session_id | Yes | ||
| slider_ref | No | Optional ref of the slider handle (.geetest_slider_button). If omitted the tool locates it by selector and registers its own ref. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden, and it does so well: it discloses the canvas extraction, diff-based hole detection algorithm, slider drag calculation, and that it performs a humanized mouse drag. It does not mention side effects or failure conditions, but the core behavioral surface is clearly communicated.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a dense, front-loaded paragraph with a clear action statement followed by implementation details, return shape, and call timing. It is longer than the minimal case but every sentence adds meaningful information for a complex automated solver; it earns its length without becoming bloated.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the complexity of the tool and the absence of an output schema, the description does a good job explaining the return JSON shape and the high-level algorithm. It is missing explicit semantics for the required parameters and does not mention edge cases, but for an agent selecting and invoking this captcha solver, the essential context is present.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 33%, and the description does not compensate for the undocumented required session_id and page_id parameters. It gives useful context about the live page and slider, but it never explains how the parameters map to the tool's operation, leaving the agent to guess the role of the required inputs.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: it 'solve[s] the GeeTest v3-style slide puzzle from the live page, then perform[s] the human drag itself.' It clearly differentiates from sibling tools like page_geetest_click by focusing on the slide drag behavior and explicitly calling out that it uses 'pure pixel math' rather than a vision model.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear timing guidance with 'Call AFTER the challenge popup is open,' and implies this tool handles the slide drag, not other captcha types. It does not explicitly name alternatives or say when not to use it, but the context is strong enough for an agent to know when this tool is appropriate.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
page_hcaptchaA
HCAPTCHA SOLVER: solve hCaptcha image challenges from the live page with the host vision model (Cloudflare Workers AI GLM). Handles ALL challenge variants — tile grids, pattern-break icon fields, any image-pick — by asking the model for the exact pixel coordinates of every element to click, then clicking with the humanized mouse. Multi-round: loops until the page's h-captcha-response field receives a token (up to max_rounds, default 4). Credentials come from CLOUDFLARE_API_KEY / CLOUDFLARE_ACCOUNT_ID env or the params. Returns JSON {success, rounds, response_len}.
| Name | Required | Description | Default |
|---|---|---|---|
| page_id | Yes | ||
| cf_api_key | No | Cloudflare API key for the host vision model (defaults to CLOUDFLARE_API_KEY env). | |
| max_rounds | No | Max challenge rounds (hCaptcha chains 2-4). Default 4. | |
| session_id | Yes | ||
| checkbox_ref | No | Optional ref of the hCaptcha anchor iframe / checkbox. If omitted the tool auto-locates the anchor iframe by src pattern. | |
| cf_account_id | No | Cloudflare account id (defaults to CLOUDFLARE_ACCOUNT_ID env). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It does this well by disclosing the vision model (Cloudflare Workers AI GLM), the clicking mechanism ('humanized mouse'), the multi-round loop until a token appears, and credential sourcing from env vars or params. It does not describe failure or timeout behavior, but the core behavioral profile is unusually explicit.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but efficient: the purpose is front-loaded, and each clause adds useful detail about model, variants, clicking behavior, loop termination, credentials, and return shape. It is long, but not bloated, and the JSON return shape is cleanly included.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has no output schema and no annotations, the description compensates well by stating the return JSON, auth mechanism, loop cap, and challenge scope. The main gaps are the missing semantics of the required page_id/session_id parameters and lack of error-condition details, but the description is otherwise sufficient for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 67%, with page_id and session_id completely undocumented, and the description does not clarify those two required parameters. It mostly repeats what the schema already says about cf_api_key, cf_account_id, and max_rounds (env defaults and default 4), adding little new meaning for parameter selection.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with 'HCAPTCHA SOLVER' and states a specific verb and resource: 'solve hCaptcha image challenges from the live page.' It goes beyond a generic captcha label by naming the host vision model and the specific challenge variants it handles, clearly distinguishing it from sibling tools like captcha_solve, page_captcha_ocr, and page_captcha_rotate.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage through phrases like 'from the live page' and 'Handles ALL challenge variants,' which tells an agent this is the hCaptcha-specific solver. However, it never explicitly states when to prefer this tool over sibling captcha tools or when not to use it, leaving the routing decision to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
page_init_scriptA
INIT SCRIPTS: run your JS at DOCUMENT START on every navigation — BEFORE any page script loads. The deepest interception layer in Ghostfox: hooks installed here are what page libraries (gt.js, analytics, frameworks) capture, so their stored references are YOURS. Use for: API response interception (JSONP callback swaps), state capture, anti-detection instrumentation, environment patches. Applies to subsequent navigations.
| Name | Required | Description | Default |
|---|---|---|---|
| source | Yes | JavaScript source to run at document start on every navigation. | |
| page_id | Yes | ||
| session_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It discloses important behavior: runs at document start, before page scripts, on every navigation, and applies to subsequent navigations. However, it does not mention potential side effects like page breakage or how to remove the script, which is a notable gap for a script-injection tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with a clear, bolded summary ('INIT SCRIPTS: run your JS at DOCUMENT START on every navigation') followed by relevant use cases. It is somewhat verbose but each sentence adds value, explaining the interception layer and why it matters. No filler or repetition exists.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description adequately covers the core purpose, timing, and use cases for the tool. However, given that there is no output schema and no annotations, it misses key context: what session_id and page_id are for, how to undo or stop the script, and what happens on errors. These gaps leave an agent with uncertainty about full usage.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 33% (only 'source' is described). The description explains the purpose of 'source' (run JS at document start) but provides no additional context for 'session_id' or 'page_id' beyond what the schema shows (which is nothing). It does not compensate for the low coverage, as these parameters are left ambiguous.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('run') and resource ('your JS at DOCUMENT START on every navigation') with explicit timing. It clearly distinguishes itself from sibling tools like page_eval by emphasizing it runs before any page script loads and applies to subsequent navigations, making its purpose unmistakable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear use cases ('Use for: API response interception, state capture, anti-detection instrumentation, environment patches') and timing context. However, it does not explicitly mention when not to use this tool or name alternative tools (e.g., page_eval for one-time execution), so it lacks explicit exclusions/alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
page_match_imageA
REAL TEMPLATE MATCHING: multi-scale normalized cross-correlation of a needle against a haystack, computed IN-PAGE at full grayscale resolution (no grid loss). Returns top matches [{x, y, score, scale}] — needle-center positions in haystack-image pixels. needle_rect optionally crops the needle (e.g. an instruction glyph band from the same image — GeeTest icon-click pattern). Images are cached per URL: single-use challenge URLs fetch exactly once. THE tool for: captcha piece->gap, glyph->character, logo->page.
| Name | Required | Description | Default |
|---|---|---|---|
| hay_ref | Yes | Ref of the HAYSTACK element to search in. | |
| page_id | Yes | ||
| hay_rect | No | Optional search-region crop of the haystack (x, y, w, h in haystack-image px). Omit for the full image. USE THIS to exclude the needle's own area (self-match guard). | |
| needle_ref | Yes | Ref of the NEEDLE element (canvas / img / background-image). | |
| session_id | Yes | ||
| needle_rect | No | Optional crop of the needle image (x, y, w, h in needle-image px). Omit to use the whole image. Use this to match a sub-region (e.g. an instruction glyph band inside the same image). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, and it does substantial work: it discloses in-page computation, full-grayscale/no-grid-loss behavior, return format, coordinate semantics, caching behavior, and single-fetch behavior for challenge URLs. It does not explicitly assert read-only/side-effect-free status, but the matching operation itself strongly implies a non-destructive read.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is longer than minimal, but nearly every sentence carries distinct useful information: algorithm, output shape, crop behavior, caching, and use cases. The marketing-style caps ('REAL', 'THE tool') add noise, but the structure is front-loaded and information-dense.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 6-parameter tool with no output schema and no annotations, the description is notably complete: it provides the return structure, coordinate meaning, algorithm, cache semantics, and intended use cases. It does not specify how many top matches are returned or answer edge cases like no-match, but those are minor gaps for an image-matching tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 67%, and the description enriches parameter meaning beyond the schema by giving a concrete needle_rect use case (instruction glyph band in a GeeTest pattern) and explaining output coordinates in haystack-image pixels. Not every parameter is expanded, but session_id/page_id are standard context IDs and the schema already documents the crop parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific operation: multi-scale normalized cross-correlation of a needle against a haystack, and specifies the output as matched needle-center positions. It also names concrete use cases (captcha piece->gap, glyph->character, logo->page), which clearly differentiates it from generic vision or OCR siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description states 'THE tool for: captcha piece->gap, glyph->character, logo->page,' giving clear when-to-use context. It does not explicitly name excluded alternatives, but the use-case framing is strong enough to route an agent toward this tool for template-matching-style problems.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
page_move_toA
Move the mouse onto an element (hover) along a HUMAN-LIKE path: bezier arc, ease-in-out velocity, sub-pixel tremor, overshoot+correction. Triggers hover menus and tooltips. Pass the ref from page_a11y.
| Name | Required | Description | Default |
|---|---|---|---|
| ref | Yes | Element ref from page_a11y (e.g. "e12"). | |
| page_id | Yes | ||
| session_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the behavioral burden. It discloses meaningful movement traits: bezier arc, ease-in-out velocity, sub-pixel tremor, overshoot+correction, and the hover effect. It does not mention timing, failure cases, or whether actual system cursor movement occurs, so it is not a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded: the core action appears first, followed by the human-like movement detail and the practical purpose. Every sentence earns its place and there is no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple three-parameter hover action, the description covers what the tool does, why it is used, and where the required ref comes from. It omits return/error behavior and leaves the standard session/page identifiers undocumented, but these are minor gaps given the tool's low complexity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 33%, and the description does little to compensate. It repeats that ref comes 'from page_a11y', matching the schema, but session_id and page_id remain completely unexplained. The description adds essentially no meaning beyond the input schema for two of the three required parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb and resource: 'Move the mouse onto an element (hover)'. It also adds the human-like path and the hover menu/tooltip effect, which clearly distinguishes it from sibling actions like page_click_ref or page_drag.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It clearly indicates when this tool is useful: to trigger hover menus and tooltips, and it gives the prerequisite to pass a ref from page_a11y. It does not explicitly name alternatives or state when not to use it, so it falls just short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
page_network_bodyC
Fetch the BODY of a captured response by requestId — Network.getResponseBody BY PROTOCOL. The JSONP/API answer data that page JS can't see, served from the engine. THIS is the 0s and 1s layer.
| Name | Required | Description | Default |
|---|---|---|---|
| page_id | Yes | ||
| request_id | Yes | requestId from page_network_read. | |
| session_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, but it only presents the body as 'the 0s and 1s layer' and doesn't disclose operational behavior such as base64 encoding, availability only for captured responses, or error cases when the body wasn't saved. It adds a bit of context about engine-side data but no actionable behavioral detail.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The first sentence is direct and front-loaded, but the remaining sentences — 'THIS is the 0s and 1s layer' — are stylistic filler that adds little semantic value. Not bloated, but not every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a network-body tool with no output schema and no annotations, the definition omits the capture prerequisite, where requestId comes from (except a schema hint), and what the response contains or how it is encoded. An agent would need cross-tool inference to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 33%: only request_id has a description. The prose mentions requestId and 'captured response', but it doesn't explain session_id or page_id or the expected format/type of any parameter, so the description does not compensate for the low schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening sentence names a specific action ('Fetch the BODY'), a specific resource ('a captured response'), and the selection key ('requestId'), and even names the underlying protocol call. It does not explicitly distinguish itself from sibling tools like page_network_read, but the 'BODY... page JS can't see' phrasing implies the raw-response-body niche.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use it — when you need the raw JSONP/API response data that page JS conceals — but never states a prerequisite (network capture started, requestId obtained from page_network_read) or contrasts it with alternatives. The only cross-reference to page_network_read lives in the schema's request_id description, not in the tool description.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
page_network_readA
List captured HTTP responses [{url, requestId}] since capture start. Optionally filter by URL substring (e.g. 'geetest' or 'api.'). Pair with page_network_body to read the response content BY PROTOCOL.
| Name | Required | Description | Default |
|---|---|---|---|
| filter | No | Optional URL substring filter (e.g. "geetest"). | |
| page_id | Yes | ||
| session_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden of behavioral disclosure. It conveys a read-only listing operation and the 'since capture start' temporal scope, but it does not explicitly state that it never mutates state, nor does it disclose potential pitfalls like empty results or capture prerequisites (e.g., needing page_network_start). It is not misleading, but it lacks some explicit safety context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with zero fluff. The primary action and output are front-loaded, followed by the filter usage and the pairing hint. Every word earns its place, and the structure is easy to scan.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers the core behavior and filter, but it omits crucial context like the requirement to start a capture session via page_network_start (implied by 'since capture start' but not stated), and it doesn't mention error scenarios or how empty results are represented. For a tool with no output schema and no annotations, this is a notable gap in making it fully self-sufficient.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 33% (only 'filter' has a description). The description compensates partially by explaining the filter parameter with examples ('geetest' or 'api.'), but it does not clarify the required session_id and page_id, leaving those two parameters undocumented. It adds value for the filter but falls short of fully compensating for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('List'), a clear resource ('captured HTTP responses'), and the exact output format ([{url, requestId}]). It also distinguishes this tool from its sibling page_network_body by noting the pairing protocol, making its role unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly pairs this tool with page_network_body for reading response content, which signals when to use this vs. the sibling. However, it does not cover exclusions (e.g., when not to use it) or mention other network-related alternatives like page_console or page_errors, so it stops short of full routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
page_network_startA
PROTOCOL-LEVEL NETWORK CAPTURE: start recording every HTTP response for this page, BELOW the page (invisible to page JS, unpatchable). The answer data of any captcha/API travels here. Start before triggering the flow you want to see.
| Name | Required | Description | Default |
|---|---|---|---|
| filter | No | Optional URL substring filter (e.g. "geetest"). | |
| page_id | Yes | ||
| session_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral burden. It discloses that capture happens below the page, is invisible to page JS, is unpatchable, and that captcha/API answer data travels through it. It does not mention side effects, whether repeated starts reset capture, or how recording ends, but the disclosed traits are meaningful and non-obvious.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded with a clear protocol-level capture label. The two explanatory sentences add important context without padding. Minor redundancy between 'PROTOCOL-LEVEL NETWORK CAPTURE' and 'BELOW the page' keeps it from a 5.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description explains what, why, and when to start, but it omits how to stop or retrieve the recorded data, and does not address the optional filter parameter. Given there is no output schema and no annotations, an agent would still need to infer the full capture workflow from sibling tool names.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 33%: session_id and page_id have no descriptions in the schema, and the tool description does not explain their meaning or relationship. The description only generically refers to 'this page' without mapping it to page_id. The optional filter is documented in the schema but not reconciled with the claim of recording 'every HTTP response.'
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('start recording') and resource ('every HTTP response for this page'), and explains the protocol-level scope ('BELOW the page, invisible to page JS, unpatchable'). This clearly distinguishes it from sibling network-read tools and other page operations.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives a clear timing instruction: 'Start before triggering the flow you want to see.' This tells the agent when to invoke the tool relative to other actions. It does not explicitly name alternatives or say when not to use it, so it stops short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
page_ocrA
TIER 2 VISION — LOCAL OCR: extract text from an element (or the whole viewport if no ref). 100% local (pure-Rust ML models auto-download once to GHOSTFOX_HOME/models). Answers 'what text is written there' — for image captchas, canvas text, scanned UI. Models download on first call (~12MB, once).
| Name | Required | Description | Default |
|---|---|---|---|
| ref | No | Ref of the element to OCR. Omit for the whole viewport. | |
| page_id | Yes | ||
| session_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral burden and does well: it discloses that processing is '100% local', that models auto-download to GHOSTFOX_HOME/models, and that first call downloads ~12MB. It does not mention return format or failure behavior, but the core operational traits are covered.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded with the key action. There is slight redundancy between 'auto-download once to GHOSTFOX_HOME/models' and 'Models download on first call (~12MB, once)', which costs a point, but overall every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with no output schema and no annotations, the description covers purpose, scope, local execution, and the download side effect. It stops short of specifying the return value format or error cases, which an agent might need, but the information required to select and invoke the tool is largely present.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 33%: only 'ref' has a description. The tool description reinforces the ref semantics ('or the whole viewport if no ref') but does not add new meaning for 'session_id' or 'page_id'. These are standard context parameters, so the gap is minor, but the description does not fully compensate for the low schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action: 'extract text from an element (or the whole viewport if no ref)'. It also names concrete use cases ('image captchas, canvas text, scanned UI'). However, it does not explicitly differentiate from sibling tools like page_captcha_ocr or page_vision, so the boundary is implied rather than stated.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear when-to-use context: 'Answers "what text is written there"' and lists typical OCR scenarios. It does not provide explicit when-not-to-use guidance or name alternatives, but the intended use is evident from the phrasing and the 'LOCAL OCR' label.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
page_openA
Navigate to a URL in an existing session. Returns a page_id (string) that must be passed to all subsequent page tools. Waits for the page to load. If the page has iframes or shadow DOM, use page_a11y instead of guessing CSS selectors. Example: page_open(session_id, 'https://example.com') returns a page_id like 'abc123'.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | ||
| session_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden and adds key behavior: it returns a page_id, waits for the page to load, and clarifies that page_id must be reused. It could also mention timeouts or errors, but the critical operational behavior is disclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three purposeful sentences with no filler. The most important fact—navigation and the returned page_id—is front-loaded, and the example adds clarity without bloating the description.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter tool, the description provides enough to call it correctly: required inputs, return value, and how that return value connects to sibling tools. It does not cover failure modes or timeout behavior, but this is a minor gap given the tool's simplicity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must explain the two parameters. It does so by explaining the relationship between session_id and URL and providing a concrete example. Some parameter-level detail is still implicit rather than explicit.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly identifies the action as navigating to a URL within an existing session, linking it to the sibling tools by returning a page_id used by subsequent page tools. It is specific and distinguishes page_open from session_create and other page tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It states the context ('existing session') and gives conditional guidance to use page_a11y when iframes or shadow DOM are present. It could be more explicit about when not to use page_open, but the routing to an alternative is present.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
page_pixelsA
SUPERMAN GLASSES: render an element (canvas / img / background-image) as a compact luminance GRID of digits 0-9 the agent READS directly — see shapes, holes, object orientation, image layout WITHOUT a vision model. 0=black, 9=white. Captcha gaps appear as darker cells, upright skies are bright rows on top. Pass the ref from page_a11y. Returns JSON {w, h, grid:[rows of digits]}.
| Name | Required | Description | Default |
|---|---|---|---|
| ref | Yes | Ref of the element to render (canvas / img / background-image). | |
| grid_h | No | Grid height in cells (default 21). | |
| grid_w | No | Grid width in cells (default 32). | |
| page_id | Yes | ||
| session_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It fully discloses the output format (JSON with w, h, grid), the meaning of digit values (0=black, 9=white), and how to interpret visual artifacts. It also mentions the prerequisite of obtaining a ref from page_a11y. This is a read-only rendering tool with no side effects mentioned, which is appropriate. It does not mention error handling or auth, but the behavior is well-transparent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, information-dense paragraph. It front-loads the core purpose with the metaphor, then packs in output format, interpretation rules, and a source hint. Every sentence contributes value: the metaphor brands the concept, the digit-grid explanation clarifies use, and the interpretation examples aid the agent. It is concise relative to the amount of useful information conveyed.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given there is no output schema, the description correctly explains the return structure ({w, h, grid}). It also provides enough context for correct invocation: the required ref source, and grid dimension defaults are left to the schema (which documents default values). It does not mention error conditions or edge cases, but for a read-only grid renderer with clear input requirements, the description is largely complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 60% (descriptions for ref, grid_h, grid_w; none for session_id, page_id). The description adds meaningful context for the ref parameter ('Pass the ref from page_a11y') and explains the return format, which relates to grid dimensions. However, it does not elaborate on session_id or page_id, which are likely standard session/page identifiers. It compensates partially but not fully for the schema gaps.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('render') and resource (element as canvas/img/background-image), and specifies the exact output: a compact luminance grid of digits. It clearly distinguishes itself from vision-model tools by saying 'WITHOUT a vision model', and from siblings like page_screenshot and page_vision by its unique digit-grid format. The purpose is unambiguous and not a tautology.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives a concrete usage hint: 'Pass the ref from page_a11y', which tells the agent where to get the required ref. It also explains when the tool is useful (seeing shapes, holes, orientation without a vision model) and provides interpretative cues (captcha gaps darker, skies bright rows on top). It doesn't explicitly name alternatives or when not to use it, but the 'WITHOUT a vision model' phrase implies an alternative to vision-based tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
page_pressA
Press a named key (Enter, Tab, Escape, ArrowDown, ...) — e.g. Enter to submit a search box.
| Name | Required | Description | Default |
|---|---|---|---|
| key | Yes | Key name: Enter, Tab, Escape, Backspace, Delete, ArrowUp/Down/Left/Right, Home, End, PageUp, PageDown. | |
| page_id | Yes | ||
| session_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden of behavioral disclosure, but it only restates the action ('press') already implied by the tool name. It doesn't mention whether the press triggers navigation, key-up/key-down semantics, focus requirements, or any side effects — genuine gaps for an input-mutation tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence states the action and example with zero filler. Every word earns its place by either defining the key scope or showing a realistic use case.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a straightforward 3-param action tool, the description covers the core semantics but omits details like return behavior, errors, or whether key presses require a focused element. Against the rich sibling set, an agent could still succeed by relying on its understanding of key names and page context, but the lack of any behavioral detail keeps it at average.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 33%, so the description must compensate. It adds value by clarifying the key parameter with a concrete example and the ellipsis implying a known set, but it leaves session_id and page_id entirely unexplained, relying on sibling-tool convention for their meaning.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states a specific verb ('Press') and resource ('a named key'), immediately distinguishing it from mouse-action siblings like page_click and text-entry tools like page_type or page_fill. The example ('Enter to submit a search box') reinforces the intended action without ambiguity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear contextual usage via the search-box example, implying this tool is for discrete key presses rather than typing or clicking. It does not explicitly mention alternatives or when not to use it, so it stops short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
page_read_refA
Read the FULL value of an element by its page_a11y ref — no truncation. Use when the snapshot's 200-char preview isn't enough (body text, long input fields).
| Name | Required | Description | Default |
|---|---|---|---|
| ref | Yes | Element ref from page_a11y (e.g. "e12"). | |
| page_id | Yes | ||
| session_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the behavioral burden. It clearly identifies this as a read operation and adds the key trait 'no truncation', which is behaviorally meaningful. It does not describe error cases or exact return formatting, but for a read-only ref-based tool the core behavior is well exposed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences with no filler, and the most important information (read full value, no truncation) is front-loaded. The usage guidance is compact and directly actionable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read-by-ref tool with three parameters and no output schema, the description gives enough context: what it does, what sets it apart, and when to use it. It could mention return value shape or invalid-ref behavior, but these are minor given how clearly the core use case is defined.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 33%, with only 'ref' documented. The description adds meaning by explaining that the ref comes from page_a11y and that the point is reading the full, untruncated value. However, it does not clarify session_id or page_id, though their names are relatively self-evident from context.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a precise verb and resource: read the full value of an element using its page_a11y ref. It also differentiates itself from the snapshot tool by emphasizing no truncation, so an agent can distinguish it from page_a11y, page_snapshot, and the other page_* siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives an explicit when-to-use condition: when the snapshot's 200-char preview is insufficient, such as for body text or long input fields. This directly contrasts with the snapshot preview and makes the selection decision clear without needing to inspect sibling schemas.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
page_screenshotA
Capture a PNG screenshot of a page (viewport by default, full page with full_page=true). Saved under the session recordings dir; returns the file path. Feeds the live view when enabled.
| Name | Required | Description | Default |
|---|---|---|---|
| page_id | Yes | ||
| full_page | No | Capture the whole scrollable document instead of the viewport (optional). | |
| session_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full behavioral burden. It usefully discloses side effects and outputs: the screenshot is saved under the session recordings directory, returns the file path, and feeds the live view when enabled. This goes well beyond a bare 'capture screenshot' description, though it could still mention failure modes or prerequisites.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences with no filler. The key action is front-loaded, the parameter-dependent mode is stated clearly, and side effects are summarized in the remaining sentences. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers the action, output, and side effects reasonably well. However, with no output schema and no annotations, it leaves gaps around required parameters, prerequisite page/session setup, and error conditions, so an agent may not fully understand how to construct a correct invocation beyond the obvious session_id/page_id names.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 33%: full_page is documented in the schema, but the required session_id and page_id have no schema descriptions. The description does not explain what these identifiers refer to beyond the general 'page' and 'session recordings' context, so it fails to compensate for the low coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: capture a PNG screenshot of a page. It also clarifies the critical scope options (viewport by default, full page with full_page=true) and distinguishes this from sibling tools like page_snapshot by explicitly covering the image-capture behavior.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use the tool: whenever a page screenshot is needed, especially a visual record. However, it does not explicitly discuss alternatives such as page_snapshot, nor when screenshot might be inappropriate, leaving selection between siblings partially to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
page_snapshotA
Extract the visible text content of a page as plain text (token-friendly). Returns: url, title, and content (all visible text, no HTML). For semantic element data with refs and values, use page_a11y instead — it gives you interactive elements with roles and names. Use this when you just need to READ page content without needing to interact with elements.
| Name | Required | Description | Default |
|---|---|---|---|
| page_id | Yes | ||
| session_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It discloses that the operation is read-only ('READ page content'), specifies output format (url, title, content, no HTML), and implies no interaction with elements. It could mention whether it waits for page load or handles dynamic content, but for a simple snapshot it is adequately transparent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is brief and front-loaded with the primary action and output. Every sentence adds value: it states the output, clarifies the no-HTML format, and differentiates from a sibling. No filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's simplicity and the absence of an output schema, the description fully explains what is returned (url, title, content) and the nature of that content (visible text, no HTML). It also clarifies the use case and differentiation from a sibling, making it complete for an agent to call correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It does not mention session_id or page_id at all. While the parameter names are self-explanatory, the description fails to clarify their roles or any required format, leaving the agent to infer from names alone. This is a notable gap given the low coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb 'extract' and the resource 'visible text content of a page', and explicitly lists what it returns (url, title, content). It also differentiates from page_a11y by contrasting text-only output with semantic element data, making the tool's purpose unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit when-to-use guidance ('when you just need to READ page content') and when-not-to-use (for interactivity, use page_a11y instead). It names the alternative tool and the condition that selects it, leaving no ambiguity.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
page_typeA
Type text character-by-character into an element by CSS selector (human-like key events). Prefer page_type_ref when you have a page_a11y ref — it handles rich editors (Lexical/Draft/ProseMirror) and returns a verified receipt. Use this only when you only have a CSS selector and don't need rich editor support. Returns 'ok' on success.
| Name | Required | Description | Default |
|---|---|---|---|
| text | Yes | ||
| page_id | Yes | ||
| selector | Yes | CSS selector. | |
| session_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden and does disclose human-like key events and the success return value. However, it does not explain failure behavior, whether text is appended or replaces existing content, or any focus/visibility requirements.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three sentences with no filler. It front-loads the core operation, then gives routing guidance and the return value, so every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers the core purpose, the main alternative, and the return value, which is useful given no output schema. But with no annotations, it leaves operational gaps around error handling, content replacement, and preconditions that would be needed for fully confident invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 25%, and the description does not sufficiently compensate. It adds minimal meaning for 'text' and 'selector', but session_id and page_id remain unexplained, and selector semantics mostly repeat the schema's existing 'CSS selector' note.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action: typing text character-by-character into an element by CSS selector with human-like key events. It explicitly distinguishes this tool from page_type_ref, making the purpose and scope clear even without a title.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives explicit guidance: 'Prefer page_type_ref when you have a page_a11y ref' and 'Use this only when you only have a CSS selector and don't need rich editor support.' This clearly tells the agent when to use this tool versus the main alternative.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
page_type_refA
Type text into the element a page_a11y ref points at (inputs and rich editors). Returns landed chars as a receipt.
| Name | Required | Description | Default |
|---|---|---|---|
| ref | Yes | Element ref from page_a11y (e.g. "e12"). | |
| text | Yes | ||
| page_id | Yes | ||
| session_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral burden. It discloses the core write operation and the 'landed chars' receipt, which is genuinely useful. However, it omits side effects like whether existing text is replaced or appended, whether the element must be visible/editable, and what happens on failure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single sentence with high signal-to-noise: it states the target, allowed element types, and return receipt. No wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers target and return but omits operational details such as whether typing appends or replaces, focus/wait requirements, and failure modes. For a mutation tool with no annotations and no output schema, this leaves the agent to guess some important behaviors.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 25%, so the description must compensate. It gives meaning to ref (page_a11y element) and text (typed text), but session_id and page_id are left undocumented, and text format/behavior is not explained.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states it types text into the element a page_a11y ref targets, restricted to inputs and rich editors. The ref-based targeting differentiates it from sibling page_type, though it does not name the alternative explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It specifies the context: use this tool when you have a page_a11y ref and need to type into inputs or rich editors. No explicit when-not or alternative is named, but the scope is clear enough for an agent to route itself.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
page_upload_fileA
Upload a file to an input[type=file] by CSS selector. The file must exist on the machine running the engine.
| Name | Required | Description | Default |
|---|---|---|---|
| page_id | Yes | ||
| selector | Yes | CSS selector for the file input (input[type=file]). | |
| file_path | Yes | Absolute path to the file to upload. | |
| session_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden of behavioral disclosure. It adds a useful constraint about file existence and implies a mutating action, but it does not describe what happens after upload, whether the input is cleared first, or how failures are surfaced. No annotation contradiction exists.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences with no filler. The main action is front-loaded, and the key prerequisite is stated immediately afterward.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with no annotations and no output schema, the description is under-specified. It lacks guidance on session/page identification, return values, post-upload behavior, and failure handling, leaving the agent to infer several important details.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema already documents selector and file_path, and the description only reinforces that file_path must exist locally. session_id and page_id remain undocumented in both schema and description, which is a gap given only 50% schema description coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific operation: upload a file to an input[type=file] by CSS selector. This clearly distinguishes it from all sibling tools, none of which handle file uploads.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly conveys when to use the tool (to upload a file into a file input) and gives an important prerequisite: the file must exist on the engine machine. It does not explicitly mention when not to use it, but no sibling tool offers a comparable alternative.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
page_visionA
TIER 4 VISION — VISUAL CORTEX: detect word bounding boxes in an element (or viewport) via the built-in text-detection ML model. Returns JSON [{x, y, w, h}, ...] in image pixels. The first shipped visual-cortex model — see where text lives in ANY image (click-order captchas, canvas text, scanned layout). 100% local.
| Name | Required | Description | Default |
|---|---|---|---|
| ref | No | Ref of the element to analyze. Omit for the whole viewport. | |
| page_id | Yes | ||
| session_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries some burden. It states '100% local' (privacy), the input requirement (ref or viewport), and output format in image pixels. However, it does not disclose potential performance characteristics, model limitations, or whether the operation is read-only. The absence of annotations is mitigated somewhat by the '100% local' note, but a clear read-only hint is missing.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is fairly concise and front-loaded: the main purpose, output, and usage are in the first two lines. The subsequent sentences provide context on model capabilities and use cases. It's not overly verbose and every sentence adds value, though the 'TIER 4 VISION — VISUAL CORTEX' could be seen as redundant marketing fluff, but it helps set context for a specialized tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the complexity (vision model) and no output schema or annotations, the description covers the essential aspects: what it does, what it returns, how to invoke (ref or viewport), and key use cases. It doesn't fully explain input format details (e.g., how to specify a specific element vs viewport other than 'ref'), but it's adequate for an agent to understand and call the tool correctly. The example use cases aid in routing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 33%, with 'ref' having a description; page_id and session_id are undocumented. The description explains that 'ref' can be omitted for viewport, which adds meaning. For page_id and session_id, the description doesn't add anything—they are common identifiers, but given the low coverage, some compensation is expected. It's borderline acceptable because the main behavioral parameter (ref) is explained.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description explicitly states the tool detects word bounding boxes via ML model and returns JSON array. It names the resource (element or viewport) and the output format in image pixels. The 'visual cortex' metaphor and use cases (captchas, canvas text) make it clear what it is for. It's clearly distinct from OCR tools in the sibling list since it focuses on bounding boxes, not text extraction.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description mentions 'see where text lives in ANY image' and gives example use cases (click-order captchas, canvas text, scanned layout), implying when to use it. However, it does not explicitly state alternatives or when NOT to use it. For example, it doesn't say 'use page_ocr for text content, not bounding boxes.' It's implied but not explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
page_wait_forA
Wait until a CSS selector becomes visible on the page (replaces manual sleeps). Returns true if found, false on timeout.
| Name | Required | Description | Default |
|---|---|---|---|
| page_id | Yes | ||
| selector | Yes | CSS selector to wait for. | |
| session_id | Yes | ||
| timeout_ms | No | Timeout in milliseconds (default 10000). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the behavioral burden. It explicitly says the tool waits for visibility, returns true when found, and returns false on timeout, which tells an agent the operation does not throw on a negative result. It does not specify polling behavior or the exact definition of visibility, but the core behavior is transparent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences deliver the core behavior and return contract with no filler. The main action is front-loaded and the timeout outcome is stated immediately after.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a low-complexity wait utility, the description covers what the tool does, when it is useful, and what it returns on success and timeout. It omits edge-case behavior and relies on the schema for parameters, but nothing critical is missing for selecting and invoking it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 50%, with selector and timeout_ms described; page_id and session_id are left to convention. The description reinforces that the selector is the CSS selector being waited on but does not add meaning beyond the schema. This puts it at the baseline rather than adding parametric insight.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific action and target: 'Wait until a CSS selector becomes visible on the page.' It also states the drop-in purpose ('replaces manual sleeps') and distinguishes this from the mutation-oriented sibling page tools by focusing on waiting rather than acting.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The parenthetical 'replaces manual sleeps' gives clear context for when to use this tool instead of hard-coded delays. It does not explicitly name alternative tools or conditions to avoid using it, but the intended use case is reasonably clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
session_createA
Create a new browsing session: launches the engine with a fresh coherent identity. Returns session_id.
| Name | Required | Description | Default |
|---|---|---|---|
| proxy | No | Proxy URL, e.g. socks5://user:pass@host:port (optional). | |
| headful | No | Run with a visible browser window instead of headless (optional) — for humans who want to watch the agent work. | |
| platform | No | Restrict identity platform: "windows" | "macos" | "linux" | "android" (optional). | |
| profile_dir | No | Reuse a persistent profile directory (optional). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It does disclose that the tool launches the engine, creates a fresh coherent identity, and returns a session_id, which is useful. However, it omits session lifecycle details, cleanup expectations, resource cost, or side effects of reusing a profile_dir.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two short sentences, front-loaded with the action, and each sentence contributes: one defines what the tool does, the other states the essential output. There is no filler or repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a session-creation tool with no output schema and no annotations, mentioning session_id is valuable. Still, the description lacks lifecycle context, usage ordering relative to other tools, and any error or cleanup behavior, making it minimally complete rather than fully actionable.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, and the input schema already explains proxy, headful, platform, and profile_dir with helpful detail. The description adds no parameter-level meaning, but the schema carries that burden fully, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb and resource ('Create a new browsing session'), states the engine-launching behavior, and notes the key return value (session_id). It also distinguishes itself from siblings like page_open and identity_generate by framing this as the session-level creation step.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no guidance on when to use this tool versus alternatives such as identity_generate or page_open, and no prerequisites or ordering constraints are mentioned. An agent must infer that creating a session should precede page operations.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
session_evidenceA
READ-ONLY: Retrieve the complete audit trail for a session. No side effects — does not modify the session or browser. Returns JSON with: session_id, dir (evidence directory path), identity_toml (path to the identity file used), event_count (total tool calls recorded), events (array of all events with timestamps), and snapshots (list of page snapshot file paths). The session_id must come from a previous session_create call. For a nonexistent session, returns an error. Evidence persists after the session ends and includes: append-only events.jsonl, page snapshots (markdown), screenshots (PNG), and identity.toml. Use session_create first if you need a session.
| Name | Required | Description | Default |
|---|---|---|---|
| session_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full behavioral burden and handles it well. It explicitly states the operation is READ-ONLY, has no side effects, does not modify the session or browser, returns an error for nonexistent sessions, and notes that evidence persists after the session ends. It also discloses the append-only nature of the event log and the file types involved.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the most important signal ('READ-ONLY') and then provides a structured list of return fields. It is somewhat long but each section adds value: purpose, return shape, prerequisite, error behavior, and persistence details. No filler or repetition is present.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter read-only tool with no output schema and no annotations, the description is complete. It covers what the tool does, what it returns, the source of the required parameter, the error behavior, and what happens to evidence after the session ends. An agent has enough information to call it correctly and interpret its output.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% for the single session_id parameter, but the description compensates by stating that session_id must come from a previous session_create call. It does not detail the exact format or type beyond what the schema already provides, but the provenance requirement is the most critical semantic information and is clearly conveyed.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with 'READ-ONLY: Retrieve the complete audit trail for a session', giving a specific verb and resource. It clearly distinguishes this from sibling session tools like session_pages or session_me by focusing on the full audit trail, and it states the operation is non-mutating.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit prerequisite context: the session_id must come from a previous session_create call, and it advises 'Use session_create first if you need a session.' It also notes the error case for a nonexistent session. It does not explicitly name alternative tools for when this should not be used, but the prerequisite and error behavior provide clear usage context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
session_meA
Detect the currently logged-in username on the page (Reddit, X, GitHub, HN). Returns username or 'unknown'. Use before commenting to avoid duplicates.
| Name | Required | Description | Default |
|---|---|---|---|
| page_id | No | ||
| session_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden of behavioral disclosure. It explains that the tool returns 'username' or 'unknown', implying a read-only detection with a fallback value, and lists supported sites. It does not disclose detailed behavior around sessions, page state, or error handling, so some gaps remain.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two short sentences with no filler. The primary purpose is front-loaded, the return behavior is included, and the usage note earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool lacks annotations, output schema, and parameter explanations. While the description covers purpose and return value, it leaves the required session_id undefined and offers no details about the page context or limitations. This is insufficient for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description provides no parameter information. The required 'session_id' parameter is completely unexplained, and the optional 'page_id' is only vaguely hinted at by 'on the page'. This is a major gap for an agent trying to call the tool correctly.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Detect'), a clear resource ('currently logged-in username on the page'), and lists supported sites (Reddit, X, GitHub, HN). It also states the return value ('username' or 'unknown') and ties the tool to the commenting workflow, distinguishing it from page_comment and other siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is clear usage guidance: 'Use before commenting to avoid duplicates.' This tells an agent when to invoke the tool. However, it does not explicitly mention when not to use it or provide alternative tool comparisons.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
session_pagesA
POPUP/TAB VISION: list EVERY browser target — pages you opened AND popups the site opened (OAuth windows, payment flows). Popups are auto-attached and given a page_id usable with every page_* tool immediately. Returns JSON [{page_id, target_id, url}].
| Name | Required | Description | Default |
|---|---|---|---|
| session_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the transparency burden. It discloses non-obvious behavior: popups are automatically attached, get a page_id, and are immediately usable by all page_* tools, and it specifies the JSON return shape. It does not explicitly declare read-only status or failure behavior, but 'list' and the return-only description make the operation's nature clear.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, no filler. The first sentence labels and scopes the operation; the second adds the key behavioral detail and return format. The most important information is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter list tool with no output schema, the description gives the essential output format and a crucial behavioral detail (popups are immediately usable). It lacks a small amount of context around session_id provenance and empty/error cases, but overall an agent can call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema exposes only a required session_id string with no description, and schema description coverage is 0%. The tool description never mentions session_id, where to get it, or how it scopes the listed targets, so it fails to compensate for the schema's missing parameter documentation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description uses a specific verb and resource: it lists every browser target in the session, including user-opened pages and site-opened popups. It explicitly distinguishes itself from the page_* siblings by being the enumeration tool that returns page_id for all targets rather than operating on a single page.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Clear context is provided: use it whenever you need a complete inventory of pages and popups, especially OAuth/payment flows. It implies this should precede use of page_* tools because popups are auto-attached with usable page_id, but it does not explicitly name alternatives or give when-not conditions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
23 tool updates
v0.7.3- Added
confirm_action - Added
page_captcha_ocr - Added
page_captcha_rotate - Added
page_comment - Added
page_console - Added
page_contrast - Added
page_dismiss_modal - Added
page_drag - Added
page_errors - Added
page_geetest_click - Added
page_geetest_slide - Added
page_hcaptcha - Added
page_init_script - Added
page_match_image - Added
page_move_to - Added
page_network_body - Added
page_network_read - Added
page_network_start - Added
page_ocr - Added
page_pixels - Added
page_vision - Added
session_me - Added
session_pages
19 tool updates
v0.1.0- First observed
captcha_solve - First observed
identity_audit - First observed
identity_generate - First observed
page_a11y - First observed
page_click - First observed
page_click_ref - First observed
page_eval - First observed
page_fill - First observed
page_open - First observed
page_press - First observed
page_read_ref - First observed
page_screenshot - First observed
page_snapshot - First observed
page_type - First observed
page_type_ref - First observed
page_upload_file - First observed
page_wait_for - First observed
session_create - First observed
session_evidence
TDQS
Scored across 42 tools
Some overlap exists among the many vision, input, and captcha tools (e.g., page_snapshot vs page_a11y vs page_ocr, or page_type vs page_type_ref vs page_fill), though the descriptions do a good job of explaining when to prefer one over another. The highly specialized captcha solvers are each tied to a specific challenge type, but the sheer number of adjacent tools still creates selection ambiguity.
The vast majority of tools follow a clear page_<verb_or_noun> snake_case convention, with session_ and identity_ prefixes for lifecycle and identity concerns. Minor deviations like captcha_solve, confirm_action, and the noun-style page_console/page_a11y keep it from being perfectly uniform, but the overall pattern is predictable.
At 42 tools, this is well above the 25+ threshold and feels too heavy even for a broad stealth-browser and captcha-solving purpose. Many vision and captcha variants could likely be consolidated or surfaced through a smaller capability-oriented set. The count is not chaotic enough for a 1, but it will tax an agent's ability to choose efficiently.
The toolset covers the core lifecycle thoroughly: session creation, identity generation/audit, page navigation, interaction, reading, network capture, console/error introspection, captcha solving, and evidence retrieval. Minor gaps like explicit session teardown, page reload/back-forward navigation, or cookie management are missing, but agents can work around those with page_eval and navigation to a URL.
Maintenance
Related MCP Connectors
- openhelmOAuthai.openhelm
Autonomous cloud agent tasks: real browser + your tools, structured evidence-backed results.
Live browser debugging for AI assistants — DOM, console, network via MCP.
Stealth web browser for agents: search, fetch, click, download and type in persistent MCP sessions.
Stealth web automation for AI agents. Login, signup, navigate, screenshot.
Related MCP Servers
- AlicenseAqualityCmaintenanceEnables cloud browser automation through Browserbase and Stagehand, allowing LLMs to interact with web pages, take screenshots, extract data, and perform automated actions with support for proxies, stealth mode, and parallel sessions.141,768 npmApache 2.0
- AlicenseNot gradedqualityCmaintenanceThe Zero-Setup Local Browser MCP. Enables AI agents to control web browsers via CDP with zero vision tokens and high-speed DOM mapping.17 npmMIT
- AlicenseNot gradedqualityCmaintenanceMCP-first browser-control toolkit enabling AI agents to safely automate browser actions using redacted snapshots and ref-based interactions.Apache 2.0
- AlicenseNot gradedqualityCmaintenanceA local MCP server for Firefox that gives AI agents full control over the browser via a Unix socket, enabling automation of tabs, pages, cookies, and more without exposing any network ports.16 npmMIT