AginxBrowser
AginxBrowser is an agent-first HTTP browser service for seeing, reading, finding, acting on, and remembering the live web without Chromium.
Fetch/read pages as markdown/html/text with tiered HTTP/browser rendering, JS extraction, XHR capture, proxy/TLS stealth, and prompt-injection sanitization.
Search the web across 20 engines and 7 categories (general, news, code, packages, academic, AI, images), optionally fetching top-result content.
Act statelessly on one-off pages: click CSS selectors and evaluate JavaScript.
Run interactive sessions: create, list, clone, close persistent sessions; navigate, click, drag, type, scroll, wait, eval, set files, and read indexed page state.
Inspect sessions: cookies, storage, console, network/media sniffer, anti-bot challenges, verdict, and export replay scripts (bash/jsonl/flow JSON).
Manage named accounts: create, list, delete, verify, and log in as identities with private cookie jars and stable device personas.
Hand off login: import Chrome 'Copy as cURL' state or open a login page for human completion via live view.
Query local cache of fetched pages and searches before re-fetching.
Download files to disk with streaming, SHA-256, and resume.
Render artifacts: markdown to deterministic HTML with inline SVG diagrams; pages to PDF/PNG/PPTX/DOCX; animation timelines to MP4 with narration/audio/subtitles; and screenshots.
Automate flows: run recorded/editable flow JSON server-side with zero model tokens, composing with sessions and logins.
Provides academic search capabilities by querying the arXiv API, returning paper results with metadata.
Supports web searches through Baidu's search engine, contributing results to the meta-search feature.
Enables searching GitHub repositories via the GitHub API for code-related queries.
Allows web searches via Google, one of the engines used in the meta-search aggregation.
Integrates with Hugging Face Hub to search for AI models and related resources.
Allows operators to plug in a private Meilisearch index into the search endpoint for custom data search.
Supports searching for npm packages as part of the package search category.
Supports searching for PyPI packages as part of the package search category.
Includes Sogou as a web search engine in the meta-search feature, and also powers WeChat article search.
Provides code search via the Stack Exchange API, including Stack Overflow questions and answers.
Searches WeChat articles through Sogou's WeChat channel, with link resolution.
AginxBrowser
English | 中文
The Browser for AI Agents. See the live web. Read it. Act on it. Remember it.
A browser built for agents from the first line of code — not a human browser bolted onto automation. See the world, read it, search it, act on it, and keep what you read: one Rust binary with built-in V8, no Chromium required.
Humans have Chrome. Agents have AginxBrowser.
One binary, zero dependencies, instant service. The HTTP API is the whole interface — agents plug in and go.
The star that got our attention: Pierre Tachoire, co-founder of Lightpanda — the headless browser our bench measures against — starred the repo. 90 seconds on why that mattered to us.
Real pages rendered by AginxBrowser's diting engine (no Chromium) — Wikipedia, this repo, Rust. Screenshot it yourself →

Why Agents Need Their Own Browser
Measured against headless Chrome on the same 20 pages, same network (bench, 2026-08-28): 7.6× faster to agent-usable text (p50 532 ms vs 4 053 ms), ~10× less memory (227 MB for the whole process vs ~2.1 GB per Chrome page), and 0 hard failures where Chrome's --dump-dom produced no DOM on 5 of 40 loads. Re-run on v0.5.21 (2026-09-29) on a degraded-network day held the ratio — 5.4× p50, 468 MB whole-run vs 1.75 GB per page — both raw TSVs are committed. An agent's total cost is browser efficiency × model efficiency — this is the browser half.
Existing "browser automation" was built for humans or for one-shot scraping — not for agents:
AginxBrowser | Puppeteer/Playwright | Firecrawl | Browser-use | |
Designed for | Agents first | Human debugging | Scraping service | LLM wrapper |
Dependencies | Single binary, no Chromium | Chromium ~500MB | Docker ~1GB | Chromium |
Sees (screenshots) | ✅ built-in diting rendering engine | Needs Chromium | ❌ | Needs Chromium |
Reads | markdown + js_extract + fetch receipts | DIY | markdown | DIY |
Writes documents | ✅ | ❌ | ❌ | ❌ |
Finds (search) | ✅ 20 engines, 7 categories, merged | ❌ | ❌ | ❌ |
Acts | indexed session interaction | DevTools API | ❌ | LLM-driven |
Remembers | ✅ local fetch/search cache (SQLite FTS5) | ❌ | crawl cache | ❌ |
Protocol | HTTP | Node API | HTTP | Python |
TLS fingerprints | ✅ Chrome/Firefox/Safari/Edge | Plugin required | ❌ | ❌ |
CAPTCHA | ✅ detect + auto-wait + optional 2captcha | DIY | ❌ | ❌ |
Interactive sessions | ✅ persistent | ✅ | ❌ | ✅ |
Same-tier engines, not the tools in the table above. Cells are capabilities, not speed.
AginxBrowser | Obscura | Blitz | Lightpanda | |
What it is | Rust browser, V8, diting CSS paint | Rust headless browser, V8 | HTML/CSS engine (Stylo). Not an agent browser | Zig headless browser, V8 |
Screenshot | built-in paint, opt-in build | screenshots, screencast, PDF | paints a window | Hermes integration falls back to Chrome for screenshots |
Public CSS suite | none in CI | not published as WPT | WPT in CI, including SVG | not published as WPT |
License | Apache-2.0 | Apache-2.0 | Apache-2.0 and MIT | AGPL-3.0 |
An agent needs five things from a browser: see, read, find, act, remember. One binary covers them all — systemd-friendly, zero dependencies.
Core advantage: no Chromium. AginxBrowser inlines a full browser engine (V8 + Rust HTTP stack + the diting CSS/layout/paint rendering engine, with the Blitz/Stylo/Taffy lineage as its reference implementation). No Puppeteer, no Chrome, no Docker. One Rust binary under systemd is your agent browsing infrastructure.
Related MCP server: Pilot
Two Things Stateless Renderers Can't Do
Most new "agent browsers" are stateless, fingerprint-less one-shot renderers — fine for public pages, dead on arrival against Cloudflare or login flows. AginxBrowser goes the opposite way:
🔐 Real TLS fingerprints — stealth mode replicates the complete Chrome145 / Firefox133 / Safari / Edge TLS handshakes via BoringSSL (not just a UA string), switchable per request; Cloudflare Turnstile challenges wait automatically for
cf_clearance. Fingerprint-less engines eat 403s — we get through.🤝 Stateful interactive sessions — login state injectable and exportable (
session_create(cookies=...)↔session_cookies), surviving pagination and multi-step flows;persistent: trueeven survives idle eviction and server restarts — the same session id comes back logged in. One-shot engines throw state away.
Reference point: Cloudflare's Kitesurf explicitly ships neither real TLS-fingerprint negotiation nor persistent auth sessions — anti-bot and login territory is exactly where AginxBrowser plays.
Apache-2.0 open source, single binary — self-host today, no cloud lock-in.
Every Fetch Is a Receipt
Agents act on what a browser tells them, so the response reports what actually happened — not just "got a 200":
tier— which path served the page: plain HTTP (~100 ms) or the V8-rendered browser tier. An agent can see why a fetch was fast or slow.redirected_from— the full redirect trail.redirected_from[0]is the URL you asked for,urlis where the content actually came from — requested paired with effective, every hop visible.content_hash+changed_since_prev— every fetch is hashed; consecutive samples of the same URL can be diffed. A rate-limited origin serving the same frozen 200 body for days reads aschanged_since_prev: false— the cheapest drift detector there is.captcha_event— when a challenge page was detected (and solved, if a solver is configured), the response says so instead of handing over a challenge page as if it were content.
The local cache builds on the same idea: search hits come back with [§ heading] section prefixes so an agent knows where on the page a hit landed, and ranking fuses keyword relevance with freshness.
Capabilities
Tiered rendering: static pages over plain HTTP (~100ms); V8 spins up only when JS rendering is needed (~1-2s) — 90% of the bench page set served without spinning up V8 at all; every response reports which tier served it (
tierfield)Multi-engine meta-search: general web (Baidu / Bing / Sogou / WeChat / DuckDuckGo / Wikipedia / Hacker News), news (Bing News), code (Stack Overflow, GitHub, MDN), packages (npm, PyPI, RubyGems), academic (arXiv, OpenAlex), AI models (Hugging Face) — 20 engines across 7 categories, queried concurrently, merged and deduplicated. Operators can plug a private Meilisearch index into the same
/search. Search → read in one stepImage search:
categories=imageshits Baidu/Bing image indexes and returns direct binaryimage_urllinks (downloadable straight to jpg/png) plussource_urlprovenanceInteractive sessions: persistent browser sessions with indexed interaction (
state/click/input/scroll/eval) — agents browse like humans do, andsession_exportturns what an agent figured out into a runnable curl replay script (zero model tokens on re-run) — or, withformat=json, into a flow document (flow_runreplays it server-side with{{var}}substitution,wait/expectgates and saved outputs; installed flows live inworkflow/<name>/flow.json, dropped in without a rebuild). Session tools also cover the acting part:session_viewportsimulates device viewports (media queries respond),session_waitblocks on a selector or predicate with a timeout,session_screenshotrenders the live state,session_consolereplays the page's console ring, andsession_storageexports/restores cookies plus localStorage for login hand-offPlayback-link sniffer:
session_network(filter=media)extracts the m3u8/mp4/dash URLs a page's player actually requested at runtime — links found only in page HTML are often decoys, so the request log is the source of truth.GET /session/{id}/harexports the same traffic as HAR 1.2 (retained bodies included)File download: streaming to disk (no memory buffering), SHA-256 integrity, resume of interrupted transfers — for binaries, archives, datasets
Local cache that remembers: every fetch/search lands in SQLite (FTS5) at
~/.aginxbrowser/cache.db— a re-fetch inside the TTL answers from what the agent already read instead of re-paying network time, with CJK substring matching,[§ heading]section-aware snippets, per-URL content hashes for drift detection, and per-session scoping for shared deploymentsCAPTCHA handling: type detection with automatic Cloudflare challenge wait and optional 2captcha integration — search never stalls on verification pages
JS data extraction:
js_extractpullswindow.__INITIAL_STATE__and other structured data out of SPAsDocument generation:
POST /render_markdownturns markdown into a deterministic, self-contained HTML artifact — the document layer, so agents never write HTML by hand. Prose rides a plain offline shell (no fonts, no scripts); fencedarchifyblocks carry typed zero-coordinate diagram JSON (sequence / workflow / architecture / dataflow / lifecycle families) and render to inline SVG via the layout engine. Same input, same bytes — the receipt carries the sha256 so determinism is verifiable.theme(light/dark) andpreset(classic / signal-flow / blueprint / editorial) bake colors at generation time;quality: "showcase"is the delivery gate, grading route crossings, label clearance and rhythm without touching the artifact bytes. Guided-view tabs pluswindow.agxViewer(focus/ ego /route/reach) make the artifact interactive. Mermaid sources are the agent's job to translate into archify JSON, not the engine's. Diagram vocabulary adapted from archify (MIT)Screenshot rendering:
/screenshotendpoint (opt-in--features screenshot) paints the JS-rendered DOM with the diting rendering engine — pure CPU, no Chromium — to PNG. Vision input for agentsTimeline video:
/videorenders a page's animation timelines to MP4 — the page's scripts register GSAP-style timelines inwindow.__timelines(duration()+pause(t)), each frame seeks tot=i/fpsand paints the viewport, and the frames pipe into ffmpeg (H.264, yuv420p). Deterministic by construction: no wall clock in the pixel values, same render twice = same MP4. Needs ffmpeg on PATHPage set (PDF/PNG/PPTX/DOCX):
/pdfcuts a rendered page into pages and packages them — print mode paginates at top-level block boundaries (default A4 @96dpi, no half-cut text where a break can land on a block edge), slides mode makes one page per CSS-selector match sized to the element (an HTML deck with one.slideper page exports as a real deck). Image-based PDF: per-page JPEG via DCTDecode, hand-rolled PDF 1.4 writer, zero new dependencies. PPTX packages the same pages as one slide per page; DOCX as one page-sized section per page — both hand-rolled OOXML (stored-ZIP writer, fixed timestamps), byte-deterministic, zero new dependenciesTLS fingerprint spoofing: stealth mode impersonates Chrome145/Firefox133/Safari/Edge, switchable per request
Firecrawl compatible:
/v1/scrapeendpoint — existing Firecrawl clients migrate by changing the base URLDNS rebinding protection: built-in SSRF guard + post-resolution IP validation
A Browser, Not a Crawler
AginxBrowser exists for real-time retrieval: an agent arrives with a question, reads a handful of pages, leaves with the answer. It is not a crawling tool — and the product is shaped so it can't quietly become one:
robots.txt is not our gate. The RFC 9309 checker ships built in, but a real-time lookup layer isn't a crawler and doesn't do crawler etiquette by default; operators who want it set
AGINXBROWSER_HONOR_ROBOTS=1.No site-walking API. There is no crawl endpoint and no link-following recursion — every page load happens because an agent asked for that page.
Built-in budgets. Per-domain: 20 pages/minute. Per interactive session: 200 pages. Toggled via
AGINXBROWSER_DOMAIN_RATE_PER_MIN/AGINXBROWSER_SESSION_PAGE_LIMIT(0disables on your own instance). Generous for an agent grinding through docs or a console; fatal to the page-after-page crawl pattern, including subdomain rotation (one registrable domain, one budget).The hosted instance (browser.aginx.net) runs tighter budgets. Every user shares one egress IP, and keeping sites comfortable with that IP is part of the service. Self-host if you want different numbers.
Need to bulk-crawl a site? Use a crawler. This isn't one, and it won't become one.
What It's For
Not demos — real jobs agent browsers are doing today:
Grind through admin consoles — AWS / App Store Connect / Google Play, dozens of menu layers per task. Let the agent click; it comes back only when authorization is needed.
Batch actions behind login — fill carts, dig through order history, check pages that only render while logged in. Inject cookies, operate, export for reuse.
Past anti-bot walls — Cloudflare protection, Turnstile challenges, TLS fingerprint checks. Stealth mode pushes through instead of retreating at 403.
The Chinese internet — Baidu / Sogou / WeChat meta-search across 5 engines, correct Chinese page rendering. Not English-web-only.
On-the-spot scripting — agent reads the page, writes JS, evals it: highlighted comparison tables, reflowed content, product filters on hidden parameters. GreaseMonkey-on-steroids.
Multimodal vision — screenshots as visual input for look-and-judge flows: picking seats, recognizing layouts, verifying rendering.
Where It Sits in the Computer-Use Stack
Computer-use agents come in two layers. GUI-layer stacks (Cua, desktop CUA agents) drive a whole machine: screenshots of a display in, X11 mouse/keyboard events out, a VM or container per session. Engine-layer browsers skip the desktop — the page itself is the machine. AginxBrowser is the engine layer:
GUI layer (desktop CUA) | AginxBrowser (engine layer) | |
Action space | screen pixels → OS input events | DOM: click by coordinates or selector, real event dispatch |
State readout | screenshot of a display | structured DOM / accessibility tree, screenshot optional |
Per-session cost | Xvfb + VM/container | one process, no display server |
Frames | flattened into pixels | first-class: hit-testing descends into iframes, frame-scoped evaluate |
Replay | best-effort (compositor timing) | action logs export as runnable curl / flow JSON |
Quick Start
One-command full install (SKILL.md trigger surface + verification):
# Download -> inspect the contents -> run only after review (never blind-run network scripts)
curl -fsSL https://raw.githubusercontent.com/yinnho/aginxbrowser/main/skill.sh -o skill.sh
less skill.sh
bash skill.shInstall the skill trigger surface via skills.sh:
npx skills add yinnho/aginxbrowserOr just hand this block to your AI agent — it sets itself up:
Set up aginxbrowser as your web access layer:
1. Install it locally and verify it is healthy:
brew install yinnho/aginxbrowser/aginxbrowser
aginxbrowser doctor # then start: aginxbrowser (listens on 0.0.0.0:8089)
curl http://127.0.0.1:8089/health
2. From now on, route web tasks through http://127.0.0.1:8089 instead of raw curl:
- read a page / scrape a URL -> POST /fetch {"url":...} (markdown output, Cloudflare bypassed by default)
- search the web -> POST /search {"q":...,"fetch_top":3}
- see a page -> POST /screenshot {"url":...}
- login / form / click-through-> POST /session/create -> /session/{id}/state -> /input /click -> /closeSelf-hosting:
# macOS / Linux via Homebrew
brew install yinnho/aginxbrowser/aginxbrowser
aginxbrowser doctor # features + fonts + egress self-check
# Docker (Docker Hub, mirrored on GHCR)
docker run -p 8089:8089 yinnho/aginxbrowser:latest
# (or ghcr.io/yinnho/aginxbrowser:latest)
# Or the prebuilt binary (platform detect + sha256 + mirror fallback + doctor self-check)
# macOS / Linux / Windows (git-bash; prebuilt Windows ships from v0.3.1, full `stealth`+`screenshot` feature set from v0.4.0)
# Cautious: download -> inspect -> run (never blind-run network scripts)
curl -fsSL https://browser.aginx.net/install.sh -o install.sh
less install.sh && bash install.sh
# Or straight in, if you trust the repo:
# curl -fsSL https://browser.aginx.net/install.sh | sh
# GitHub slow/blocked? AGINXBROWSER_GH_PROXY=https://ghfast.top/ bash install.sh
aginxbrowser doctor # features + fonts + egress self-check
# Or build from source (--features stealth,screenshot or you lose both)
cargo build --release --features stealth,screenshot
# Start the service
./target/release/aginxbrowser
# → Listening on 0.0.0.0:8089
# Verify
curl http://127.0.0.1:8089/health
# → {"status":"ok","engine":"diting"}
# Fetch a page
curl -sS -X POST http://127.0.0.1:8089/fetch \
-H "Content-Type: application/json" \
-d '{"url":"https://example.com"}'
# Search (snippets only by default; fetch_top grabs page bodies for the top N)
curl -sS -X POST http://127.0.0.1:8089/search \
-H "Content-Type: application/json" \
-d '{"q":"macbook price","max_results":5,"fetch_top":2,"max_chars_per":2000}'
# Create an interactive session
curl -sS -X POST http://127.0.0.1:8089/session/create \
-H "Content-Type: application/json" \
-d '{"url":"https://example.com"}'
# → {"session_id":"s_1","url":"https://example.com/"}REST Routes
Every capability is plain HTTP — no SDK required. There is no /openapi.json (the routes are few enough to list here); full request/response fields for each route are in docs/API.md.
Method | Path | What it does |
GET |
| Liveness + build commit, UA, TLS, capabilities |
GET |
| Deep self-check: engines live, fonts, egress |
GET |
| Search engine catalog with live suspension state |
POST |
| Fetch a page → text/markdown/html. Params: |
POST |
| Multi-engine meta-search. Params: |
POST |
| Click a CSS selector on a page |
POST |
| Evaluate JavaScript on a page |
POST |
| Stream a file to disk (sha256, resume) |
POST |
| Render page → PNG ( |
POST |
| Render animation timelines → MP4 ( |
POST |
| Paginate page → PDF/PNG/PPTX/DOCX ( |
POST |
| Firecrawl-compatible scrape (+ |
POST |
| Run a recorded/edited flow JSON to completion — zero model tokens ( |
POST |
| Start an interactive session (cookies/UA carried across calls) |
POST |
| Create a logged-in session from a DevTools "Copy as cURL" |
GET |
| Live sessions |
POST |
| Session actions |
GET |
| Session inspection |
POST |
| Markdown → deterministic self-contained HTML artifact (+ inline-SVG diagrams) |
Every capability is one POST away — no SDK, no protocol adapter.
Project Layout
aginxbrowser/
├── Cargo.toml
├── build.rs # V8 snapshot generation
├── js/
│ └── bootstrap.js # V8 bootstrap script
├── workflow/ # Flow assets: <name>/flow.json replayed by flow_run (drop-in, no rebuild)
├── README.md
├── docs/
│ └── API.md # Full API reference (HTTP)
├── bench/ # Benchmark harness + results (vs headless Chrome)
│ ├── README.md # methodology + numbers
│ ├── pages.txt # fixed 20-page set
│ ├── run.py # harness
│ ├── summarize.py # TSV → results table
│ └── results/ # raw run data
└── src/
├── main.rs # HTTP service entry & routing
├── server.rs # Business layer (fetch/click/eval/search)
├── session.rs # Interactive browser sessions
├── docgen/ # Document layer: markdown → deterministic HTML + inline-SVG diagrams
├── render.rs # Tiered rendering (HTTP direct → diting browser engine)
├── store.rs # Local fetch/search cache (SQLite FTS5, drift hashes)
├── download.rs # Streaming file download (sha256, resume)
├── robots.rs # RFC 9309 robots.txt checker (opt-in gate)
├── rate.rs # Per-domain + per-session budgets
├── captcha.rs # CAPTCHA detection & auto-solve
├── firecrawl_compat.rs # Firecrawl-compatible /v1/scrape endpoint
├── video.rs # Timeline video pump (__timelines seek → ffmpeg → MP4)
├── pages.rs # Page pump (print/slides pagination → PDF/PNG)
├── ooxml.rs # OOXML containers (image-based PPTX/DOCX, stored-ZIP writer)
├── doctor_cli.rs # `aginxbrowser doctor` self-check
├── browser.rs # Top-level API: Browser, BrowserBuilder
├── page.rs # Top-level API: Page, Element
├── config.rs # BrowserConfig
├── cookie.rs # CookieStore
├── error.rs # Error types
├── search/ # 20 native search engines, 7 categories
│ ├── mod.rs # SearchEngine trait, Registry, merge/dedupe, progressive backoff
│ ├── baidu.rs # Baidu (JSON API, wreq stealth)
│ ├── baidu_images.rs # Baidu Images (acjson API, images category)
│ ├── bing.rs # Bing (HTML parsing, plain reqwest)
│ ├── bing_images.rs # Bing Images (images/async endpoint, images category)
│ ├── bing_news.rs # Bing News infinite-scroll fragment (news category; direct-first/proxy-retry)
│ ├── sogou.rs # Sogou web (HTML parsing, plain reqwest)
│ ├── sogou_wechat.rs # Sogou WeChat (HTML parsing + /link resolution)
│ ├── duckduckgo.rs # DuckDuckGo (html.duckduckgo.com, general; direct-first)
│ ├── wikipedia.rs # Wikipedia (MediaWiki search API, general; direct-first/proxy-retry)
│ ├── hn.rs # Hacker News (Algolia API, general; time_range filters created_at)
│ ├── stackexchange.rs # Stack Overflow (SE API v2.3, code category)
│ ├── mdn.rs # MDN Web Docs (v1 search API, code category only)
│ ├── github_repos.rs # GitHub repos (api.github.com, code category)
│ ├── arxiv.rs # arXiv (Atom API, academic category)
│ ├── openalex.rs # OpenAlex works (academic; DOI links, inverted-index abstracts)
│ ├── huggingface.rs # HF Hub models/datasets/spaces (ai category)
│ ├── npm.rs # npm packages (npms.io API, packages category)
│ ├── pypi.rs # PyPI name resolution (JSON API, packages)
│ ├── rubygems.rs # RubyGems gems (packages; direct-first/proxy-retry)
│ └── meilisearch.rs # Private-index adapter (env-configured)
│
├── diting_dom/ # HTML parsing, DOM tree, CSS selectors
├── diting_css/ # CSS parsing + cascade
├── diting_net/ # HTTP client, cookies, encoding, proxies
├── diting_js/ # V8 runtime, JS ops, module loading
├── diting_layout/ # Taffy-based layout, floats, hit-testing
├── diting_fonts/ # Bundled CJK font subset, fallback
└── diting_browser/ # Page navigation, lifecycle, browser contextBuild
# Standard build (no stealth; TLS fingerprint features inactive)
cargo build --release
# With stealth (requires go + cmake + C++ toolchain; enables TLS fingerprint spoofing)
cargo build --release --features stealth
# With screenshot rendering (enables /screenshot; adds the rendering stack, +30-40MB)
cargo build --release --features screenshot
# Full featured (recommended for production)
cargo build --release --features stealth,screenshotRequirements: Rust 1.78+; the V8 static library downloads automatically on first build. The stealth feature additionally needs go, cmake, and a C++ compiler. The screenshot feature ships with a bundled CJK font subset (GB2312 + common symbols) — no system fonts required for correct Chinese rendering.
If your network can't reach the rusty_v8 CDN (build hangs with zero progress after "downloading v8"), pre-fill ~/.cache/rusty_v8 with the librusty_v8.a.gz for your version (fetch it from any reachable mirror/host and gunzip into place) and the build script skips the download.
Runtime Environment Variables
Variable | Default | Description |
|
| Listen address |
| enabled |
|
| Linux Chrome145 | Spoofed User-Agent |
|
| Accept-Language header |
| none | Optional fallback proxy. Blocked-source engines (Wikipedia, Bing News, Hugging Face, RubyGems) connect directly first and fall through to this proxy only when the direct attempt fails — overseas deployments need no proxy at all; per-request |
|
| JS navigation-chain cap: documents a page may chain via |
|
|
|
| unset | robots.txt is not consulted by default on |
| unset | Opt in to |
| unset | Opt in to loopback/RFC1918/link-local fetches (the SSRF gate). Same as the |
| unset | Scoped alternative: comma-separated CIDR allowlist (e.g. |
| unset | Directory of extra fonts ( |
|
| Per-host robots.txt policy cache TTL |
|
| Per-registrable-domain page budget per minute (subdomains share one budget); over-budget requests get 429 with the stance message. |
|
| Total pages one interactive session may walk (navigation-causing clicks count); over-budget navigations are refused, the current page stays interactive. |
| on | Local fetch/search cache; |
|
| SQLite database location (created 0600) |
|
| Cached page TTL |
|
| Cached search-result-set TTL |
|
|
|
| none | 2captcha API key; enables CAPTCHA auto-solving |
|
| CAPTCHA solving provider |
| none | Meilisearch base URL; set to enable the private-index engine |
| none | Meilisearch index uid to query |
| none | Optional Bearer key for the Meilisearch instance |
API Documentation
Full API reference → docs/API.md
Security audit notes → docs/skills-sh-audit.md — why skills.sh shows "Critical Risk", and which real product feature each warning corresponds to
Covers:
All HTTP endpoints (
/fetch,/search,/screenshot,/video,/pdf,/download,/render_markdown,/v1/scrape,/flow/run,/doctor, the session endpoints)Environment variables, error codes, per-site scraping examples
Plugging Into Other Systems
AginxBrowser is pure attach-alongside infrastructure — like a real browser, it runs as an independent service that anything can call, without embedding host code or polluting host config. Deploy one instance per machine (under systemd) and every app needing "render + scrape" capability shares it.
The attach point is HTTP — /fetch, /search, /screenshot, /download, /render_markdown for any language with an HTTP client.
Integration: read the environment variable AGINXBROWSER_URL=http://127.0.0.1:8089. Unset → behavior unchanged; set → risk-controlled sites automatically route through AginxBrowser for rendering, falling back gracefully on failure.
Known Limitations
Screenshots are opt-in:
/screenshotrequirescargo build --release --features screenshot(adds the diting rendering stack). The default (and only) render engine in that build is diting — our own CSS+layout+paint stack, zero Blitz/Stylo code. The pinned-rev Blitz reference pipeline is a separate opt-in,--features blitz-reference, for comparison renders and the dual-engine cross-check tests. Complex-site CSS is approximate on both (not pixel-perfect like Chromium)Element coordinates supported:
/screenshotwithselectorreturns element page coordinates (selector_rects, CSS px);selectoralone crops directly to that element. Inline elements (<a>text</a>) get a rect too on the default diting engine — a union of their flattened inline content, strut-expanded to the element's ownline-heightlike Chrome reports for replaced-only inlines (<a><img></a>→ line-box height, not the image height). Empty inlines still have no rect — pick a block ancestor thereJS interaction broadly works; heavy-fingerprint pages may still fail: React/Vue event delegation works normally (URL-reflection attributes like
src/hrefresolve to absolute URLs so Next.js/webpack hydrate and clicks trigger handlers). Heavy-fingerprint auth pages (WorkOS/Cloudflare) probingnavigator.plugins, WebGL canvas etc. may still break until stealth fingerprint coverage completesProxy support: HTTP/HTTPS/SOCKS5 via
AGINXBROWSER_PROXYHard risk-controlled sites: Baidu Wenku unsupported; Zhihu articles need a valid
__zse_ck
Star History
If AginxBrowser saved you a headless-Chrome fleet or a scraping headache, a star is how other agents (and their humans) find the project.
License
Apache-2.0.
Available Tools
40 toolsaccount_deleteAInspect
Delete a named login identity: stored record AND live jar. Cookie values are credentials — delete means gone. Sessions currently running as the account keep their in-process jar handle, but nothing writes back. Returns {deleted: name}, or an error naming the account if it does not exist.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | The account to delete: stored record AND live jar. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With only a title annotation, the description carries the full behavioral burden and meets it excellently. It discloses that both the stored record and live jar are removed, that cookies are credentials and deletion is permanent, that running sessions retain an in-process handle but nothing writes back, and the exact return/error shape.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, each earning its place: the first defines scope, the second conveys permanence and gravity, the third covers runtime behavior and return values. Information is front-loaded and there is no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter destructive tool with no output schema, the description is remarkably complete. It covers what is deleted, the irreversible nature, session edge-case behavior, and the expected response and error format. An agent has everything needed to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the parameter description already states 'The account to delete: stored record AND live jar.' The tool description adds the context that this is a 'named login identity' but does not provide additional format, constraints, or syntax beyond the schema. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Delete') and a specific resource ('a named login identity'), and clarifies scope by naming both the stored record and the live jar. This clearly distinguishes it from siblings like account_list and account_verify.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The destructive intent is unmistakable: 'delete means gone' and 'Cookie values are credentials' make clear this is the permanent removal tool, not list or verify. It lacks an explicit 'use this when...' or 'instead of...' statement, but the context is strong enough that an agent would not confuse it with alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
account_listARead-onlyInspect
List named login identities (the multi-account layer) with metadata only: name, cookie domains, cookie count, updated_at, the last account_verify verdict, and the identity's persona User-Agent (each account is one stable device: its own UA and hardware fingerprint, drawn once and reused). Cookie values are credentials and never leave the server. Use to see which identities exist before session_create {account} picks one.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, but the description adds meaningful behavior beyond that: metadata-only results, cookie values are credentials that never leave the server, and each account is a stable device with reused UA and fingerprint. These details help an agent understand data sensitivity and account identity semantics.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a compact, front-loaded paragraph that first states the action, then the returned metadata, then a security note, then usage. Every sentence contributes a distinct, useful piece of information without repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, but the description lists the returned fields explicitly: name, cookie domains, cookie count, updated_at, last account_verify verdict, and persona User-Agent. Combined with the readOnly annotation and zero-parameter schema, this gives the agent enough context to invoke the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so there are no parameter semantics to explain. The description does not need to compensate for missing input schema documentation, and the baseline for a no-parameter tool is appropriate here.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Starts with a specific verb and resource: 'List named login identities (the multi-account layer)'. It then enumerates the exact metadata scope, including fields like cookie domains, updated_at, and account_verify verdict, making it clearly distinguishable from related account tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides explicit usage context: 'Use to see which identities exist before session_create {account} picks one.' This tells the agent when to call it, though it does not state when not to use it or name alternative listing tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
account_loginAInspect
Open a site's login page AS a named account and close the login loop. Creates the session as the account (private jar, device persona), navigates to url, and reports which generic login gates the page shows — needs: password | sms | qr | slider (QR scan like xiaohongshu, SMS code, password form, slider/captcha) — plus a session_verdict classification and, when a human step is needed, the session_id and a /live handoff so a person can finish it in the live view (the engine detects and describes; it never fills credentials or solves challenges). With predicate and no human gates detected, waits for the automatic bounce, then teaches + stamps the account's verify spec — later account_verify calls re-check it bare. Cookies write back after every action, so a login finished in /live is already persisted. status is "logged_in" (verify stamped), "waiting" (drive session_wait/session_input on session_id, or re-call account_login after the human finishes — the account's shared jar already holds it), or "opened" (no predicate given).
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | The login page URL to open as this account. | |
| name | Yes | The account to log in as (created implicitly on first use; 1-64 chars of [a-zA-Z0-9_-]). | |
| predicate | No | A JS expression truthy on the page the site lands on AFTER login, e.g. `!!document.querySelector('.user-nick')`. With it and no human step detected, the call waits for the automatic login bounce and stamps the account's verify spec on success. Without it, the call just opens the page and reports what kind of login it sees. | |
| use_proxy | No | Route through the engine proxy. Seeds a fresh account; an account with an existing record reuses its recorded egress. | |
| timeout_ms | No | Wait budget in ms for the automatic-login bounce (default 60000, clamped 1000..120000). Never spent while a human step (QR/SMS/password/slider) is outstanding — those return immediately with a session handoff. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations provide only a title, so the description carries the full burden of behavioral disclosure. It thoroughly explains that the tool creates a session, never fills credentials or solves challenges, writes cookies back after every action, and reports statuses ('logged_in', 'waiting', 'opened'). It also discloses the handoff mechanism and verify spec stamping. No contradictions with annotations exist.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long but densely packed with necessary information. It front-loads the core purpose, then covers predicate behavior, cookie persistence, and status outcomes. While it could be better structured with bullet points, every sentence adds value and avoids fluff. It is appropriately sized for the tool's complexity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity, the description covers the essential behaviors: login gate detection, session handoff, verify spec stamping, and status values. It mentions the session_verdict classification and the /live handoff but does not detail every edge case, such as error handling or exact return structure. Still, it provides enough for an agent to call the tool correctly in most scenarios.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3. The description adds meaningful context beyond the schema: it explains how 'predicate' drives the automatic-login wait and stamps the verify spec, how 'use_proxy' seeds a fresh account, and how 'timeout_ms' is never spent during human steps. These details help an agent choose and configure parameters correctly, exceeding the baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Open a site's login page AS a named account and close the login loop.' It clearly states the tool's primary function and distinguishes it from siblings like account_verify (which re-checks a verify spec) and session_* tools. The purpose is unambiguous and actionable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context for when to use the tool, including the predicate parameter for waiting on automatic login bounce, and how to handle human gates by using session_wait/session_input or re-calling account_login. It references account_verify as a follow-up step, but does not explicitly state when to prefer alternatives like session_navigate or session_create. This is strong guidance but not fully explicit about exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
account_verifyAInspect
Check whether a named account is still logged in. Teach-once: the first call passes url + predicate (a JS expression truthy on a logged-in page, e.g. !!document.querySelector('.user-nick')); the spec is remembered and later calls can be bare. Runs in a scratch session AS the account (private jar), so the probe doubles as a cookie refresh. Returns {name, logged_in, url, checked_at}.
| Name | Required | Description | Default |
|---|---|---|---|
| url | No | Teach-once: the page that shows login state (its login wall if the account is logged out). Remembered after the first call. | |
| name | Yes | The account to check. | |
| predicate | No | Teach-once: a JS expression that is truthy when logged in, e.g. `!!document.querySelector('.user-nick')`. Remembered after the first call — later calls can pass neither and rerun the spec. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no readOnly or destructive annotations, the description carries the full burden and excels: it discloses the private-jar scratch session, the fact that the probe doubles as a cookie refresh, and the remembered-spec statefulness. These are exactly the behavioral traits an agent needs beyond the schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three dense sentences, each earning its place: primary purpose, teach-once behavior, and side-effect/return information. The main verb and resource are front-loaded, with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Although there is no output schema, the description explicitly gives the return shape {name, logged_in, url, checked_at}. Combined with the side-effect, statefulness, and session context, nothing essential is missing for an agent to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the schema already documents url and predicate thoroughly, including the teach-once behavior and example predicate. The description reinforces this but adds no meaning beyond the schema, so it stays at the baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Check whether a named account is still logged in.' This clearly differentiates it from siblings like account_list and account_delete, and the teach-once detail further scopes what the tool does.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives clear operational context: first call passes url and predicate, later calls can be bare, and the check runs in a scratch session as the account. It does not explicitly name alternatives or exclusion conditions, but the tool's unique purpose makes the intended usage unambiguous.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
cacheAInspect
Query the LOCAL CACHE of every page this server has fetched and every search it has run. Check here BEFORE re-fetching or re-searching — a hit is instant and free while a fresh fetch costs 5-60s. Use query for full-text search (works for Chinese substrings and English words), get to pull a page's full cached content, stats for counts, clear to delete rows.
| Name | Required | Description | Default |
|---|---|---|---|
| all | No | With clear: delete everything cached for this caller | |
| get | No | Return the FULL cached content of this exact URL instead of listing hits | |
| url | No | Only rows whose URL contains this substring | |
| kind | No | Which rows to search: "auto" (default, pages + searches), "pages", or "searches" | |
| clear | No | Delete matching rows instead of returning them (requires url, since_hours, or all) | |
| limit | No | Maximum rows returned (default: 10, max 100) | |
| query | No | Full-text search over cached page contents, titles, URLs and past search queries. Omit to list the latest rows. | |
| stats | No | Return row counts and database size instead of rows | |
| since_hours | No | Only rows stored within the last N hours |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only provide readOnlyHint=false, so the description carries most of the behavioral burden. The description clearly discloses the destructive 'clear to delete rows' path, scopes the tool to cached pages and searches, and explains the operational modes (query, get, stats, clear). It does not cover irreversibility or eviction/persistence details, but the destructive behavior is explicit.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three sentences, front-loaded with the tool's purpose, and every clause earns its place. It communicates scope, usage timing, cost rationale, and mode selection without redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 9-parameter tool with no output schema, the description covers the core invocation model, the main modes, the destructive path, and the decision to check the cache before more expensive operations. Field-level details such as defaults, limit, since_hours, and clear requirements are fully covered by the schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description adds extra semantic value beyond the schema by grouping modes ('Use query for full-text search, get to pull full content, stats for counts, clear to delete rows') and by adding operational details such as Chinese-substring and English-word search behavior.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening sentence names a specific verb and resource: 'Query the LOCAL CACHE of every page this server has fetched and every search it has run.' It also differentiates from the sibling fetch/search tools by telling the agent to check here before re-fetching or re-searching.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit usage guidance: 'Check here BEFORE re-fetching or re-searching — a hit is instant and free while a fresh fetch costs 5-60s.' This tells the agent when the cache is the right choice, and why, which is strong routing guidance relative to alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
clickAInspect
Click an element on a one-off page: loads url in a fresh browser context (stateless — no cookies unless passed, no shared state with other calls), waits wait_secs after load before clicking, then fires a DOM click on the first CSS-selector match. The click may trigger navigation (link, form submit) — the response url and text_after are read after that navigation lands. Returns clicked:false when the selector matches nothing. For multi-step interaction on a shared page use session_click instead.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | The URL to load | |
| selector | Yes | CSS selector of element to click | |
| wait_secs | No | Seconds to wait for the page to settle after load, before clicking |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Discloses key runtime behavior beyond the schema: a fresh stateless browser context, waiting after load, DOM click implementation, possible navigation, and response read after navigation. Also documents the clicked:false miss case. Since there are no behavioral annotations, this description carries the burden and does so thoroughly.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four dense sentences with no filler; the main action comes first, then side effects, error case, and alternative are sequenced for efficient parsing. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For three simple parameters, the description covers statelessness, clicking behavior, navigation, and the no-match return value. It is nearly complete, but without an output schema it leaves the exact success shape implicit beyond url/text_after and clicked:false.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3. The description adds semantics on top: the selector matches the first element, wait_secs governs time after page load before clicking, and url is loaded in a fresh context. This extra context justifies a 4 rather than baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific operation ('Click an element on a one-off page') with a clear resource (the page URL) and mechanism (CSS-selector match). It explicitly contrasts itself with session_click, so an agent can distinguish it from sibling click tools without inferring.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives the exact usage context: one-off, stateless page interaction. It closes with an explicit routing condition: 'For multi-step interaction on a shared page use session_click instead,' which tells the agent when to prefer the sibling.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
downloadAInspect
Download a file over HTTP(S) with streaming to disk (no memory buffering), SHA-256 integrity hash, and optional resume of interrupted transfers. Filename resolution: explicit param → Content-Disposition → URL tail. Use for binaries, archives, datasets, documents - anything where the agent wants the FILE saved, not its text content read.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | URL of the file to download (http/https) | |
| resume | No | Resume an interrupted download when a local partial file exists | |
| cookies | No | Cookies to send with the request: `"name=value"` strings or CDP-style objects `{"name","value","domain",...}` for gated downloads | |
| filename | No | Explicit output filename. When omitted: Content-Disposition → URL tail → "download" | |
| use_proxy | No | Route through proxy (default: false; auto-enabled for known blocked domains) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With only a title annotation, the description carries the behavioral disclosure burden. It does well by revealing no-memory-buffering streaming, SHA-256 hashing, resume behavior, and the filename resolution precedence. It falls short of 5 because it does not disclose the return value, where the file lands on disk, or how failures are surfaced.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences with no filler. The core behavioral constraint is front-loaded, filename resolution is compressed into a clear chain, and the usage guidance closes the description efficiently. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a download tool, the description is largely complete: purpose, major behavioral traits, filename handling, and use cases are all present. However, without an output schema or supporting annotations, the absence of return-value details, target path, and failure-mode expectations leaves a moderate gap for an agent that must verify the tool's result.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3; each parameter is already described in the schema. The description adds meaningful context for 'filename' via the resolution precedence chain (explicit → Content-Disposition → URL tail), but it does not deepen semantics for url, resume, cookies, or use_proxy beyond what the schema provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource ('Download a file over HTTP(S)') and immediately adds distinguishing mechanics: streaming to disk, SHA-256 integrity hash, resume, and filename resolution. It also explicitly contrasts with reading text content, which differentiates it from sibling tools like fetch.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The final sentence gives a clear when-to-use rule: binaries, archives, datasets, and documents where the agent wants the FILE saved. It also states the when-not condition ('not its text content read'), effectively routing the agent toward a fetch-style alternative for text retrieval.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
evalAInspect
Execute JavaScript on a one-off page: loads url in a fresh browser context, optionally waits wait_secs for the page to settle, evaluates script (async/Promise supported) and returns {url, result}. Script-driven navigation (location.href, form submit) is drained and reflected in the returned url. Stateless — no cookies or page state shared with other calls; when the script needs prior page state or a login, use session_eval.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | The URL to load | |
| script | Yes | JavaScript code to execute (supports async/Promise) | |
| wait_secs | No | Seconds to wait before executing |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations provide only a title, so the description carries the full behavioral burden. It discloses fresh browser context, optional wait, async/Promise support, statelessness, no cookie/page-state sharing, and navigation being drained into the returned `url`. It does not mention error handling, timeouts, or broader side effects of executing arbitrary JavaScript, but the core execution model is clearly explained.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three efficient sentences with the main action and return value front-loaded, followed by a navigation edge case and a statelessness caveat with the sibling alternative. No filler or redundant restating of the schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no output schema and minimal annotations, the description covers the essential invocation details: return shape, stateless boundary, navigation handling, and when to choose `session_eval`. Minor gaps remain around error/timeout behavior and exact script execution scope, but the description is sufficient for correct use.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, but the description adds meaning beyond the schema: `url` is loaded in a fresh browser context, `wait_secs` is for the page to settle, and `script` supports async/Promise. This contextualizes the parameters better than the terse JSON Schema descriptions alone.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a precise action ('Execute JavaScript'), a specific resource ('a one-off page' loaded from `url`), and a clear return shape (`{url, result}`). It also distinguishes itself from sibling `session_eval` by explicitly noting that this tool is stateless.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly names `session_eval` as the alternative when prior page state or login is needed, giving a concrete when-not-to-use condition. It also implies the appropriate use case: one-off, stateless JavaScript evaluation on a fresh page.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
fetchARead-onlyInspect
Fetch a webpage and return clean markdown/html/text. Use whenever the agent needs to READ any web page - blogs, docs, articles, JS-rendered SPAs, Cloudflare-protected sites. Static pages are served over plain HTTP (~100ms tier:"http"); pages that need JS get the full browser (tier:"browser"). render_tier selects auto (default) / http (pure HTTP, refuses the upgrade) / browser (always the JS browser).
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | The URL to fetch | |
| format | No | Output format: "markdown", "html", or "text" (default: markdown) | markdown |
| sanitize | No | Strip prompt-injection payloads from the text output (default true): zero-width/steganographic characters, instruction-shaped lines ("ignore previous instructions", chat markup tokens, CJK variants), and text hidden via opacity:0 / tiny fonts. A `sanitize_report` field counts what was removed — stripping is observable, never silent. Set false for raw output. | |
| selector | No | CSS selector to extract specific content | |
| max_chars | No | Maximum characters to return (default: 50000) | |
| use_proxy | No | Route through proxy (for blocked foreign sites) | |
| wait_secs | No | Seconds to wait for JS rendering | |
| js_extract | No | JS expression to extract from the page after rendering | |
| capture_xhr | No | Capture script-initiated API responses: a list of URL substrings (e.g. ["/api/"]) whose matching fetch/XHR bodies come back in an `xhr` array; an empty list captures every XHR/Fetch. Forces browser rendering (script-initiated requests only exist after JS runs). | |
| render_tier | No | Rendering strategy: "auto" (default), "http", or "browser" | auto |
| tls_fingerprint | No | TLS fingerprint override (stealth mode only): "chrome145", "firefox133", etc. | |
| auto_bypass_challenge | No | Auto-detect and bypass Cloudflare Turnstile challenges (default: true) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnlyHint annotation, the description discloses meaningful behavior: static pages use plain HTTP, JS pages use a full browser, and render_tier controls auto/http/browser with the nuance that http 'refuses the upgrade.' This gives the agent a real sense of how the tool behaves at runtime, though some details like sanitization are left to the schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three sentences, front-loaded with the core purpose, and each sentence adds useful information. The rendering-tier explanation is a bit dense and contains awkward formatting, but there is no wasted prose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with 12 parameters and no output schema, the description covers the essential invocation context: what it returns, when to use it, and how rendering tiers work. It does not enumerate every parameter, but the schema provides that detail, so the description is sufficiently complete for an agent to select and call the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description adds a little extra color around render_tier (e.g., 'refuses the upgrade') but mostly restates what the schema already documents. It does not materially improve understanding of the other 11 parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Fetch a webpage and return clean markdown/html/text.' It clearly positions the tool as the read-only web-fetching option among siblings like download, search, and session_navigate, and the mention of JS-rendered SPAs and Cloudflare-protected sites further distinguishes its scope.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly says 'Use whenever the agent needs to READ any web page' and gives concrete examples (blogs, docs, articles, JS-rendered SPAs, Cloudflare-protected sites). It does not explicitly name alternative tools or exclusion cases, but the context is clear enough for an agent to select this tool over session-based or download siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
flow_runAInspect
Run a flow — a recorded, editable JSON browser-session script — deterministically, with zero model tokens. Steps are {op, args, expect?, save?}: ops cover navigate/click/click_xy/input/scroll/eval/wait/screenshot/state/cookies; {{var}} placeholders in args are filled from vars; expect asserts (url_contains | selector | text_contains | eval_truthy) abort with evidence on failure; save collects a step's output into the receipt. Source the flow inline via "flow", or by "name" from the server's workflow//flow.json (unknown name → error lists installed workflows). Pass session_id to reuse a live session (e.g. from import_curl) so login state and flows compose. The receipt carries status ok/failed, saved outputs, the session_id (kept alive), and on failure the failing step, reason and a diagnostic screenshot — fix the flow or take the session over from there.
| Name | Required | Description | Default |
|---|---|---|---|
| flow | No | Inline flow document: {create?, vars?, steps:[{op, args, expect?, save?}]} | |
| name | No | Or run a server-side workflow/<name>/flow.json asset. An unknown name errors back with the list of installed workflows — that error is the discovery call. | |
| vars | No | Values for {{placeholders}} in step args; wins over the flow's own vars defaults. | |
| max_steps | No | Override the run's step-execution budget (branch loops re-run steps, so every revisit counts). Default 1000, clamped 1..=100000. A flow document may also declare its own max_steps; this wins. | |
| session_id | No | Reuse a live session (e.g. from import_curl) instead of creating a fresh one — that's how login state and flows compose. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations provide only a title, so the description carries the full behavioral burden. It discloses that execution is deterministic and uses zero model tokens, that expect assertions abort with evidence on failure, that save collects step output into the receipt, and that on failure a diagnostic screenshot is included. It also mentions session_id is kept alive after the run, implying side effects on session state. This is rich, honest disclosure of execution behavior and failure modes, though it doesn't detail permission needs or reversibility, which are less relevant for a sandboxed browser script.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single dense paragraph but front-loads the core concept, then expands into operational details. Every sentence serves a purpose: defining the script shape, ops, placeholders, expect/save semantics, sourcing options, session reuse, and receipt behavior. It avoids filler, though it could be slightly better structured with bullet points for ops or expectations. As is, it is efficient and scannable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool of this complexity (5 optional params, no output schema), the description covers the critical usage aspects: how to provide a flow, how to name it, how placeholders and vars work, how assertions and saving behave, and what the receipt contains on both success and failure. It also explains the session_id reuse mechanism. The only minor missing piece is an explicit list of supported 'op' values, though the description enumerates them inline (navigate/click/click_xy/etc.), so it is adequately complete for an agent to understand what it can do.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so a baseline 3 applies, but the description adds substantial meaning beyond the schema. It explains how 'flow' and 'name' are alternatives, what happens on unknown names, how 'vars' overrides defaults, and how 'max_steps' interacts with branch loops and document-level declarations. The schema descriptors are terse; the description resolves their operational intent (e.g., 'max_steps' becomes a budget with branch re-counting). This goes well beyond the raw JSON schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with 'Run a flow — a recorded, editable JSON browser-session script — deterministically, with zero model tokens.' This gives a precise verb (run), a specific resource (flow) and its characterization. It also distinguishes itself from sibling session_* tools by framing flows as composed scripts, not single operations. The explanation of ops, placeholders, and expect/save further clarifies what the tool does.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explains how to source flows (inline via 'flow' or by 'name' from server assets) and explicitly recommends passing session_id to reuse a live session (e.g., from import_curl) so login state and flows compose. It also states that an unknown name errors with a list of installed workflows — turning an error into a discovery call. While it doesn't explicitly say 'when not to use this versus session_* tools', the composition and determinism cues imply it is for scripted multi-step runs, which is enough context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
import_curlAInspect
Import login state from a real browser in one paste. The human logs into a site in their own Chrome (solving the CAPTCHA/SMS once), opens DevTools → Network, right-clicks any authenticated request → "Copy as cURL", and passes the command here. Returns a live session_id already carrying that site's cookies and sitting on the copied request's URL — the agent continues from where the human left off, no password or second login needed. Works with bash, PowerShell and cmd copy flavors.
| Name | Required | Description | Default |
|---|---|---|---|
| curl | Yes | A "Copy as cURL" command pasted from Chrome DevTools (Network panel → right-click any authenticated request). bash, PowerShell and cmd flavors all parse; the cookie set is injected and the session navigates to the copied request's URL. | |
| account | No | Attach the session to a named account: the imported login lands in the account's private jar and is written back under its name after every action — one import per identity, no clobbering. | |
| use_proxy | No | Route the session's traffic through the engine proxy. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only provide a title, so the description carries the behavioral burden. It discloses that the tool imports cookies, navigates to the copied request's URL, returns a session_id, and supports multiple shell flavors. It does not mention side effects like session replacement or security warnings, but the core mutation behavior is transparent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, moderately long paragraph but every sentence earns its place. It front-loads the core purpose and then details the workflow and output. Slightly run-on but not wasteful; could be split into bullet points for clarity, but it is effective.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 3-param tool with no output schema, the description covers the input format, the expected usage flow, the return value (session_id), and account attachment behavior. It lacks edge-case handling (invalid cURL, overwrite rules) but is complete enough for an agent to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3. The description adds value beyond the schema by explaining the curl param's multi-flavor parsing and the account param's 'private jar' and 'no clobbering' behavior, which clarifies intent beyond the raw property descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb (import) and resource (login state from a browser via cURL), and distinguishes it from siblings like session_create by framing it as 'Import login state from a real browser in one paste.' It precisely explains what the tool does and how it fits into the session workflow.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives a concrete scenario (human solves CAPTCHA/SMS in Chrome, copies cURL, agent continues) and implicitly positions it as the alternative to manual authentication. However, it does not explicitly name when not to use it or mention alternatives like session_create for fresh sessions, so it stops short of full exclusion guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
render_markdownAInspect
Render a markdown document into a deterministic, self-contained HTML artifact - the document layer, so the agent never writes HTML by hand. Prose rides a plain offline shell (no fonts, no scripts); archify fenced code blocks carry typed zero-coordinate diagram JSON (sequence, workflow, architecture, dataflow, lifecycle families) and render to inline SVG via the layout engine. Same input, same bytes: the receipt carries the sha256 so determinism is verifiable. theme picks light (default) or dark; preset picks the palette family — classic (default), signal-flow, blueprint, editorial — orthogonal to theme; colors bake at generation time (presentation attributes, not CSS variables), and the receipt records both preset and theme. quality picks the composition audit profile — standard (default) or showcase, the delivery gate: the receipt's diagrams[].composition grades route crossings, ambiguous corridors, label clearance (2px standard / 4px showcase), route rhythm, and node text projected to the 930px reader width; the audit never changes the artifact bytes. Mermaid sources are the agent's job to translate, not the engine's: flowchart/graph → workflow (lanes + columns), sequenceDiagram → sequence, stateDiagram-v2 → lifecycle (bands), erDiagram/class → architecture (grid + boundaries) — read the topology and emit the matching zero-coordinate archify JSON; the engine accepts only archify JSON. A broken diagram degrades to a visible code block and lands in receipt.diagnostics; an authored route preset that cannot be honored is self-repaired to a verified semantic substitute and disclosed in receipt diagrams[].repairs - the document still renders. A fence may also carry views: [{id,label,nodes,note?}] (node ids of the active family), emitted as guided-view tabs above the diagram plus an inlined viewer script - clicking a tab lights the member nodes and the routes between them (subgraph), clicking a node lights it with its direct neighbors (ego graph), everything else dims; a view's optional note shows as a caption while it is active (the story layer). window.agxViewer in a session drives and reads the same state programmatically: {focus,view,state} as before, plus route(i,from,to) which returns and lights the shortest authored directed path between two nodes (null when unreachable, state untouched), and reach(i,id,down|up) which returns and lights the authored downstream/upstream closure ({nodes,links}); both dim the rest of the diagram. diagrams[].views in the receipt lists the tabs. motion: true bakes an entrance choreography into the artifact: pure-declarative CSS animation with zero scripts - headings split into per-glyph (CJK) / per-word (latin) spans that rise in with expo easing, prose blocks stagger up an nth-child delay ladder, diagram figures grow in with a back ease (GSAP's easing math as public cubic-bezier equivalents, nothing embedded); the diagrams themselves play a flow story on the same clock - nodes land beat by beat, solid edges draw in (dash-offset), dashed returns fade, sequence messages arrive as sent - with a timed caption strip under each figure as the subtitles, which becomes a static transcript under prefers-reduced-motion; the file itself animates in any browser and the receipt records motion plus diagrams[].story (beat times and captions - the hook for muxing voice later). With session_id the artifact is also loaded into that session (local, free) and the reply carries viewport acceptance: scroll extents measured in the live session and graded fits/tall/wide/oversized, telling the agent how to read the page back. Diagram vocabulary adapted from archify (MIT).
| Name | Required | Description | Default |
|---|---|---|---|
| theme | No | Color theme: "light" (default) or "dark" — the shell background/ foreground and every SVG palette slot swap together; the receipt records which theme produced the bytes | |
| motion | No | Bake the entrance choreography into the artifact (default false): pure-declarative CSS animation — headings split into per-glyph/per- word spans that rise in with expo easing, prose blocks stagger up a nth-child delay ladder, and diagram figures grow in with a back ease (GSAP's easing math as public cubic-bezier equivalents). The diagrams animate too, on one story clock: nodes pop in one beat at a time, solid edges draw themselves (dash-offset drain), dashed returns fade, sequence messages land as they are "sent", and a timed caption strip under each figure subtitles the beats — under prefers-reduced-motion the strip becomes a static transcript. Zero scripts: the file itself animates in any browser, subtitles and all; the receipt records motion (plus diagrams[].story with the beat times, the hook for muxing voice later) so a cached artifact is never mistaken for the static one | |
| preset | No | Visual preset: "classic" (default), "signal-flow", "blueprint", or "editorial" — a palette family orthogonal to theme (each preset exists in both light and dark). The receipt records preset and theme separately | |
| quality | No | Quality profile for the composition audit: "standard" (default) or "showcase" — the delivery gate. The audit grades route crossings, corridors, label clearance, rhythm, and projected text size in the receipt (diagrams[].composition); it never changes the artifact bytes, only how findings are severity-rated | |
| markdown | Yes | Full markdown document. Prose rides a plain offline shell; archify fenced code blocks carry typed zero-coordinate diagram JSON and render to inline SVG. | |
| session_id | No | Optional session ID: also load the rendered HTML into that live session (local and free) so session_screenshot / session_state can verify the artifact |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotations carry only a title, so the description bears the full disclosure burden, and it delivers extensively: deterministic bytes with a sha256 receipt for verification, a fully offline shell (no fonts/scripts), graceful degradation of broken diagrams to visible code blocks with diagnostics, self-repair of unhonorable route presets disclosed in repairs, reduced-motion fallback to a static transcript, and session loading described as local/free. This far exceeds the transparency expectations for an unannotated tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The first sentence is well front-loaded, but the description runs roughly 700 words and buries actionable content under implementation minutiae an agent does not need to invoke the tool: GSAP easing math as cubic-bezier equivalents, per-glyph CJK span splitting, an nth-child delay ladder, and full window.agxViewer API signatures (route(i,from,to), reach(i,id,down|up)). Much of the motion prose is duplicated nearly word-for-word in the schema's motion description, so those sentences do not earn their place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's high complexity (6 parameters, a demanding archify JSON input contract, no output schema) and a title-only annotation, the description is exceptionally complete: it specifies defaults for every option, input format requirements, failure and repair semantics, determinism verification via the receipt, post-call viewport grading (fits/tall/wide/oversized), and even licensing provenance. An agent has everything needed to call and verify this tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, and the description adds meaning above it: theme/preset are orthogonal with colors baked at generation time as presentation attributes rather than CSS variables, quality thresholds are quantified (2px standard / 4px showcase, 930px reader width), and failure/repair behavior is tied to each option. Some marginal value is lost because the schema's motion parameter description already repeats nearly the full animation behavior verbatim.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening sentence is specific: 'Render a markdown document into a deterministic, self-contained HTML artifact - the document layer, so the agent never writes HTML by hand.' The verb (render), resource (markdown document), and output (self-contained HTML) are unambiguous. However, unlike the strongest examples, it never names its siblings render_pdf/render_video or explains the boundary between them, so the differentiation is implicit rather than explicit.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives strong in-tool input guidance ('Mermaid sources are the agent's job to translate, not the engine's... read the topology and emit the matching zero-coordinate archify JSON; the engine accepts only archify JSON'), which tells the agent what to feed it and what not to feed it. But there is no explicit when-to-use-this-vs-alternatives guidance for tool selection against render_pdf/render_video; the only usage framing is the implied 'use this instead of writing HTML by hand.'
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
render_pdfAInspect
Cut a rendered page into pages and package as PDF, PNGs, PPTX or DOCX. Print mode (no selector) paginates the document into fixed-height pages (default 794x1123, A4 @96dpi), breaking at top-level block boundaries — no half-cut text where a break can land on a block edge. Slides mode (selector set) makes one page per match, sized to that element — generate an HTML deck with one .slide per page and each becomes a deck page. format "pdf" (default) returns base64 image-based PDF; "png" returns one base64 PNG per page in pages_base64; "pptx" returns a base64 PPTX (one slide per page, deck-sized to the largest page); "docx" returns a base64 DOCX (one page-sized section per page, each section keeps its own height). Returns page count and packaging.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | Page URL to cut into pages. | |
| width | No | Page width in CSS pixels. Default 794 (A4 @96dpi). | |
| format | No | Output format: "pdf" (default), "png" (one base64 PNG per page), "pptx" (one slide per page, image-based), "pptx-native" (editable: element-level DrawingML — real text runs, gradient shapes, image parts; requires `selector`), or "docx" (one page-sized section per page). | |
| height | No | Page height in CSS pixels — print pagination only. Default 1123. | |
| selector | No | CSS selector; present → slides mode (one page per match, sized to the element). Absent → print mode (fixed-height pages at block boundaries). | |
| max_pages | No | Safety cap on emitted pages. Default 50. | |
| use_proxy | No | Route through proxy (for blocked foreign sites) | |
| jpeg_quality | No | JPEG quality for PDF page embedding (1-100). Default 90. | |
| tls_fingerprint | No | TLS fingerprint override (stealth mode only) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With only a title annotation and no read-only/destructive hints, the description carries the behavioral burden. It discloses pagination boundaries, slide-per-match behavior, image-based PDFs, per-format packaging details, and the return of page count. It stops short of describing side effects like network fetching or permission requirements, but these are largely non-mutating.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but well-structured: purpose first, then mode behavior, then format-specific output details, then return summary. Every clause adds information, though the long middle sentence packs many behaviors together and could be slightly hard to scan.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with 9 parameters, two modes, and five output formats, the description covers the essential decisions and return characteristics. With no output schema, it supplies high-level return information ('page count and packaging') and names pages_base64 for PNGs, but does not fully specify the response envelope.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description adds real meaning beyond the schema by explaining mode-dependent behavior (block-boundary breaks vs element-sized pages), the deck-size behavior of PPTX, and the section-height behavior of DOCX. This is more than a restatement of parameter names.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('cut') and resource ('rendered page') and enumerates the exact output formats (PDF, PNGs, PPTX, DOCX). It clearly distinguishes itself from sibling rendering tools like render_video and render_markdown by describing page-based packaging.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear when-to-use guidance for the two modes: print mode when no selector is set, and slides mode when a selector is provided. It does not explicitly contrast the tool with sibling alternatives, but the mode-level guidance is strong and actionable.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
render_videoAInspect
Render a page's animation timelines to an MP4 video. The page's scripts must expose window.__timelines — objects with duration() and pause(t) (a paused gsap.timeline registered there works as-is). Each frame seeks every timeline to t=i/fps and paints the viewport, so the output is deterministic — no wall clock in the pixel values. Audio: narration[] places TTS/voice clips at start times (mixed into one AAC track), audio adds looped background music, and subtitles_srt muxes an SRT as a soft mov_text track and (by default, burn_subtitles: false to opt out) burns the same cues into the frame pixels — QuickTime, WeChat and most social embeds ignore the soft track. Requires ffmpeg on the server. Returns base64 MP4 (H.264, yuv420p) plus frame count and durations.
| Name | Required | Description | Default |
|---|---|---|---|
| fps | No | Frames per second. Default 24. | |
| url | Yes | Page URL whose scripts register timelines in `window.__timelines` (GSAP-style objects with `duration()` + `pause(t)`). | |
| audio | No | Background music: looped to cover the video, volume-scaled, faded out at the tail. | |
| width | No | Viewport width in CSS pixels (floored to even — yuv420p). Default 1280. | |
| height | No | Viewport height in CSS pixels. Default 720. | |
| narration | No | Voiceover clips, each starting at its own time (any TTS output; mixed into one AAC track). | |
| use_proxy | No | Route through proxy (for blocked foreign sites) | |
| subtitles_srt | No | Inline SRT subtitles muxed as a soft (toggleable) mov_text track. | |
| burn_subtitles | No | Burn the cues into the frame pixels too (hardsub) — on by default when `subtitles_srt` is present; QuickTime, WeChat and most social embeds ignore the soft mov_text track. `false` keeps the soft track only. | |
| hold_tail_secs | No | Freeze the final timeline state for this many extra seconds. Default 0.5. | |
| tls_fingerprint | No | TLS fingerprint override (stealth mode only) | |
| max_duration_secs | No | Safety cap on timeline + hold tail, seconds. Default 120. | |
| wait_timelines_ms | No | How long to wait for `window.__timelines` to appear, ms. Default 10000. | |
| subtitles_language | No | ISO language tag for the subtitle track, e.g. "eng" / "zh". |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With only a title annotation, the description carries the full burden and does an excellent job: it discloses the deterministic frame-seeking behavior, the audio mixing rules, the subtitle muxing and default burn behavior, the ffmpeg dependency, and the exact return value (base64 MP4 plus frame count and durations).
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is multi-sentence but every clause earns its place. It is logically organized: core mechanism first, then audio, then subtitles, then requirements and output. No filler or repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given 14 parameters, no output schema, and the tool's complexity, the description covers all essential aspects an agent needs to invoke it correctly: what it does, how it works, the audio and subtitle behaviors, the server dependency, and the return format. Nothing critical is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Although the schema already covers all 14 parameters (100% coverage), the description adds critical meaning beyond it: how timelines are sought, how audio clips are mixed, why burn_subtitles defaults to true (social embeds ignore soft track), and the hold-tail freeze behavior. This is far more than the schema provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb-resource pairing: 'Render a page's animation timelines to an MP4 video.' It then details the exact mechanism (seeking timelines per frame) and the output format, which fully distinguishes it from sibling tools like render_pdf and render_markdown.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly implies when to use it — when a page exposes window.__timelines and you need a deterministic MP4 of the animation. It also states the ffmpeg requirement. However, it does not explicitly name alternative rendering tools or state when not to use it, so it falls just short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
searchARead-onlyInspect
Search the web across Baidu/Bing/Sogou/WeChat/Google (aggregated + deduped) and optionally fetch the top results' full content. Use when the agent needs to FIND information online - replaces a search API. Supports image search returning direct image URLs. Optional engines: ["baidu"]-style filter by engine name (invalid names error with the valid list; /doctor lists them with live health). Optional time_range day/week/month/year for news freshness (engines without dated results ignore it). Response carries engine_errors explaining any engine that contributed nothing (CAPTCHA suspension, transient failure).
| Name | Required | Description | Default |
|---|---|---|---|
| q | Yes | Search query | |
| engines | No | Restrict to these engine names (e.g. ["baidu"], ["sogou_wechat"]). Empty = all engines serving `categories`. Invalid names return an error listing the valid ones. | |
| fetch_top | No | Fetch content for top N results | |
| categories | No | Search categories (default: general) | general |
| time_range | No | Freshness window: "day" | "week" | "month" | "year". Honored by engines with dated results (e.g. bing_news filters by pubDate); others ignore it. | |
| max_results | No | Maximum number of results (default: 10) | |
| max_chars_per | No | Max characters per result content |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=true, so safety is covered. The description adds meaningful behavioral details beyond that: aggregation and dedup, engine-filter error behavior, engine_errors field explaining partial failures (CAPTCHA/transient), and time_range being ignored by some engines. This is valuable runtime behavior that the schema alone would not convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but front-loaded with the core action and ends with error-handling behavior. It is longer than ideal, and 'Optional engines: [...]' wording is slightly awkward, but every sentence carries information not present in the schema, so it earns its length.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no output schema, the description helpfully explains engine_errors and image-search return behavior, filling an otherwise empty return-value story. It doesn't describe dedup mechanics or result shape in detail, but for a 7-parameter tool with full schema coverage, readOnlyHint, and error-behavior notes, it is reasonably complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description does add a little cross-parameter meaning — e.g., engines work with `categories`, invalid engine names error with a valid list, and `time_range` is only honored by dated-result engines — but most parameter meaning already lives in the schema's property descriptions. It doesn't rise above baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb-resource pair ('Search the web'), enumerates the aggregated engines, and distinguishes itself from sibling tools like `fetch` and `click` by stating it replaces a search API and supports image search. This is a clear, specific purpose statement that differentiates it from the surrounding navigation and session tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly says 'Use when the agent needs to FIND information online' and mentions fetching full content as an optional follow-on, which implies when `fetch_top` is the right parameter. It does not explicitly name sibling alternatives or state when not to use this tool (e.g., use `fetch` for a single known URL), so it loses a point on exclusions, but the core usage context is clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
session_challengesARead-onlyInspect
One-call risk-control report: did this session hit an anti-bot wall? Taobao/tmall's x5 risk control answers 200 like a normal response — either a redirect onto a punish page (tmd/punish, punish.taobao.com) or an MTop API body carrying FAIL_SYS_USER_VALIDATE / RGV587 / x5secdata. Returns {total, events:[{url,method,status,kind,via}]} where via says whether the wall was navigated into ("url") or swallowed by an API response ("body"). When there are hits, the response also carries the account name (which identity got walled) and a handoff instruction: the engine detects and surfaces but does not auto-bypass — a human opens the live view (/live?session= on the engine's HTTP port), solves the challenge in this session, and the retry rides the cookie that solving sets. Detection only; no automated solving or bypass.
| Name | Required | Description | Default |
|---|---|---|---|
| session_id | Yes | Session ID |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnlyHint annotation, the description thoroughly discloses behavior: the detection signals, the two possible wall locations via 'url' or 'body', the output structure, the account-name/handoff addition, and the explicit statement that the engine only detects and surfaces, never solves or bypasses. This is rich, accurate behavioral context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long but information-dense: every component, from purpose to failure indicators to output shape to human handoff, earns its place. It is also front-loaded with the primary question the tool answers.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema present, the description compensates fully by specifying the return shape and hit-specific fields. It also explains the follow-up human workflow and the tool's non-bypass boundary, making the definition complete enough for correct invocation and interpretation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema already documents the single session_id parameter with 100% coverage. The description only loosely ties it to 'this session' and does not add format, constraints, or lifecycle context. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a clear, specific purpose: a 'one-call risk-control report' that tells whether a session hit an anti-bot wall. It names concrete failure indicators and the expected output shape, and it distinguishes itself from generic session tools by being detection-only.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context for when to use it: whenever you need to check whether a session was walled, and it explicitly notes that the tool does not auto-bypass. However, it does not name sibling alternatives or state when to prefer a different session inspection tool, leaving some routing to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
session_clickAInspect
Click an interactive element by its index (from session_state output) inside a live browser session: scrolls it into view and fires a DOM click on the session's current page. Before clicking it re-verifies the element in the same frame — if the page changed since session_state (element detached, disabled, hidden, or covered by an overlay), it returns clicked:false with a reason ("detached"/"disabled"/"not_visible"/"covered_by") and, when covered, a covered_by description of the element that would eat the click — never a silent no-op. A submit click may navigate the session — the returned url/text_after reflect the page after the action, and session state (cookies, localStorage, globals) persists for follow-up calls. Indexes come from the most recent session_state; re-list after navigation.
| Name | Required | Description | Default |
|---|---|---|---|
| index | Yes | Element index (from /state output) | |
| session_id | Yes | Session ID |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only provide a title, so the description carries the full burden. It thoroughly discloses behavior: scrolls into view, fires DOM click, re-verifies element, returns clicked:false with reasons on failure, never silent no-op, may navigate, and session state persists. This is exceptionally transparent and goes beyond basic expectations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single dense paragraph but well-structured: it opens with the core action and source, then details pre-conditions, failure modes, side effects, and persistence. It is not overly verbose for the amount of critical information it conveys, though breaking it into bullets could improve readability. It earns its length by covering essential behavioral details.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite lacking an output schema, the description explains return values (clicked:false, reason, covered_by, url, text_after), failure modes, navigation side effects, and session persistence. For a two-parameter tool, this is remarkably complete—an agent has everything needed to invoke it correctly and interpret outcomes.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% with both parameters described (index from /state output, session_id). The description adds value by clarifying that index must come from the most recent session_state and re-list after navigation, plus explaining how the index is used (re-verification). This enriches the schema meaning beyond the basic descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Click an interactive element'), the resource (element by index in a live browser session), and the source of the index (session_state output). It also distinguishes itself from sibling session_click_xy by specifying index-based clicking, making it easy to select the right tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context: indexes come from the most recent session_state, and it advises re-listing after navigation. It implicitly differentiates from session_click_xy (coordinate-based) by emphasizing index-based selection, though it doesn't explicitly name the alternative. It covers when the tool may not work (page changed) and how to handle it, giving sufficient usage guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
session_click_xyAInspect
Click at viewport coordinates (CSS pixels) via real mouse events — pointerdown/mousedown, pointerup/mouseup, then click on whatever element is hit there. For canvas/map surfaces with no DOM element to index. click_count 2 adds dblclick.
| Name | Required | Description | Default |
|---|---|---|---|
| x | Yes | Viewport X coordinate in CSS pixels | |
| y | Yes | Viewport Y coordinate in CSS pixels | |
| button | No | Mouse button: "left" (default), "right", "middle" | |
| session_id | Yes | Session ID | |
| click_count | No | Click count: 1 single (default), 2 adds dblclick, 3+ sets detail |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only contain a title, so the description carries the full burden. It discloses the event sequence (pointerdown/mousedown, pointerup/mouseup, then click) and the click_count behavior, which adds value beyond the schema. It does not mention side effects like navigation or page state, but the core behavior is well covered.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with zero filler. The core action and event sequence are front-loaded, the use case is stated clearly, and click_count behavior is appended without unnecessary detail.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a coordinate-based click tool with no output schema and moderate parameter count, the description covers purpose, event behavior, and primary use case. It doesn't address error conditions or prerequisites like page loaded state, but these are not essential for a simple click operation. The distinction from element-based clicking is implied, not fully elaborated.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so baseline is 3. The description adds semantics for click_count ('2 adds dblclick') and reinforces x/y as viewport coordinates, but it does not add meaning for button or session_id beyond what the schema already provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Click at viewport coordinates'), explains the real mouse event sequence, and explicitly scopes to canvas/map surfaces with no DOM element to index. This clearly differentiates it from sibling click/session_click tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides clear context for when to use: 'For canvas/map surfaces with no DOM element to index.' It implies the alternative (element-based clicking) but does not name specific sibling tools or explicitly say 'use session_click instead', so it falls short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
session_cloneAInspect
Derive a new browser session from a live one, carrying the full login state: cookies, localStorage/sessionStorage, viewport pin, dialog policy, proxy and keepalive flags. The source session stays untouched. Use to snapshot a logged-in state before risky actions, or to run the same login in parallel tabs. Returns {session_id (new), cloned_from, url, viewport}.
| Name | Required | Description | Default |
|---|---|---|---|
| session_id | Yes | Session ID to derive from (stays alive and untouched) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With only a title annotation, the description carries the behavioral burden. It explicitly guarantees 'The source session stays untouched,' lists the state that is cloned, and gives the return shape. This gives an agent meaningful expectations without contradicting annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three dense sentences with no filler: the operation, the carried state, the source-safety guarantee, the use cases, and the return object are all front-loaded. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter tool with no output schema, the description is complete: it covers purpose, exact behavioral effect, when to use it, and the return fields. An agent can select and invoke it correctly without needing further context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, and the single parameter is already well documented as 'Session ID to derive from (stays alive and untouched).' The description adds no new parameter-level detail beyond contextualizing that the new session is derived from this one, so the schema carries the weight.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Derive') with a clear resource ('a new browser session') and names exactly what is carried over, including cookies and storage. It is distinct from sibling tools like session_create or session_export, so an agent can tell when session_clone is the right call.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly states when to use the tool: to snapshot a logged-in state before risky actions or run the same login in parallel tabs. It does not mention when not to use it or name alternatives, so it stops short of the full 5, but the use context is clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
session_closeAInspect
Close a browser session and free its resources. For a persistent session this also drops the on-disk login snapshot - idle expiry keeps it, an explicit close does not.
| Name | Required | Description | Default |
|---|---|---|---|
| session_id | Yes | Session ID |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no readOnly or destructive hints in the annotations, the description carries the behavioral disclosure burden. It does so well by revealing that an explicit close drops the on-disk login snapshot for persistent sessions, while idle expiry keeps it. This is a meaningful side effect beyond what the schema or annotations convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences with no filler. The core action is front-loaded, and the second sentence adds only the critical distinction about persistent session snapshots. Every word earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a low-complexity, one-parameter close tool with no output schema, the description covers the action, resource impact, and a key side effect. It does not mention error behavior or idempotency, but these are not essential for an agent to invoke the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema already provides 100% coverage for the single session_id parameter with the description 'Session ID', so the baseline is 3. The description does not add parameter-level detail, but none is needed given the simple schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Close') with a clear resource ('browser session') and a concrete consequence ('free its resources'). It clearly differentiates this from sibling tools like session_create or session_list, which serve different lifecycle stages.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The intended usage is implied clearly: call this when a browser session should be ended and its resources released. The description adds useful context by contrasting explicit close with idle expiry, but it does not explicitly name alternatives or state when not to use the tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
session_consoleARead-onlyInspect
Read the session's recent page console output (log/info/warn/error) as {url, total, matched, messages:[{ts_ms, level, text, url}]}, newest last. Ring buffer of 500 entries; captures output from page scripts, clicks, evals and navigation alike. Optional filters: level (exact, e.g. "error"), since_ts (epoch ms), url_contains (page URL substring), limit (most recent N matches). The fastest way to see WHY a page misbehaves: click the button, call this, read the error.
| Name | Required | Description | Default |
|---|---|---|---|
| level | No | Only entries at this level: "log" | "info" | "warn" | "error" | |
| limit | No | Keep only the most recent N matching entries | |
| since_ts | No | Only entries logged at or after this Unix epoch millisecond timestamp | |
| session_id | Yes | Session ID | |
| url_contains | No | Only entries whose page URL contains this substring |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, and the description adds meaningful behavioral context: it is a ring buffer of 500 entries, newest last, and captures output from multiple sources. This goes beyond the annotation and helps the agent understand the tool's memory and ordering behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact, front-loaded with the return shape, and every sentence adds value. The final sentence is a memorable usage heuristic, and there is no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only tool with a fully documented schema, the description covers the return shape, ordering, buffer size, capture scope, and filters. It doesn't explain pagination or what happens when the buffer overflows, but those are minor gaps given the tool's simplicity and the readOnlyHint annotation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all five parameters. The description adds a little extra context by grouping filters and giving an example value for level, but it doesn't substantially extend the schema's meaning.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Read'), a clear resource ('the session's recent page console output'), and the exact return shape. It also distinguishes itself from siblings by focusing on console output, not network, state, or storage.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly says this is 'The fastest way to see WHY a page misbehaves' and gives a concrete workflow: 'click the button, call this, read the error.' It also clarifies that it captures output from page scripts, clicks, evals, and navigation, which helps an agent know when to use it over alternatives like session_network or session_state.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
session_cookiesARead-onlyInspect
Export the session's current cookies as ["name=value", ...] for the page's URL. Use to persist a logged-in session and replay it later via session_create with cookies. Round-trips with session_create's cookies field.
| Name | Required | Description | Default |
|---|---|---|---|
| session_id | Yes | Session ID |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, so the safety profile is known. The description adds valuable behavioral context beyond that: cookies are scoped to the page's URL, returned in a specific serialized format, and designed to round-trip with session_create's cookies field. This is meaningful context without contradicting the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences with no filler. The main action and output format come first, followed by the use case and round-trip relationship. Every sentence contributes necessary information for correct invocation.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple one-parameter read-only tool, the description is complete: it states the input (session_id), the exact output format, the URL scoping, and the intended replay workflow with session_create. There is no output schema, but the description adequately covers what the agent needs to know.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the only parameter, session_id, is already described as 'Session ID'. The description focuses on the output and use case rather than adding new parameter-level detail, which is fine given the schema already carries the full burden.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Export'), names the exact resource ('the session's current cookies'), and specifies the output format ('["name=value", ...] for the page's URL'). It also distinguishes itself from siblings by explicitly connecting to session_create's cookies field, making the tool's role clear relative to related operations.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly states when to use the tool: to persist a logged-in session and replay it later via session_create with cookies. It does not explicitly list exclusions or alternatives like session_export, but the intended use case is unambiguous and actionable.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
session_createAInspect
Create a persistent interactive browser session for multi-step interaction - clicking, typing, scrolling, reading state across page transitions. Use when the agent must INTERACT with a page (login flows, forms, pagination, click-through) rather than read it once. Returns session_id; persists 8 min idle. With persistent:true the login state survives idle eviction and server restarts - the same session_id revives logged-in.
| Name | Required | Description | Default |
|---|---|---|---|
| url | No | Initial URL to navigate to (optional). `start_url` is honored as an alias (#115) — callers guessing that name must not land on about:blank. | |
| width | No | Initial viewport width in CSS pixels. Pinned for the session's life (survives navigation) so element rects and media queries anchor to the same layout across every page of the visit. | |
| height | No | Initial viewport height in CSS pixels. | |
| mobile | No | Mobile device emulation (coarse pointer, no hover) for the initial viewport. | |
| account | No | Run as a named login identity (the multi-account layer): a private cookie jar seeded from the account record, write-back to the account store after every action. Concurrent logins (`taobao-scraper` vs `taobao-publisher`) never clobber each other. The account record survives the session — a later create with the same name picks up the warm jar. 1-64 chars of [a-zA-Z0-9_-]. | |
| cookies | No | Cookies to inject before navigation: `"name=value"` strings or CDP-style objects `{"name","value","domain",...}`. Lets the session start already logged-in. Round-trips with session_cookies. | |
| storage | No | Web Storage to inject after the initial navigation lands: {"local_storage": {"k":"v"}, "session_storage": {"k":"v"}}. For login states that live in localStorage rather than the cookie jar. Round-trips with session_storage. | |
| ttl_secs | No | Idle time-to-live in seconds before the session is evicted (default: 480, clamped 60..3600). Raise it for long workflows. | |
| keepalive | No | Exempt the session from the idle reaper: it lives until session_close or server exit, so a workflow interrupted by long non-browser steps keeps its login state. | |
| use_proxy | No | Route through proxy (default: false) | |
| persistent | No | Persist the login state (cookies + localStorage/sessionStorage + viewport + dialog policy) to the server's local store after every action. If the session idles out — or the whole server restarts — the next call with the same session_id revives it logged-in (storageState-style recovery, no re-login). Explicit session_close drops the snapshot. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations provide only a title with no readOnly/destructive hints, so the description carries the full burden. It discloses key behavioral traits: the session is persistent and interactive, idles out after 8 minutes, and with persistent:true the login state survives both idle eviction and server restarts with the same session_id reviving logged-in. It does not mention rate limits, resource cleanup, or explicit side-effects beyond persistence, but covers the most important behavior for an agent deciding to repeat or retain sessions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three sentences and every sentence earns its place: the first defines the tool, the second gives the usage rule, and the third explains the persistence option. It is front-loaded with purpose and usage. Slight redundancy with schema defaults (8 min vs ttl_secs default 480) keeps it from being as tight as the two-sentence high-calibration example, but it remains efficient.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is complex (11 optional parameters, no output schema, no annotations except title), yet the description covers the essential top-level context: what the session is for, when to use it, idle timeout, persistence, and that it returns a session_id. All parameter details are in the fully covered schema, so nothing critical is missing for an agent to invoke the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the input schema already documents all 11 parameters in detail. The description adds minimal parameter-specific value beyond echoing persistent behavior (persistent:true, 8 min idle), which is already expressed in the ttl_secs and persistent parameter descriptions. Baseline 3 is appropriate because the schema carries the parameter semantics completely.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Create a persistent interactive browser session for multi-step interaction - clicking, typing, scrolling, reading state across page transitions.' It then explicitly contrasts with read-only use cases ('rather than read it once'), which distinguishes it from sibling tools like fetch, search, and read-only session tools without needing to open their schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description states precisely when to use the tool: 'Use when the agent must INTERACT with a page (login flows, forms, pagination, click-through) rather than read it once.' It also gives conditional guidance for persistent:true, explaining when the login state survives eviction and server restarts. This provides clear context and an explicit exclusion (not for one-time reads), even without naming a specific alternative tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
session_dialogAInspect
Inspect or flip the session's dialog policy for window.alert/confirm/prompt. Dialogs never block the page: each is auto-answered (default dismiss) and logged into session_console at level "dialog". action "list" reports {policy, prompt_text, dialogs}; "accept" makes subsequent confirm() true and prompt() return prompt_text (or the call's default argument); "dismiss" restores the default.
| Name | Required | Description | Default |
|---|---|---|---|
| action | Yes | "list" reports the policy and dialog history; "accept"/"dismiss" set the answer applied to subsequent window.confirm/prompt calls (alert is always logged, never blocking). | |
| session_id | Yes | Session ID | |
| prompt_text | No | With action "accept": text window.prompt returns once accepted (omitted keeps the current text). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With minimal annotations (only title), the description carries the full burden. It discloses that dialogs never block the page, are auto-answered with default dismiss, and are logged to session_console at level 'dialog'. It also details the exact behavioral changes for 'accept' and 'dismiss' actions, including how prompt() returns prompt_text.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two well-structured sentences pack all necessary information with zero fluff. The purpose is front-loaded, and behavioral details follow logically. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers the tool's purpose, all actions, the non-blocking behavior, logging, and the return structure for 'list'. While it doesn't detail error handling or the exact contents of 'dialogs', the provided information is sufficient for an agent to call the tool correctly given the lack of output schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description adds meaningful context by explaining how 'action' values behave and how 'prompt_text' interacts with the 'accept' action. This goes beyond the schema's terse descriptions, justifying a slightly higher score.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states a specific verb (inspect or flip) and resource (the session's dialog policy for window.alert/confirm/prompt). It distinguishes itself from siblings by focusing solely on dialogs, which no other sibling handles.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explains the three actions and their effects, making it clear when to use this tool (to inspect or change dialog policy). It doesn't explicitly mention alternatives because none exist, but it does refer to session_console for logging, which provides context without contradicting usage.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
session_dragAInspect
Drag the mouse from one viewport position to another: press at from, steps mousemove events, release at to. The trajectory is humanized by default (eased velocity, wobble, jittered timing, overshoot) — the shapes anti-bot checks score for; pass humanize:false for exact linear interpolation. Moves AMarker-style drag targets, canvas selections and captcha sliders that only track while the pointer travels.
| Name | Required | Description | Default |
|---|---|---|---|
| to | Yes | Where to release it | |
| from | Yes | Where to press the mouse button down | |
| steps | No | Interpolated mousemove events between from and to. Default: 24 with humanize on, 10 without | |
| delay_ms | No | Mean delay between moves in ms (default 18 humanized / 30 linear) — per-step timing is jittered around this when humanizing | |
| humanize | No | Humanize the trajectory: minimum-jerk easing, perpendicular wobble, timing jitter, grip/settle pauses, occasional hesitation and overshoot-and-correct. Set false when a test/tool needs exact linear interpolation. Default: true | |
| session_id | Yes | Session ID |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations are minimal (only a title), so the description carries the full burden. It discloses the humanized trajectory behavior (eased velocity, wobble, jittered timing, overshoot) and explains that it exists to pass anti-bot checks. It also clarifies the behavior of the `humanize` flag. This is substantial behavioral context beyond what the schema provides.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences with zero filler. The first sentence front-loads the core action and mechanics; the second adds crucial behavioral nuance (humanization) and use-case examples. Every word earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex action with six parameters, the description covers the purpose, the humanization behavior, and typical use cases. It does not explicitly mention the session context, but the `session_id` parameter and tool name make that clear. The lack of an output schema is acceptable since the action is a gesture with no meaningful return value. The description is sufficiently complete for an agent to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all six parameters are already documented in the schema. The description adds some contextual meaning (e.g., that `steps` and `humanize` relate to trajectory smoothing), but it does not significantly enhance parameter understanding beyond what the schema provides. Baseline 3 is appropriate given full coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Drag'), a specific resource ('the mouse'), and precise mechanics (press at `from`, mousemove events, release at `to`). It also names concrete target types (AMarker-style drag targets, canvas selections, captcha sliders), which clearly distinguishes it from sibling click tools like session_click and session_click_xy.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context for when to use the tool—when dragging is needed—and even gives examples of drag targets that only track while the pointer travels. However, it does not explicitly mention when NOT to use it or point to alternative tools (e.g., using session_click for simple clicks). Since the sibling list includes click tools, this guidance is implied but not explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
session_evalAInspect
Execute arbitrary JavaScript in a live browser session and return the result. Runs in the session's current page, so DOM mutations, globals and storage persist across calls — unlike the stateless eval tool, which loads its own throwaway page each call. Script-driven navigation moves the session's URL. JS exceptions are reported with name, line/column and stack.
| Name | Required | Description | Default |
|---|---|---|---|
| script | Yes | JavaScript code to execute | |
| session_id | Yes | Session ID | |
| timeout_ms | No | Await budget for the script's promise in ms (default 5000, clamped 100..120000). Pass a larger budget for slow page-side work such as uploads through the page's own fetch; on expiry the tool errors with EVAL_TIMEOUT (the script may still be running) instead of returning a null result. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
There are no relevant annotations beyond the title, so the description carries the full burden. It discloses persistence of DOM mutations, globals, and storage across calls, the navigation side effect, and exception reporting details. It does not explicitly warn that arbitrary JavaScript can be destructive to the session, but the persistence statement effectively conveys the risk.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is four concise sentences, each adding distinct value: core action, persistence behavior, contrast with eval, and error reporting. There is no filler or redundancy, and the most important information is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity, the description covers execution context, side effects, and error behavior, which is substantial for a 3-parameter tool with no output schema. It does not detail the exact return serialization format, but for an eval tool that returns a JavaScript result, this is a minor gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents script, session_id, and timeout_ms including defaults and clamping behavior. The description adds execution-environment context but not new parameter-level semantics, so a baseline score of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: it executes arbitrary JavaScript in a live browser session and returns the result. It also explicitly distinguishes itself from the sibling eval tool by contrasting persistent session state with a throwaway page, making the tool's purpose unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly tells the agent when to prefer this tool over the alternative eval: use session_eval when running in the current session page where mutations, globals, and storage persist, and use eval when a stateless isolated page is desired. It also notes that script-driven navigation changes the session URL, which is critical for choosing the tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
session_exportARead-onlyInspect
Export a browser session's recorded action log. Format "bash" (default) returns a runnable curl script that replays every recorded action (navigate/click/input/scroll/eval) against a fresh session on this server — hand it to a shell or cron, zero model tokens. Format "jsonl" returns the raw action log, one JSON object per line. Format "json" returns a flow.json document — the same recording as editable ops ({op, args}) with cookies/storage stripped — that flow_run replays server-side.
| Name | Required | Description | Default |
|---|---|---|---|
| format | No | Output format: "bash" (default) renders a runnable curl script that replays every recorded action against a fresh session; "jsonl" returns the raw action log, one JSON object per line; "json" returns a flow.json document (editable ops, cookies/storage stripped) for replay via flow_run | |
| session_id | Yes | Session ID |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=true and no destructive consequences, so the description already starts with a safety baseline. It goes further by disclosing format-specific behaviors: bash produces a runnable curl script that replays against a fresh session, and json strips cookies/storage into editable ops — adding meaningful behavioral context beyond the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The scope statement is front-loaded and each subsequent sentence covers a distinct format with its purpose and output shape — no filler. The content is dense but appropriately sized for a tool that needs to explain three formats.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the burden of explaining return values and does so for all three formats, including the replay mechanics and stripped data. It doesn't cover error cases (e.g., missing session), but for a read-only export tool the essential usage guidance is present.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the schema already documents session_id and format in detail. The description reinforces the meaning of each format value (bash default and runnable, jsonl raw log, json editable for flow_run), adding semantic color that helps the agent select the right value without needing to read the full schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States the verb 'Export' with the specific resource 'a browser session's recorded action log,' which clearly distinguishes the tool from the many session_* siblings. It also contrasts with flow_run by explaining that json output is replayed server-side, so an agent can tell export apart from the sibling that consumes it.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides clear context on when each format is appropriate: bash for shell/cron with zero model tokens, json for replay via flow_run, and jsonl for the raw action log. It names flow_run as the alternative tool for replay, though it doesn't exhaustively cover when not to use each format.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
session_inputAInspect
Type text into an input/textarea element by its index (from session_state output), focusing it and dispatching input/change events. events:"full" is the complete human typing gesture: per-character keydown/keypress/input/keyup cycles, trailing change, then blur — the tail blur commits on forms that save in onBlur (React capture listeners, #100). A disabled, readonly, or detached field answers filled:false with a reason instead of a silent write. Hidden inputs are legitimate targets and are filled normally.
| Name | Required | Description | Default |
|---|---|---|---|
| text | Yes | Text to type into the input field | |
| index | Yes | Element index (from /state output) | |
| events | No | Event fidelity: "full" types one character at a time with a keydown/keypress/input/keyup cycle per character, for pages whose listeners key on keyboard events (e.g. keypress-Enter login forms). Default fires a single input+change pair after the value is set. | |
| session_id | Yes | Session ID |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only provide a title, so the description carries the full behavioral burden and does so thoroughly. It discloses focusing, event cycles, trailing blur semantics, the filled:false with reason response for disabled/readonly/detached fields, and the special case for hidden inputs. This is far beyond the minimal schema information.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but every sentence earns its place: purpose, event fidelity, failure behavior, and hidden-input edge case. It loses a point for awkward formatting ('events:"full"' with stray punctuation) and the unexplained '#100' reference, which slightly muddies an otherwise efficient structure.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, yet the description provides the key response signal (filled:false with a reason) and covers important edge cases. It does not describe what a successful call returns, but the operation and failure behavior are specified well enough for an agent to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema already covers all parameter descriptions at 100%, giving a baseline of 3. The description adds meaning beyond the schema by clarifying that the index comes from session_state output and by explaining the practical consequences of the events parameter, especially the blur behavior that commits onBlur forms.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource ('Type text into an input/textarea element by its index') and clearly scopes the operation to focused typing plus event dispatch. This distinguishes it from siblings like session_click or session_set_files without needing to inspect the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives useful contextual usage guidance: index comes from session_state output, 'full' events are for pages with keyboard listeners, and hidden inputs are valid targets. However, it never says when to prefer this tool over nearby alternatives such as session_eval or session_click, so the when-vs-alternatives guidance is only implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
session_listARead-onlyInspect
List live browser sessions with idle age and the time left before auto-eviction. Use to discover a session to reuse instead of creating a new one; sessions expire after 8 min idle.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotations already declare readOnlyHint=true, so the read-only nature is covered. The description adds meaningful context about live sessions, idle age, auto-eviction, and the 8-minute idle limit, which helps the agent understand the lifecycle without contradicting annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two purposeful sentences. The first states the core function and output fields, and the second gives usage guidance and expiration policy. No filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter, read-only list tool with no output schema, the description is complete. It tells the agent what the tool returns (idle age, eviction time), when to use it, and a key behavioral constraint (8-minute idle expiry). Nothing essential is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so the schema is trivially complete at 100% coverage. Baseline for 0 params is 4; the description adds no parameter details because none are needed, and that is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('List') and resource ('live browser sessions') and clearly states what information is returned ('idle age and the time left before auto-eviction'). It distinguishes itself from session_create by framing the tool as a discovery mechanism for reuse.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly states when to use the tool: 'discover a session to reuse instead of creating a new one'. It also provides a critical operational condition ('sessions expire after 8 min idle'), giving the agent a clear decision rule.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
session_networkARead-onlyInspect
Read the session's network request log. filter="media" extracts playback/stream URLs (m3u8/HLS, mp4, dash, flv...) actually requested by the page's player at runtime - the reliable way to get a real video link, since links embedded in page HTML are often decoys. Media elements and player iframes the engine never fetches (video/audio/source/iframe src) are merged in as candidates: via="network" entries are confirmed requests, via="dom" entries are candidates carrying their tag (iframes = kind "iframe", navigate into them to sniff). Default returns every request as compact rows (method/url/status/type/size). Navigate to the video page first, let it load, then call this.
| Name | Required | Description | Default |
|---|---|---|---|
| filter | No | "media" extracts playback/stream links (m3u8/HLS, mp4, dash, ...) from the requests the page actually issued - the reliable way to get a real video link, since URLs embedded in page HTML are often decoys. Media elements and player iframes the engine never fetches (video/audio/source src, iframe src) are merged in as candidates: entries carry via="network" (confirmed requests) or via="dom" (candidates, with their tag; iframes surface as kind "iframe" - player pages to navigate or sniff inside, not playable URLs). Omit to list every request as compact rows. | |
| session_id | Yes | Session ID | |
| url_contains | No | Narrow the `xhr` array to URLs containing this substring. | |
| body_max_chars | No | Per-body character cap for the `xhr` array (default 4000). | |
| include_bodies | No | Add an `xhr` array of background API responses (the page's own fetch/XHR traffic with retained bodies) alongside the request rows — the page's API face is often the cleanest structured read of its data. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=true, and the description adds substantial behavioral detail: it explains the via='network' vs. via='dom' distinction, merges media elements and iframes as candidates, describes the default output format (compact rows with method/url/status/type/size), and clarifies that iframes are navigable rather than playable. This goes beyond annotations to disclose runtime behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but organized, leading with the core purpose and then layering details about the filter, candidates, and output format. It repeats some of the filter explanation from the schema but remains readable. It could be slightly more concise without losing nuance, but it's efficient.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers the primary use case, prerequisite action, output structure, edge cases (iframe candidates), and parameter behavior. Since there is no output schema, it adequately describes the returned fields (via, kind, rows). This is complete for an agent to call correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has 100% description coverage, so all parameters are documented there. The description adds contextual explanation (e.g., why filter='media' is reliable) but doesn't introduce new parameter meaning beyond what the schema already provides. Per the baseline for full schema coverage, this is a 3.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb and resource: 'Read the session's network request log.' It explains the filter parameter's purpose (extract media URLs) and differentiates from page HTML links as decoys. This distinguishes it from siblings like session_navigate or fetch, which operate differently.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides explicit usage context: 'Navigate to the video page first, let it load, then call this.' It also explains when to use filter='media' vs. the default, and mentions the reliability advantage over embedded links. This effectively guides when to choose this tool over alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
session_preloadAInspect
Replace the session's document-start preload group (empty array clears). Sources run before each new document's own scripts — including inline tags — which is the only hook that beats pages whose signing layer captures window.fetch/XHR natives at parse time (xhs's inline jsvmp). Set before the first navigate; applies to every navigation from then on.
| Name | Required | Description | Default |
|---|---|---|---|
| scripts | No | Full JS sources, in order. Sources run before each new document's own scripts (including inline ones) — the only hook that beats pages whose signing layer captures window.fetch/XHR natives at parse time. `[]` clears the group. Set before the session's first navigate and it applies to every navigation from then on. | |
| session_id | Yes | Session ID |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only provide a title, so the description carries the burden. It discloses the timing (document-start, before inline scripts), persistence across navigations, and clearing semantics. It doesn't mention error cases or whether scripts are re-run on every navigation, but the core behavioral traits are well covered.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded with the core action, then explains the timing and use case. It is slightly dense with technical detail (xhs's inline jsvmp) but every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with two parameters and no output schema, the description covers the essential context: what it does, when to set it, and how it behaves across navigations. It could mention whether the preload applies to the current page if set mid-session, but the guidance to set before first navigate is clear.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents both parameters. The description adds context about the scripts array's ordering and clearing behavior, but this is largely duplicated in the parameter description. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'Replace the session's document-start preload group' and clarifies that an empty array clears it. It also explains the timing ('before each new document's own scripts') and the unique advantage over sibling tools, distinguishing it from session_eval and session_navigate.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly says when to use it: 'Set before the first navigate; applies to every navigation from then on.' It also explains the behavioral context (beats pages whose signing layer captures window.fetch/XHR natives at parse time) and the clearing behavior, giving clear guidance on how to use it versus alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
session_screenshotAInspect
Screenshot the session's CURRENT DOM state (mutations from clicks/evals included) as a base64 PNG via the built-in renderer. Width/height default to the session's viewport, so session_viewport + session_screenshot shows the responsive layout. Returns {url, width, height, image_base64, format}.
| Name | Required | Description | Default |
|---|---|---|---|
| width | No | Render width in CSS pixels; defaults to the session's current viewport | |
| height | No | Render height in CSS pixels; defaults to the session's current viewport | |
| selector | No | CSS selector: capture only that element's box | |
| full_page | No | Capture the full scrollable page instead of the viewport (default: false) | |
| session_id | Yes | Session ID | |
| selector_all | No | With selector, capture every match (default: first match only) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description discloses that it captures the CURRENT DOM state including mutations, which is a key behavioral trait beyond what annotations provide. It also mentions the built-in renderer and the return shape. Annotations only provide a title, so the description carries the burden and does so well, though it doesn't mention potential side effects or performance implications.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences with no waste. The key behavioral trait (current DOM state) is front-loaded, followed by the default behavior and return format. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a screenshot tool with 6 parameters and no output schema, the description covers the essential context: what is captured, defaults, and return shape. It doesn't explain the selector/full_page/selector_all options in detail, but the schema covers those. The description is complete enough for an agent to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all parameters. The description adds context about defaults (viewport) and the return format, but doesn't add significant meaning beyond the schema. Baseline 3 is appropriate since the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Screenshot'), a precise resource ('the session's CURRENT DOM state'), and clarifies that mutations from clicks/evals are included. It also distinguishes itself from generic screenshot tools by mentioning the built-in renderer and the return format. This is clearly differentiated from siblings like session_viewport and render_pdf.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly notes that width/height default to the session's viewport and suggests pairing with session_viewport to show responsive layout. It doesn't explicitly state when NOT to use it versus alternatives like render_pdf or render_video, but the context is clear enough for an agent to infer appropriate usage.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
session_scrollAInspect
Scroll the page up or down by a number of viewport-heights.
| Name | Required | Description | Default |
|---|---|---|---|
| amount | No | Scroll amount in viewport-heights (default: 3) | |
| direction | No | Scroll direction: "up" or "down" (default: down) | down |
| session_id | Yes | Session ID |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With only a title annotation and no readOnly/destructive hints, the description carries the burden of behavioral disclosure. It does state the core behavior (scrolling up/down by viewport-heights), but it omits side effects or limitations such as behavior at the top/bottom of the page, whether scrolling is relative to the current position, or whether the action waits for rendering to settle. This is a reasonable but not thorough disclosure for a simple scroll operation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, front-loaded sentence that conveys the essential action and unit of measurement with zero filler. Every word earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple tool with fully documented parameters, the description covers the primary action but leaves gaps: no mention of return behavior or output (no output schema exists), no mention of session applicability beyond the parameter, and no guidance on repeated or large scroll amounts. It is adequate but not complete for an agent with no other context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all parameters (amount, direction, session_id) are already documented with defaults and formats. The description adds no new parameter-level meaning beyond restating 'up or down' and 'viewport-heights,' which the schema already captures. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Scroll'), a clear resource ('the page'), and a precise unit of measure ('viewport-heights'), which fully captures the tool's function. It is readily distinguishable from sibling tools like session_navigate, session_click, and session_viewport, which do not involve scrolling by viewport heights.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage: use this tool when you need to scroll the currently active page vertically. However, it provides no explicit guidance on when to prefer this tool over alternatives, no mention of prerequisites (e.g., an active session), and no exclusions or edge-case conditions. Usage is self-evident from the action but not explicitly framed.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
session_set_filesAInspect
Select files on a file input programmatically (Playwright setInputFiles semantics): builds File objects from base64 content, assigns them to input.files, then dispatches input+change so framework onChange handlers fire. Selector-addressed because file inputs are often hidden and absent from the session_state index.
| Name | Required | Description | Default |
|---|---|---|---|
| files | Yes | Files to select | |
| selector | Yes | CSS selector for the file input, e.g. "input[type=file]". File inputs are often hidden, so this is selector-addressed rather than using the /state index. | |
| session_id | Yes | Session ID |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations provide only a title, so the description carries the behavioral burden. It transparently discloses the full mechanism: builds File objects from base64, assigns them to input.files, and dispatches input+change events so framework onChange handlers fire. It does not mention failure modes or return behavior, but the core side effects are clear.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two compact sentences front-load the purpose, then efficiently explain the mechanism and the selector rationale. Every clause adds value, with no redundant or filler content.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a three-parameter tool with complete schema coverage and no output schema, this description gives enough behavioral context for correct invocation: the selector approach, event dispatch, and framework-handler implications. It does not cover return values or failure conditions, but that is a minor gap given the detailed mechanism already provided.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all parameters and nested fields. The description adds a useful rationale for the selector parameter and confirms the base64 construction path, but it does not materially extend the schema's parameter documentation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific action—'Select files on a file input programmatically'—and anchors it to Playwright setInputFiles semantics. It also distinguishes itself from state-indexed sibling tools by explaining that file inputs are often hidden and absent from the session_state index, so selector addressing is required.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Clear context is provided: this tool is for attaching files to file inputs, and it explains why selector addressing is used instead of the session_state index. It does not explicitly name an alternative or give when-not-to-use conditions, but the intended scenario is readily inferable.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
session_stateARead-onlyInspect
Get the current page state as an indexed list of interactive elements. Returns compact text with [N] indexes for use with click/input tools.
| Name | Required | Description | Default |
|---|---|---|---|
| session_id | Yes | Session ID |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, so the safety profile is covered. The description adds useful behavioral context by specifying the return format: 'compact text with [N] indexes.' It does not describe edge cases or limitations, but for a simple read-only tool with annotation support, this is adequate.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two concise sentences with no filler. The core behavior ('Get the current page state') is front-loaded, and the return-format detail is provided efficiently. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter read-only tool with no output schema, the description provides enough information to call it correctly: what it returns, the format ([N] indexes), and its intended downstream use. It does not describe the exact structure of an element entry, but this is a minor gap given the simplicity of the operation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The only parameter, session_id, is already documented in the schema with a description, giving 100% schema description coverage. The tool description does not add additional meaning about the parameter. Baseline 3 applies because the schema carries the parameter documentation burden.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'Get the current page state as an indexed list of interactive elements.' It clearly differentiates this from the action-oriented sibling tools (session_click, session_input) by positioning itself as a read-only state retrieval tool. The mention of 'for use with click/input tools' strengthens its identity as the preparatory inspection tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description conveys when to use the tool: before invoking click/input tools, to obtain the [N] indexes needed for those actions. It does not explicitly name alternative tools or state exclusions, but the context is clear and sufficient for an agent to select it over mutation-focused siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
session_storageARead-onlyInspect
Snapshot the session's localStorage/sessionStorage for the current origin: {url, local_storage, session_storage}. Feed it back via session_create's storage field to restore a logged-in state in a new session — the half of login state that cookies can't carry (many sites keep the session token in localStorage). Call before the session idles out.
| Name | Required | Description | Default |
|---|---|---|---|
| session_id | Yes | Session ID |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint: true, and the description reinforces this with 'Snapshot', which implies no mutation. The description adds valuable behavioral context: the exact data returned (url, local_storage, session_storage), its purpose in restoring login state, and a timing constraint (before idle). No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences that are information-dense and free of redundancy. It front-loads the main action and return shape, then provides the use case and a timing note. Each clause earns its place, though it could be slightly shorter without losing critical context.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given there is no output schema, the description compensates by specifying the return fields. It also explains the intended workflow (session_create) and why the tool matters. While it doesn't cover edge cases like empty storage or error conditions, for a simple read-only snapshot tool this is sufficient for an agent to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has a single parameter 'session_id' with 100% coverage (description: 'Session ID'). The tool description does not add any additional meaning or constraints for this parameter beyond what the schema provides. Per the baseline rule for high schema coverage, a score of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb 'Snapshot' and clearly identifies the resource (localStorage/sessionStorage for the current origin). It also specifies the return shape {url, local_storage, session_storage}, which makes the tool's function unambiguous and distinct from generic session tools like session_state or session_export.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides concrete usage guidance: feed the result back via session_create's 'storage' field to restore a logged-in state, and call it before the session idles out. It also explains why it's needed (cookies can't carry this half of login state). It doesn't explicitly mention alternatives or exclusions, but the use case is well-defined.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
session_verdictARead-onlyInspect
One call answers "where did this session land": verdict is one of challenge (risk control engaged — punish page or a 200-status API body that swallowed the wall; the response carries a handoff instruction for a human to solve it in the live view), captcha (explicit CAPTCHA interstitial), login (bounced to a login form — auth expired), empty, landed (normal 2xx content page), or unknown (couldn't classify — read facts; a non-2xx main document lands here with facts.doc_status carrying the number, so a zhihu-style burst 403 is branchable). The facts sheet also carries challenge_events, requests, console_errors and the fired signals. Pure code over signals the engine already holds (current URL, risk-control rows, main document status/size, console errors) — no screenshots, no page evals, single-digit milliseconds. Verdict observes, it never bypasses.
| Name | Required | Description | Default |
|---|---|---|---|
| session_id | Yes | Session ID |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The readOnlyHint annotation is reinforced and extended with valuable behavioral detail: it is pure code over existing signals, takes no screenshots, runs no page evals, executes in single-digit milliseconds, and never bypasses protections. The description also discloses that challenge responses carry a handoff instruction for human resolution, adding context beyond the annotation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the core purpose ('One call answers where did this session land') and then delivers a dense but relevant enumeration of verdicts, facts, and behavior. It is long but every clause earns its place; a list format would improve scannability, but the content is not wasteful.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the burden of explaining return semantics, and it does so thoroughly: all verdict categories, the facts sheet contents, handling of non-2xx documents, and the safety/performance profile. For a one-parameter classification tool, an agent has enough context to call it correctly and interpret the result.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage for the single parameter is 100%, so the schema already documents session_id. The description does not add any additional meaning about the parameter itself, which matches the baseline of 3 when structured schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's job: classify where a session landed, and enumerates the possible verdicts (challenge, captcha, login, empty, landed, unknown) with concrete meanings. It distinguishes itself from action tools by saying it 'observes, never bypasses' and uses 'no screenshots, no page evals,' but it does not explicitly differentiate itself from similar observation tools like session_state or session_network.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context for when to call it: after a session action, to answer 'where did this session land' and to branch on the resulting verdict/facts, including handling non-2xx documents via facts.doc_status. It does not explicitly name alternative tools or state when not to use it, but the intended use case is strongly implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
session_viewportAInspect
Set the session's viewport (device emulation): scripts see innerWidth/innerHeight move, media queries like (max-width: 600px) re-evaluate, element rects re-anchor, and mobile=true flips pointer/hover matchMedia answers to coarse/none. Omitted width/height keeps the current value.
| Name | Required | Description | Default |
|---|---|---|---|
| width | No | Viewport width in CSS pixels; omit to keep the current width | |
| height | No | Viewport height in CSS pixels; omit to keep the current height | |
| mobile | No | Mobile emulation: matchMedia answers pointer:coarse / hover:none and navigator.maxTouchPoints reports 5 (default: false) | |
| session_id | Yes | Session ID |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With annotations only providing a title, the description carries the behavioral burden and does so well: it reveals that scripts observe innerWidth/innerHeight changes, media queries re-evaluate, element rects re-anchor, and mobile flips matchMedia pointer/hover answers. It does not mention session prerequisites or persistence, but the core state-changing behavior is transparent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two dense sentences deliver high-signal information with no filler. The main action is front-loaded, followed by the most important behavioral effects and the omitted-parameter rule, making it easy for an agent to parse quickly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a state-mutating tool with a complete input schema and no output schema, the description explains the behavioral consequences well enough to invoke it correctly. It lacks explicit return-value or error/edge-case information, but those are not essential for a setter of this kind.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and each parameter already has a clear description. The tool description reinforces the semantics of omitted width/height preserving current values and describes the mobile flag's effect, but it does not add meaning substantially beyond what the schema already provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Set') and resource ('the session's viewport'), and clearly distinguishes device emulation from a simple resize by enumerating the observable effects on scripts, media queries, and element rects. It also clarifies the mobile mode behavior, leaving no ambiguity about what the tool does.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context for when to use this tool: when the session needs device emulation or viewport resizing, including media-query re-evaluation and pointer/hover emulation. It does not explicitly name alternatives or state when not to use it, but no direct sibling appears to overlap with this capability.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
session_waitARead-onlyInspect
Wait until a CSS selector matches or a JS predicate turns truthy, with a timeout. The page's event loop keeps running while waiting (fetches, timers, promise chains progress), so this replaces blind sleeps for async content: navigate, session_wait for '.price-card', then click/read. Returns {matched, elapsed_ms, detail:{tag,text} or the predicate value}; errors with timeout ... naming the selector/predicate on expiry. Exactly one of selector/predicate.
| Name | Required | Description | Default |
|---|---|---|---|
| selector | No | CSS selector to wait for (e.g. ".price-card") | |
| predicate | No | JS expression polled until truthy (e.g. "document.querySelectorAll('.card').length >= 3") | |
| session_id | Yes | Session ID | |
| timeout_ms | No | Give up after this many milliseconds (default: 10000, max: 120000) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnlyHint annotation, the description reveals that the page's event loop continues running (fetches, timers, promise chains progress), which is important behavioral context. It also discloses the return shape, error text format, and the exactly-one-of constraint. This is substantial added value beyond the annotation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences, tightly packed with purpose, behavior, usage example, return format, error behavior, and the exclusivity constraint. No filler words; key information is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with no output schema, the description fully covers what an agent needs: what it waits for, the async behavior, a usage example, the return shape, error format, and the parameter constraint. The only minor omission is polling interval, but that is not essential for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3. The description adds a critical non-schema constraint: 'Exactly one of selector/predicate.' It also explains the meaning of selector and predicate in the wait context and what the return detail field contains depending on which is used, going beyond the schema's basic definitions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states a specific action: wait until a CSS selector matches or a JS predicate is truthy, with a timeout. It gives a concrete example workflow (navigate, session_wait for '.price-card'), which distinguishes it from blind sleeps, but it does not explicitly compare itself to sibling tools like session_eval or session_console.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly positions the tool as a replacement for blind sleeps in async scenarios: 'so this replaces blind sleeps for async content'. It provides a typical usage pattern, which gives clear context for when to use it. It stops short of stating exclusions (e.g., when not to use) and does not name alternative tools, so it loses the fifth point.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
1 tool update
v0.5.11- Changed
fetch2 fields changed- changed
Input schema / $defs / RenderTier / oneOfPrevious value: -[ - { - "const": "auto", - "description": "HTTP-direct first, fall back to diting browser. (default)", - "type": "string" - }, - { - "const": "http", - "description": "Pure HTTP, no V8/JS. Fastest; misses JS-rendered content.", - "type": "string" - }, - { - "const": "obscura", - "description": "Always use the diting browser (current behaviour pre-tiering).\n\"browser\" is accepted as an alias — agents guess it before \"obscura\".", - "type": "string" - } -]New value: +[ + { + "const": "auto", + "description": "HTTP-direct first, fall back to diting browser. (default)", + "type": "string" + }, + { + "const": "http", + "description": "Pure HTTP, no V8/JS. Fastest; misses JS-rendered content.", + "type": "string" + }, + { + "const": "browser", + "description": "Always use the diting browser (current behaviour pre-tiering).\nWire name is \"browser\". \"obscura\" is still accepted and not advertised.", + "type": "string" + } +] - changed
Input schema / properties / render_tier / descriptionPrevious value: -"Rendering strategy: \"auto\" (default), \"http\", or \"obscura\""New value: +"Rendering strategy: \"auto\" (default), \"http\", or \"browser\""
1 tool update
v0.5.8- Changed
session_create1 field changed- changed
Input schema / properties / url / descriptionPrevious value: -"Initial URL to navigate to (optional)"New value: +"Initial URL to navigate to (optional). `start_url` is honored as an\nalias (#115) — callers guessing that name must not land on about:blank."
1 tool update
v0.5.6- Added
session_preload
3 tool updates
v0.5.5- Added
account_login - Changed
flow_run1 field changed- added
Input schema / properties / max_stepsAdded value: +{ + "default": null, + "description": "Override the run's step-execution budget (branch loops re-run steps,\nso every revisit counts). Default 1000, clamped 1..=100000. A flow\ndocument may also declare its own max_steps; this wins.", + "format": "uint64", + "minimum": 0, + "type": [ + "integer", + "null" + ] +}
- Added
session_verdict
36 tool updates
v0.5.1- Added
account_delete - Added
account_list - Added
account_verify - Changed
cache1 field changed- removed
Input schema / titleRemoved value: -"CacheParams"
- Changed
click1 field changed- removed
Input schema / titleRemoved value: -"ClickParams"
- Changed
download1 field changed- removed
Input schema / titleRemoved value: -"DownloadParams"
- Changed
eval1 field changed- removed
Input schema / titleRemoved value: -"EvalParams"
- Changed
fetch1 field changed- removed
Input schema / titleRemoved value: -"FetchParams"
- Added
flow_run - Changed
import_curl2 fields changed- added
Input schema / properties / accountAdded value: +{ + "default": null, + "description": "Attach the session to a named account: the imported login lands in\nthe account's private jar and is written back under its name after\nevery action — one import per identity, no clobbering.", + "type": [ + "string", + "null" + ] +} - removed
Input schema / titleRemoved value: -"ImportCurlParams"
- Changed
render_markdown2 fields changed- added
Input schema / properties / motionAdded value: +{ + "description": "Bake the entrance choreography into the artifact (default false):\npure-declarative CSS animation — headings split into per-glyph/per-\nword spans that rise in with expo easing, prose blocks stagger up a\nnth-child delay ladder, and diagram figures grow in with a back\nease (GSAP's easing math as public cubic-bezier equivalents). The\ndiagrams animate too, on one story clock: nodes pop in one beat at\na time, solid edges draw themselves (dash-offset drain), dashed\nreturns fade, sequence messages land as they are \"sent\", and a\ntimed caption strip under each figure subtitles the beats — under\nprefers-reduced-motion the strip becomes a static transcript.\nZero scripts: the file itself animates in any browser, subtitles\nand all; the receipt records motion (plus diagrams[].story with\nthe beat times, the hook for muxing voice later) so a cached\nartifact is never mistaken for the static one", + "type": [ + "boolean", + "null" + ] +} - removed
Input schema / titleRemoved value: -"RenderMarkdownParams"
- Added
render_pdf - Added
render_video - Changed
search1 field changed- removed
Input schema / titleRemoved value: -"SearchParams"
- Added
session_challenges - Changed
session_click1 field changed- removed
Input schema / titleRemoved value: -"SessionClickParams"
- Changed
session_click_xy1 field changed- removed
Input schema / titleRemoved value: -"SessionClickXyParams"
- Changed
session_clone1 field changed- removed
Input schema / titleRemoved value: -"SessionCloneParams"
- Changed
session_close1 field changed- removed
Input schema / titleRemoved value: -"SessionCloseParams"
- Changed
session_console1 field changed- removed
Input schema / titleRemoved value: -"SessionConsoleParams"
- Changed
session_cookies1 field changed- removed
Input schema / titleRemoved value: -"SessionCookiesParams"
- Changed
session_create2 fields changed- added
Input schema / properties / accountAdded value: +{ + "default": null, + "description": "Run as a named login identity (the multi-account layer): a private\ncookie jar seeded from the account record, write-back to the account\nstore after every action. Concurrent logins (`taobao-scraper` vs\n`taobao-publisher`) never clobber each other. The account record\nsurvives the session — a later create with the same name picks up\nthe warm jar. 1-64 chars of [a-zA-Z0-9_-].", + "type": [ + "string", + "null" + ] +} - removed
Input schema / titleRemoved value: -"SessionCreateParams"
- Changed
session_dialog1 field changed- removed
Input schema / titleRemoved value: -"SessionDialogParams"
- Changed
session_drag4 fields changed- changed
Input schema / properties / delay_ms / descriptionPrevious value: -"Delay between moves in ms (default 30) — gives mousemove-driven\nwidgets time to react per step"New value: +"Mean delay between moves in ms (default 18 humanized / 30 linear) —\nper-step timing is jittered around this when humanizing" - added
Input schema / properties / humanizeAdded value: +{ + "default": null, + "description": "Humanize the trajectory: minimum-jerk easing, perpendicular wobble,\ntiming jitter, grip/settle pauses, occasional hesitation and\novershoot-and-correct. Set false when a test/tool needs exact linear\ninterpolation. Default: true", + "type": [ + "boolean", + "null" + ] +} - changed
Input schema / properties / steps / descriptionPrevious value: -"Interpolated mousemove events between from and to (default 10)"New value: +"Interpolated mousemove events between from and to. Default: 24 with\nhumanize on, 10 without" - removed
Input schema / titleRemoved value: -"SessionDragParams"
- Changed
session_eval2 fields changed- added
Input schema / properties / timeout_msAdded value: +{ + "default": null, + "description": "Await budget for the script's promise in ms (default 5000, clamped\n100..120000). Pass a larger budget for slow page-side work such as\nuploads through the page's own fetch; on expiry the tool errors with\nEVAL_TIMEOUT (the script may still be running) instead of returning\na null result.", + "format": "uint64", + "minimum": 0, + "type": [ + "integer", + "null" + ] +} - removed
Input schema / titleRemoved value: -"SessionEvalParams"
- Changed
session_export2 fields changed- changed
Input schema / properties / format / descriptionPrevious value: -"Output format: \"bash\" (default) renders a runnable curl script that\nreplays every recorded action against a fresh session; \"jsonl\" returns\nthe raw action log, one JSON object per line"New value: +"Output format: \"bash\" (default) renders a runnable curl script that\nreplays every recorded action against a fresh session; \"jsonl\" returns\nthe raw action log, one JSON object per line; \"json\" returns a\nflow.json document (editable ops, cookies/storage stripped) for\nreplay via flow_run" - removed
Input schema / titleRemoved value: -"SessionExportParams"
- Changed
session_input1 field changed- removed
Input schema / titleRemoved value: -"SessionInputParams"
- Changed
session_navigate1 field changed- removed
Input schema / titleRemoved value: -"SessionNavigateParams"
- Changed
session_network1 field changed- removed
Input schema / titleRemoved value: -"SessionNetworkParams"
- Changed
session_screenshot1 field changed- removed
Input schema / titleRemoved value: -"SessionScreenshotParams"
- Changed
session_scroll1 field changed- removed
Input schema / titleRemoved value: -"SessionScrollParams"
- Added
session_set_files - Changed
session_state1 field changed- removed
Input schema / titleRemoved value: -"SessionStateParams"
- Changed
session_storage1 field changed- removed
Input schema / titleRemoved value: -"SessionCookiesParams"
- Changed
session_viewport1 field changed- removed
Input schema / titleRemoved value: -"SessionViewportParams"
- Changed
session_wait1 field changed- removed
Input schema / titleRemoved value: -"SessionWaitParams"
3 tool updates
v0.3.3-rc1- Changed
fetch3 fields changed- changed
Input schema / $defs / RenderTier / oneOfPrevious value: -[ - { - "const": "auto", - "description": "HTTP-direct first, fall back to diting browser. (default)", - "type": "string" - }, - { - "const": "http", - "description": "Pure HTTP, no V8/JS. Fastest; misses JS-rendered content.", - "type": "string" - }, - { - "const": "obscura", - "description": "Always use the diting browser (current behaviour pre-tiering).", - "type": "string" - } -]New value: +[ + { + "const": "auto", + "description": "HTTP-direct first, fall back to diting browser. (default)", + "type": "string" + }, + { + "const": "http", + "description": "Pure HTTP, no V8/JS. Fastest; misses JS-rendered content.", + "type": "string" + }, + { + "const": "obscura", + "description": "Always use the diting browser (current behaviour pre-tiering).\n\"browser\" is accepted as an alias — agents guess it before \"obscura\".", + "type": "string" + } +] - added
Input schema / properties / capture_xhrAdded value: +{ + "default": null, + "description": "Capture script-initiated API responses: a list of URL substrings\n(e.g. [\"/api/\"]) whose matching fetch/XHR bodies come back in an\n`xhr` array; an empty list captures every XHR/Fetch. Forces browser\nrendering (script-initiated requests only exist after JS runs).", + "items": { + "type": "string" + }, + "type": [ + "array", + "null" + ] +} - added
Input schema / properties / sanitizeAdded value: +{ + "default": true, + "description": "Strip prompt-injection payloads from the text output (default true):\nzero-width/steganographic characters, instruction-shaped lines\n(\"ignore previous instructions\", chat markup tokens, CJK variants),\nand text hidden via opacity:0 / tiny fonts. A `sanitize_report`\nfield counts what was removed — stripping is observable, never\nsilent. Set false for raw output.", + "type": "boolean" +}
- Changed
search2 fields changed- added
Input schema / properties / enginesAdded value: +{ + "default": [], + "description": "Restrict to these engine names (e.g. [\"baidu\"], [\"sogou_wechat\"]).\nEmpty = all engines serving `categories`. Invalid names return an\nerror listing the valid ones.", + "items": { + "type": "string" + }, + "type": "array" +} - added
Input schema / properties / time_rangeAdded value: +{ + "default": null, + "description": "Freshness window: \"day\" | \"week\" | \"month\" | \"year\". Honored by\nengines with dated results (e.g. bing_news filters by pubDate);\nothers ignore it.", + "type": [ + "string", + "null" + ] +}
- Changed
session_network3 fields changed- added
Input schema / properties / body_max_charsAdded value: +{ + "default": null, + "description": "Per-body character cap for the `xhr` array (default 4000).", + "format": "uint", + "minimum": 0, + "type": [ + "integer", + "null" + ] +} - added
Input schema / properties / include_bodiesAdded value: +{ + "default": null, + "description": "Add an `xhr` array of background API responses (the page's own fetch/XHR\ntraffic with retained bodies) alongside the request rows — the page's\nAPI face is often the cleanest structured read of its data.", + "type": [ + "boolean", + "null" + ] +} - added
Input schema / properties / url_containsAdded value: +{ + "default": null, + "description": "Narrow the `xhr` array to URLs containing this substring.", + "type": [ + "string", + "null" + ] +}
TDQS
Scored across 40 tools
Tool families are cleanly separated: stateless fetch/eval/click are explicitly distinguished from their session_* counterparts, and account_*/render_*/session_* names signal intent. The main fuzzy boundaries are session_challenges vs session_verdict (both report anti-bot/risk status) and the three render_* tools, though descriptions usually resolve them.
Most tools follow predictable families—account_<verb>, session_<operation>, render_<format>—and bare stateless verbs (fetch, search, eval, click) are easy to read. Deviations such as flow_run instead of run_flow and noun-style read tools like session_state/session_network/session_cookies keep it from being perfectly uniform.
40 tools is a heavy surface, especially with 26 session_* variants; even though each is a real browser action, the set will tax an agent's selection and context budget. The server's broad scope explains the count, but it still exceeds what most MCP clients handle gracefully.
The surface covers the full browser workflow: read/search/download, session lifecycle and input, login-state persistence, risk-control verdicts, flow replay, and render/export targets. Minor gaps such as no explicit multi-tab/window management or standalone proxy/user-agent configuration are workaroundable.
Maintenance
Related MCP Connectors
Headless browser primitives for AI agents when sites need real JS rendering.
Hosted browser for AI agents: screenshots, post-JS DOM, console, WCAG. No install, no API key.
Web search, browser automation, scraping, crawling and CAPTCHA solving for AI agents.
- TabfleetOAuthcom.tabfleet
Launch, inspect, control, and share isolated cloud browsers for your agents.
Related MCP Servers
- AlicenseNot gradedqualityNot gradedmaintenanceA Playwright-based MCP server that exposes a live browser as a traceable, inspectable, debuggable and controllable execution environment for AI agents.5,334 npm57-
- AlicenseAqualityBmaintenanceA high-performance browser automation MCP server that provides AI agents with a fast, persistent Chromium instance via Playwright. It features reference-based element interaction, snapshot diffing, and manual handoff capabilities to handle complex tasks like CAPTCHAs.6137 npm32MIT
- AlicenseAqualityDmaintenanceMulti-session browser MCP server that gives AI agents up to 15 fully-isolated browsers running in parallel. 36 tools including navigation, extraction, network intercept, stealth, and self-improvement. Each session has its own cookies, storage, and fingerprint so agents never collide.3762 npm5MIT
- FlicenseBqualityDmaintenanceA minimalist browser control engine that allows LLM agents to visually perceive and interact with web pages through the Chrome DevTools Protocol and MCP standard.411-