Auto-Browser
Auto Browser is an MCP server for controlling a shared Playwright browser with human-in-the-loop safeguards, session/auth reuse, and page inspection tools.
Create, list, get, fork, and close browser sessions, optionally navigating to a start URL.
Observe pages as text, screenshots, OCR, or rich summaries, and take standalone screenshots.
Read full page HTML or visible text in paged chunks, and find elements by CSS selector or text/regex query.
Execute browser actions like navigate, click, type, hover, select, scroll, wait, reload, go back/forward, and upload.
Manage tabs: list, activate, and close open tabs/pages.
Save and reuse named auth profiles, or start sessions from saved sign-in state.
List and read downloaded files as text, with binary files refused but accessible via artifact URL.
Wait for CSS selectors to reach visible, hidden, attached, or detached states.
Evaluate JavaScript expressions, requiring operator approval for each exact expression.
Request human takeover of the shared browser desktop.
Use safety features such as policy checks, approvals, audit trails, operator identity, PII scrubbing, and Witness receipts when configured.
Work with MCP clients over HTTP or via the bundled stdio bridge.
Auto Browser

Give your AI agent a real browser, with a human in the loop.
Auto Browser is an MCP-native browser control plane for authorized workflows. It gives MCP clients, LLM agents, and operators a shared Playwright browser with human takeover, reusable auth profiles, approvals, audit trails, and local-first deployment.
Works with:
Claude Code, Codex (CLI, IDE extension, and app), Google Antigravity, Cursor, and VS Code, connected directly over HTTP
Claude Desktop, through the bundled stdio bridge
any MCP client that can talk HTTP or stdio
direct REST callers when you want curl-first control
Why Auto Browser
MCP-native from day one. The browser surface is already packaged as an MCP server instead of bolted on after the fact.
Human takeover when the web gets brittle. noVNC keeps the same live session available when a person needs to step in.
Login once, reuse later. Save named auth profiles and reopen fresh sessions that are already signed in.
Local-first by default. Run the full stack on your own box with Docker Compose, or use Codespaces for a quick hosted demo.
Safety rails built in. Approvals, operator identity, PII scrubbing, Witness receipts, and policy presets are all part of the product surface.
Evidence you can hand to someone else. Witness receipt chains are Ed25519-signed, and an exported bundle verifies with
scripts/verify_witness_bundle.py— which imports nothing from this project, so a recipient need not run or trust this controller to check it.We audit ourselves in public.
docs/audits/2026-08-execution-audit.mddocuments an adversarial audit of this repo that found safety controls which reported success while doing nothing, with reproductions, the fixes, and the gates that close the class.Governed skill induction. Verified browser traces can become staged skill candidates with provenance that is signed when a mesh identity is configured — and checked on read, not just produced — plus verifier adapters and review-only graduation — agents that prove they can repeat themselves correctly, not just act once.
Related MCP server: Browser-Use MCP Server
Good Fits
internal dashboards and admin tools
operator-assisted QA and browser debugging
login-once, reuse-later account workflows
brittle sites where a human may need to recover the flow
MCP-powered agent workflows that need a real browser, not just HTML fetches
Not the Goal
CAPTCHA solving
unauthorized scraping or account automation
deceptive identity shaping or bypass tooling
What You Get
Browser Control | Operator Safety | Deployment and Integration |
Playwright-backed sessions with screenshots, DOM summaries, OCR excerpts, tab controls, downloads, and network inspection | approval gates, operator identity headers, audit events, PII scrubbing, Witness receipts, and protection profiles | MCP over HTTP, bundled stdio bridge, REST API, Docker Compose, Codespaces, auth profiles, and optional per-session isolation |
Quickstart
git clone https://github.com/LvcidPsyche/auto-browser.git
cd auto-browser
docker compose up --buildThat is enough for local development with the default settings.
Optional:
cp .env.example .env
make doctorRun make doctor from a normal terminal with local Docker access and permission to open localhost sockets.
Open:
API docs:
http://127.0.0.1:8000/docsOperator dashboard:
http://127.0.0.1:8000/dashboard, which includes the queue of actions waiting on your approvalVisual takeover:
http://127.0.0.1:6080/vnc.html?autoconnect=true&resize=scale
All published ports bind to 127.0.0.1 by default.
Try It in Codespaces
Codespaces provisions the stack automatically. The dashboard and noVNC tabs are usually ready in about 90 seconds.
First Useful Demo
The highest-signal flow in this repo is:
create a session
log in manually if the site needs a human
save the session as a named auth profile
open a new session from that auth profile
continue work without reauthing
Start here:
Minimal session creation:
curl -s http://127.0.0.1:8000/sessions \
-X POST \
-H 'content-type: application/json' \
-d '{"name":"demo","start_url":"https://example.com"}' | jqMinimal observation:
curl -s http://127.0.0.1:8000/sessions/<session-id>/observe | jqRecent Changes
1.8.1
Sturdier sessions and stores. Actions that ran are no longer reported as failed when the page navigates right after, an approved action runs at most once,
MAX_SESSIONSholds under concurrent creates, and sessions, tabs, cron jobs and the JSON stores no longer leak or lose writes.The stdio MCP bridge survives a controller restart and reports auth and rate-limit errors instead of hanging. The SDK, bridge and LangChain adapters can send an operator id for controllers with
REQUIRE_OPERATOR_ID=true.Faster actions and observations, fewer browser round trips per observation, and a controller image without the test tooling.
Security fixes for the noVNC socket, TOTP autofill, auth-profile paths, approvals and witness receipts.
create_sessionwithtotp_secretnow needstotp_hostsor astart_url.
1.8.0
Smaller results for agents. MCP results refer to sessions instead of repeating the full session record, and
execute_actionno longer returns the pre-action snapshot (detail="full"restores both). An action result is less than half its old size. The default tool list carries the 20 tools a browsing agent needs, and the rest are oneMCP_TOOL_PROFILE=fullaway.Agents can see and read.
browser.screenshotand observe'sfastpreset return the screenshot as MCP image content.browser.read_downloadreads a downloaded CSV, JSON or text file.browser.get_html(text_only=true)is paged and keeps line breaks and table cells.Approvals you can find and trust. The dashboard has a pending-approvals queue with Approve and Reject. A governed tool call such as
browser.eval_jsis approved for its exact arguments, which the operator sees.Observations name things as a person reads them. Fields are labelled from their
<label>, never from what was typed into them. The accessibility outline works again on current Playwright.A security pass. The navigation allowlist now matches how Chromium parses URLs. A tokenless controller refuses DNS-rebinding Host headers, downloaded artifacts are served sandboxed, share links are scoped, and the browser runs as an unprivileged user.
A leaner image. The controller leaves out the provider CLIs (about 750 MB) unless built with
INSTALL_AGENT_CLIS=true.
1.7.0 closed the fail-open design issues from GHSA-xmh3-cw7j-9gp5. A reachable API now needs a token (API_BIND_SCOPE), operator identity can be proven by a named credential, auth profiles belong to the operator who saved them, and staged skills are signature-checked on read.
See CHANGELOG.md for the full release history.
MCP Clients
Auto Browser exposes:
an HTTP MCP endpoint at
http://127.0.0.1:8000/mcpconvenience endpoints at
http://127.0.0.1:8000/mcp/toolsandhttp://127.0.0.1:8000/mcp/tools/calla stdio bridge:
uvx auto-browser-mcpfrom PyPI, orscripts/mcp_stdio_bridge.pyin a repo checkout
Clients that speak MCP over HTTP connect directly. With Claude Code or Codex:
claude mcp add --transport http auto-browser http://127.0.0.1:8000/mcp
codex mcp add auto-browser --url http://127.0.0.1:8000/mcpGoogle Antigravity takes the server in its raw MCP config (MCP Servers, then Manage MCP
Servers, then View raw config), or in ~/.gemini/config/mcp_config.json. It reads
serverUrl, not url:
{
"mcpServers": {
"auto-browser": { "serverUrl": "http://127.0.0.1:8000/mcp" }
}
}Cursor, VS Code, bearer tokens, and pairing Auto Browser with a web-search MCP
server are covered in docs/mcp-clients.md.
The default MCP tool profile is curated: the 20 tools a browsing agent needs (sessions, observe and screenshot, execute_action, page and download reading, tabs, auth profiles, human takeover). Every listed tool costs the model context on every request, so the diagnostics, audit, harness, and admin tools are in the full profile. To expose them, set:
MCP_TOOL_PROFILE=fullRaw tool-call example:
curl -s http://127.0.0.1:8000/mcp/tools/call \
-X POST \
-H 'content-type: application/json' \
-d '{
"name":"browser.create_session",
"arguments":{
"name":"demo",
"start_url":"https://example.com"
}
}' | jqClient setup guides:
For resource listing, resource reads, and subscription-style update examples,
see docs/mcp-clients.md#resources-and-subscriptions.
Convergence Harness
Auto Browser ships a Stage 0 convergence harness for Agent Skill Induction. It runs a structured task contract, records tamper-checked traces, verifies completion, and writes a staged skill candidate carrying provenance. With a mesh identity configured that provenance is signed, and the registry verifies the signature before serving a candidate — a candidate that fails the check, or that was dropped into the staging directory unsigned, is refused. Candidates induced from a mock run are marked simulated so they cannot pass as converged. Generated skills are staged only — promotion stays explicit and reviewed.
The harness tools — convergence runs, run status and traces, drift checks, candidate management, and graduation — are in the full MCP tool profile (MCP_TOOL_PROFILE=full), or can be invoked directly over REST.
Start with docs/convergence-harness.md. A deterministic local smoke is:
python -m controller.harness.run --contract evals/contracts/example_read.json --mock-final-url https://example.com --mock-final-text "Example Domain"For MCP clients, set MCP_TOOL_PROFILE=full to expose the harness.* tools.
Security and Compliance
For a real private deployment, set at least:
APP_ENV=production
API_BIND_SCOPE=exposed
API_BEARER_TOKEN=<strong-random-secret>
REQUIRE_OPERATOR_ID=true
AUTH_STATE_ENCRYPTION_KEY=<44-char-fernet-key>
REQUIRE_AUTH_STATE_ENCRYPTION=true
SHARE_TOKEN_SECRET=<strong-random-secret>
CONTROLLER_ALLOWED_HOSTS=<controller hostname>
ALLOWED_HOSTS=<sites the browser may visit>
REQUEST_RATE_LIMIT_ENABLED=true
METRICS_ENABLED=true
STEALTH_ENABLED=falseWith APP_ENV=production the controller refuses to start while any of these is missing, and the startup log names the missing ones.
By default every session shares one Chromium process and one noVNC desktop (SESSION_ISOLATION_MODE=shared_browser_node). Cookies and storage stay separate per session, but a human taking over one session can see the others' windows. When sessions belong to different people, accounts or trust domains, start with make up-isolation to give each session its own browser container and takeover surface (docker_ephemeral). docs/session-isolation-audit.md has the details.
COMPLIANCE_TEMPLATE can apply a preconfigured posture at startup:
Preset | Auth Encryption | Operator ID | PII Scrub | Isolation | Max Session Age |
| required | required | all layers |
| 4h |
| - | required | network + text | shared | 24h |
Both presets require upload approvals and enable Witness receipts. Startup writes the applied policy to /data/compliance-manifest.json. The legacy names (HIPAA, SOC2, GDPR, PCI-DSS) still work as deprecated aliases and emit a warning at startup.
Example:
COMPLIANCE_TEMPLATE=strict docker compose upFor deployment details, hosted Witness notes, CLI auth modes, and reverse-SSH guidance, see:
Architecture at a Glance
flowchart LR
User[Human operator] -->|watch / takeover| noVNC[noVNC]
LLM[Any model: OpenAI / Claude / Gemini / OpenRouter / Grok / DeepSeek / MiniMax / local] -->|shared tools| Controller[Controller API]
Controller -->|Playwright protocol| Browser[Browser node]
noVNC --> Browser
Browser --> Artifacts[(screenshots / traces / auth state)]
Controller --> Artifacts
Controller --> Policy[Allowlist + approval gates]Model providers
First-class adapters for OpenAI, Claude, and Gemini (API or CLI). Beyond those,
a single generic OpenAI-compatible adapter drives any model reachable over an OpenAI
/chat/completions endpoint — set an API key to enable it:
Provider | Reaches |
| one key → ~every frontier model (Claude, GPT, Gemini, Grok, DeepSeek, Llama, Mistral, Qwen, …) |
| Grok |
| DeepSeek (text-only; driven from the DOM/accessibility outline) |
| MiniMax |
| any custom base URL — self-hosted Ollama / vLLM / LM Studio, Azure OpenAI, Together, Groq, Fireworks, … |
Vision (screenshots) is used for every provider except text-only ones. See .env.example for
the *_API_KEY / *_BASE_URL / *_MODEL settings.
Core components:
browser-node/runs Chromium, Xvfb, x11vnc, and noVNCcontroller/exposes the FastAPI controller, MCP transport, policy rails, and orchestration endpointsdata/holds runtime artifacts, auth state, approvals, audit logs, and optional CLI cachesscripts/contains local helpers for doctor, smoke tests, bridges, and release checks
Repo Guide
Path | What It Contains |
controller API, MCP transport, tests, and packaging | |
browser runtime and Playwright connection layer | |
copy-paste flows and MCP client setup | |
LangChain, LangGraph, and CrewAI adapters | |
architecture, deployment, hardening, and launch docs | |
doctor, smoke harnesses, stdio bridge, and auth helpers | |
supporting service templates and operational assets |
Common Commands
Command | Purpose |
| list available repo commands |
| run Ruff checks across the whole repo |
| run controller tests in Docker |
| run controller tests on host Python 3.11+ |
| run deterministic provider/profile eval scoring |
| run the local readiness smoke |
| run the fuller release-validation pass |
| verify per-session Docker isolation |
| verify reverse-SSH remote access |
Documentation Map
If You Want To... | Start Here |
understand the system shape | |
connect Claude, Codex, Antigravity, Cursor, or VS Code | |
run the curl-first examples | |
deploy on a trusted host | |
review production constraints | |
run the convergence harness | |
inspect release history | |
see where the project is headed |
Contributing
If you want to help, start with:
If Auto Browser is useful, a star helps other people find it. Sponsorship and tip options live in TIPS.md.
Available Tools
20 toolsbrowser.activate_tabA
Bring one tab to the foreground so subsequent observations and actions target it. Tab indexes come from browser.list_tabs. Returns the activated index, the updated session summary, and the current tab list.
| Name | Required | Description | Default |
|---|---|---|---|
| index | Yes | Zero-based index of the target tab, as reported by browser.list_tabs. | |
| session_id | No | Target session id, from browser.create_session or browser.list_sessions. Optional when exactly one session is live; observe and act tools create one when none is. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish the mutation profile (readOnlyHint=false, idempotentHint=false), so the bar is lower. The description adds useful behavior beyond annotations: it changes the foreground target for later calls and documents the return payload (index, session summary, tab list). No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences with the main purpose front-loaded and no filler. Every clause earns its place: what it does, why, where the parameter comes from, and what it returns.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a low-complexity tool with two well-described parameters, no output schema, and annotations, the description's mention of return values compensates for the missing output schema. The only omission, error behavior on invalid index, is minor and inferable from the schema's minimum constraint.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the schema fully documents index and session_id. The description reinforces only that indexes come from browser.list_tabs and adds no new semantics beyond what the schema already provides, so baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States exactly what it does ('Bring one tab to the foreground') and why ('so subsequent observations and actions target it'), clearly distinguishing it from siblings like browser.list_tabs and browser.close_tab. The verb-resource pairing is specific and unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit workflow: tab indexes come from browser.list_tabs, and activation is for targeting subsequent observations and actions. It does not enumerate exclusions or name alternative tools, but the context is clear enough to choose it over tab-listing and tab-closing siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser.close_sessionBDestructive
Close a session and finalize its trace/artifacts.
| Name | Required | Description | Default |
|---|---|---|---|
| session_id | No | Target session id, from browser.create_session or browser.list_sessions. Optional when exactly one session is live; observe and act tools create one when none is. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already convey destructiveHint=true, readOnlyHint=false, and non-idempotence. The description adds the useful detail that trace/artifacts are finalized, but it does not clarify irreversibility, effect on open tabs, or whether closing returns any confirmation. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single compact sentence with no filler. The primary action is front-loaded, and the trace/artifact finalization detail earns its place by clarifying an important side effect.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with one optional parameter and a rich schema description, the description is largely sufficient: it states the action and the key finalization behavior. The main gap is the absence of any expected return or error information, but given the annotations and simple scope this is a minor omission.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The single parameter session_id is already fully described in the schema, including its optionality and how to obtain the id. The tool description itself adds no parameter-level meaning, so the baseline of 3 applies because schema coverage is 100%.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly names the verb ('Close') and the resource ('a session'), and adds a specific secondary effect ('finalize its trace/artifacts'). It implicitly differentiates from browser.close_tab by targeting the session level, though it does not explicitly distinguish itself from other session-management siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no explicit guidance on when to prefer close_session over alternatives like close_tab, nor any prerequisites such as having a live session. Usage is only implied by the action name and the 'finalize' phrasing, not stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser.close_tabADestructive
Close one tab index if more than one tab is open.
| Name | Required | Description | Default |
|---|---|---|---|
| index | Yes | Zero-based index of the target tab, as reported by browser.list_tabs. | |
| session_id | No | Target session id, from browser.create_session or browser.list_sessions. Optional when exactly one session is live; observe and act tools create one when none is. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true and readOnlyHint=false, so the description's 'Close' action aligns with those. The description adds the guardrail that it only works when more than one tab is open, which is useful behavioral context. However, it doesn't disclose what happens if the index is invalid or if only one tab is open, nor does it mention whether the active tab changes after closing.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single sentence that is front-loaded with the action and resource, and includes the key condition. Zero wasted words. It earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple destructive action with full schema coverage and annotations indicating destructiveness, the description is mostly adequate. However, it doesn't mention edge-case behavior (e.g., what happens if index is out of range or if only one tab is open), and there's no output schema to clarify return values. The description is sufficient for basic use but leaves some behavioral gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents both parameters thoroughly. The description adds no additional parameter-level meaning beyond what the schema provides. Baseline 3 is appropriate since the schema carries the full burden.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Close') and resource ('one tab index') with a clear condition ('if more than one tab is open'). It distinguishes itself from sibling tools like browser.activate_tab and browser.list_tabs, though it doesn't explicitly name them. The purpose is clear and actionable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use it (when a tab needs closing and more than one is open) but doesn't explicitly state when not to use it or mention alternatives like browser.close_session. The condition 'if more than one tab is open' provides some usage context, but there's no explicit routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser.create_sessionA
Create a new browser session and optionally navigate to a start URL.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | Label shown in session lists. | |
| start_url | No | Page to open once the session starts. | |
| totp_hosts | No | Hosts where one-time codes from totp_secret may be typed ("*.example.com" also matches subdomains). Defaults to the start_url host. | |
| user_agent | No | User-Agent header to send instead of the default. | |
| totp_secret | No | Base32 TOTP secret for this session's one-time codes, typed only on totp_hosts. | |
| auth_profile | No | Start signed in from a saved auth profile (browser.list_auth_profiles). | |
| proxy_server | No | Proxy URL such as http://host:port. Omitted: the deployment default, if any. | |
| proxy_persona | No | Named proxy configuration to route through. Not with proxy_server. | |
| memory_profile | No | Load a named memory profile into this session. | |
| proxy_password | No | Password for proxy_server. | |
| proxy_username | No | Username for proxy_server. | |
| protection_mode | No | Witness evidence mode; 'confidential' is stricter. Omitted: the deployment default. | |
| storage_state_path | No | Playwright storage-state file, relative to the auth directory. Not with auth_profile. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already indicate this is a non-read-only, non-idempotent, non-destructive operation, so the description doesn't need to restate that. The description adds the 'optionally navigate' behavior, which is useful. However, it doesn't disclose side effects like session limits, resource cleanup, or whether creating a session with the same name replaces or errors, which would add value beyond the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence that states the core action and the optional behavior. Every word earns its place; there is no fluff or repetition of schema details.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 13-parameter creation tool with no output schema, the description is minimal but the schema is rich and fully covers parameters. The main missing context is what the tool returns (e.g., session ID) and how it relates to fork_session, which an agent would need to know to use it correctly. The annotations cover safety, so the description doesn't need to repeat that.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all 13 parameters. The description adds no parameter-level meaning beyond the schema. Baseline 3 is appropriate because the schema carries the full burden and the description doesn't need to compensate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a clear verb and resource: 'Create a new browser session and optionally navigate to a start URL.' This distinguishes it from session-listing, closing, and observation siblings. However, it doesn't explicitly contrast with fork_session, which also creates a session, so it falls short of full sibling differentiation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies the primary use case (starting a new session) and the optional start URL, but it does not state when to prefer this over fork_session or when not to use it. The schema hints at exclusions (e.g., proxy_persona not with proxy_server, storage_state_path not with auth_profile), but the description itself provides no usage guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser.eval_jsA
Execute a JavaScript expression in the current page context and return the result. Every call needs operator approval for that exact expression: the first call returns status approval_required with an approval_id; retry with approval_id once it is approved. Prefer find_elements, get_html or observe for reading the page — they need no approval.
| Name | Required | Description | Default |
|---|---|---|---|
| expression | Yes | JavaScript expression or function source, evaluated in the page; returns its JSON value. | |
| session_id | No | Target session id, from browser.create_session or browser.list_sessions. Optional when exactly one session is live; observe and act tools create one when none is. | |
| approval_id | No | ID of the approval granted for exactly this expression. The first call returns status approval_required with an approval_id; call again with it once an operator approves. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond annotations, the description discloses the critical approval behavior: every call requires operator approval for that exact expression, the first call returns approval_required with an approval_id, and the call must be retried with that ID. This is essential operational context that is not present in the annotations or schema and directly affects how an agent invokes the tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three focused sentences: purpose, approval protocol, and alternative tools. It front-loads the core behavior and approval constraint, with no filler or redundant content.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation-capable eval tool with no output schema, the description plus schema covers all necessary invocation details: the expression, session handling, approval flow, result expectation, and less-privileged alternatives. An agent has enough information to call it correctly and recover from the approval_required status.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, and each parameter already has a meaningful schema description, particularly expression and approval_id. The tool description reinforces the approval flow but adds no parameter semantics beyond what the schema provides, so the baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific operation: execute a JavaScript expression in the current page context and return the result. It also distinguishes itself from sibling read tools by explicitly naming find_elements, get_html, and observe as preferred alternatives. This leaves no ambiguity about what the tool does or how it differs.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit when-not guidance: prefer find_elements, get_html, or observe for reading because they need no approval. It also explains the approval retry workflow, so an agent knows the exact conditions under which eval_js should be used and how to proceed after an approval_required response.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser.execute_actionA
Execute one browser action (navigate, click, hover, type, press, select_option, scroll, …) in a session, using the same action schema the agent planner emits. Actions are policy-checked and audited, and governed actions may require a granted approval_id. Call browser.observe first to get targetable element IDs and selectors. Returns the action's verification (what changed) and an observation of the page after it (up to 20 interactables).
| Name | Required | Description | Default |
|---|---|---|---|
| action | Yes | The browser action to execute, in the shared action schema: an action type (navigate, click, type, scroll, wait, done, …) plus its arguments, e.g. a URL for navigate or an element selector/index for click and type. | |
| detail | No | 'compact' (default) leaves out what choosing the next step does not need: the full session record (see browser.get_session), remote-access diagnostics and, for actions, the pre-action snapshot. 'full' returns the complete payload. | compact |
| session_id | No | Target session id, from browser.create_session or browser.list_sessions. Optional when exactly one session is live; observe and act tools create one when none is. | |
| approval_id | No | ID of a granted approval that authorizes this action when the active policy requires one. Omit when the action does not need approval. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations provide basic flags (readOnlyHint=false, openWorldHint=true, idempotentHint=false, destructiveHint=false) but are not highly informative. The description adds valuable behavioral context: actions are policy-checked and audited, governed actions may require approval, and the return value includes verification and a post-action observation. This goes beyond the annotations and clarifies side effects and authorization needs.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, well-structured paragraph with three sentences. It front-loads the main purpose, then covers policy/approval, the prerequisite, and the return format. No filler or repetition, though it could be slightly more compact by removing the ellipsis in the action list, which is already clear from the schema enum.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the complexity (nested action schema, many parameters, no output schema), the description is sufficiently complete. It explains the workflow (observe first), the approval mechanism, and the return value. It does not enumerate all action types, but the schema covers that. It also doesn't mention session creation, but that is handled in the session_id parameter description.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so parameters are well-documented in the schema. The description adds semantic value by explaining the action schema is the same one the planner emits and by advising to call browser.observe for element IDs and selectors, which directly aids correct use of element_id and selector. This exceeds the baseline for full schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Clearly states it executes one browser action, enumerates action types, and distinguishes itself from observation tools by specifying the action-execution role. Mentions the shared action schema, making its purpose unambiguous and distinct from siblings like browser.observe or browser.wait_for_selector.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly instructs to call browser.observe first to obtain element IDs and selectors, which is a critical prerequisite. Also mentions the approval_id for governed actions, giving context on when to include it. Does not explicitly state when to avoid this tool in favor of alternatives, but the observe-first hint is strong guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser.find_elementsA
Find elements either by CSS selector or by a text/regex query across the page's text content. selector mode returns matching elements' text, href, value, bounding box, and visibility -- useful before clicking or scraping multiple items. query mode (set query, optionally regex and context) returns the matched text plus surrounding context for each hit -- a cheap way to locate a specific string on the page without a full observe.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Most matches to return. | |
| query | No | Text or regex to search for across the page's text content, as an alternative to selector. Matching is case-insensitive in both modes. Returns matched elements plus surrounding text. | |
| regex | No | Treat query as a regular expression (JavaScript syntax, 'gi' flags) instead of a plain-text substring match. | |
| context | No | Characters of surrounding text to include around each match when query is used. | |
| selector | No | CSS or Playwright selector. Give this or query. | |
| session_id | No | Target session id, from browser.create_session or browser.list_sessions. Optional when exactly one session is live; observe and act tools create one when none is. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotations say little beyond abstract hints, so the description carries much of the behavioral burden. It discloses what each mode returns: element text/href/value/bounding box/visibility for selector mode, and matched text with surrounding context for query mode. It does not discuss side effects or navigation behavior, but for a find-style tool the output disclosure is the most important behavioral trait.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact—three sentences—with the core purpose in the first clause. Each subsequent clause earns its place by describing a specific return mode, use case, or sibling alternative; there is no repetitive filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description compensates by specifying return fields for both modes and their practical use cases. The input schema already documents all six parameters, including constraints and defaults, so nothing needed for correct invocation is left unaddressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3. The description adds value by organizing the otherwise flat parameter list into the two modes and tying 'context' and 'regex' to query mode and 'selector' to the element-attribute return path. This helps the agent assemble valid parameter combinations beyond what individual schema descriptions already provide.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb and resource ('Find elements') and immediately distinguishes the two operating modes: CSS selector and text/regex query. It also identifies distinct outputs for each mode, and the note about not needing a 'full observe' separates it from the browser.observe sibling.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives concrete usage context for both modes: selector mode is 'useful before clicking or scraping multiple items,' and query mode is positioned as a 'cheap way to locate a specific string on the page without a full observe.' This explicitly names the observe sibling as an alternative and the condition for choosing the query path.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser.fork_sessionA
Fork a session: snapshot its cookies, storage state, and current URL, then create a new independent session with that state. Useful for branching workflows or running parallel variants.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | Omitted: fork-of-<source name>. | |
| start_url | No | Omitted: the source session's current URL. | |
| session_id | No | Target session id, from browser.create_session or browser.list_sessions. Optional when exactly one session is live; observe and act tools create one when none is. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations include openWorldHint=true and destructiveHint=false, indicating it creates a new state without destroying existing ones. The description adds that it snapshots cookies/storage/URL and creates an 'independent' session, clarifying that the original is unaffected. This is useful context, but it doesn't mention error scenarios or edge cases (e.g., if session_id is invalid), so it doesn't go beyond the annotations' baseline significantly.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences, front-loading the core action and how it works, then adding the use case. Every sentence adds value; no fluff or redundancy. Ideal length for a tool of this complexity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has no output schema and moderate complexity (3 params, all optional), the description covers the core function and expected use. It could mention potential errors or post-conditions (e.g., new session's status), but the description is adequate for an agent to decide to use it and understand the result. The high schema coverage compensates for missing param details.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema provides 100% description coverage for all three parameters (name, start_url, session_id), each with clear defaults. However, the description adds no additional parameter-level detail, but since the schema is thorough, a baseline of 3 would be standard; the description's overall context about forking helps interpret the parameters, justifying a 4.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Fork a session'), the resource ('session'), and the mechanism ('snapshot its cookies, storage state, and current URL, then create a new independent session'). It also distinguishes from siblings by specifying it is for branching workflows, which aligns with the purpose of create/edit/delete operations, unlike list/get/observe tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context for when to use the tool ('branching workflows or running parallel variants') and implies it builds on create_session/list_sessions. It doesn't explicitly exclude alternatives like create_session, but the use case is well-defined. The 'openWorldHint' annotation suggests side effects are expected, which the description aligns with.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser.get_htmlARead-onlyIdempotent
Read the current page: its serialized DOM (the whole document, not just the viewport), or with text_only=true its visible text — the cheapest way to read a page's content. Returns up to max_chars (default 20,000) from offset; when truncated is true, call again with offset=next_offset for the rest.
| Name | Required | Description | Default |
|---|---|---|---|
| offset | No | Character to start from. A result with truncated=true gives the next_offset to continue from. | |
| max_chars | No | Most characters to return. | |
| text_only | No | Return the page's visible text instead of HTML: line breaks kept, table cells tab-separated. | |
| session_id | No | Target session id, from browser.create_session or browser.list_sessions. Optional when exactly one session is live; observe and act tools create one when none is. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description adds meaningful behavior beyond the annotations: it returns serialized DOM (not just viewport), text_only strips to visible text, output is subject to max_chars/offset, and truncation is handled by calling again with offset=next_offset. This gives the agent a clear mental model of pagination and output modes that annotations do not convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with no wasted words. The core purpose is front-loaded, the text_only variant is introduced immediately, and the pagination contract is stated compactly in one sentence. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the burden of explaining return behavior, and it does: max_chars, offset, truncated, and next_offset. The text_only behavior is also described well enough for an agent to know exactly what it will receive. Session_id requirements are already covered by the schema, so nothing essential is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the schema already describes all four parameters clearly. The description adds value by explaining the relationship between offset, max_chars, truncated, and next_offset, and by detailing what text_only does (visible text with line breaks and tab-separated cells). This goes beyond the baseline schema documentation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb ('Read') and resource ('current page'), and immediately clarifies scope: serialized DOM of the whole document versus visible text with text_only=true. It also differentiates itself from viewport-only reads, which helps distinguish it from browser.screenshot and browser.observe.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives clear context for when to use the tool ('the cheapest way to read a page's content') and explains the trade-off between DOM and text-only modes. However, it does not explicitly mention when to prefer siblings like browser.observe, browser.eval_js, or browser.find_elements, so exclusion guidance is left implicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser.get_sessionARead-onlyIdempotent
Get the full record for one browser session by ID: the live session summary (status, current page, tabs) when the session is active, or the persisted session record when it has been closed. Use browser.list_sessions to discover session IDs.
| Name | Required | Description | Default |
|---|---|---|---|
| session_id | No | Target session id, from browser.create_session or browser.list_sessions. Optional when exactly one session is live; observe and act tools create one when none is. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is covered. The description adds valuable behavioral context: it explains that the tool returns a live summary when the session is active versus a persisted record when closed. This goes beyond the annotations and helps the agent understand what to expect.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with no wasted words. The core behavior is front-loaded, the live-vs-closed distinction is explained, and the pointer to the sibling tool is placed at the end. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter read tool with rich annotations, the description is nearly complete. It explains the return behavior for both active and closed sessions and points to the discovery mechanism. It doesn't describe the exact structure of the returned record, but with no output schema and a simple read operation, this is a minor gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already fully documents the session_id parameter, including its optionality and how to obtain it. The description reinforces that the ID comes from browser.create_session or browser.list_sessions, but this is already in the schema. Baseline 3 is appropriate since the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: retrieving the full record for one browser session by ID. It distinguishes between live and closed session records, which adds precision. It also names the sibling tool (browser.list_sessions) for discovering session IDs, helping differentiate it from related tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly says to use browser.list_sessions to discover session IDs, providing clear context for when this tool is appropriate. It doesn't explicitly state when not to use it or mention alternatives like browser.observe, but the guidance is sufficient for an agent to select it correctly.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser.list_auth_profilesARead-onlyIdempotent
List reusable saved auth profiles that can be loaded into a new session.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already provide readOnlyHint, destructiveHint, idempotentHint. Description adds no behavioral details beyond stating profiles are reusable, which is implied by 'saved'.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Single 11-word sentence, no fluff, purpose front-loaded. Every word earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Low complexity tool with no params, annotated safety, and clear purpose. Description is adequate but doesn't mention return format or empty list scenario.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
No parameters exist; schema coverage is 100%. Description adds no parameter info but doesn't need to. Baseline 4 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description uses specific verb 'List' and identifies resource 'reusable saved auth profiles', distinguishing from related tools like get_auth_profile (single) and save_auth_profile (create).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Implicitly indicates usage context: listing profiles for loading into a session. No explicit when-not-to-use or alternatives, but clear enough for an agent.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser.list_downloadsARead-onlyIdempotent
List files captured from browser downloads for one session.
| Name | Required | Description | Default |
|---|---|---|---|
| session_id | No | Target session id, from browser.create_session or browser.list_sessions. Optional when exactly one session is live; observe and act tools create one when none is. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is covered. The description adds the 'captured from browser downloads' context and session scoping, but doesn't disclose details like whether the list includes file paths, sizes, or statuses, or whether it only lists files captured so far.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, clear sentence that front-loads the action and resource. No wasted words; every part earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read-only list tool with one optional parameter and full schema coverage, the description is mostly adequate. However, with no output schema, it doesn't hint at what the returned list contains (e.g., file names, paths, metadata), which an agent might need to know to decide whether to call read_download next.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents the session_id parameter well, including its optionality and how to obtain it. The description adds no additional parameter semantics beyond what the schema provides, so baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('List') and resource ('files captured from browser downloads') scoped to 'one session'. It clearly distinguishes from siblings like browser.list_sessions and browser.read_download, though it doesn't explicitly name them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for listing downloads within a session, and the schema clarifies the session_id is optional when exactly one session is live. However, it doesn't explicitly state when to prefer this over browser.read_download or how it relates to browser.list_sessions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser.list_sessionsARead-onlyIdempotent
List live and persisted browser sessions, one reference each (id, name, status, live, current page, takeover URL). browser.get_session returns a full record.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover readOnlyHint, idempotentHint, and destructiveHint, so the description doesn't need to restate safety. It adds behavioral context by specifying that the tool returns 'one reference each' with a defined set of fields (id, name, status, live, current page, takeover URL) and that both live and persisted sessions are included. This goes beyond the bare read-only annotation, though it stops short of discussing edge cases like empty lists, ordering, or pagination.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences with zero filler. It front-loads the action and resource, packs the essential field list into a concise parenthetical, and then immediately points to the main alternative. Every word contributes value.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a no-parameter, no-output-schema list tool, the description covers the core needs: what it lists (live and persisted sessions), what each entry contains, and where to get a fuller record. It is complete enough for an agent to execute the tool correctly. The only minor gaps are unspecified edge-case behaviors (e.g., empty result set, session ordering) and the exact meaning of 'current page' and 'takeover URL', but these are not critical for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema is empty (0 parameters), so the baseline for this dimension is 4. The description adds no parameter-specific meaning because there are no parameters to explain, and the empty schema already communicates that no arguments are required. The description instead focuses on the output, which is appropriate for a no-argument tool.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('List') and resource ('live and persisted browser sessions'), names the exact reference fields, and contrasts itself with browser.get_session ('returns a full record'). This clearly distinguishes it from sibling tools like list_tabs or list_downloads without needing to inspect schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly names the key alternative (browser.get_session) and indicates the difference in return granularity: this tool returns one summary reference per session, while get_session provides a full record. An agent can infer when to use this list tool vs. the detail tool. The context of other sibling tools (create, close, fork) is also implicitly separated by the focus on listing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser.list_tabsARead-onlyIdempotent
List currently open tabs/pages for one session.
| Name | Required | Description | Default |
|---|---|---|---|
| session_id | No | Target session id, from browser.create_session or browser.list_sessions. Optional when exactly one session is live; observe and act tools create one when none is. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is fully covered. The description adds the behavioral traits 'currently open' (indicating a snapshot) and 'for one session' (scoping the operation), which go slightly beyond annotations but do not disclose details such as return format, pagination, or how tabs are represented. Given the annotations, this is adequate but not rich.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single nine-word sentence that is front-loaded with the action and resource, contains zero filler, and communicates the core scoping constraint. There is nothing extraneous; it is as concise as a description for this simple tool can be.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a low-complexity, read-only tool with one well-documented optional parameter, the description plus annotations and schema provide everything an agent needs to invoke it correctly. The absence of an output schema is acceptable because the tool's purpose strongly implies a list result, and the schema explains how to get a session_id. No critical gaps remain.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the session_id parameter is described in detail (source, optionality, fallback to observe/act tools). The tool description itself does not add parameter semantics beyond the phrase 'for one session', which merely restates the parameter's role. Since the schema carries the weight, the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('List'), a clear resource ('currently open tabs/pages'), and a scoping constraint ('for one session'). This distinguishes it from siblings like list_sessions (which lists sessions, not tabs) and activate_tab/close_tab (which mutate tabs rather than list them). An agent can tell exactly what this tool does without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description does not explicitly state when to use this tool versus alternatives, nor does it name a sibling like list_sessions as a prerequisite. The parameter description hints that session ids come from browser.create_session or browser.list_sessions and that observe/act tools create one when none exists, which implies the session-management workflow, but this is not framed as a when-to-use rule for list_tabs itself. The usage context is implied rather than explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser.observeARead-only
Capture the current browser observation: interactables, tabs, console, and a perception summary. Presets: 'text' — no screenshot or OCR: interactables, accessibility tree and the first 2,000 characters of page text; for text-only models. 'fast' — screenshot only, returned as an image, no text/accessibility extraction; for vision models. 'normal' (default) — text plus a screenshot URL and OCR. 'rich' — normal with twice the interactables and 4,000 characters of text. To read a whole page, use browser.get_html with text_only=true.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Most interactable elements to return. | |
| detail | No | 'compact' (default) leaves out what choosing the next step does not need: the full session record (see browser.get_session), remote-access diagnostics and, for actions, the pre-action snapshot. 'full' returns the complete payload. | compact |
| preset | No | text: page text and interactables, no screenshot. fast: screenshot and title only. normal: both. rich: more text and twice the interactables. Omitted: the deployment default. | |
| session_id | No | Target session id, from browser.create_session or browser.list_sessions. Optional when exactly one session is live; observe and act tools create one when none is. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description goes well beyond the readOnlyHint/openWorldHint annotations by disclosing what each preset returns, including OCR, accessibility tree, screenshot URL vs. image-only output, and character limits. It also explains the relationship between text and rich modes, giving the agent accurate expectations about output content without contradicting the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description front-loads the core purpose and then uses a dense but scannable preset breakdown; every sentence carries decision-relevant information. The alternative routing is placed at the end without redundant restatement of schema fields.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite no output schema, the description specifies the return composition for each preset, including whether output is an image, a URL, OCR, or text, and it covers the main decisions an agent needs to make. Given the 100% parameter schema coverage, nothing essential to selecting or invoking the tool is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3; the description adds real value by explaining the semantics of the preset enum and the trade-offs among text/fast/normal/rich, plus the recommendation to use get_html for full-page reading. Limit and detail are already self-documenting in the schema, so no further compensation is needed.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Opens with a specific verb and resource ('Capture the current browser observation') and enumerates its contents (interactables, tabs, console, perception summary), so an agent can tell it apart from sibling read tools. The closing sentence explicitly distinguishes it from browser.get_html for full-page reading.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives concrete selection cues by mapping presets to model types ('text' for text-only models, 'fast' for vision models) and names browser.get_html as the alternative for reading a whole page. It does not explicitly state when observe should be avoided in favor of browser.screenshot or browser.find_elements, so it falls just short of full when/when-not coverage.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser.read_downloadA
Read a downloaded file as text (CSV, JSON, TXT, HTML, ...), by download_id from browser.list_downloads or, if omitted, the latest completed download. Paged like browser.get_html. Binary files (PDF, XLSX, images) are refused with their artifact URL.
| Name | Required | Description | Default |
|---|---|---|---|
| offset | No | Character to start from. A result with truncated=true gives the next_offset to continue from. | |
| max_chars | No | Most characters to return. | |
| session_id | No | Target session id, from browser.create_session or browser.list_sessions. Optional when exactly one session is live; observe and act tools create one when none is. | |
| download_id | No | id from browser.list_downloads. Omit to read the most recent completed download. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations are sparse (readOnlyHint=false, destructiveHint=false) so the description carries the burden. It discloses that binary files are refused and returns an artifact URL, and that pagination behaves like get_html. It does not detail error handling or side effects, but given the low annotation coverage, the description adds meaningful behavioral context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two concise sentences with the purpose front-loaded and no redundant wording. The first sentence covers purpose and selection, the second covers pagination and binary handling. Every word earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers the main use case, pagination, and binary refusal, and references get_html for behavioral similarity. Since there is no output schema, the return format is implied ('as text') but not explicitly structured. The mention of pagination and refusal covers key edge cases, making it reasonably complete for an agent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and parameter descriptions are already detailed. The description repeats the download_id semantics (omit for latest) and adds the file-type context, but does not significantly enhance parameter meaning beyond the schema. The reference to get_html for pagination is helpful but not parameter-specific.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool reads a downloaded file as text, specifies supported file types (CSV, JSON, TXT, HTML), and identifies the download via download_id or latest completed. It distinguishes itself from siblings like list_downloads (which lists) and get_html (which gets page HTML) by focusing on reading download content.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explains how to select a download (by id or omit for latest) and references list_downloads as the source of download_id. It also mentions pagination like browser.get_html, giving usage context. It does not explicitly exclude alternatives beyond noting binary files are refused, but the usage guidance is clear and actionable.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser.request_human_takeoverB
Ask for a human to take over the shared browser desktop.
| Name | Required | Description | Default |
|---|---|---|---|
| reason | No | What the human should do, shown to them. | Manual review requested |
| session_id | No | Target session id, from browser.create_session or browser.list_sessions. Optional when exactly one session is live; observe and act tools create one when none is. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already indicate non-read-only, non-idempotent, open-world behavior, but the description adds no detail about what happens after the request: whether it blocks, pauses automation, notifies a human, or returns a confirmation. The description contributes minimal behavioral context beyond the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence with no filler or repetition. It is concise and easy to parse, though it is so brief that it omits important usage and behavioral context.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool that requests human intervention, an agent would benefit from knowing whether the call is blocking, how the request is surfaced, and under what circumstances it should be used. With no output schema and no usage guidance, the description is not fully complete for an agent selecting or invoking the tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, and the schema richly documents both parameters, including the optional session_id fallback rule. The tool description adds no additional parameter meaning, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a clear, specific action: 'Ask for a human to take over the shared browser desktop.' This uniquely identifies the tool's purpose and distinguishes it from the sibling tools, none of which cover human takeover.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to request a human takeover or when to prefer alternatives like browser.observe, browser.execute_action, or other automation tools. The intent is implied by the name and description, but no explicit when/when-not conditions are provided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser.save_auth_profileA
Save the current session storage state into a reusable named auth profile.
| Name | Required | Description | Default |
|---|---|---|---|
| session_id | No | Target session id, from browser.create_session or browser.list_sessions. Optional when exactly one session is live; observe and act tools create one when none is. | |
| profile_name | Yes | Name to save the signed-in state under; pass it to browser.create_session as auth_profile. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description states what is captured (current session storage state) and where it is stored (a reusable named auth profile), which is the core behavioral contract. However, it does not disclose overwrite behavior, whether an authenticated session is required, or what the tool returns; the annotations are all false flags and do not fill that gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One sentence with no filler; it front-loads the action and destination. Every word contributes to the meaning.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter save action, the required inputs are clear and the schema covers them well. However, there is no output schema and the description does not mention what happens on success, what the call returns, or what occurs when an existing profile name is reused, leaving moderate gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, and the schema already documents both parameters clearly, including the optionality of session_id and the relationship of profile_name to browser.create_session. The description adds no parameter-level detail beyond the schema, so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific action and resource: save the current session storage state into a named auth profile. This clearly differentiates it from siblings like list_auth_profiles, create_session, and close_session.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies the use case by calling the profile 'reusable,' and the schema explains that the resulting profile is passed to browser.create_session, but the description itself gives no explicit when-to-use guidance or exclusions. An agent can infer the intended workflow but is not told directly when to prefer this over alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser.screenshotARead-only
Capture the current viewport and return it as an image (plus its artifact URL), without the full observe payload.
| Name | Required | Description | Default |
|---|---|---|---|
| label | No | Word added to the screenshot's file name. | manual |
| session_id | No | Target session id, from browser.create_session or browser.list_sessions. Optional when exactly one session is live; observe and act tools create one when none is. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and destructiveHint=false, so the safety profile is covered. The description adds useful behavioral context: it captures only the current viewport and returns an image plus artifact URL rather than the full observe payload. It does not discuss auth or rate limits, but that is acceptable for a simple read-only capture tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence with no filler. It front-loads the action and return value, then adds the key differentiator from observe. Every part earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read-only tool with two optional parameters fully documented in the schema, the description is complete. It states what is captured, what is returned, and how it differs from the main sibling. The absence of an output schema is mitigated by explicitly naming the return type.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, with both label and session_id already documented in the input schema. The description adds no parameter-specific meaning beyond the schema, so the baseline score of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Capture') and resource ('the current viewport'), and states the return value ('an image plus its artifact URL'). It also distinguishes itself from browser.observe by noting it lacks the full observe payload, so an agent can tell it apart from siblings without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'without the full observe payload' clearly positions this as the lightweight image-only alternative to browser.observe. It does not explicitly say 'use this when you only need a screenshot' or list exclusions, but the contrast with observe gives sufficient contextual guidance for selection.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser.wait_for_selectorARead-only
Wait for a CSS selector to reach a specific state (visible, hidden, attached, detached). Returns when the condition is met or raises on timeout.
| Name | Required | Description | Default |
|---|---|---|---|
| state | No | Condition to wait for, as in Playwright's wait_for_selector. | visible |
| selector | Yes | CSS or Playwright selector. | |
| session_id | No | Target session id, from browser.create_session or browser.list_sessions. Optional when exactly one session is live; observe and act tools create one when none is. | |
| timeout_ms | No | How long to wait before failing. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description explicitly states the completion condition (selector reaches the requested state) and the failure behavior (raises on timeout), which is useful beyond the annotations. Annotations already cover readOnlyHint=true and destructiveHint=false, so the safety profile is clear. It does not mention possible session creation, but that detail is available in the session_id parameter description.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence that front-loads the verb and resource, lists the relevant states, and states the timeout behavior. There is no filler or repetition of schema fields, making it highly efficient.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple synchronization tool, the description covers the core contract: what condition to wait for and what happens on timeout. The schema fully documents parameters and defaults, and annotations cover the read-only/destructive profile. The main residual gap is that the return value is not described, but the absence of an output schema makes this less critical.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline applies and the description does not need to compensate. The description adds no substantive parameter meaning beyond saying 'CSS selector', which is actually narrower than the schema's 'CSS or Playwright selector'. State, timeout, and session_id are all already well documented in the input schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action (wait), the resource (selector), and the specific completion states (visible, hidden, attached, detached). It is easy to distinguish from siblings like list_sessions or screenshot, though it does not explicitly contrast with find_elements or observe. The description calls it a CSS selector while the schema says 'CSS or Playwright selector', a minor narrowing that keeps it from being fully precise.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this tool versus siblings like find_elements or observe, and no exclusions or alternatives are mentioned. The verb 'Wait' implies a synchronization use case, but the description leaves the decision entirely to the agent's inference. Session-related prerequisites are documented in the schema, not in the description.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
21 tool updates
v1.8.1- Added
browser.activate_tab - Added
browser.close_session - Added
browser.close_tab - Added
browser.create_session - Added
browser.eval_js - Added
browser.execute_action - Changed
browser.find_elements17 fields changed- added
Input schema / properties / contextAdded value: +{ + "default": 0, + "description": "Characters of surrounding text to include around each match when query is used.", + "maximum": 500, + "minimum": 0, + "title": "Context", + "type": "integer" +} - added
Input schema / properties / limit / descriptionAdded value: +"Most matches to return." - added
Input schema / properties / queryAdded value: +{ + "anyOf": [ + { + "maxLength": 500, + "minLength": 1, + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "description": "Text or regex to search for across the page's text content, as an alternative to selector. Matching is case-insensitive in both modes. Returns matched elements plus surrounding text.", + "title": "Query" +} - added
Input schema / properties / regexAdded value: +{ + "default": false, + "description": "Treat query as a regular expression (JavaScript syntax, 'gi' flags) instead of a plain-text substring match.", + "title": "Regex", + "type": "boolean" +} - added
Input schema / properties / selector / anyOfAdded value: +[ + { + "maxLength": 2000, + "minLength": 1, + "type": "string" + }, + { + "type": "null" + } +] - added
Input schema / properties / selector / defaultAdded value: +null - added
Input schema / properties / selector / descriptionAdded value: +"CSS or Playwright selector. Give this or query." - removed
Input schema / properties / selector / maxLengthRemoved value: -2000 - removed
Input schema / properties / selector / minLengthRemoved value: -1 - removed
Input schema / properties / selector / typeRemoved value: -"string" - added
Input schema / properties / session_id / anyOfAdded value: +[ + { + "maxLength": 120, + "minLength": 1, + "type": "string" + }, + { + "type": "null" + } +] - added
Input schema / properties / session_id / defaultAdded value: +null - changed
Input schema / properties / session_id / descriptionPrevious value: -"ID of the target browser session, as returned by browser.create_session or listed by browser.list_sessions."New value: +"Target session id, from browser.create_session or browser.list_sessions. Optional when exactly one session is live; observe and act tools create one when none is." - removed
Input schema / properties / session_id / maxLengthRemoved value: -120 - removed
Input schema / properties / session_id / minLengthRemoved value: -1 - removed
Input schema / properties / session_id / typeRemoved value: -"string" - removed
Input schema / requiredRemoved value: -[ - "session_id", - "selector" -]
- Added
browser.fork_session - Added
browser.get_html - Added
browser.get_session - Added
browser.list_auth_profiles - Added
browser.list_downloads - Removed
browser.list_memory_profiles - Added
browser.list_sessions - Added
browser.list_tabs - Added
browser.observe - Added
browser.read_download - Added
browser.request_human_takeover - Added
browser.save_auth_profile - Added
browser.screenshot - Added
browser.wait_for_selector
13 tool updates
v1.4.2- Removed
browser.activate_tab - Removed
browser.close_session - Removed
browser.close_tab - Removed
browser.eval_js - Removed
browser.execute_action - Added
browser.find_elements - Removed
browser.get_auth_profile - Removed
browser.get_console - Removed
browser.get_session - Removed
browser.list_sessions - Removed
browser.observe - Removed
browser.save_auth_profile - Removed
harness.get_status
33 tool updates
v1.4.2- Changed
browser.activate_tab2 fields changed- added
Input schema / properties / index / descriptionAdded value: +"Zero-based index of the target tab, as reported by browser.list_tabs." - added
Input schema / properties / session_id / descriptionAdded value: +"ID of the target browser session, as returned by browser.create_session or listed by browser.list_sessions."
- Changed
browser.close_session1 field changed- added
Input schema / properties / session_id / descriptionAdded value: +"ID of the target browser session, as returned by browser.create_session or listed by browser.list_sessions."
- Changed
browser.close_tab2 fields changed- added
Input schema / properties / index / descriptionAdded value: +"Zero-based index of the target tab, as reported by browser.list_tabs." - added
Input schema / properties / session_id / descriptionAdded value: +"ID of the target browser session, as returned by browser.create_session or listed by browser.list_sessions."
- Removed
browser.create_session - Removed
browser.drag_drop - Changed
browser.eval_js1 field changed- added
Input schema / properties / session_id / descriptionAdded value: +"ID of the target browser session, as returned by browser.create_session or listed by browser.list_sessions."
- Changed
browser.execute_action3 fields changed- added
Input schema / properties / action / descriptionAdded value: +"The browser action to execute, in the shared action schema: an action type (navigate, click, type, scroll, wait, done, …) plus its arguments, e.g. a URL for navigate or an element selector/index for click and type." - added
Input schema / properties / approval_id / descriptionAdded value: +"ID of a granted approval that authorizes this action when the active policy requires one. Omit when the action does not need approval." - added
Input schema / properties / session_id / descriptionAdded value: +"ID of the target browser session, as returned by browser.create_session or listed by browser.list_sessions."
- Removed
browser.find_elements - Removed
browser.fork_session - Changed
browser.get_auth_profile1 field changed- added
Input schema / properties / profile_name / descriptionAdded value: +"Name of the saved auth profile to inspect, as listed by browser.list_auth_profiles."
- Changed
browser.get_console1 field changed- added
Input schema / properties / session_id / descriptionAdded value: +"ID of the target browser session, as returned by browser.create_session or listed by browser.list_sessions."
- Removed
browser.get_html - Removed
browser.get_memory_profile - Removed
browser.get_network_log - Removed
browser.get_page_errors - Removed
browser.get_request_failures - Changed
browser.get_session1 field changed- added
Input schema / properties / session_id / descriptionAdded value: +"ID of the target browser session, as returned by browser.create_session or listed by browser.list_sessions."
- Removed
browser.list_auth_profiles - Removed
browser.list_downloads - Removed
browser.list_tabs - Changed
browser.observe1 field changed- added
Input schema / properties / session_id / descriptionAdded value: +"ID of the target browser session, as returned by browser.create_session or listed by browser.list_sessions."
- Removed
browser.readiness_check - Removed
browser.request_human_takeover - Changed
browser.save_auth_profile1 field changed- added
Input schema / properties / session_id / descriptionAdded value: +"ID of the target browser session, as returned by browser.create_session or listed by browser.list_sessions."
- Removed
browser.save_memory_profile - Removed
browser.screenshot - Removed
browser.set_viewport - Removed
browser.stop_trace - Removed
browser.verify_witness - Removed
browser.wait_for_selector - Changed
harness.get_status1 field changed- added
Input schema / properties / run_id / descriptionAdded value: +"ID of the convergence run, as returned by harness.start_convergence or listed by harness.list_runs."
- Removed
harness.get_trace - Removed
harness.list_runs
1 tool update
v1.3.1- Added
browser.verify_witness
3 tool updates
v1.1.1- Added
harness.get_status - Added
harness.get_trace - Added
harness.list_runs
5 tool updates
v1.0.5- Changed
browser.execute_action1 field changed- changed
Input schema / $defs / BrowserActionDecision / properties / action / enumPrevious value: -[ - "navigate", - "click", - "hover", - "select_option", - "type", - "press", - "scroll", - "wait", - "reload", - "go_back", - "go_forward", - "upload", - "social_login", - "social_post", - "social_comment", - "social_like", - "social_follow", - "social_unfollow", - "social_repost", - "social_dm", - "request_human_takeover", - "done" -]New value: +[ + "navigate", + "click", + "hover", + "select_option", + "type", + "press", + "scroll", + "wait", + "reload", + "go_back", + "go_forward", + "upload", + "request_human_takeover", + "done" +]
- Removed
social.extract_posts - Removed
social.extract_profile - Removed
social.login - Removed
social.search
36 tool updates
v1.0.3- Changed
browser.activate_tab3 fields changed- added
Input schema / additionalPropertiesAdded value: +false - added
Input schema / properties / session_id / maxLengthAdded value: +120 - added
Input schema / properties / session_id / minLengthAdded value: +1
- Changed
browser.close_session3 fields changed- added
Input schema / additionalPropertiesAdded value: +false - added
Input schema / properties / session_id / maxLengthAdded value: +120 - added
Input schema / properties / session_id / minLengthAdded value: +1
- Changed
browser.close_tab3 fields changed- added
Input schema / additionalPropertiesAdded value: +false - added
Input schema / properties / session_id / maxLengthAdded value: +120 - added
Input schema / properties / session_id / minLengthAdded value: +1
- Changed
browser.create_session13 fields changed- added
Input schema / additionalPropertiesAdded value: +false - changed
Input schema / properties / auth_profile / anyOfPrevious value: -[ - { - "type": "string" - }, - { - "type": "null" - } -]New value: +[ + { + "maxLength": 120, + "minLength": 1, + "type": "string" + }, + { + "type": "null" + } +] - added
Input schema / properties / memory_profileAdded value: +{ + "anyOf": [ + { + "maxLength": 120, + "minLength": 1, + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "description": "Load a named memory profile into this session.", + "title": "Memory Profile" +} - changed
Input schema / properties / name / anyOfPrevious value: -[ - { - "type": "string" - }, - { - "type": "null" - } -]New value: +[ + { + "maxLength": 200, + "minLength": 1, + "type": "string" + }, + { + "type": "null" + } +] - added
Input schema / properties / protection_modeAdded value: +{ + "anyOf": [ + { + "enum": [ + "normal", + "confidential" + ], + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "title": "Protection Mode" +} - changed
Input schema / properties / proxy_password / anyOfPrevious value: -[ - { - "type": "string" - }, - { - "type": "null" - } -]New value: +[ + { + "maxLength": 500, + "type": "string" + }, + { + "type": "null" + } +] - added
Input schema / properties / proxy_personaAdded value: +{ + "anyOf": [ + { + "maxLength": 200, + "minLength": 1, + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "title": "Proxy Persona" +} - changed
Input schema / properties / proxy_server / anyOfPrevious value: -[ - { - "type": "string" - }, - { - "type": "null" - } -]New value: +[ + { + "maxLength": 500, + "minLength": 1, + "type": "string" + }, + { + "type": "null" + } +] - changed
Input schema / properties / proxy_username / anyOfPrevious value: -[ - { - "type": "string" - }, - { - "type": "null" - } -]New value: +[ + { + "maxLength": 200, + "type": "string" + }, + { + "type": "null" + } +] - changed
Input schema / properties / start_url / anyOfPrevious value: -[ - { - "type": "string" - }, - { - "type": "null" - } -]New value: +[ + { + "maxLength": 2000, + "minLength": 1, + "type": "string" + }, + { + "type": "null" + } +] - changed
Input schema / properties / storage_state_path / anyOfPrevious value: -[ - { - "type": "string" - }, - { - "type": "null" - } -]New value: +[ + { + "maxLength": 500, + "minLength": 1, + "type": "string" + }, + { + "type": "null" + } +] - changed
Input schema / properties / totp_secret / anyOfPrevious value: -[ - { - "type": "string" - }, - { - "type": "null" - } -]New value: +[ + { + "maxLength": 500, + "type": "string" + }, + { + "type": "null" + } +] - changed
Input schema / properties / user_agent / anyOfPrevious value: -[ - { - "type": "string" - }, - { - "type": "null" - } -]New value: +[ + { + "maxLength": 2000, + "minLength": 1, + "type": "string" + }, + { + "type": "null" + } +]
- Changed
browser.drag_drop3 fields changed- added
Input schema / additionalPropertiesAdded value: +false - added
Input schema / properties / session_id / maxLengthAdded value: +120 - added
Input schema / properties / session_id / minLengthAdded value: +1
- Changed
browser.eval_js3 fields changed- added
Input schema / additionalPropertiesAdded value: +false - added
Input schema / properties / session_id / maxLengthAdded value: +120 - added
Input schema / properties / session_id / minLengthAdded value: +1
- Changed
browser.execute_action14 fields changed- changed
Input schema / $defs / BrowserActionDecision / properties / element_id / anyOfPrevious value: -[ - { - "type": "string" - }, - { - "type": "null" - } -]New value: +[ + { + "maxLength": 500, + "minLength": 1, + "type": "string" + }, + { + "type": "null" + } +] - changed
Input schema / $defs / BrowserActionDecision / properties / file_path / anyOfPrevious value: -[ - { - "type": "string" - }, - { - "type": "null" - } -]New value: +[ + { + "maxLength": 500, + "minLength": 1, + "type": "string" + }, + { + "type": "null" + } +] - changed
Input schema / $defs / BrowserActionDecision / properties / key / anyOfPrevious value: -[ - { - "type": "string" - }, - { - "type": "null" - } -]New value: +[ + { + "maxLength": 120, + "minLength": 1, + "type": "string" + }, + { + "type": "null" + } +] - changed
Input schema / $defs / BrowserActionDecision / properties / label / anyOfPrevious value: -[ - { - "type": "string" - }, - { - "type": "null" - } -]New value: +[ + { + "maxLength": 1000, + "type": "string" + }, + { + "type": "null" + } +] - changed
Input schema / $defs / BrowserActionDecision / properties / platform / anyOfPrevious value: -[ - { - "type": "string" - }, - { - "type": "null" - } -]New value: +[ + { + "maxLength": 120, + "minLength": 1, + "type": "string" + }, + { + "type": "null" + } +] - changed
Input schema / $defs / BrowserActionDecision / properties / recipient / anyOfPrevious value: -[ - { - "type": "string" - }, - { - "type": "null" - } -]New value: +[ + { + "maxLength": 200, + "minLength": 1, + "type": "string" + }, + { + "type": "null" + } +] - changed
Input schema / $defs / BrowserActionDecision / properties / selector / anyOfPrevious value: -[ - { - "type": "string" - }, - { - "type": "null" - } -]New value: +[ + { + "maxLength": 2000, + "minLength": 1, + "type": "string" + }, + { + "type": "null" + } +] - changed
Input schema / $defs / BrowserActionDecision / properties / text / anyOfPrevious value: -[ - { - "type": "string" - }, - { - "type": "null" - } -]New value: +[ + { + "maxLength": 5000, + "minLength": 1, + "type": "string" + }, + { + "type": "null" + } +] - changed
Input schema / $defs / BrowserActionDecision / properties / url / anyOfPrevious value: -[ - { - "type": "string" - }, - { - "type": "null" - } -]New value: +[ + { + "maxLength": 2000, + "minLength": 1, + "type": "string" + }, + { + "type": "null" + } +] - changed
Input schema / $defs / BrowserActionDecision / properties / username / anyOfPrevious value: -[ - { - "type": "string" - }, - { - "type": "null" - } -]New value: +[ + { + "maxLength": 500, + "minLength": 1, + "type": "string" + }, + { + "type": "null" + } +] - changed
Input schema / $defs / BrowserActionDecision / properties / value / anyOfPrevious value: -[ - { - "type": "string" - }, - { - "type": "null" - } -]New value: +[ + { + "maxLength": 1000, + "type": "string" + }, + { + "type": "null" + } +] - added
Input schema / additionalPropertiesAdded value: +false - added
Input schema / properties / session_id / maxLengthAdded value: +120 - added
Input schema / properties / session_id / minLengthAdded value: +1
- Removed
browser.find_by_vision - Changed
browser.find_elements3 fields changed- added
Input schema / additionalPropertiesAdded value: +false - added
Input schema / properties / session_id / maxLengthAdded value: +120 - added
Input schema / properties / session_id / minLengthAdded value: +1
- Changed
browser.fork_session3 fields changed- added
Input schema / additionalPropertiesAdded value: +false - added
Input schema / properties / session_id / maxLengthAdded value: +120 - added
Input schema / properties / session_id / minLengthAdded value: +1
- Changed
browser.get_auth_profile1 field changed- added
Input schema / additionalPropertiesAdded value: +false
- Changed
browser.get_console3 fields changed- added
Input schema / additionalPropertiesAdded value: +false - added
Input schema / properties / session_id / maxLengthAdded value: +120 - added
Input schema / properties / session_id / minLengthAdded value: +1
- Changed
browser.get_html3 fields changed- added
Input schema / additionalPropertiesAdded value: +false - added
Input schema / properties / session_id / maxLengthAdded value: +120 - added
Input schema / properties / session_id / minLengthAdded value: +1
- Added
browser.get_memory_profile - Changed
browser.get_network_log3 fields changed- added
Input schema / additionalPropertiesAdded value: +false - added
Input schema / properties / session_id / maxLengthAdded value: +120 - added
Input schema / properties / session_id / minLengthAdded value: +1
- Changed
browser.get_page_errors3 fields changed- added
Input schema / additionalPropertiesAdded value: +false - added
Input schema / properties / session_id / maxLengthAdded value: +120 - added
Input schema / properties / session_id / minLengthAdded value: +1
- Changed
browser.get_request_failures3 fields changed- added
Input schema / additionalPropertiesAdded value: +false - added
Input schema / properties / session_id / maxLengthAdded value: +120 - added
Input schema / properties / session_id / minLengthAdded value: +1
- Changed
browser.get_session3 fields changed- added
Input schema / additionalPropertiesAdded value: +false - added
Input schema / properties / session_id / maxLengthAdded value: +120 - added
Input schema / properties / session_id / minLengthAdded value: +1
- Changed
browser.list_auth_profiles1 field changed- added
Input schema / additionalPropertiesAdded value: +false
- Changed
browser.list_downloads3 fields changed- added
Input schema / additionalPropertiesAdded value: +false - added
Input schema / properties / session_id / maxLengthAdded value: +120 - added
Input schema / properties / session_id / minLengthAdded value: +1
- Added
browser.list_memory_profiles - Changed
browser.list_sessions1 field changed- added
Input schema / additionalPropertiesAdded value: +false
- Changed
browser.list_tabs3 fields changed- added
Input schema / additionalPropertiesAdded value: +false - added
Input schema / properties / session_id / maxLengthAdded value: +120 - added
Input schema / properties / session_id / minLengthAdded value: +1
- Changed
browser.observe5 fields changed- added
Input schema / additionalPropertiesAdded value: +false - changed
Input schema / properties / limit / maximumPrevious value: -100New value: +200 - added
Input schema / properties / presetAdded value: +{ + "default": "normal", + "enum": [ + "fast", + "normal", + "rich" + ], + "title": "Preset", + "type": "string" +} - added
Input schema / properties / session_id / maxLengthAdded value: +120 - added
Input schema / properties / session_id / minLengthAdded value: +1
- Added
browser.readiness_check - Changed
browser.request_human_takeover3 fields changed- added
Input schema / additionalPropertiesAdded value: +false - added
Input schema / properties / session_id / maxLengthAdded value: +120 - added
Input schema / properties / session_id / minLengthAdded value: +1
- Changed
browser.save_auth_profile3 fields changed- added
Input schema / additionalPropertiesAdded value: +false - added
Input schema / properties / session_id / maxLengthAdded value: +120 - added
Input schema / properties / session_id / minLengthAdded value: +1
- Added
browser.save_memory_profile - Changed
browser.screenshot3 fields changed- added
Input schema / additionalPropertiesAdded value: +false - added
Input schema / properties / session_id / maxLengthAdded value: +120 - added
Input schema / properties / session_id / minLengthAdded value: +1
- Changed
browser.set_viewport3 fields changed- added
Input schema / additionalPropertiesAdded value: +false - added
Input schema / properties / session_id / maxLengthAdded value: +120 - added
Input schema / properties / session_id / minLengthAdded value: +1
- Changed
browser.stop_trace3 fields changed- added
Input schema / additionalPropertiesAdded value: +false - added
Input schema / properties / session_id / maxLengthAdded value: +120 - added
Input schema / properties / session_id / minLengthAdded value: +1
- Changed
browser.wait_for_selector3 fields changed- added
Input schema / additionalPropertiesAdded value: +false - added
Input schema / properties / session_id / maxLengthAdded value: +120 - added
Input schema / properties / session_id / minLengthAdded value: +1
- Changed
social.extract_posts3 fields changed- added
Input schema / additionalPropertiesAdded value: +false - added
Input schema / properties / session_id / maxLengthAdded value: +120 - added
Input schema / properties / session_id / minLengthAdded value: +1
- Changed
social.extract_profile3 fields changed- added
Input schema / additionalPropertiesAdded value: +false - added
Input schema / properties / session_id / maxLengthAdded value: +120 - added
Input schema / properties / session_id / minLengthAdded value: +1
- Changed
social.login3 fields changed- added
Input schema / additionalPropertiesAdded value: +false - added
Input schema / properties / session_id / maxLengthAdded value: +120 - added
Input schema / properties / session_id / minLengthAdded value: +1
- Changed
social.search3 fields changed- added
Input schema / additionalPropertiesAdded value: +false - added
Input schema / properties / session_id / maxLengthAdded value: +120 - added
Input schema / properties / session_id / minLengthAdded value: +1
32 tool updates
v0.5.0- First observed
browser.activate_tab - First observed
browser.close_session - First observed
browser.close_tab - First observed
browser.create_session - First observed
browser.drag_drop - First observed
browser.eval_js - First observed
browser.execute_action - First observed
browser.find_by_vision - First observed
browser.find_elements - First observed
browser.fork_session - First observed
browser.get_auth_profile - First observed
browser.get_console - First observed
browser.get_html - First observed
browser.get_network_log - First observed
browser.get_page_errors - First observed
browser.get_request_failures - First observed
browser.get_session - First observed
browser.list_auth_profiles - First observed
browser.list_downloads - First observed
browser.list_sessions - First observed
browser.list_tabs - First observed
browser.observe - First observed
browser.request_human_takeover - First observed
browser.save_auth_profile - First observed
browser.screenshot - First observed
browser.set_viewport - First observed
browser.stop_trace - First observed
browser.wait_for_selector - First observed
social.extract_posts - First observed
social.extract_profile - First observed
social.login - First observed
social.search
TDQS
Scored across 20 tools
Every tool targets a distinct resource or action: sessions, tabs, downloads, auth profiles, observation, reading content, and execution. Even similar tools like observe, screenshot, get_html, and find_elements have clearly different purposes and output formats, eliminating ambiguity.
All tools follow a consistent browser. prefix with verb_noun snake_case naming (list_sessions, create_session, activate_tab, save_auth_profile). Even single-word verbs like observe and screenshot are acceptable and maintain predictability.
With 20 tools, the set is on the heavier side but justified by the breadth of browser automation: session management, tab control, page reading, actions, downloads, auth profiles, and advanced features like forking and JS evaluation. Each tool serves a distinct purpose, so the count feels appropriate for a full-featured server.
The surface covers the full lifecycle: session CRUD (create, list, get, close), observation (multiple modes), navigation and interaction (execute_action), tab management, downloads, auth profiles, and even human takeover and forking. There are no obvious dead ends; every action has a complementary read or verification tool.
Maintenance
Related MCP Connectors
Run multi-step tasks in a real Chrome browser: persistent environments, live view, human takeover.
Hosted AI agents and workflows with app OAuth, human approval gates, and a run ledger.
Human-in-the-loop API for AI agents. CAPTCHA, OTP, KYC, and approvals by real humans.
Human approvals, notifications, inbound webhooks, wake-ups for headless agents. Free trial: /try
Related MCP Servers
- -licenseNot gradedqualityFmaintenanceThis server provides cloud browser automation capabilities using Browserbase, Puppeteer, and Stagehand. This server enables LLMs to interact with web pages, take screenshots, and execute JavaScript in a cloud browser environment.1,999 npm3,412MIT
- AlicenseNot gradedqualityFmaintenanceFacilitates browser automation with custom capabilities and agent-based interactions, integrated through the browser-use library.1968MIT
- AlicenseAqualityBmaintenanceEnables LLMs like Claude to navigate the web through Puppeteer-based tools and Steel. Based on the Web Voyager framework, it provides tools for all the standard web actions click clicking/scrolling/typing/etc and taking screenshots.1656MIT

gotoHuman MCPofficial
AlicenseAqualityFmaintenanceAdd human approval steps to your AI agents and automations with gotoHuman.346 npm52MIT