ascended-browser
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@ascended-browserSearch Wikipedia for the capital of Australia and summarize it"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
A real browser for AI agents, as an MCP server. Your agent opens pages, reads them, and acts on them through verified actions: it fills a whole form in one call, picks "Mrs." from a custom React dropdown by name, and is told what the page did in response, instead of clicking coordinates and hoping.
It is the browser from Ascended, packaged on its own: Camoufox (a hardened Firefox that looks like a person's browser to the sites it visits) behind the same tool dispatcher, page reading and result formatting Ascended's own agent uses.
See it work
Three real runs of Claude Code with only this server attached, on live sites.
The panel on the side is Claude Code's own transcript: every tool call it made
and what came back, errors included, with the real elapsed time. Waiting is
cut and tool time plays faster; the pointer, device frames, devtools panels
and outlines are drawn afterwards from the server's event log
(demo/record_demo.py), so they sit on what the agent really touched and read.
The full unedited screen recordings are in videos/unedited/
(the browser is driven through Playwright, so no OS pointer appears in them).
Front-end QA. react.dev on a phone and a tablet, dark/light, before/after screenshots, an audit outlining the real offending elements, console and network.
Browsing. Types into Wikipedia's search, finds a fact in the article, then searches YouTube and plays the first video.
Real forms. booking.com: popup, destination autocomplete, date picker, search, sort by price.
Click any clip for the full-quality MP4.
Related MCP server: browser-control-mcp-server
Install
From Python (3.11 or newer) or from npm; both run the same server.
uvx ascended-browser doctor # checks the machine; downloads nothing
uvx ascended-browser fetch # downloads the browser now (~700 MB; otherwise on first use)
npx -y ascended-browser doctor # the same, from npmThe browser is Camoufox 135.0.1-beta.24, the build every test here ran on; it is
pinned, so a newer Camoufox release never changes what your agent drives, and
any other Camoufox you have installed is left as it is. Run fetch once before
adding the server to an agent, so its first tool call does not wait for the
download.
The npm package is a small launcher: it runs the Python package with uvx
when uv is installed (uv brings its own Python),
else pipx, else a private venv made with your Python 3.11+. Use whichever
command you prefer in the configs below (npx -y ascended-browser in place of
uvx ascended-browser).
Linux: install xvfb to keep the browser on a virtual display (closest to a
real screen); without it the browser runs headless. On a display,
browser_viewport resizes the real window, so phone and tablet checks reflow
the page at true breakpoints.
Add it to your agent
Claude Code
claude mcp add ascended-browser -- uvx ascended-browser
# or: claude mcp add ascended-browser -- npx -y ascended-browserCodex
codex mcp add ascended-browser -- uvx ascended-browserWhether Codex asks before each browser action follows the permission mode you
pick in Codex: under Full access it just runs them; under Ask for approval
it asks first. In a sandboxed mode, codex exec cannot ask and fails every
call: use Full access, or pre-approve this server's tools by adding
default_tools_approval_mode = "approve" under [mcp_servers.ascended-browser]
in ~/.codex/config.toml.
opencode (opencode.json)
{
"mcp": {
"ascended-browser": { "type": "local", "command": ["uvx", "ascended-browser"], "enabled": true }
}
}Cursor, Windsurf, Claude Desktop and other clients: a stdio server with
command uvx and args ["ascended-browser"], or command npx and args
["-y", "ascended-browser"].
Tools
Tool | What it does |
| Open a URL (or several at once) and return what is on the page, each control with a ref |
| Look at the page again, or narrow to a region, a query or a filter |
|
|
| Read the page's text and field state, query by CSS selector, pull repeated items into a JSON shape, or |
| A picture, when an observation cannot describe it: canvas, charts, visual layout; |
| Resize to phone/tablet/desktop (Linux), emulate dark mode, reduced motion, forced colors or offline |
| Read-only JavaScript, policy-checked |
| List, close or sleep tabs |
| Record a task once, replay it on the next page with new values |
| Wait out a "checking your browser" page; press a Turnstile/reCAPTCHA checkbox if one blocks a form |
Long results come back clipped, with an evidence_ref that
browser_extract pages through, so a huge page cannot flood the agent's
context.
Settings
Variable | Default | |
| hidden |
|
|
| Browser profile (sign-ins persist), session files |
|
| Longer results are clipped with an |
| Any browser setting, e.g. | |
|
| Logs go to stderr |
Limits (0.1)
Resizing the window (phone/tablet presets, multi-size screenshot grids) runs through Ascended's live view, which this package does not ship yet. Emulation (dark mode and the rest) works.
Saved logins (
browser_login) and schema extraction backed by a model need the Ascended app.One server process is one browser session: tabs and refs last until your client disconnects; the profile (cookies, sign-ins) lasts across sessions.
How it is built
src/ascended_browser/_app is generated from Ascended by
scripts/sync_from_ascended.py: the browser modules copied as they are, the
browser tool dispatcher and result formatter extracted by reachability, and
every import of the rest of the app rewritten to runtime/ (small standalone
stand-ins). The sync refuses any app import it cannot map.
Tested with Ascended's own stress harnesses run against this package
(tests/stress/), a client-side MCP smoke test (tests/smoke_mcp.py), and
live-website tasks given to real agents (tests/agents/).
What has been verified so far: Linux (Python 3.11, 3.12 and 3.14), with Claude Code and Codex (0.160) on live-site tasks, opencode on a navigation task, and the npm launcher through uvx and through its own venv. macOS and Windows should work headless or with a visible window, but are untested, and window resizing for phone/tablet checks is Linux-only for now.
License
MIT. The bundled axe-core (_app/browser_vendor/axe-core) is MPL-2.0 and keeps
its notice in the file.
Available Tools
10 toolsbrowser_actA
Perform one verified action in a leased tab. action.kind must be navigate|click|fill|fill_form|sequence|select|check|date|press|upload|scroll|wait (action.type is accepted as an alias for kind; type=type means fill). ref must be an exact ref=e… token from the newest page/observe/extract result for this tab — labels and row numbers are rejected; CSS/XPath/text selectors are not an agent-facing escape hatch. Reuse the fresh refs returned in an action's page; do not observe again after every field. A successful act returns either page (the page it produced, with fresh refs) or page_unchanged: true, meaning the URL and visible text are what your last observation showed and its refs still apply — observe again only for a different view (query, selector, region). Batch stable fields with fill_form, which reports each field separately. For anything that picks from a list (native selects, custom ARIA comboboxes and dropdowns, type-to-search boxes) use kind=select with option= and optional query=; this performs and verifies open/filter/choose in one call. For anything with a checked or selected state, including list items, use kind=check with checked=true|false. Independent calls for distinct tab_ids may be issued together; calls targeting one tab remain ordered. Use sequence for a short reversible multi-step workflow when later-page targets can be named semantically. Every step is re-resolved from current page state, authority is rechecked, and the sequence stops before dispatch when an exact expect_before guard or a unique target is missing; use expect_after to guard a planned transition. Do not place a consequential final submit in a sequence. until waits inside the call for a late effect (redirect, async save, hot reload); kind=wait only waits. The same select fields can be included in fill_form. For a reversible draft/save action, a click receipt with state=network_effect_observed, a successful same-origin non-read response, matching control readback, and the page's save status is sufficient browser evidence. Do not activate unrelated controls, inspect hidden databases, or use shell/page scripts solely to obtain stronger private persistence proof. upload attaches files to an input[type=file] ref (or file_chooser=true with an observed attachment button ref) and is the ONLY way to do so — fill rejects file inputs. For a request such as 'keep scrolling until I stop you', use one scroll action with duration_seconds instead of spending repeated model rounds; cancelling/replacing the run stops it.
| Name | Required | Description | Default |
|---|---|---|---|
| action | Yes | ||
| tab_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does: it discloses the two success return shapes (a fresh `page` vs `page_unchanged: true`), ref staleness/verification rules, that sequence stops before dispatch on ambiguity or failed guards, that upload is the only file path and fill rejects file inputs, and what counts as sufficient evidence for a save action. This is unusually complete behavioral disclosure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The core purpose is front-loaded, but the body is a dense wall of run-on prose mixing workflow guidance, evidence policy, and parameter semantics with little structure; several clauses (save-evidence proof, private-persistence warnings) could be tightened or moved. Length is partly justified by the large nested schema, but structure is weak.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool of this complexity with no output schema and many nested action variants, the description covers return values, verification, guard semantics, batching, uploads, and waiting behavior — everything an agent needs to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description explains the semantics of the hard parameters beyond a bare schema: kind/type alias handling, the exact-ref requirement, option+query for select, checked state, fill_form batching with depends_on, and expect_before/expect_after guards. It does not explain tab_id or the top-level container, and the reported schema coverage is low, so a 4 rather than 5.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Opens with a specific verb+resource+scope ('Perform one verified action in a leased tab') and enumerates the exact action.kind values, so the agent knows precisely what the tool does. It also differentiates from siblings by referring to page/observe/extract results and browser_extract, so an agent can route correctly without opening other schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives rich when-to-use guidance: select for anything list-picking, check for any checked/selected state, sequence for short reversible multi-step flows (with an explicit 'do not place a consequential final submit in a sequence'), fill_form for stable fields, scroll+duration_seconds for continuous scrolling, until for late effects. It also states when NOT to act (no unrelated controls, no inspecting hidden databases, no shell scripts for stronger proof).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_evaluateA
Run policy-gated read-only JavaScript in a tab only when browser_extract or browser_observe cannot return the required fact. For visible page text, labels, controls, or a region owned by aria-controls, use browser_extract instead; for computed styles, colors, fonts, sizes, icons or backgrounds, browser_extract read=design (the page, or one component with target_ref) or read=inspect (one element's box and styles).
| Name | Required | Description | Default |
|---|---|---|---|
| arg | No | Optional argument passed to the function. | |
| tab_id | Yes | ||
| function | Yes | JS function/expression (aliases: expression, code, js, script). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description says the JavaScript runs 'policy-gated' and 'read-only,' which are meaningful behavioral disclosures: the agent knows it cannot mutate state and that some policy restricts execution. However, it doesn't detail what the policy prohibits, how failures appear, or whether the return is serialized, so some gaps remain for an unannotated tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two dense sentences that front-load the core purpose and then immediately provide alternative routing. It is well-structured and every clause adds routing information, though slightly long for a single-breath read.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 3-parameter, no-output-schema, no-annotation tool, the description covers purpose, constraints, and alternatives well. It does not explain what the function returns or how results are structured, which is acceptable given no output schema, but it also doesn't mention error handling or serialization limits for read-only evaluation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 67%. The description does not mention the 'arg', 'tab_id', or 'function' parameters, leaving those to the schema. The schema already documents 'arg' and 'function' partially; the missing third parameter description is not compensated in the description. Baseline 3 for moderate coverage with no extra parameter guidance.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description delivers a precise verb ('Run policy-gated read-only JavaScript'), the resource ('in a tab'), and an explicit scope constraint ('only when browser_extract or browser_observe cannot return the required fact'). It clearly distinguishes itself from sibling tools by naming them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides explicit when-to-use (only when other tools fail) and when-not-to-use, with concrete alternative routing: 'use browser_extract instead' for text/labels/controls, and 'browser_extract read=design or read=inspect' for styles/visual properties. No inference is left to the agent.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_extractA
Extract the complete current visible content and structured field/page state from a tab. Pass target_ref to scope a fresh read to an observed content block, control or the ARIA region it owns. This is a fresh high-level read, not a change-only observation. If a large result returns evidence_ref and next_cursor from browser_extract, continue it with this same tool. Do not continue a managed-output reference from browser_open, browser_observe, or browser_act; start a fresh extraction with tab_id and instruction instead. Pass selector for an instant CSS query instead of a full read: it counts and lists matching elements (tag, text, requested attributes, and the ref of any match already in the current observation) without building a snapshot, e.g. every product link's href, how many rows a table has, or each card's price. read=console|network reads the tab's logs (with same-site error bodies); read=inspect with target_ref says why a click is refused or a control will not take a value: what receives the click there, what covers or clips it, disabled state, styles. read=design returns how the page looks in CSS terms: variables, colors by use, type scale, radii, shadows, spacing and layout regions with sizes; target_ref limits it to one component. read=audit audits the page you are on (a deployed or localhost site): accessibility violations (axe-core) with the ref of each element, load timings and weight, title/lang/description/headings/alt gaps, broken same-origin links and console errors, worst first.
| Name | Required | Description | Default |
|---|---|---|---|
| find | No | Return every passage that mentions this, with the words around it, instead of the page's text. Use it when the page is long and the question is narrow (a deadline, a price, a name). | |
| read | No | ||
| level | No | ||
| limit | No | Continuation only: maximum characters for this slice; defaults to 8000. | |
| types | No | read=network: e.g. ['fetch','xhr']. | |
| checks | No | read=audit: which checks to run; omit for all four. | |
| cursor | No | With evidence_ref: character offset from next_cursor. With selector or log reads: entry offset from next_cursor. Defaults to 0. | |
| filter | No | read=network: URL regex, e.g. '/api/'. | |
| tab_id | Yes | ||
| from_end | No | Read the end of the page rather than the beginning. Footers, totals and closing dates live there, and paging forward charges for everything above them. | |
| selector | No | CSS selector to query instead of reading the page (e.g. 'table tbody tr', 'a.product-link', 'article h2'). Returns total, and for each match up to max_results: tag, text, attrs, children_count, visible, and ref when the element is in the current observation. An invalid selector is an error; no match is an empty result, not an error. | |
| attributes | No | selector only: attributes to return per match, e.g. ['href'] or ['src', 'alt']. href/src are complete absolute URLs, or explicitly omitted when longer than 8192 characters. Other attributes are bounded previews. value is the live field value (secrets read as [redacted]). | |
| target_ref | No | Fresh extraction, read=inspect or read=design: exact ref from the current observation (read=design also takes a region id r…). Reads that element, or the uniquely related region named by aria-controls, without executing page JavaScript in the model loop. | |
| failed_only | No | read=network: failed or 4xx/5xx only. | |
| instruction | No | Fresh extraction: optional objective used to focus field state and report an explicit no-match result. Continuation: accepted as a harmless repeated hint but archived evidence is returned unchanged. | |
| max_results | No | selector or log reads: entries per page (default 50). total always counts every match; continue with cursor=next_cursor. | |
| navigations | No | ||
| evidence_ref | No | Continuation only: opaque evidence_ref returned by an earlier browser_extract result in this session, accompanied by next_cursor. | |
| include_text | No | selector only: include each match's text (default true). Set false when only attributes or the count matter. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden and does so well: it distinguishes fresh reads from change-only observations, explains continuation semantics (evidence_ref + next_cursor, cursor offset, limit default 8000), notes that selector queries avoid building a snapshot, states invalid selector is an error while no match is an empty result, and discloses redaction of secret values and bounded attribute previews.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core purpose, then ordered into scoping, continuation, selector, and mode-specific paragraphs. It is dense and almost every sentence adds operative detail, though the sheer length (especially the read=audit and read=design clauses) makes it slightly sprawling for an agent to scan.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex 19-param multi-mode tool with no output schema and no annotations, the description covers the major modes, continuation behavior, and error semantics well. It is not fully complete: the level and navigations parameters remain unexplained, and the default page-read return shape is only partially implied through continuation limits.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 79%, and the description adds substantial meaning for target_ref, selector, attributes, find, from_end, evidence_ref, cursor, instruction, and each read mode. However, params level and navigations have empty schema descriptions and are never addressed in the description, leaving real gaps for a 19-parameter tool.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource (extract visible content and structured state from a tab) and explicitly differentiates itself from siblings: 'fresh high-level read, not a change-only observation' versus browser_observe, and it warns not to continue managed-output refs from browser_open/browser_observe/browser_act. An agent can distinguish this tool's role without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Extremely explicit routing: it says when to pass target_ref, when to use selector instead of a full read, when to use find, from_end, and each read= mode (console/network/inspect/design/audit). It also gives a clear when-not instruction ('Do not continue a managed-output reference from browser_open, browser_observe, or browser_act; start a fresh extraction').
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_flowA
Record and replay browser work. Every verified browser_act on a tab is recorded (targets by what they are, not by ref; passwords and other secrets never). After a task succeeds once, action=save turns that tab's recorded steps into a flow; action=run replays it on a tab with no model call between steps, re-finding each element on the live page, so repeating the task (the next job application on the same site, the next record in the same form) takes seconds. Pass start_url for the new page and variables for the values that change; save reports the detected variables. A run stops at the first step it cannot find unambiguously or verify, and says which steps are done and which remain; finish that step yourself, then run again with from_step. The same review and authority apply to a replayed submit as to one you click yourself: run only a flow whose steps the user's request covers.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | save: a short name, e.g. 'Apply on careers.example.com'. | |
| action | Yes | recorded: numbered verified steps on tab_id so far. save: store them (or from_step..to_step) as a flow. list: saved flows. show: one flow's steps and variables. run: replay flow_id on tab_id. delete: remove flow_id. | |
| tab_id | No | recorded, save, run: the tab. | |
| flow_id | No | show, run, delete: from save or list. | |
| to_step | No | save: last recorded step to include (default: the latest). | |
| from_step | No | save: first recorded step to include (default 1). run: step to resume from after finishing a stopped step yourself. | |
| start_url | No | run: page to start on, e.g. the next job's apply page. Omit to start where the flow was recorded, or on the current page when the tab is already there. | |
| variables | No | run: values for this run by variable name, e.g. {"email": "a@b.com"}. Unnamed variables keep the recorded value; an unknown name is an error. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does so richly: secrets are never recorded, targets are stored by identity not ref, replayed steps involve no model call, elements are re-found live, and a run halts at the first unverifiable step and reports done/remaining with from_step resume. This is exactly the behavioral context an agent needs before invoking a replay tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The purpose is front-loaded and every sentence contributes information (recording rules, save/run semantics, variable handling, stop/resume behavior, authority constraint). It is dense and fairly long, but there is no filler or repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an 8-parameter, no-annotation, no-output-schema tool, the description covers the important behavioral ground including failure/resume semantics. The main remaining gap is that it doesn't describe what recorded/list/show return, though those are read-only inspection modes where the schema enum hints at intent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is already 100%, so the baseline is 3; the description nonetheless adds value by explaining the record->save->run lifecycle, what start_url means for a new page, what variables carry, and that save reports detected variables. It doesn't elaborate on from_step/to_step beyond what the schema already documents.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening sentence names a specific verb+resource ('Record and replay browser work') and the body specifies that recording happens from verified browser_act steps, which cleanly distinguishes it from siblings like browser_act and browser_open. An agent can tell what this tool is for without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives concrete trigger conditions ('After a task succeeds once, action=save... action=run replays it') and a real-world motivation (next job application, next record in the same form). It also states a usage boundary ('run only a flow whose steps the user's request covers'), though it never explicitly names a sibling alternative for one-off actions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_observeA
Observe a live tab and return generation-scoped element refs, on-screen controls first, current viewport content_blocks and measured scroll_containers. Treat auth_state as page evidence for your next decision; a visible login form does not itself stop the run. diagnostics lists page errors and failed requests since your previous result. When the result says elements were not shown, narrow the view instead of reading the DOM with browser_evaluate: query finds controls by their words, within scopes to a region, frame or element, and cursor continues where the last response stopped. Every ref returned is valid for browser_act.
| Name | Required | Description | Default |
|---|---|---|---|
| query | No | Words to find. Returns only controls whose role, label, text, value or option labels contain every word, searched across the whole page, not just what an earlier observation showed. | |
| cursor | No | The next_cursor an earlier observation printed; continues the same ordering. | |
| filter | No | interactive (default) lists controls; all also lists the page's headings; links lists where the page can take you as text and URL, which fits a whole index in one answer. query applies to links too and matches the URL, so filter:'links' with query:'/jobs/' returns that section of the site. Every ref is still clickable. | |
| tab_id | Yes | ||
| within | No | Scope to one region id (r…) or frame id (f…) from an earlier observation, or to an element ref (e.g. a dialog, listbox or form) and its descendants. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does well: it discloses that refs are generation-scoped, that they remain valid for browser_act, that auth_state is page evidence rather than a run-stopping condition, and that diagnostics reports errors and failed requests since the previous result. It never explicitly states the operation is read-only or mentions cost/rate characteristics, leaving a small gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The core return shape is front-loaded in the first sentence, followed by behavioral notes and then the routing advice, so an agent gets the essentials immediately. It is dense but every sentence carries information; a few clauses (auth_state, diagnostics) could be tightened without loss.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description must describe results, and it does: element refs, content_blocks, scroll_containers, auth_state, diagnostics, and the 'elements were not shown' signal. Combined with 80% schema coverage and one enum, the definition is nearly self-sufficient; only explicit read-only framing and pagination/volume expectations are missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 80%, so the baseline is 3, but the description adds genuine meaning beyond the schema: query searches the whole page (not just the last observation), within scopes to a region, frame or element, and cursor continues the prior ordering. It frames these as one coherent narrowing strategy rather than restating field definitions, though it does not add format/syntax detail for within or cursor values.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource ('Observe a live tab') and enumerates exactly what is returned: generation-scoped element refs, on-screen controls first, content_blocks and scroll_containers. That level of specificity separates it cleanly from browser_evaluate, browser_extract and browser_screenshot without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives explicit routing: 'When the result says elements were not shown, narrow the view instead of reading the DOM with browser_evaluate', and points to query/within/cursor as the narrowing path. It also tells the agent how to treat auth_state. It stops short of contrasting with the other siblings (browser_extract, browser_screenshot, browser_flow), so it is strong but not exhaustive.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_openA
Wake the workspace and open or navigate your current browser tab. Pass either url or a bounded urls list (maximum 5). To check several sites, open them in ONE call with urls: it opens independent tabs in the same authenticated browser context and returns one compact, ordered result per URL; that results list is the authoritative source for each tab_id. A URL that fails (DNS, block) is reported on its own row and the other tabs stay open. Batch results omit page bodies: next, call browser_extract/observe for every returned tab_id together in one response, since calls on different tabs run in parallel. Returns once the page has settled. By default it reuses your current actor-owned tab, including across a resumed task; set reuse=false only for a genuinely independent second tab (the result carries reused=true when reused). The single-URL result includes a settled page snapshot; use it directly instead of immediately calling browser_observe. Only a returned attention_required=true represents an explicit handoff. A blocked=true page result does not suspend the run: decide whether to take another route or tell the user plainly what could not be completed. For broad public discovery use web_search rather than navigating this tab to a search engine; browser search remains valid when explicitly requested or interaction/personalization matters.
| Name | Required | Description | Default |
|---|---|---|---|
| url | No | Single-page form. Set this URL and leave urls empty/omitted. | |
| read | No | Return the page's readable text with this result, so opening a page in order to read it is one call rather than open then browser_extract. | |
| urls | No | Batch form only. Leave url empty/omitted. Each URL opens its own tab; a failed URL does not close the others. | |
| reuse | No | Single-page form only: default true. URL batches always use independent tabs so atomic rollback is possible. | |
| atomic | No | Rarely needed. true closes every batch tab if any URL fails (all-or-nothing). Omit for the normal independent behavior. | |
| disposition | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden and does so: returns once the page has settled, batch results omit page bodies, failed URLs are reported per-row without closing siblings, atomic rollback is opt-in, reuse persists across a resumed task, and attention_required/blocked semantics are spelled out rather than suspending the run.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core purpose and dense throughout, with each clause carrying operational detail. It runs long and some sentences (the web_search aside, the blocked/attention handoff) could be tightened, but there is very little filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, so the description must describe returns, and it does: the ordered results list as authoritative for tab_id, the settled page snapshot for single-URL, reused=true signaling, and attention_required/blocked fields. Nothing needed to call or interpret the tool correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is already 83%, so the schema documents most parameters. The description still adds real meaning beyond it: reuse defaults and cross-resumed-task behavior, the atomic all-or-nothing tradeoff, the url/urls mutual exclusivity, and the read flag's one-call shortcut. Only 'disposition' is left without prose explanation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('open or navigate your current browser tab') plus the secondary effect of waking the workspace. It also implicitly separates itself from siblings by directing content retrieval to browser_extract/observe and discovery to web_search.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit when-to-use rules: one call with urls for several sites, single url otherwise, reuse=false only for a genuinely independent tab, and web_search for broad public discovery. It names the alternative tools and the conditions that select them.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_screenshotA
Take a picture of an agent-owned tab. This is a last resort, not a way to find or operate controls: browser_observe returns the refs browser_act needs, and a picture returns none. Use it only for what an observation cannot describe (canvas, charts, images, visual layout). Password and payment fields are masked. The result states the viewport size at capture; a person may resize the live window, so trust it over sizes you saw earlier. state captures a ref's :hover/:focus styling; compare_with diffs against an earlier picture or another tab and boxes what changed. To check a page you are building against an original, open both and take one with compare_with='tab:': one picture of both.
| Name | Required | Description | Default |
|---|---|---|---|
| ref | No | scope=element or state: a ref from the latest browser_observe of this tab. | |
| label | No | Short caption naming the picture, e.g. 'Application form before submit'; used in the evidence file name. | |
| scope | No | viewport (default) is what is on screen now; full_page is the scrollable document, cut at 8000 px; element is one ref. | |
| state | No | ||
| tab_id | Yes | The agent-owned tab to capture. | |
| compare_with | No | 'previous' (this tab's last picture), 'tab:<tab_id>' (another tab, e.g. the original site beside your build; pages of different heights come back side by side), or an evidence path; same scope and size. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations the description carries the full burden and discloses real behavioral traits: password/payment fields are masked, the result reports the capture viewport size, and the live window may be resized so the reported size should be trusted over earlier observations. It does not cover permissions/ownership rules or how the image is delivered, but the disclosure is well beyond the schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the core purpose and the 'last resort' caveat before elaborating on parameters, and every sentence carries information. It is dense and a touch sprawling, but nothing is filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a screenshot tool with no output schema and no annotations, the description covers the key behaviors an agent needs: masking, returned viewport size, comparison semantics, full_page truncation. Only the delivery format of the image itself is unaddressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 83% (baseline 3), and the description adds genuine meaning: 'state captures a ref's :hover/:focus styling', compare_with 'diffs against an earlier picture or another tab and boxes what changed', and the tab:<tab_id> side-by-side example. The 8000px full_page cut is also surfaced. It stops short of explaining ref/label futher than the schema already does.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a concrete verb+resource ('Take a picture of an agent-owned tab') and immediately positions it against siblings: browser_observe returns refs, browser_act operates, a picture returns none. An agent can distinguish this from every sibling without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says it is 'a last resort, not a way to find or operate controls' and names the exact condition for use ('only for what an observation cannot describe: canvas, charts, images, visual layout'), routing the agent to browser_observe/browser_act otherwise. It even adds the build-vs-original workflow via compare_with.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_tabsA
List, close, or sleep workspace tabs. Closing takes tab_id for one tab or tab_ids for several in a single call (at most 20); each tab's outcome is reported separately, so one unknown id does not strand the rest.
| Name | Required | Description | Default |
|---|---|---|---|
| action | Yes | ||
| tab_id | Yes | ||
| tab_ids | No | Several tabs to close in one call. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It usefully discloses batch semantics (max 20, per-tab outcomes, one bad id does not strand the rest), but stays silent on whether 'close' is destructive/reversible, how 'sleep' differs in effect, and what permissions are needed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two dense sentences, front-loaded with the three actions, with the batch constraint and failure-isolation behavior stated compactly. No filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation-capable tool with no annotations and no output schema, the description covers purpose and batch behavior but omits action-specific requirements (whether 'list' ignores tab_id) and the effect/risk profile of close vs sleep.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 33% (tab_id has no description), yet the description compensates by mapping tab_id to one tab and tab_ids to several, and adds the non-schema 20-item cap. It does not clarify why tab_id is required even for 'list', which leaves a small gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States specific verbs (list/close/sleep) against a specific resource (workspace tabs), matching the action enum exactly. An agent can distinguish this from browser_open, browser_act, or browser_observe without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied by the action verbs; there is no guidance on when to choose 'sleep' over 'close', nor when to prefer this tool over browser_open/browser_observe for tab state. Adequate but gaps remain for an agent deciding between actions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_viewportA
Check a responsive layout in the current agent-owned tab. action=set with preset phone (390x844), tablet (768x1024) or desktop (1365x768), or width and height in CSS pixels, resizes the real browser window so the page reflows as it would on that screen (Linux; elsewhere the window keeps its size). action=restore returns to the launch size and clears emulation; call it when the check is done. action=get reads the current size. action=emulate sets color_scheme (light or dark), reduced_motion, forced_colors or offline. A successful set or restore includes a fresh page snapshot; earlier element positions are stale after a resize, so use the new refs.
| Name | Required | Description | Default |
|---|---|---|---|
| clear | No | emulate: clear all emulation. | |
| width | No | set only, without preset: requested page-content width in CSS pixels. | |
| action | Yes | get reads size and emulation; set applies a preset or width and height; restore returns to dynamic sizing and clears emulation. | |
| height | No | set only, without preset: requested page-content height in CSS pixels. | |
| preset | No | set only: the live view's profile sizes, phone 390x844, tablet 768x1024, desktop 1365x768. | |
| tab_id | Yes | The current agent-owned browser tab whose shared native window is targeted. | |
| offline | No | emulate: fail every request in this browser context. | |
| color_scheme | No | ||
| forced_colors | No | ||
| reduced_motion | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does so: it discloses that resizing only affects the real window on Linux and elsewhere the window keeps its size, that a successful set/restore returns a fresh page snapshot, that earlier element positions become stale and new refs must be used, and that restore clears emulation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
It is front-loaded with the primary purpose and walks through actions in order, but the single long paragraph is dense and the clause about Linux vs elsewhere is easy to miss.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 10-parameter emulation tool with no output schema, the description covers the important behavioral details an agent needs: tab ownership context, per-action intent, platform-specific behavior, snapshot return, and ref staleness. Nothing material is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema already documents most parameters, but the description adds specifics beyond it: the concrete preset dimensions were listed in the description body, and it clarifies semantics of set vs emulate vs restore. Schema coverage is 70% with 5 enums, so it complements rather than repeats the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names the specific operation (resize the real browser window for responsive-layout checks) and enumerates each action verb with its resource. It differentiates itself from layout-inspection siblings like browser_screenshot and browser_observe by scoping to viewport sizing and emulation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives clear contextual guidance for each action (set for resizing, restore when the check is done, get for reading). It does not name alternative sibling tools or explain when not to use this one, so it falls short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
wait_for_bot_wallA
Wait out a bot wall on a tab, then attempt to solve it if it does not clear on its own. Polls the tab while the wall is auto-verifying (e.g. a Cloudflare 'Just a moment…' check that clears itself). If it becomes a user-required challenge it runs the configured solver, and if that fails it returns a structured verdict. Also call it before submitting a form that carries an embedded Turnstile, reCAPTCHA or hCaptcha checkbox: it waits for the widget to pass, presses the checkbox if needed, and reports whether a token was issued (image puzzles are never solved). When the page is required for the user's goal, stop and tell them to solve it (or hand off if you are a browser sub-agent); otherwise continue to the next task.
| Name | Required | Description | Default |
|---|---|---|---|
| tab_id | Yes | ||
| max_wait_ms | No | How long to wait for the wall to clear on its own, in milliseconds (default 60000). | |
| apply_solver | No | Default true. Run the configured bot-wall solver after the wait budget if the wall is still present. | |
| poll_interval_ms | No | How often to re-observe the tab while waiting, in milliseconds (default 5000). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and delivers: it polls while auto-verifying, runs the configured solver for user-required challenges, presses the checkbox if needed, and enforces a hard limit ('image puzzles are never solved'). Failure behavior (structured verdict) and side-effect (checkbox press) are disclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the core action, then layers in the captcha-before-submit case and the handoff rule. It is dense and long, but each sentence contributes distinct behavioral information rather than restating the name.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, and the description compensates by describing the return shape ('structured verdict', 'reports whether a token was issued'). Combined with the parameter context this is nearly complete, though it could say a bit more about the verdict's fields.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 75%, with defaults and meanings already documented on max_wait_ms, apply_solver and poll_interval_ms. The description reinforces these by referring to the 'wait budget' and the 'configured solver', adding contextual meaning for how the parameters interact without contradicting the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('wait out a bot wall on a tab, then attempt to solve it') and goes further by naming the concrete scenarios (Cloudflare 'Just a moment…', embedded Turnstile/reCAPTCHA/hCaptcha). None of the sibling browser_* tools handle bot walls, so an agent can distinguish it immediately.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit when-to-call conditions: before submitting a form with an embedded captcha widget, and whenever a bot wall may auto-clear. It also states when to stop and hand off to the user vs continue, which is exactly the routing decision an agent needs.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
10 tool updates
v0.1.1- First observed
browser_act - First observed
browser_evaluate - First observed
browser_extract - First observed
browser_flow - First observed
browser_observe - First observed
browser_open - First observed
browser_screenshot - First observed
browser_tabs - First observed
browser_viewport - First observed
wait_for_bot_wall
TDQS
Scored across 10 tools
Most tools target clearly different browser operations, but the read-side tools (browser_observe, browser_extract, browser_evaluate, browser_screenshot) overlap in purpose and require the long descriptions to separate them. The descriptions do provide strong guidance, so an agent can usually pick correctly, but some ambiguity remains.
Nine tools use a consistent snake_case browser_ prefix, and the naming is readable and predictable. The lone exception is wait_for_bot_wall, and some browser_ tools are noun-based rather than verb-based, which are minor deviations.
Ten tools is well-scoped for a browser automation server, with each tool covering a distinct capability such as opening, acting, observing, extracting, screenshotting, viewport control, flow replay, and bot-wall handling. No tool feels redundant or missing at the count level.
The surface covers the browser lifecycle: opening/navigating, tab management, interaction, observation, extraction, evaluation, screenshots, viewport emulation, workflow replay, and bot-wall handling. There are no obvious dead ends for typical browser automation tasks.
Maintenance
Related MCP Connectors
AI-powered browser automation — navigate, click, fill forms, and extract data from any website.
Web search, browser automation, scraping, crawling and CAPTCHA solving for AI agents.
Live web access for agents: scrape, SERP search, crawl/map, 100+ collectors, datasets, proxies.
AI-powered web automation. Navigate websites using AI agents for one page or a thousand
Related MCP Servers
- AlicenseNot gradedqualityCmaintenanceEnables AI agents to interact with web browsers using natural language, featuring automated browsing, form filling, vision-based element detection, and structured JSON responses for systematic browser control.62MIT
- AlicenseAqualityCmaintenanceEnables AI agents to fully control a browser for web automation, including navigation, clicking, typing, scrolling, screenshots, and DOM inspection, with session persistence and anti-bot bypass.4024 npmMIT
- FlicenseNot gradedqualityCmaintenanceEnables AI assistants to control a browser through a set of tools, allowing them to perform web automation tasks like navigation, typing, clicking, and taking screenshots.-
- AlicenseNot gradedqualityBmaintenanceEnables AI agents to search the web, read and extract content from webpages, fetch JSON from REST APIs, and collect links while bypassing anti-bot protections.50 npmISC