agent-browser
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@agent-browseropen my bank portal — I'll take over for the 2FA, then read my balance"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Why this exists
Most browser tooling for agents hands the model a browser you cannot see. When it misreads a page or stalls on a login, all you get is a transcript and a guess.
Agent Browser runs one headed Chrome session. Your agent drives it through a small JSON API, and a live view of that same session sits open in front of you. When the agent gets stuck, you take over in the window it is already using — then hand it back.
Start the runtime:
docker compose up --buildThen reach it however you like:
npm install aether-browser # TypeScript
pip install aether-browser # Python
claude mcp add agent-browser -- npx -y aether-browser mcp # any MCP clientRelated MCP server: monkeysee
Quickstart
One command, on Docker Engine for Linux with Compose v2. It builds from the source checkout and starts Xvfb, x11vnc, noVNC, and the API — the stack that owns a single headed Chrome session. Both user-facing listeners bind to numeric loopback.
docker compose up --buildOnce health responds, open the live view at
127.0.0.1:6080/vnc.html and check the API from another terminal:
curl -fsS http://127.0.0.1:8092/browser/health | jq .Ctrl+C stops it.
The first build takes several minutes. It installs the hash-locked Python environment, then uses Patchright to install the current Google Chrome Stable package. The exact browser version is captured with each accepted image, so rebuilding the same source later may pick up a newer Stable.
Linux host networking is deliberate. It keeps the unauthenticated v0.x noVNC surface on numeric loopback. Docker Desktop and remote-host deployment are outside this quickstart.
The API in four calls
With the runtime healthy and curl plus jq installed. Local loopback needs no bearer token by
design — read the authority contract before you change the
deployment shape.
# Open the session. This is the browser you are about to watch.
SESSION_ID="$(curl -fsS -X POST http://127.0.0.1:8092/browser/session/create \
-H 'Content-Type: application/json' \
-d '{"api_version":"v1"}' | jq -er '.session_id')"
# Go somewhere. The page changes in the live view as this runs.
curl -fsS -X POST http://127.0.0.1:8092/browser/navigate \
-H 'Content-Type: application/json' \
-d "{\"api_version\":\"v1\",\"session_id\":\"${SESSION_ID}\",\"url\":\"https://example.com\"}" \
| jq '{status, final_url, title, readable_text}'
# Read it back as structure, not pixels.
curl -fsS -X POST http://127.0.0.1:8092/browser/snapshot \
-H 'Content-Type: application/json' \
-d "{\"api_version\":\"v1\",\"session_id\":\"${SESSION_ID}\"}" \
| jq '{status, url, title, sequence, vision_steps_remaining}'
# Give the browser back.
curl -fsS -X POST http://127.0.0.1:8092/browser/session/end \
-H 'Content-Type: application/json' \
-d "{\"api_version\":\"v1\",\"session_id\":\"${SESSION_ID}\"}" | jq .What that did: session/create started one headed Chrome session and returned its UUID plus the
local view URL. navigate validated the destination, pinned the allowed addresses, and changed the
page on the shared display. snapshot returned bounded text, accessibility state, viewport
metadata, counters, and a PNG of that same page. At no point did the API and the human view create
competing browsers — they met at one owned session.
The same flow ships as examples/curl.sh.
Use it from an MCP client
Any MCP client — Claude Code, Claude Desktop, Cursor, Windsurf — can drive the session while you watch, and take over when it gets stuck.
claude mcp add agent-browser -- npx -y aether-browser mcpAlready have the Python client? aether-browser mcp serves the same nine tools.
{
"mcpServers": {
"agent-browser": {
"command": "npx",
"args": ["-y", "aether-browser", "mcp"]
}
}
}Nine tools — browser_open · browser_navigate · browser_read · browser_click ·
browser_type · browser_press · browser_scroll · browser_status · browser_close
browser_open hands back the live view URL, and every response after it repeats that URL, so you
always know where to look. On connect, the server tells the model to stop and ask for a takeover at
a login, a payment, or a 2FA prompt rather than guessing — the loop in the recording above.
Setup and limits: docs/MCP.md.
Node and TypeScript client
The same API from Node, with types and cleanup you cannot forget:
npm install aether-browserimport { AgentBrowser, withSession } from 'aether-browser'
const browser = new AgentBrowser({ controllerToken: process.env.AGENT_BROWSER_CONTROLLER_TOKEN })
await withSession(browser, async (session) => {
const page = await session.navigate('https://example.com')
await session.click({ selector: '#login' })
await session.type({ selector: '#user', text: 'ada' })
console.log(page.title, session.viewUrl)
})withSession always ends the session, including when your callback throws, so a crash cannot leave
the single slot occupied. No runtime dependencies, and it runs anywhere that can reach the server.
It carries a small CLI too: npx aether-browser doctor tells you what is missing before a first
run, and up builds and starts the runtime on a Linux host.
Python client
The same client, same name, same commands, released version for version with the npm package:
pip install aether-browserimport os
from aether_browser import AgentBrowser, session
browser = AgentBrowser(controller_token=os.environ["AGENT_BROWSER_CONTROLLER_TOKEN"])
with session(browser) as live:
page = live.navigate("https://example.com")
live.click(selector="#login")
live.type("ada", selector="#user")
print(page["title"], live.view_url)The session context manager always ends the session, including when the body raises. No runtime
dependencies — the transport is urllib from the standard library — ships type hints, and runs on
Python 3.10 or newer, anywhere that can reach the server. Same CLI as the npm package:
aether-browser doctor, up, status, open, down.
Optional Jev web decisions (Python)
The Python client has an opt-in aether_browser.jev layer for Jev → your selected model.
Jev reads a caller-selected text excerpt and chooses between up to three caller-approved URLs,
handoff, or human takeover. The browser navigates only to a URL the caller supplied, and its
normal destination checks still apply. On handoff, your callback receives the current page
evidence and Jev's request IDs, token counts, and costs; your application invokes its chosen
reasoning model to write, analyze, or continue browsing. Jev returns decisions, not prose or
visual judgments, so screenshots and complex page interactions belong to that selected model or
the human watching the live browser.
import os
from aether_browser import AgentBrowser, session
from aether_browser.jev import HumanTakeover, JevWebAgent, NavigationOption, OpenRouterJev
browser = AgentBrowser(controller_token=os.environ["AGENT_BROWSER_CONTROLLER_TOKEN"])
with session(browser) as live:
first_page = live.navigate("https://example.com")
result = JevWebAgent(OpenRouterJev(os.environ["OPENROUTER_API_KEY"])).run(
live,
goal="Read the site's documentation",
initial_page=first_page,
options_for=lambda page: (
[NavigationOption("https://example.com/docs", "Official documentation")]
if page["final_url"] == "https://example.com/"
else []
),
excerpt_for=lambda page: page["readable_text"][:6000],
selected_model=lambda handoff: your_model(handoff), # Supply your own model callback.
)
if isinstance(result, HumanTakeover):
print("Take control at", result.view_url)
input("Press Enter when the human is done to close this session: ")
else:
print(result)The example's your_model is an application-defined function, not part of this package. The
excerpt_for callback explicitly chooses text sent to OpenRouter, and options_for explicitly
chooses candidate URLs. Never return private page content or signed URLs unless your application
intends to send them to that provider. The OpenRouter key stays in the calling process; the
browser server does not store it. No Jev request occurs unless you call this optional layer.
If Jev fails or returns an invalid decision, the selected-model callback receives a handoff
with reason="jev_unavailable" and the browser takes no additional action. Recognized authentication
or payment pages prompt human takeover before Jev receives their text. The session remains owned
by your code; keep it open while the human takes over. See docs/JEV.md for
the complete contract and limitations.
What makes it different
One session, two participants. The agent acts through JSON. You watch the same display, and take the controls whenever you want them.
Structure before pixels. Readable text and a bounded accessibility tree come back before you spend a vision step on a screenshot.
A small control surface. The v0.x API exposes explicit browser actions — not a shell, not arbitrary JavaScript, not raw DevTools.
Model-agnostic and self-hosted. Bring the framework you already use, and keep the browser on hardware you control.
What it does today
Capability | v0.x contract |
Browser | One headed Google Chrome Stable session launched through Patchright |
State | URL, title, readable text, bounded accessibility nodes, viewport, and PNG snapshot |
Actions | Navigate, click, type, scroll, and allowlisted key presses |
Human view | The same Xvfb display through loopback-only x11vnc and noVNC |
Ownership | One explicit UUID session with expiry, vision budget, and idempotent cleanup |
Authority | Observer/controller separation when authenticated; strict local loopback mode otherwise |
MCP | Nine stdio tools from either client ( |
Navigation | HTTP(S)-only validation across requested, redirected, and browser-initiated navigation |
Request and response shapes, limits, and stable error codes: docs/API.md.
Architecture
flowchart LR
Agent["Agent client"] -->|bounded JSON API| API["FastAPI"]
API --> Guard["authority + navigation policy"]
Guard --> Session["single-session manager"]
Session --> Chrome["Patchright + headed Google Chrome"]
Chrome --> State["text · accessibility · PNG"]
State --> Agent
Chrome --> Display["shared Xvfb display"]
Display -->|loopback noVNC| Human["Human observer / takeover"]The session manager owns the page, browser context, temporary profile, timers, counters, and
cleanup. The API and the live view are different interfaces to that shared resource, not two
independent automation paths. See docs/ARCHITECTURE.md.
Agent Browser v0.2.3 is source-first and self-hosted. It is not a hosted service, and the v0.x noVNC surface is unauthenticated and meant for numeric loopback on a machine you control. No Chrome-containing image, image tar, or public layer cache is distributed unless separate redistribution authorization is documented.
Security boundary
API and noVNC listen on numeric loopback by default; noVNC remains loopback-only in the v0.x line.
Remote API clients require a separately operated same-host HTTPS reverse proxy, an exact trusted loopback peer, strict Host validation, and distinct strong observer/controller tokens.
Destination validation rejects credentials, unsupported schemes, blocked address classes, unsafe redirects, and DNS rebinding. Browser egress is pinned through an owned TCP proxy.
Non-proxied WebRTC UDP is disabled so it cannot silently bypass the TCP egress boundary.
Inputs, outputs, interactions, timeouts, lifetimes, and screenshot budgets are bounded.
Cleanup converges on session end, expiry, launch failure, application shutdown, and process failure.
Trust assumptions and residual risks are spelled out in docs/SECURITY.md.
Report vulnerabilities privately through SECURITY.md — please do not open a public
security issue.
Source recovery and exclusions
The source-recovery rule is reuse general browser behavior, not private domain code.
Lifecycle, structured-state, interaction, and cleanup patterns may be adapted from authorized
references; ATS/trading integrations, broker or account selectors, order actions, secrets,
and credential injection are excluded from the public core. Provenance status is tracked in
docs/SOURCE-RECOVERY.md.
What it does not do
No hosted cloud service, cloud control plane, or production remote-hosting claim.
No bundled LLM or default model calls, account system, dashboard, credential vault, or credential injection.
No CAPTCHA bypass, anti-detection guarantee, stealth claim, or proxy rotation.
No arbitrary JavaScript, shell, filesystem, upload, clipboard, download, or raw CDP API.
No multi-session pool, ATS integration, trading integration, or brokerage behavior.
Roadmap
Shipped. Both clients and their CLI are published as aether-browser, version for version,
on npm from clients/node and
on PyPI from clients/python. Since
0.2.0 both also serve the MCP server.
Two tracks are open, each with an issue, and each is a good first contribution:
Multi-session worker pool — explicit isolation and capacity semantics.
Session trace and recording export — with clear privacy controls.
These are candidates, not shipped features.
Contributing
Start with CONTRIBUTING.md, the
Code of Conduct, and the current API contract. Small,
well-tested changes that keep the authority boundary narrow are very welcome — the two roadmap
issues above are the best place to start. Security reports go through the private process in
SECURITY.md, never a public issue.
License and third-party notices
Aether-owned source code is licensed under the Apache License 2.0. Google Chrome is
separately licensed under Google's Chrome terms and is not
covered by Aether's Apache license; dependencies, system packages, fonts, and web assets also
remain under their respective terms. See THIRD_PARTY_NOTICES.md.
Aether is not affiliated with or endorsed by Google. The v0.x distribution target is source that
builds locally; this repository does not distribute a prebuilt Chrome-containing image.
Available Tools
9 toolsbrowser_clickB
Click a CSS selector, or an x/y point in the viewport. Provide one, not both.
| Name | Required | Description | Default |
|---|---|---|---|
| x | No | Viewport x coordinate. | |
| y | No | Viewport y coordinate. | |
| selector | No | CSS selector to click. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full behavioral burden and largely fails: it says nothing about whether the click waits for the element to exist, scrolls it into view, what happens if the selector matches nothing or the page is not loaded, or whether a browser session must already be open. It only restates the mutation action implied by the name.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, the primary mode listed first, and the argument constraint stated immediately after with no filler. Every word earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with no annotations and no output schema, the description leaves the agent without error/failure semantics, session prerequisites, or any indication of what a successful click returns. A simple tool, but the missing behavioral context is exactly the gap the description should have filled.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already explains x, y, and selector, setting the baseline at 3. The description adds genuine meaning beyond the schema by declaring mutual exclusivity between the selector mode and the coordinate mode, a constraint the schema itself does not encode (zero required parameters).
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (click) and the exact resource it acts on (CSS selector or viewport x/y point), so an agent can tell it apart from browser_type, browser_press, and browser_scroll. It does not explicitly name sibling tools or say how it differs from e.g. browser_press, but the action itself is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
"Provide one, not both" gives a real constraint on argument selection, which is useful since no parameter is marked required. However, it gives no guidance on when to prefer a selector versus coordinates (selector is normally the reliable path, coordinates the fallback for non-CSS targets), nor any precondition such as the page already being open.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_closeA
End the session and release the browser. Idempotent. Call this when the task is finished so the single session slot is free.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, and it does disclose two non-obvious traits: idempotency and the existence of a single session slot that must be freed. It does not say whether an unclosed session blocks future browser_open calls or what an error looks like if no session exists, but the key behavioral context is present.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with what the tool does, then the idempotency guarantee, then the calling condition. No filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter teardown tool with no output schema and no annotations, the description covers what it does, when to call it, the idempotency edge case, and the resource constraint. Nothing an agent needs in order to call it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so the baseline is 4. There is nothing to document and the description does not pretend otherwise.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('End the session and release the browser'), which is unambiguous against siblings like browser_open or browser_read. It also names the shared resource being affected (the browser session), so an agent can tell it apart from navigational tools without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly gives the trigger condition: 'Call this when the task is finished so the single session slot is free.' That is clear when-to-use guidance. It stops short of naming browser_open as the counterpart that must precede it, or warning what happens if the session is still mid-task, so it is not a full 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_openA
Start the one headed Chrome session and return the live view URL. Give that URL to the human: they can watch this exact session, and take over in it, while you drive. Safe to call twice — it reuses the session that is already open.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does well: it discloses idempotency ('Safe to call twice — it reuses the session that is already open'), the headed/single-session nature, and human takeover semantics. It omits failure modes (e.g., what happens if Chrome cannot launch) and any auth or timing behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences: action and return value first, then human-facing usage, then idempotency. Every sentence carries distinct information with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter tool with no output schema, the description adequately covers the return value (live view URL), idempotency, and human cooperation model. It could be slightly more complete by stating that this must be opened before the other browser_* tools can drive the session.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There are zero parameters with 100% schema coverage, so the baseline is 4. The description correctly adds no parameter detail because none is needed, though it does not mention a configurable session name or similar implied knobs.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource — starting the single headed Chrome session and returning the live view URL — and the phrase 'the one headed Chrome session' plus 'reuses the session that is already open' carves out its singleton role versus siblings like browser_close and browser_status. An agent immediately knows this is the session entry point.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives clear operational context: the returned URL goes to the human, who can watch or take over while the agent drives, and it is safe to call twice. It never explicitly says when not to call it or names a sibling alternative, so it falls short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_pressC
Press one allowlisted key or combination. Allowed: Enter, Escape, Tab, Backspace, Delete, Space, ArrowUp, ArrowDown, ArrowLeft, ArrowRight, Home, End, PageUp, PageDown, Control+A, Control+Z, Control+Shift+Z, Meta+A, Meta+Z, Meta+Shift+Z.
| Name | Required | Description | Default |
|---|---|---|---|
| key | Yes | Key to press. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It does disclose the allowlist restriction, which is genuine behavioral context, but says nothing about what pressing does (e.g., acting on the focused element), whether an error is returned for unlisted keys, or the post-press effect.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The bulk of the text is a 20-item enumeration that redundantly duplicates the input schema's enum, which does not earn its place. The useful content ('press one allowlisted key or combination') is a single short clause buried ahead of the redundant list.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple one-parameter tool with no output schema, the description covers the what and the allowed values. It is nonetheless missing focus/prerequisite context and error behavior that an agent would need to invoke it reliably.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the single parameter is a fully enumerated enum, so the schema already documents every valid value. The description's list adds no format or syntax meaning beyond what the enum provides, giving the baseline 3.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('press') and resource ('one allowlisted key or combination'), which clearly separates it from siblings like browser_type and browser_click. It is clear, though it never explicitly names an alternative sibling to differentiate against.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to reach for this tool versus browser_type (text entry) or browser_click, and no prerequisites such as an element needing focus. Usage is only weakly implied by the word 'press'.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_readA
Read the current page: title, URL, readable text and a bounded accessibility tree. Prefer this over a screenshot — it is structure, not pixels, and it does not spend a vision step unless you ask for the image.
| Name | Required | Description | Default |
|---|---|---|---|
| include_screenshot | No | Also return the PNG as base64. Costs one vision step from the session budget. Default false. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden and does disclose a real behavioral trait: the accessibility tree is bounded, and the vision-step cost only applies when the image is requested. It doesn't cover permissions, pagination, or truncation limits of the bounded tree, which is why it isn't a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences, front-loaded with what is returned and immediately followed by the routing rationale. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, so the description must describe return values and it does, listing title, URL, text, and the accessibility tree. 'Bounded' is left undefined, so an agent doesn't know the truncation limits, but the description is otherwise sufficient for a zero-required-param read tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the single parameter is fully documented in the schema, so the schema does the heavy lifting. The description's 'unless you ask for the image' only loosely echoes include_screenshot without adding format or default detail beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (read) and resource (current page) and enumerates exactly what comes back: title, URL, readable text, and a bounded accessibility tree. It also distinguishes itself from the screenshot path, so an agent can tell it apart from sibling browser_* tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says to prefer this over a screenshot and explains why (structure not pixels, no vision cost unless requested), which routes the agent correctly. It gives no guidance on when NOT to use it or how it relates to siblings like browser_status, so it falls short of a full 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_scrollC
Scroll the page by a nonzero pixel delta.
| Name | Required | Description | Default |
|---|---|---|---|
| delta_x | No | Horizontal pixels. Positive scrolls right. | |
| delta_y | No | Vertical pixels. Positive scrolls down. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and delivers little: it discloses only that a zero delta is invalid, but says nothing about whether the scroll is instant or animated, whether it waits for load, what state it mutates, or what it returns.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single zero-waste sentence with the key constraint front-loaded. Brevity is appropriate for a trivial operation, though the terseness leaves gaps that are penalized elsewhere.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-optional-parameter utility with no output schema, the description is minimally adequate, but an agent still lacks any statement of preconditions or post-scroll state, which matters when orchestrating a browser session.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and both deltas are documented with direction and sign conventions, so the schema does the heavy lifting. The description adds only the 'nonzero' constraint, which is a marginal but real addition.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (scroll) and resource (the page) with the delta semantics of the operation. It is clearly distinguishable from siblings like browser_navigate or browser_click, though it never explicitly contrasts itself with them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use guidance, no prerequisites (e.g., must a page be open first?), and no mention of how scroll interacts with browser_navigate or browser_read. The only inferred guidance is the 'nonzero' wording.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_statusA
Report whether the runtime is up, whether a session is open, how many vision steps remain, and the live view URL to hand to a human.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full disclosure burden. It usefully enumerates the four pieces of state returned (runtime, session, vision steps, view URL), but omits any hint about permissions, side effects, or when a session must exist before calling.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single well-formed sentence that front-loads the most important outputs (runtime and session state) before the ancillary URL detail. No wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, so the description properly enumerates the return contents, which is what an agent most needs here. For a zero-parameter, side-effect-free status probe it is essentially complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so the schema is trivially complete and there is nothing for the description to compensate for. Baseline 4 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Report') and enumerates exactly what is reported: runtime state, session presence, remaining vision steps, and the live view URL. This clearly contrasts with the action-oriented siblings (browser_open, browser_click, etc.), though it does not name a sibling explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied rather than stated: the mention of the live view URL 'to hand to a human' hints at one use case, but there is no explicit when-to-use/when-not guidance or routing against the other browser_* tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_typeA
Type text into a CSS selector, or at an x/y point. Text is sent byte for byte. Do not use this for a secret a human should enter — ask them to take over in the live view instead.
| Name | Required | Description | Default |
|---|---|---|---|
| x | No | Viewport x coordinate. | |
| y | No | Viewport y coordinate. | |
| text | Yes | Text to type. | |
| selector | No | CSS selector to type into. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It does disclose two useful traits — 'Text is sent byte for byte' (no key interpretation) and the secret-handling safety rule — but says nothing about focus behavior, error/failure results, whether existing field content is cleared, or permission requirements.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with purpose, then the byte-for-byte behavior, then the safety exclusion. No filler and nothing buried.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Covers purpose, targeting modes, and a safety caveat adequately for a 4-param, no-output-schema tool. With zero annotations, though, the behavioral profile is thin: an agent cannot tell what happens on a bad selector, whether the element is focused first, or whether text is appended or replaced.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds meaning the schema lacks: selector and x/y are presented as alternative targeting modes, implying a mutual-exclusion relationship the flat schema never states. It stops short of explaining precedence if both are supplied.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Names a specific verb (Type) and target ('CSS selector, or at an x/y point'), which is enough to distinguish it from browser_click and browser_press. It does not explicitly name the sibling it differs from (browser_press for keys), so it falls just short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit negative case and a concrete alternative: 'Do not use this for a secret a human should enter — ask them to take over in the live view instead.' That is real when-not-to-use guidance with a fallback, though the positive when-to-use case versus browser_press is left implicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
9 tool updates
v0.2.3- First observed
browser_click - First observed
browser_close - First observed
browser_navigate - First observed
browser_open - First observed
browser_press - First observed
browser_read - First observed
browser_scroll - First observed
browser_status - First observed
browser_type
TDQS
Scored across 9 tools
Each tool has a distinct purpose, but browser_open and browser_navigate could be confused since both initiate a session and navigate to a URL. The descriptions clarify the difference (open creates/reuses session, navigate goes to URL), but an agent might still misfire.
All tool names follow a consistent snake_case verb_noun pattern with the 'browser_' prefix. This is predictable and readable.
Nine tools is well-scoped for a browser automation server, covering essential actions like open, navigate, read, click, type, press, scroll, status, and close. No obvious bloat or missing core operations.
The surface covers the full browser lifecycle and common interactions (open, navigate, read, click, type, press, scroll, status, close). Minor gaps include no explicit wait or screenshot tool (though read can return an image), and no form submission or download handling, but these are workaroundable.
Maintenance
Related MCP Connectors
Hosted real Google Chrome MCP with per-user persistent state. Navigate, click, type, screenshot.
Browser MCP for logged-in tasks. Uses your Chrome — credentials stay local. Zero-token replay.
Run multi-step tasks in a real Chrome browser: persistent environments, live view, human takeover.
- TabfleetOAuthcom.tabfleet
Launch, inspect, control, and share isolated cloud browsers for your agents.
Related MCP Servers
- AlicenseNot gradedqualityBmaintenanceEnables MCP clients to control and interact with the user's real Chrome browser session, leveraging existing logins, cookies, and extensions for AI-driven automation.5MIT
- AlicenseNot gradedqualityDmaintenanceEnables MCP clients to drive a real, logged-in Chrome browser for web automation tasks like navigation, clicking, typing, and screenshotting.2 npmMIT
- AlicenseBqualityCmaintenanceEnables MCP clients to drive a real Chromium browser for automation, including navigation, JavaScript execution, CDP commands, network capture, and multi-tab control.7GPL 3.0
- AlicenseNot gradedqualityBmaintenanceEnables MCP-driven browser automation from any MCP client by pairing with a Chrome extension, translating tool calls into WebSocket commands to drive Chrome without an Electron host.3,989 npmMIT