Web Search Neo
Web Search Neo is an MCP server that gives AI agents free, API-keyless web search plus visible Chrome automation through two tools (web_info and web_action).
Web search — text search with automatic fallback across engines (Brave, DuckDuckGo, Bing, etc.), no API key required, with challenge detection and manual/solve modes.
Browser automation — open pages, fill forms, click, submit, upload files, scroll, type text, work in your already-open signed-in Chrome via a companion extension, or use isolated/temporary/persistent/attach Selenium profiles.
Page perception — accessibility outline, page text, element text, semantic element lookup (
find), page elements, locators (CSS, ref handles, piercing paths) including Shadow DOM and same-origin iframes.Diagnostics — read console logs, uncaught exceptions, network requests/status/errors, response bodies, and export HAR.
Game/input control — step-mode rendering, frozen virtual time, keyboard/pointer/touch input, pointer lock, frame-by-frame control for canvas/WebGL games.
Site checks — passive security reports graded A–F, API reports, secret scanning, active probes, performance reports, HAR export, and regression test runs.
Macros — validate, preview, and replay reusable multi-step web_action flows, including guarded two-phase consequential actions.
Tab and session management — list/attach tabs, session isolation, parallel sessions, persistent parked sessions, agent labels, and safe cleanup.
HTTP requests — direct HTTP calls with headers/body/cookies, cookie jars, redirect validation, and plain-HTTP policy.
Self-describing contract —
web_info()returns capabilities, action schemas, recipes, pitfalls, limits, and examples on demand.
Provides web search capabilities through DuckDuckGo's search engine, allowing queries with configurable result limits.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Web Search Neosearch DuckDuckGo for latest AI news and fetch the top result"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Web Search Neo is an MCP server that gives a model two tools — web_info to
look, web_action to act — and behind them four capabilities:
Capability | What it is |
Search | Text search with automatic fallback across independent engines. No paid API, no provider key. |
Your own Chrome | The already-open, signed-in browser you are looking at, driven through a local companion extension: tabs, forms, uploads, games, screenshots. |
Perception | An accessibility outline, readable text, semantic element lookup, the page console, and its HTTP traffic. |
Isolation | Separate Selenium profiles — isolated (the default for a new session), temporary, persistent, or attached. |
Site checks | For your own site: a passive security report graded A–F, load metrics, HAR export, and step-by-step regression tests. |
Brave is the default search route. DuckDuckGo, Yahoo, Bing, Mojeek, and
Startpage are available as fallbacks; an engine that keeps answering "nothing"
where another finds hits is tried last and named in unreliable_engines.
Contents
Examples — copy-paste calls for the seven things people do most
Check your own site —
security_report,perf_report,har_export,test_run· the long versionExtending Web Search Neo — plugin actions, observation topics, and search providers
Why Web Search Neo — the capability table
Connect your already-open Chrome · The bridge daemon · Bridge authentication
Search behavior · HTTP requests without a browser · CAPTCHA and challenge modes
Two-tool MCP contract · Optional agent skill · Command-line client
Related MCP server: MCP Web Tools Server
Quick start
Requirements: Python 3.10–3.13 on PATH and Google Chrome 116+ for rendered browser tools.
git clone https://github.com/NeoXider/web-search-neo.git
cd web-search-neo
python -m venv .venv
.\.venv\Scripts\Activate.ps1
python -m pip install --upgrade pip
python -m pip install -r requirements.txt
python main.pyThe server uses MCP over stdio. Keep stdout reserved for MCP messages; rotating diagnostic logs are written to %LOCALAPPDATA%\web-search-neo\logs\msp_server.log on Windows, or $XDG_STATE_HOME/web-search-neo/logs/msp_server.log — by default under ~/.local/state — elsewhere; WEB_SEARCH_NEO_LOG_FILE overrides the path. The companion bridge runs as its own process and keeps a separate rotating log at %LOCALAPPDATA%\WebSearchNeo\bridge-daemon.log, or $XDG_DATA_HOME/WebSearchNeo/bridge-daemon.log — by default ~/.local/share/WebSearchNeo/bridge-daemon.log — because two processes rotating one file collide on Windows.
For Linux/macOS activation and detailed setup/troubleshooting, see INSTALL.md.
Connect to LM Studio
Open LM Studio's MCP configuration and add:
{
"mcpServers": {
"web-search-neo": {
"command": "python",
"args": ["main.py"],
"cwd": "C:/path/to/web-search-neo"
}
}
}The Python executable is intentionally resolved through PATH, not pinned to a machine-specific absolute interpreter path. Restart or toggle the MCP server after changing the configuration. A ready-to-edit example is included in mcp_servers.json.
Examples
Every example below is the argument object of one of the two tools:
web_action takes {"actions": [...]} and performs 1–32 ordered actions;
web_info takes {"topic": ..., "params": {...}} and reads state without
changing anything. One session_id is one page — reuse it across calls.
1. Search and read a result
{
"actions": [
{"action": "search", "query": "model context protocol servers", "num": 5},
{"action": "fetch_text", "url": "https://modelcontextprotocol.io/", "max_chars": 4000}
]
}Search first, then read the page you picked. fetch_text needs no browser at
all; fetch_many reads up to 16 URLs concurrently.
2. Open a page and read it the way the agent does
{"actions":[{"action":"open","url":"https://example.com","session_id":"demo"}]}Opens the page in a fresh isolated, headless browser — the default for a new
session since 1.20. Add "profile_mode": "current" to open it as a tab of your
own Chrome instead, in the visible 🟢 AI tab group, with your logins.
{"topic":"page_outline","params":{"session_id":"demo","limit":60}}Returns an indented tree of roles, accessible names, states, ref:<epoch>:N
handles, and on-screen boxes — the page as a screen reader sees it, not as HTML.
{"topic":"find","params":{"session_id":"demo","query":"more information link","role":"link"}}Ranked matches for a plain-language query when a CSS selector would be a guess.
{"actions":[{"action":"click","selector":"a[href*='iana.org']","session_id":"demo"}]}Clicks it. click and wait also accept a ref: handle or a piercing path in
the Selenium-backed profile modes.
3. Fill in a form and submit it
{
"actions": [
{"action": "fill", "session_id": "intake",
"fields": {"#full-name": "Ada Lovelace", "#topic": "demo", "#subscribe": true},
"files": {"#attachment": "C:/docs/report.pdf"}},
{"action": "submit", "session_id": "intake",
"form_selector": "#intake-form", "submit_selector": "#submit-button"}
]
}Text fields, <select> options, checkboxes, and file inputs in one call, then a
submit that runs native browser validation. fill reports filled and a
per-selector errors map, so a partial failure is visible rather than silent.
→ Complex forms: the long version
4. Find out why a page broke
{"topic":"console","params":{"session_id":"demo","levels":["error"],"limit":20}}{"topic":"network","params":{"session_id":"demo","only_errors":true,"limit":20}}The page's own console — including uncaught exceptions with stack frames — and
every failed or 4xx/5xx request as method status type ms size url, back to and
including the navigation that opened the session: capture is armed before the
first request, so nothing has to be reloaded to be seen. Use output="json" on
network to get the request id, then read one body with the network_body
topic.
5. Play a game frame by frame
With the game already open in session game:
{
"actions": [
{"action": "render", "mode": "step", "session_id": "game"},
{"action": "input", "session_id": "game", "target_selector": "#game",
"key_actions": [{"key": "ARROW_RIGHT", "action": "hold"}]},
{"action": "step", "frames": 3, "session_id": "game"},
{"action": "input", "session_id": "game",
"key_actions": [{"key": "SPACE", "action": "tap"}]},
{"action": "release_inputs", "session_id": "game"},
{"action": "render", "mode": "normal", "session_id": "game"}
]
}In step mode the page is frozen: page time only advances when a frame is
released, so the game measures a fixed delta instead of the model's thinking
time, and a tapped key stays down for the whole frame an engine polls.
→ Playing games: the long version
6. Work in a tab you already have open
{"topic":"browser_tabs"}{"actions":[{"action":"attach_tab","tab_id":123,"session_id":"existing"}]}Lists the tabs in your Chrome with their ids and group names, then claims one
without navigating or moving it. close later releases the session and leaves
that tab exactly where it was.
7. Check your own site before a release
{"actions":[{"action":"security_report","url":"https://staging.example.com/","summary":"min"}]}Returns the grade computed with Mozilla HTTP Observatory's tests and modifiers, and the
things to fix first. Without summary the answer carries every finding with its fix, the
arithmetic behind the score, the cookies' flags and the third parties the page
loads. scope: "site" crawls the site, scope: "hosts" checks several origins (a front end and
an API on localhost, graded in development mode). → Check your own site
Where to go next
{"topic":"capabilities"}web_info() with no arguments returns the whole agent-facing contract: every
action with its required parameters, the observation topics, recipes, pitfalls,
limits, and worked examples. web_info(topic="action_schema", params={"action": "input"}) returns one full JSON Schema on demand.
Why Web Search Neo
Strength | What it means |
Free search | Uses public search routes through the maintained DDGS library; no paid search plan or API key. |
Resilient fallback | Provider health, cooldowns, bounded retries, caching, and an overall deadline prevent one challenged engine from stalling the agent. |
Your current Chrome on request | With |
Reusable authorization | List and claim existing tabs, use the current signed-in Chrome, a persistent MCP-owned profile, or a DevTools attach window. |
Authenticated companion | Server and extension prove knowledge of a machine-local secret to each other before a single command crosses the loopback bridge. |
One bridge, many clients | The companion port belongs to a standalone bridge process, not to whichever agent happens to be running, so Claude Code and LM Studio can drive the same Chrome at the same time and the badge stays |
Two-tool MCP surface | Models see only |
Self-describing contract |
|
Semantic page reading |
|
Visible failures | The |
Deterministic frames | Gated render modes freeze |
Concurrent work | Search, HTTP fetches, and independent browser sessions run outside the MCP event loop; eight browser sessions work in parallel by default (the companion popup and |
Connect your already-open Chrome
profile_mode="current" drives this Chrome. Since 1.20 it is chosen explicitly
(or implied by current_tab_id); a new session opened without a mode gets an
isolated headless browser instead — see the migration note.
current fails with a clear setup error when the companion isn't connected; it
does not silently open a different browser.
An agent prepares the bundled companion through the compact MCP contract:
{"actions":[{"action":"setup_current_chrome"}]}That call opens no page and touches no browsing data. It writes the shared secret
into chrome-extension/bridge-token.js, checks the bundled build against
whatever is connected, and returns manual_steps — the clicks that are left,
with the absolute path of the folder to pick — plus token_ready and
token_file. The one thing it can change is the companion itself: a connected
build older than the bundled one is asked to reload itself, reported as
self_update (done, unsupported, or timeout) with replaced_version.
Chrome does not let a program add an unpacked extension to a browser that is
already open: the installed set is signed inside Secure Preferences, policy
installs need a packed CRX behind an update URL, and a Chrome started without a
DevTools port cannot be given one later. So three steps stay with the user:
Open
chrome://extensions.Switch on Developer mode, then choose Load unpacked.
Select this repository's
chrome-extensionfolder.
Keep Web Search Neo Companion enabled. Click its toolbar icon to open the
companion panel: it shows the bridge state and controlled-tab count, and offers
Reconnect, Release tabs, a GitHub release/version check, a repository link,
and a persistent on/off switch. Turning the switch
off closes the bridge and detaches every controlled tab. Its toolbar badge reads
ON while the companion is connected and authenticated to the bridge daemon,
which is between agent calls too, not only while one is running. Nobody has to
create or copy a token: it is written whenever the bridge comes up and on every
setup_current_chrome call.
Earlier revisions tried to perform those clicks through Windows UI Automation.
That code is gone: it depended on the interface language, on which window had
focus, and on a folder picker the automation backend does not enumerate. There
is no automatic substitute. If you would rather not install an extension at all,
profile_mode="temporary" and profile_mode="persistent" drive a Selenium
browser that needs no companion.
After the server is updated, the running companion keeps the code it loaded.
The server now reloads a connected companion whose manifest version differs from
its folder by itself (on browser_tabs and before open; both answers report it
as companion_refresh), but only while no agent is driving a tab - a reload
detaches every debugger session at once - at most once per five minutes across
all server processes, and never again for a build whose reload did not take (the
answer then carries the manual steps). A companion that differs only in its code
hash (same version, an edited checkout) is reported, not reloaded. When the bridge has just (re)started,
the companion reconnects on its own retry schedule, which backs off to about a
minute: browser_tabs then answers connected: false with a companion_note,
and wait_seconds (up to 90) waits it out. If it still does not connect, the
companion popup's Reconnect (Restart companion when its service worker has
stopped) is the one action; Reload on its card at chrome://extensions is the
fallback.
The bundled companion is version 1.24.0. Chrome does not refresh an unpacked
extension by itself, but from 1.3.1 the server does it instead: the worker
understands a runtime.reload command, and setup_current_chrome sends it
whenever the connected build is older than the bundled one. That only works for
a companion that can still authenticate, though: 1.15.0 changed the bridge
handshake (protocol 2), so an older build is refused before the command can reach
it. Upgrading onto 1.15.0 therefore costs one click, just as upgrading onto
1.3.1 did; the very first install still needs the three steps above, which no
program can perform.
So press Reload on the companion card at chrome://extensions once, after the
pull that crosses 1.15.0 (or 1.3.1). Skipping it leaves a build that cannot
authenticate:
a service worker from 1.2.0 or older sends no token, and one older than 1.15.0 still speaks bridge protocol 1, so the bridge closes it with code 1008 and a reason that says to update and reload the extension on
chrome://extensions;the badge then stays
OFFand the worker retries on a slow ladder, from about ten seconds up to two minutes;the bridge daemon records the rejection in
%LOCALAPPDATA%\WebSearchNeo\bridge-daemon.log, not inmsp_server.log;a 1.3.0 worker also prints the close code and reason in its own service-worker console; an older one does not, so on a stale install the badge and that daemon log are the signal.
The bridge accepts only the fixed bundled extension ID, binds only to loopback, and never reads Chrome profile files, cookies, or saved passwords directly. The extension uses Chrome's standard tabs, tab groups, and debugger APIs, plus storage for its own session state and alarms for the long waits in its reconnect backoff, with the permissions shown by Chrome.
A badge reading OFF means this companion is not connected and authenticated, which is usually — but not only — because nothing is listening on the bridge port: no daemon has been started since the machine booted, the last one exited after its idle window, or it was stopped with --bridge --stop. That is the normal state of a machine on which no MCP server has run yet, not a fault — but it is no longer the state between two agent calls, which is what it was when the port lived inside the MCP server. The badge also reads OFF in two cases where the port is very much held: when the daemon refused this companion's token, and when a companion in a second Chrome profile authenticated later and displaced it. So OFF is not by itself a reason to start a daemon — read web_info(topic="browser_status") and bridge-daemon.log before concluding the port is free. While nothing listens, the worker retries with an exponential backoff — about 1.5 seconds after the first failure, doubling to a ceiling of one minute, and dropping back to the floor the moment a handshake verifies. Chrome suspends an idle service worker after roughly thirty seconds, which is why the longer waits are chrome.alarms rather than timers: a timer that long would die with the worker. Chrome also logs every refused attempt as an extension error of ours and an extension cannot suppress that, so a red Errors button on the card of an idle machine is expected; before 1.3.1 the worker retried every two seconds and could fill that page with hundreds of identical ERR_CONNECTION_REFUSED lines. Starting the browser or reloading the extension resets the schedule, and the popup's Reconnect button retries immediately.
Keep the companion enabled in one Chrome profile at a time. The bridge holds exactly one companion connection and the most recent authenticated one wins, so a second profile with the companion enabled quietly takes the agent's tabs with it.
Dynamic companion widget
The toolbar popup (also on Alt+Shift+N, its suggested shortcut) is where the
companion is switched on, watched, and configured. It is a compact, icon-first
widget: one live status line with a pulse marker — Connected, Connecting,
Waiting plus the retry countdown, Disabled, or an error state — then three
figures: controlled tabs, the parallel-session limit against its ceiling with a
small capacity meter, and the bridge endpoint. Icon buttons carry accessible
names (aria-label/title): release check, repository link, apply/reset port,
apply/reset session limit. A version chip and a GitHub release check round off
the top; Settings holds the port and session limit. Everything renders from the
status the service worker reports each second — connected, connecting,
waiting, disabled, or error — and nothing else: no fabricated connection state,
no secrets in the UI. Opening the panel is read-only by construction: it never
navigates, submits, closes, or releases anything on its own.

While the bridge is merely unreachable, the widget says so and shows when the next attempt lands, which is what tells a stalled bridge apart from a deliberate backoff:

Opening the panel — by click or by shortcut — performs no navigation and controls nothing; it only reads status.
Under Settings:
Setting | What it does |
Bridge port | The loopback port the companion dials, 1024–65535. It must match |
Protocol access | States what the companion will forward. It is a list, not a passthrough: only the DevTools methods this MCP issues reach the tab. |
That allowlist is the answer to the obvious question about a companion holding
debugger on your signed-in tabs. Authentication proves a peer knows a
machine-local secret; it does not make that peer trusted with all of Chrome's
protocol. Any method outside the list is refused inside the extension, and a test
compares the list against the server's own call sites, so the two cannot drift
apart without the suite failing.
The bridge daemon
The listener the companion dials is not inside the MCP server any more. It is a
standalone process — bridge_daemon.py — that owns 127.0.0.1:8765, holds the
single connection to the extension, and relays commands for any number of local
MCP clients. ChromeBridge is now a client of it. Its public surface did not
change and neither did the frames the extension exchanges, so nothing an agent
calls looks any different.
Three things follow, and they are why the listener moved out:
The badge stays
ONbetween calls. The port used to exist only while an agent was running, so the companion spent most of the day disconnected and reconnected on its own backoff schedule whenever a server appeared.Two MCP clients can share one Chrome. Claude Code and LM Studio can now run at the same time; the second server to start used to fail to bind the port with
OSError: [Errno 10048]and silently loseprofile_mode: "current".The first call after a quiet stretch is not spent waiting. A freshly started server no longer has to sit out whatever reconnect delay the extension had climbed to.
An MCP server starts the daemon as it comes up, if nothing is listening yet, so the companion is connected before the first action rather than during it. It is started detached — with no console window on Windows, in its own session elsewhere, and never inheriting the stdio that carries MCP — and it outlives the server that started it. Nothing is registered for autostart: no service, no scheduled task, no login item. It exists once an MCP server has run, or once you have run one yourself, and it goes away again on its own.
On Windows the launcher deliberately bypasses a uv virtual-environment
python.exe/pythonw.exe redirector and starts the base pythonw.exe with
CREATE_NO_WINDOW, while preserving the venv through __PYVENV_LAUNCHER__.
This matters because a redirector can start the real console interpreter after
the original creation flags are gone; combining DETACHED_PROCESS with
CREATE_NO_WINDOW does not fix it, because Windows ignores the latter flag in
that combination.
Two entry points are yours:
python main.py --bridge # run the daemon in the foreground
python main.py --bridge --stop # ask a running daemon to exitFor a companion configured to a non-default port, set the same port in the
daemon's environment before starting it (or in the MCP server's env block):
$env:WEB_SEARCH_NEO_BRIDGE_PORT = "18765"; .\.venv\Scripts\python.exe -m web_search_neo.main --bridgeFor a hands-off setup, scripts/bridge_autostart.bat [port] launches the daemon
through pythonw (no console window), and running
scripts/install_bridge_autostart.bat [port] once drops that launcher into the
Windows Startup folder so the bridge is up on every logon. Delete the generated
startup file to undo. The launcher sets WEB_SEARCH_NEO_BRIDGE_IDLE_SECONDS=0,
so a logon-started bridge does not idle-exit before the next browser session.
Quiet on Windows: nothing the agent runs may pop a console onto your screen. The
generated MCP configuration points at pythonw (no console, same stdio pipes) -
and past any venv launcher shim straight at the base interpreter, since a shim
starts the real one as a visible child - the daemon starts with
CREATE_NO_WINDOW, chromedriver is launched hidden, console MCP servers can
run behind scripts/quiet_stdio.py, and the test suite spawns
node and python windowless too. If a console window still
appears, its title names the owner: node.exe is a dev/test harness, python.exe
is an MCP server started with a console interpreter (regenerate the config with
python scripts/make_mcp_config.py), and anything else belongs to whoever
started it — not to this project.
After a reboot nothing needs clicking: the bridge is already listening (logon
launcher), and opening Chrome reconnects the companion by itself — its badge
turns ON within seconds on the same backoff it uses after any outage. If it
reads OFF, the daemon is not up yet or the token the loaded copy reads is
stale; run the MCP once (it starts a daemon on demand) and press Reconnect
in the popup, or Reload the card first when the copy Chrome loads is not
the folder setup_current_chrome reports.
--bridge exits quietly if another daemon already owns the port, because that one
serves just as well. It prints nothing while it runs either: the daemon logs to
%LOCALAPPDATA%\WebSearchNeo\bridge-daemon.log, foreground or not, and never to
msp_server.log, because two processes rotating one file collide on Windows.
--bridge --stop is the one that talks back — it says whether it stopped a daemon
or found none — and it never starts one.
Left with neither a companion nor a client attached, the daemon exits after
fifteen minutes. A connected companion on its own keeps it alive indefinitely —
that is the case where the badge stays on all day. Set
WEB_SEARCH_NEO_BRIDGE_IDLE_SECONDS to change the window, or to 0 to disable it;
a daemon inherits the environment of the server that spawned it, so setting it in
the MCP client configuration is enough.
The daemon reports its version during the client handshake, and versions are
ranked by component rather than compared as text — 1.3.10 is newer than
1.3.9, which the string order gets backwards. Only a strictly older daemon
is asked to exit and replaced with one of this server's own, so a daemon started
before a git pull cannot keep serving the old code.
A daemon that is newer than the server talking to it is used, not replaced.
That is the ordinary state of an updated machine, not a fault: an MCP server
keeps running the revision it imported at startup, so after a git pull every
daemon started from main.py is ahead of the server that started it. Refusing
that was unfixable from the user's side — the server evicted the newer daemon,
started a replacement from the same updated file, was handed the same newer
version again, and after two rounds gave up with an error blaming a second
checkout that did not exist. What matters for sharing a companion is the bridge
protocol, not the release number, and that is now the only skew that stops the
server: the message names the protocol each side speaks instead of guessing at a
cause.
Replacing an older daemon still gives up after two attempts, and the error names
both versions — and, when this process is running code the checkout no longer
has, says so and tells you to restart the MCP server rather than hunt for another
clone. The error is latched so that later calls do not restart the tug-of-war,
but the latch expires after ten seconds
(WEB_SEARCH_NEO_BRIDGE_RECHECK_SECONDS) and the look that follows is passive —
it starts no daemon and evicts none — so the state clears itself once the other
side is stopped, instead of surviving until someone restarts the server by hand.
web_info(topic="browser_status") reports the link under current_chrome.daemon:
linked says whether this server currently holds one, version and pid come
from the last handshake this server completed on that port — whether or not it
went on to link — and server_version is this server's own, so a skew is visible
as two numbers side by side. version and pid are cleared when the link ends;
they used to survive it, so a startup_error about the daemon now on the port
could sit next to the version of a daemon that had already exited.
current_chrome.connected still means what it always did — the companion itself
is attached — and a daemon.linked: true with connected: false is exactly the
case of a running bridge with Chrome closed.
Bridge authentication
Loopback is not a trust boundary: every process running as the same user can open
127.0.0.1:8765, an Origin header stops web pages but not local programs, and the
extension ID is derived from the public key in the manifest, so it is not a secret
either. Before this release, a local process that reached the port first received the
full companion protocol, including cdp.send on any signed-in tab. Both sides now prove
knowledge of a machine-local secret before a single command is executed:
a 32-byte random token — 64 hex characters — is minted on first use, by the daemon as it starts or by a server as it connects, and kept in
%LOCALAPPDATA%\WebSearchNeo\bridge-tokenon Windows, or in$XDG_DATA_HOME/WebSearchNeo/bridge-token— by default~/.local/share/WebSearchNeo/bridge-token;the same secret is copied into
chrome-extension/bridge-token.jsbefore Chrome is asked to load the folder, so setup stays hands-free. That file is listed in.gitignoreand is never committed or shared between machines;the token itself never crosses the socket (bridge protocol 2). The peer's
hellocarries aroleand a fresh nonce; the daemon answers with its own nonce and anHMAC-SHA256proof over both, and only once the peer has verified that proof — with WebCrypto in the companion — does it send its own HMAC back. The proofs carry direction labels, so neither can be replayed as the other. A squatter on the port learns nothing from the companion, and a peer that sends anything else first is closed;an MCP server authenticates the same way with
role: "client"; theextensionrole additionally requires the companion'sOrigin. Being a relay is not a way around the token: a local process without it is closed on both roles alike;the token file is written atomically and restricted to the current user —
0600on POSIX, anicaclsACL on Windows — and a new secret is minted only when the file is missing. The daemon caps connections (64 open, 16 mid-handshake) and refuses a command for a tab another connected client has claimed.
A newly authenticated connection replaces the previous one, so a companion whose service worker Chrome had suspended reclaims the bridge on reconnect instead of finding it held by a stale socket.
The honest limit: the secret is a file owned by the user account, so any process running
as that same user can read it and impersonate either side. This closes the "whoever binds
the port first owns the browser" hole; it is not protection against malware already
running as you. The real fix is Chrome Native Messaging, where Chrome launches the server
itself and no port is listened on at all — it is tracked in TODO.md. The
companion forwards only the DevTools methods on its cdp.send allowlist.
The daemon did not move that boundary, but it did change how long the door stands open.
The listener used to exist only while an agent was running — minutes a day on a normal
machine, which is the assumption the paragraph above was written under. Now it is held
by a process that outlives every agent call and stays up for as long as Chrome keeps the
companion attached, which in practice means the whole working day. Nothing about the
authentication moved: the bind is still loopback-only, both roles still have to present
the token, and a process running as you could always read the token file — same-user
processes were never excluded by any of it. What is larger is simply the share of the day
during which such a process finds something listening. Close it deliberately with
python main.py --bridge --stop, or quit Chrome and let it close itself: a daemon with
neither a companion nor a client attached exits after fifteen minutes
(WEB_SEARCH_NEO_BRIDGE_IDLE_SECONDS, 0 disables the timer).
Search behavior
The normal agent call is:
{
"actions": [{
"action": "search",
"query": "best local-first MCP tools",
"num": 5,
"engine": "brave",
"fallback": true,
"challenge_mode": "fallback"
}]
}Send that object to web_action. Search state is readable separately:
{"topic":"search_status","params":{"check_live":true}}It reports configured engines, current live availability, latency, cooldown state, and detected challenges. A live probe is a diagnostic: a provider that fails during the probe is not pushed into the cooldown used by real searches, so checking status can no longer degrade the next search. Status checks are cached for five minutes; search results are cached for two minutes.
result_status says how good an answer is: ok, partial (fewer hits than
asked), empty, or off_topic — no returned row mentions any word of the
query (Bing does this for some queries, Cyrillic ones reliably). An off-topic
answer is treated like an empty one: the next engine is asked, the skipped
engine is listed in engines_off_topic, and off-topic rows are returned only
when nothing better exists, with a note. Bing's bing.com/ck/a tracking
links are unwrapped to the real page URL.
HTTP requests without a browser
http_request sends any method with headers, query, body or body_json and
never opens a browser; a 4xx/5xx comes back with its status and the server's error
body, and every Set-Cookie is listed apart in set_cookies. By default no cookie
rides from one call to the next. Since 1.20 http_session names a cookie jar that
does:
{"actions":[
{"action":"http_request","method":"POST","url":"http://127.0.0.1:8000/login",
"body_json":{"user":"demo","password":"demo"},"http_session":"api","agent_label":"qa"},
{"action":"http_request","url":"http://127.0.0.1:8000/me","http_session":"api","agent_label":"qa"}
]}The jar belongs to
(agent_label, http_session): another agent using the same name never sees it. Jars live in the server process only, expire after an idleWEB_SEARCH_NEO_HTTP_SESSION_TTL(default 1800 s, at least 60) and are capped at 32, the least recently used dropped first.Cookies are matched by domain, path,
Secureand expiry; cookies set on redirect hops are kept, a cross-origin hop still strips credentials, and a cookie the server deletes leaves the jar.http_session_clear: trueempties the jar before the request.The answer's
http_sessionblock listssent_cookiesandreceived_cookiesper hop with their flags. Values are redacted there and inset_cookies/headersunlessshow_values: true. ACookieheader andhttp_sessionin one call are refused.
CAPTCHA and challenge modes
challenge_mode="fallback"is the default. A challenged provider is skipped immediately and the search continues through another route.challenge_mode="manual"opens visible Chrome and waits up tomanual_timeout_seconds=180. If you complete the challenge, the agent receives the open browser session; otherwise the window closes and fallback continues.
Detection answers "is a challenge in the way", not "does this page contain a captcha". It looks for a rendered provider widget — reCAPTCHA, hCaptcha, Cloudflare Turnstile, Yandex SmartCaptcha, PerimeterX, DataDome or AWS WAF — at least 20x20 pixels, walking open shadow roots and same-origin frames rather than only the top document, which is where the last three and anything one frame down were being missed. A bare data-sitekey counts only when the element or a nested frame says captcha, so a chat widget carrying one is not a challenge. Wording counts only in a page heading or in a short interstitial. Every page summary carries challenge_detected, challenge_type, challenge_evidence and captcha_widgets — the last lists every captcha seen, blocking or not, so an article with an hCaptcha in its comment form is reported without being called challenged.
A challenge with no box is the other half of this. An invisible Turnstile renders no picture and no checkbox — the whole widget is a container in the DOM and a hidden field its token would go into — so a visibility test walked straight past it while the site's own submit handler waited for a token that was never coming: the button sits on "Submitting…", no request leaves the page, and the console stays empty. Every page summary therefore also carries invisible_challenge_pending, and with it invisible_challenge naming the vendor, the state (token_empty when the hidden cf-turnstile-response, g-recaptcha-response, h-captcha-response or smart-token field is empty; widget_hidden when the widget rendered before its field existed), the evidence and what to do about it. It gates the form rather than the page, so challenge_detected stays false — but captcha with op=detect reports it, mode='wait' no longer calls it resolved, and a click that made no network request at all while it is pending comes back with submit_blocked_by_challenge and a reason.
The captcha web action can also hand a widget's sitekey to a third-party solving service (2captcha-compatible, WEB_SEARCH_NEO_CAPTCHA_HOST), but only on request: mode='solve', or mode='auto' with WEB_SEARCH_NEO_CAPTCHA_AUTO_SOLVE=1, and always with WEB_SEARCH_NEO_CAPTCHA_KEY set, because every solve costs money. By default auto waits for a human. The returned token is applied only if the page URL has not changed, and the result reports success only when it was. Search challenge_mode never calls a service.
Current Chrome automation
Yes — the agent can work in your normal already-open Chrome while you watch it. New pages go into a visible tab group named 🟢 AI by default — the mascot is there so agent tabs are obvious in a crowded window, and the companion colours the group purple when it is the one creating it:
{
"actions": [{
"action": "open",
"url": "https://example.com",
"session_id": "demo"
}]
}From there, inspect with page_outline, page_text, find, or page_elements,
then send ordered fill, upload, click, scroll, and submit actions through
web_action using the same session_id. scroll defaults to the viewport centre,
uses positive delta_y for down and negative for up, and returns before/after page
scroll metrics. Screenshots are returned by the screenshot info topic: its default
is the actual current viewport, mode="full_page" captures the document, and
mode="region" captures an exact x/y/width/height CSS-pixel rectangle
without resizing Chrome. The screenshot action takes the same modes inside a
web_action batch and writes the PNG under the download directory
(WEB_SEARCH_NEO_DOWNLOAD_DIR, default ./downloads), answering with
saved_to, size_bytes and the image size; a path must end in .png, stay
inside that directory, and replaces an existing file only with overwrite: true
(all checked before the capture). wait with seconds is a plain delay that
needs no selector and no open session.
type_text has two modes. The default mode="insert" sends one CDP
Input.insertText, which React-controlled inputs see as a real edit. The focused
element is found through open shadow roots and same-origin frames (web-component
editors, TinyMCE, a game inside an iframe); a cross-origin frame counts as
unknown and the text is sent, while nothing focused or a read-only control is
refused instead of dropping the text. mode="keys" presses one key per
character (any script, Cyrillic included, at most 500) for canvas apps such as
Unity WebGL, which build text from key events and never receive an insertText.
Without a selector the keys go only to an editable control, a canvas, an element
with an explicit tabindex, or a frame; a focused button or link (Enter or Space
would activate it) or the bare page is refused, and a custom element that hides
its focus in a closed shadow root counts as unknown and is allowed. Keys into a
tabindex element are meant for canvas and game surfaces and come with a
keys_warning: on an ordinary site with keyboard shortcuts each character may
trigger one, so pass a selector for a text field there. A focused <canvas>
switches to keys by itself. fill writes text through the browser's input channel and,
when a React-style controlled input's value tracker still missed the edit,
raises one input/change event so onChange runs (framework_resynced);
typing=true stays for masked and per-keystroke inputs.
Every selector click reports what it did. verified is true when the URL or
title changed, the clicked element's own subtree or its ancestors mutated (a
ticking clock elsewhere on the page does not count, nor do the agent's own
overlays), focus moved somewhere other than the clicked element, or the element
left the document; false only for a top-document target that showed none of
that; null when it cannot be measured - a frame_selector, ref or >>>
(shadow) target. post_state carries the raw counts. Neither false nor null is
a reason to click again: the first click may already have sent a form or a
payment, so read the page first. page_elements marks a fragile selector (an
nth-of-type chain or a generated React id) with stable_selector: false and a
suggested_locator of role and accessible name for click, and surfaces
data-testid hooks.
Chrome sometimes replaces the tab under a page itself - a prerendered or
instant navigation, a restored discarded tab - and hands it a new id
(chrome.tabs.onReplaced). Companion 1.18.4 records exactly those replacements,
redirects commands addressed to the old id, and says so on the answer; the
server then moves the session (and its tab claim) to the new tab and keeps its
ownership, since it is the same page. If another session or agent already drives
the new tab, the session is dropped instead. Nothing looser is ever followed: a
tab the lost one opened (a popup, a payment window) or a tab on the same URL is
not the agent's. Read-only
topics and wait/screenshot are repeated once on the new tab; every other
step fails with a SessionTabFollowed error that says so, and a tab that is
really gone still drops the session.
open accepts persist: true (with an explicit session_id - never default).
A new session must name its mode with it: profile_mode: "current" parks one tab of
the user's Chrome, "isolated" or "temporary" the server's own browser (below);
without a mode the call is refused, and a session already open or parked keeps its
own mode. In current Chrome it parks one tab. Only tabs the server opens can be parked: a tab
claimed with attach_tab is the user's, and attach_tab has no persist. Such a
session survives the MCP client process: at exit its tab is detached and left
open, and a record in the per-user state directory keeps its tab id, Chrome run,
agent group and redacted URL; while the session is in use the record is kept
fresh, and expiry never touches a live session. A later client continues it only
explicitly, with {"action": "reattach", "session_id": ...} (or open with
persist: true, or attach_tab on that same tab); a plain call never picks it
up. Re-attaching is refused with success: false - the record dropped, the tab
left alone - unless the tab is provably the same one: the same Chrome run, still
in the agent tab group, still on the recorded origin, and not driven by another
client. A server tab whose only change is the site it shows comes back as
left_open_tab, so it can be closed rather than forgotten. An explicit close
of a parked session closes its tab and reports it under retired_parked;
records expire after WEB_SEARCH_NEO_PARKED_SESSION_TTL (default 24 hours, at
least 60 seconds) and their tabs are closed the same way. browser_status lists
parked_sessions. attach is accepted as an alias of attach_tab.
Since 1.20 persist: true also works with profile_mode: "temporary" or "isolated":
the whole browser the server launched is parked instead of one tab. Chrome and
chromedriver are started so they can outlive the server; at exit the session is
recorded (chromedriver's URL and session, the pids with their start stamps, the
profile, the window it drove) instead of quit, and a later server continues it
with reattach or open + persist: true, on the same page and cookies. A record
is never taken over while another server holds that session live, and nothing is
killed on a pid whose start stamp no longer matches. Every record has a small
detached watchdog that retires the browser (quits it, removes the profile) when
WEB_SEARCH_NEO_PARKED_SESSION_TTL runs out while it is parked, or when the
server that held it died without parking it; close retires it at once. When the
MCP client runs the server inside a Windows job object that ends every child with
the client, the open answer carries persist_warning: the browser then survives
a server restart under the same client but not the client itself. Persistent and
attach browsers are never parked.
New tabs and popups of a Selenium session (temporary, isolated, persistent,
attach) are visible since 1.20. When a click, click_text, submit, pointer,
run_script, input, press_keys, type_text, fill, scroll or wait makes
the page open a window (target="_blank", window.open), the answer carries
new_tabs with each window's handle, URL, title and opener; the session keeps
driving the tab it was on. follow_new_tab: true on click or click_text moves it
to the new window, and the tabs action lists (op: "list"), switches to ("switch",
handle or index) or closes ("close") the session's windows. A popup that
closes itself hands the session back to the window that opened it
(tab_closed_by_page). The last window is never closed by tabs - close the
session - and in an attached browser only windows its pages opened during this
session can be closed; closing the session closes them too. In the user's own
Chrome (current) new tabs stay the user's: browser_tabs, attach_tab,
close_tabs.
A domain filter on cookies means that domain and its subdomains, never a
substring, for get and clear alike. op: "clear" needs a domain - a cookie
name alone exists on every site - and deletes exactly the matching cookies
(partitioned CHIPS cookies with their partition), reporting what a fresh read
shows as gone and anything left as not_deleted. A domain without a dot or a
public suffix (com, co.uk, github.io) would match every site under it and
needs confirm_clear_all: true too. Chrome's own clear has no
filter at all and wipes every cookie of the profile - in current Chrome, every
login of the user - so clearing without a domain is refused unless
confirm_clear_all: true says that is really meant.
A session tracks whether it owns its tab, and never navigates one it borrowed.
An open on a session that claimed a tab through attach_tab opens the agent's
own tab in the 🟢 AI group and hands the borrowed one back — the debugger
detaches, and the page the user was reading is neither closed nor navigated. The
freed tab id comes back as left_claimed_tab. close removes a tab that the
agent opened itself, and leaves a claimed tab open and detached, so the group no
longer accumulates abandoned pages; pass close_tab: true to close a claimed one
anyway.
Tabs the agent never opened are closed with close_tabs, naming the ids from
web_info(topic="browser_tabs"). There is no close-everything switch: closing a
tab cannot be undone, so each one is named. Two kinds are skipped rather than
closed — pinned tabs, which are the set a person keeps on purpose, and tabs
another agent is driving, where closing one would pull the page out from under a
running session. include_pinned and include_claimed override each, and every
skip says which rule it hit. Whether a tab actually went is decided against a
fresh tab list rather than against Chrome's acknowledgement, because
chrome.tabs.remove reports failure for a tab the user had already closed by
hand; a tab that was already gone is reported as already_gone, not as an error.
A session sitting on a closed tab is forgotten and its claim released, so a later
action cannot dispatch at a tab id Chrome has since reused. Both close and close_all report what they
could not release — a tab that would not close, a debugger that would not detach
— instead of answering closed: true over the top of it; closing a session that
was never open stays a no-op and says so in note. A tab the user closed by hand
frees its session slot the next time the session cap is reached, so the
server does not refuse to open a page because of tabs that no longer exist. close_all — and the
shutdown hook that runs it — applies the same rule instead of leaving every
self-opened tab behind. close_all closes only the sessions carrying your own
agent_label unless you pass scope="all", and always reports kept_sessions
so a skipped session is never silent. The MCP doesn't read cookies or passwords; the page
simply continues using the authorization already present in Chrome.
Teardown is not unconditional, and the exception is the point of it. Tab ids
restart with Chrome, so a session that outlived a restart holds a number that now
names somebody else's tab. Every teardown path — close, close_all, the
shutdown hook, the cap sweep — asks first whether the session's browser run is
still the one it was opened in, and if it is not, sends nothing at all and forgets
the session. close reports that as browser_gone: true with a note;
close_all lists the affected ids under browser_gone and still answers
closed_all: true, because nothing of ours was left to leak. Both keys are absent
on the ordinary path. The claim is not released either: the daemon dropped its
whole registry when the run changed, so a release aimed at the old id could only
hit a claim made since. Before this, closing such a session sent tabs.remove for
the stale id to the new browser, closed a tab of the user's, and reported success.
browser_status answers the same way rather than describing a stranger's tab: a
session whose Chrome is gone is dropped and reported as session_open: false,
browser_gone: true, with a next saying to open the page again. It used to read
a page summary off whatever tab had inherited the id and answer session_open: true — the same identity bug, in the topic an agent uses to check for it.
If the companion isn't ready, read web_info(topic="browser_status") and run the
setup_current_chrome action for the exact steps. profile_mode="auto" falls
back to a separate headless Selenium session, so it does not raise another window.
Chrome profile modes
Mode | Authorization and lifetime | Best for |
| Disposable separate Selenium browser with its own profile and storage, headless by default; per-session | Tests, audits, anything that should not see the user's logins. |
| Companion extension controls the user's open Chrome; chosen explicitly (or by | Authorized sites, work alongside the user, existing tabs. |
| Prefer | Portable background clients. |
| Clean disposable profile, headless by default; cookies disappear when the session closes. | Search, scraping, isolated tests. |
| MCP owns a durable profile under | Repeated automation with a separate signed-in profile. |
| MCP connects to a Chrome process that you started with a DevTools port and does not close it on detach. | Watching the agent work in an already authorized managed Chrome window. |
Who opens and closes Chrome
open with temporary, isolated, or persistent starts an MCP-owned Chrome
through Selenium Manager, and close / close_all quits that process again —
the MCP genuinely opens and closes its own browser. attach connects to a
Chrome you started yourself (for example with
scripts\start_managed_chrome.ps1) and only detaches from it. current drives
the Chrome you already have open: the MCP never launches it and never quits it,
it only opens tabs the agent asked for and closes those same tabs back, while a
tab claimed with attach_tab is always handed back open.
Working in the same Chrome as the user
In current mode the agent takes no focus. Its tabs open in the background, in
an existing window — the group's own window when there is one, otherwise any
window that is neither focused nor minimized. It does not open a window of its
own, because on Windows a new window raises itself into the taskbar; the one
exception is a Chrome left running with no window at all, where a tab has nowhere
else to go. Navigation and keyboard input do not activate a tab or focus its OS
window either. Only web_action with the explicit show action requests the
foreground; it never minimizes, maximizes, restores, or resizes a window.
Chrome starves a tab nobody is looking at, which would make that useless: a
hidden tab gets no requestAnimationFrame callbacks at all, timers are clamped
to a second and then to a minute, and — measured on Chrome 151 — input dispatched
into a tab that has never been shown is silently dropped. So the companion
turns on focus emulation for every tab it drives, which restores all three
(49 fps, 4.5 ms timers, input delivered). The page believes it is focused and
visible while it is not, and stops believing it the moment the debugger detaches.
Two consequences worth knowing. A targeted keyboard action can change DOM focus
inside the controlled background page, but it does not take OS focus or change the
active user tab. Since Companion 1.19.0, viewport screenshots capture one fresh PNG
video frame with an 8-second frame deadline, then stop the owned recording. This
does not activate tabs, restore windows, resize the viewport, or alter emulation.
An existing recording or overlapping capture is refused. Full-page and region
screenshots still use Chrome's surface path (up to 45 seconds), which can stall
in an obscured window; use viewport capture or DOM/text instead. A timeout never
authorizes show or an automatic foreground retry. Reload the Companion and
reconnect the MCP server after updating to pick up both halves of this change.
Several agents can drive one Chrome at once, so tab ownership lives in the bridge
daemon rather than in any one server process: every tab an agent opens or claims
is registered there, an attach_tab for a tab another agent already drives is
refused with who holds it and for how long, and a claim is dropped the moment its
client's socket does — including an abrupt one, so a crashed agent cannot strand
a tab. A client that reconnects re-asserts its claims, and gives up any it lost.
When there is no daemon to ask, nothing is guarding the browser and the claim is
skipped rather than failing closed.
That register is also readable, which it had to become: a refusal arrives after an
agent has already chosen a tab, and the list it chose from said nothing. Every
client is told who else is on the bridge — browser_status reports
current_chrome.daemon.clients and a claims list naming each held tab, its
holder and whether it is ours — and browser_tabs marks a tab another agent is
driving with driven_by, so the choice is informed rather than corrected.
Agents inside one MCP server are not separated by any of that. Subagents of a
single run share one server process, so session_id is the only thing between
them, and two that both leave it at its default "default" are handed the same
session and the same tab: each one's navigation and typing lands in the other's
page. Nothing here can arbitrate — an MCP call carries no caller identity — so
the server does the one useful thing instead of pretending otherwise and says so:
sessions_in_use lists the sessions another caller is inside right now, and
shared_session with a warning appears on browser_status and on open when
this session is one of them. Give each agent its own session_id and they get a
tab each. The session cap is per process, so parallel agents share it too; it
defaults to eight, the companion popup carries a user's own number to the server,
and WEB_SEARCH_NEO_MAX_SESSIONS in the server's environment overrides both. The
refusal at the cap names the setting and lists every holder - session, agent, tab,
age, idle time, busy flag and last URL. Sessions idle past
WEB_SEARCH_NEO_SESSION_IDLE_TTL (default 30 minutes) are reaped when the cap is
hit, and close_all with scope="all" and idle_for_seconds releases orphaned
slots explicitly without touching anything in use.
Pass agent_label on open or attach_tab and the session records who opened
it. Nothing is refused without one — it is a courtesy, not a credential — but two
things become possible with it. browser_status lists every session in the
server with its owner, tab id and group, the page it was last seen on, when it
was created and last used, whether another thread is inside it, and N of M
occupancy; capabilities carries the same occupancy under limits. And
close_all defaults to scope="mine": it closes the sessions with your label
(or, with no label, the unlabelled ones), lists what it left standing under
kept_sessions, and only ends everyone's work when asked with scope="all".
Process exit still closes everything, because at that point nobody is left to own
a session.
In profile_mode="current" the label is also visible in the tab strip: open
and attach_tab prefix the tab's title with [agent_label] (or
[session_id] without a label), kept across navigations and reloads, while
page scripts and every read topic keep seeing the unlabelled title. Headless,
persistent, and attach sessions are never labelled. Pass label_tab=false, or
set WEB_SEARCH_NEO_LABEL_TABS=0 in the server's environment, to turn it off.
The label says whose tab it is; three further signals say whether anything is
happening in it. Every action marks the page it ran in: the tab's favicon gets
a small semi-transparent slime in its corner — awake and green while an agent
is working there, sleepy and amber for five minutes after its last action — on
top of the site's own icon, which comes back clean afterwards (a favicon the
page cannot read cross-origin is left untouched) — so a glance at the
tab strip separates the tab being driven right now from the one abandoned an
hour ago. In the page itself, the element the action touched flashes for a
quarter of a second, in red when the action was refused, which turns a burst of
thirty clicks into thirty visible taps instead of a page that mutates on its
own. And a ghost cursor follows the agent's virtual pointer from action to
action, gliding between points with the agent's name riding next to it, landing
a fading ring on every press — red when the press was refused. All pointer
input here is synthetic CDP events, so the operating system's mouse never moves
and agents in different tabs cannot disturb each other or the user; the cursor
is one drawing per tab, purely for the watcher. All three are shown in every
session with a window, headless excepted, and all are drawn aria-hidden and
hidden before every screenshot, so nothing an agent reads back can see them. They
are on by default; the popup's Agent presence switch turns them off for this
Chrome without touching the bridge connection — the companion answers the
server's paint requests with success-shaped stand-ins, so actions are unaffected,
and flipping the switch mid-session takes already-painted tabs down at once. Set
WEB_SEARCH_NEO_AGENT_PRESENCE=0 to turn them off everywhere instead: the server
environment wins over the popup.
Sessions are pinned to the browser run they were opened in. Tab ids restart with Chrome, so a session that outlived a restart would address whatever tab inherited its number — quite possibly one of the user's. Such a session is dropped, with an error that says to open the page again, and nothing is sent to the new browser on the way out. The companion updating itself counts as a new run, since its reload drops every debugger attachment anyway.
For isolated work, opt into Selenium explicitly. profile_mode="temporary" and
profile_mode="persistent" are headless when headless is omitted; pass
headless=false only when a visible MCP-owned Chrome is intentional. Persistent
mode keeps a durable MCP-owned profile. For attach, the
launcher — not the headless argument — determines whether the already-running
Chrome is visible or headless.
headless=true cannot be combined with current: that mode drives a Chrome the
user is looking at, so the request is refused outright rather than quietly
answered by some other browser. auto with headless=true resolves straight to
temporary without even probing for the companion — worth knowing, because it is
the one way auto reaches a Selenium profile on a machine where current works
perfectly.
Start a durable visible Chrome for attach mode (visible is the launcher default):
powershell -ExecutionPolicy Bypass -File scripts\start_managed_chrome.ps1 -ProfileId authorized -Port 9222 -WindowMode visibleSign in to the sites you need in that window, keep it open, then let the agent attach:
{
"actions": [{
"action": "open",
"url": "https://example.com/",
"session_id": "authorized",
"headless": false,
"profile_mode": "attach",
"debugger_address": "127.0.0.1:9222"
}]
}Chrome 136+ does not allow remote debugging against its normal default data directory. The included launcher therefore uses a separate durable profile. It feels like a normal visible Chrome window, keeps its logins, and remains open after MCP disconnects. See the Chrome remote debugging security change.
An ordinary Chrome window cannot accept a DevTools-port attach retroactively, which is why Web Search Neo includes its own companion extension. Codex's private extension/runtime isn't copied or required.
Attach mode can also connect to a managed Chrome running without a window:
powershell -ExecutionPolicy Bypass -File scripts\start_managed_chrome.ps1 -ProfileId automation -Port 9223 -WindowMode headlessThe headless argument cannot hide or reveal a Chrome process that is already running; for attach, -WindowMode on the launcher controls the actual window.
Reading a page
Four observation topics describe an open session, from semantic structure down to raw selectors:
Topic | Returns |
| An indented tree of roles, accessible names, states, |
| The rendered text, with headings, list items, and table cells preserved. |
| The whole content of one element — not a clipped slice of the page. |
| Ranked matches for a plain-language |
| The flat lists of links, forms, fields with |
{"topic":"page_outline","params":{"session_id":"demo","output":"json","limit":80}}The outline and find walk open shadow roots and same-origin iframes, including nested ones, and map every box into top-document coordinates so a reported center can be clicked or pointed at directly. A frame is mapped through the full transform of itself and its ancestors — transform, the individual rotate/scale/translate properties, and CSS zoom — because a frame scaled to fit its container is ordinary on a checkout page, and translating by its origin alone put the reported centre tens of pixels from the control. center is the point to aim at; page_rect is the smallest rectangle containing the mapped corners, which for a rotated frame is larger than the element itself and is not something to click. A chain that is not affine — a 3D perspective — can only be approximated flat, and those nodes say so with page_rect_approximate rather than presenting the guess as an ordinary number. A cross-origin frame appears as one node with same_origin: false; read it by passing its selector as frame_selector. Closed shadow roots are counted in closed_shadow_roots rather than entered. A node inside a frame reports the frame path that reaches it, verified to match exactly one element and written as a #host >>> #inner piercing path when the frame is itself nested or inside a shadow root; a frame with no such path is marked frame_addressable: false instead of being given a selector that would land in the wrong one.
page_text never answers with a blank page, and never hands a fragment over as if it were the page:
mode="main"on a document that is one large form — a sign-up, login, or checkout page — used to drop everything as chrome. An empty result is now retried without the noise filter and reported asfallback_used: truewithmode_used: "full".A landmark is only trusted when it holds at least half of what the body renders, so an app shell whose real content is mounted outside its
<main>no longer reads asLoading....excluded_charscounts how many of thebody_charsthe body renders are missing from the answer, andexcludedsays why: chrome dropped bymode="main", cross-origin frames, frames nested deeper than the shared frame-depth limit, a frame whose document has not parsed yet, or a clip atmax_chars.Same-origin frames are read in place,
<dialog open>androle="dialog"overlays that sit outside the chosen root are appended and counted indialogs_appended, andframesreports what was entered, skipped as cross-origin, or too deeply nested.With
include_links=truethe link index is paid for out of the same budget, so the response stays inside the requestedmax_charsinstead of overshooting it by the size of the listing.
What find returns
Each match carries two numbers, because one number cannot say both things:
Field | Meaning |
| 0–100. How well the query matched that element alone, before any ranking bonus: 100 the whole field, 62 a prefix, 45 a substring, 34 every query token present — times the field's weight. |
| The ranking score: |
Field weights, and the values matched_field can take: name 1.0, text — the
element's own visible words, scored separately when an accessible name overrode
them — 0.9, placeholder 0.8, title 0.65, testid 0.6, name_id (the name
attribute and the id) 0.5, role and value 0.4, href 0.3.
Two flags sit on top of them, and they mean different things:
low_confidence— nothing clearedmatch_threshold(25) onmatch_score, so the matches are the closest things on the page, offered as guesses. It is derived frommatch_scoreand never fromscore, which is the fix for a real failure: context alone is worth 36 points, so on any action-shaped query every visible enabled control cleared a bar of 25 and unrelated elements came back withlow_confidence: false. Arolefilter is a filter too — if nothing of that role clears the bar, the result islow_confidencerather than a confidently wrong element of another role.ambiguous— the top two matched the query equally well and ranked equally (within five points on both), so document order alone decided which came first. Both are good matches; the choice between them is not, which is why it is a separate flag. It is computed beforelimitis applied, so asking for one answer still tells you the second was just as good.
The counts account for everything the page offered: candidates were examined,
scored resembled the query at all — weak role-word brushes included — matched
cleared match_threshold, returned fit inside limit (capped at 25), and
truncated says matched did not. aria_hidden_skipped counts elements skipped
because an ancestor is aria-hidden or [hidden], by the same rule page_outline
uses, and frames_too_deep counts frames no topic enters. Under low_confidence
nothing cleared the bar, so matched is 0 while returned counts the guesses —
the one case where returned exceeds matched.
low_confidence: true is an instruction, not a footnote: re-query with different
words or a role filter. Clicking matches[0] because it is the only thing on
offer is how an agent ends up pressing Cancel when it meant Submit order.
Locators
Form | Example | Notes |
CSS selector |
| Unchanged, and still the only form accepted everywhere. |
Ref handle |
| Issued by |
Piercing path |
| The separator is a space-padded |
Ref handles carry the document they were read from — which, for a node inside an iframe, is that frame and not the top page:
element numbers restart at 1 in every document, so the earlier
ref:Nform silently resolved to an unrelated element after a navigation — aref:1saved from a search page addressed whatever happened to be the first node of the next page;a node inside a frame is numbered in that frame's own registry, so its epoch differs from the page's
dom_epoch, and resolving it enters the frame that issued it. Until 1.3.2 every ref was minted in the top document, which made the naturalpage_outline→clicksequence fail on a payment or consent frame — after a ten-second poll, blaming staleness for what was really the wrong browsing context;a handle whose epoch no longer matches now resolves to nothing, and the action fails with an explicit "read the page again with
page_outline" message. So does a handle whose element has been detached from the DOM. Both fail immediately rather than polling: a replaced document does not come back, and the two cases are reported apart, since a frame that is still open but no longer holds the element can still be reached by thecentercoordinates the outline reported;a bare
ref:Nwithout an epoch is refused outright, with an error that says to readpage_outlineagain: it carries no document identity, and answering it from whatever document happens to be loaded is the very mistake the epoch exists to prevent.
Ref handles and piercing paths are accepted by fill (both its field and its file keys),
upload, click, wait, and the form_selector of submit. click and wait took
plain CSS only until 1.3.0, which made the natural find → click sequence fail with
InvalidSelectorException; they now poll the resolved handle for the requested
present/visible/clickable state. The submit_selector still expects a plain CSS
selector. Both non-CSS forms need a live element handle, so they
resolve in the Selenium-backed modes (temporary, persistent, attach); in companion
current mode they are refused with an explicit message and CSS selectors remain the way
to address elements.
A CSS selector that contains >>> inside quotes or brackets — div[data-op='a >>> b']
— is no longer mistaken for a piercing path; the separator is only recognised outside
quoted strings and attribute brackets.
frame_selector accepts less than the locators do
A frame_selector always names exactly one frame. A CSS selector matching more than one is
refused with the count, everywhere, rather than answered with whichever came first.
What differs is which forms are accepted. render and the reading topics —
page_outline, page_text, find — take all three, so a nested frame can be named by the
#host >>> #frame path the outline reports for it, as can fill, click, wait, submit
and upload.
The input actions take plain CSS and nothing else: press_keys, pointer, touch, the
pointer entries of an input batch, pointer_lock, and game_probe. They aim by
coordinate, which needs the frame's own box measured in the top-level page, and neither a
ref: handle nor a piercing path yields one. Both are refused before a single event is sent
— not half-way through a batch, which is what an input mixing keys and pointer entries
used to do — with an error that says to pass a CSS selector matching this frame and nothing
else. pointer_lock refuses them on release and status too, though only acquire
dispatches a click: one call cannot read the same string two ways.
game_probe sits on that side too, though it dispatches nothing. It reports canvas
rectangles that are then aimed at with the same string, so a probe that read one of four
matching frames while the input went to another was the same defect wearing a different hat.
So the sequence the games section leads to has one step in it: the outline names a nested
game frame #host >>> #frame, and that path cannot be handed to game_probe or input.
Give the input actions a CSS selector that is unique by itself, and keep the piercing path
for reading.
Console and network diagnostics
The page's own console and HTTP traffic are readable without leaving the MCP contract:
Topic | Returns |
|
|
| One compact line per request in the documented |
| One response body, by |
{"topic":"network","params":{"session_id":"demo","only_errors":true}}One line per request, for example a form post the server rejected:
POST 422 Document 4ms 0.5KB https://example.com/submitonly_errors=true keeps exactly these — failed requests and everything from 400 up — so a
post the server accepted is filtered out and an empty list is itself an answer.
Use output="json" on the network topic when you need the per-request id that network_body expects; the default text lines omit it.
har_export {session_id} writes the same journal as a HAR 1.2 file into the download folder,
for DevTools, a proxy or a bug report. It carries method, URL, status, type, timing, transfer
size, the security-relevant response headers and post data; request headers and bodies are
not recorded, and the file's comment says so.
Capture is armed when the session takes its tab, not when a topic is first read. In current mode the subscription is the last step of opening the tab, while it is still about:blank, so the requests and logs of the very first navigation are in the buffer — a single open reports the document, its subresources, and a 404 favicon without anyone reloading anything. This is a fix, not a nicety: network capture used to be armed by the first network read, so an agent that opened a page, saw it fail, and then asked which requests failed was told there were none.
attach_tab is necessarily different. Capture starts at the claim, and whatever the tab did before it was claimed was recorded by nobody and cannot be recovered — reload the claimed page if you need its load traffic.
Each stream keeps up to 500 entries and reports how many older records were dropped. Two further ceilings belong to companion current mode alone, because there the buffers live inside the extension: a 512 KB budget shared between console and network, and in-flight requests capped at 1000. Those limits bite on a chatty page: one that fires 700 requests keeps the newest 500 and reports a little over 200 dropped — a few more than the arithmetic, because the shared byte budget evicts as well — and eviction is oldest-first, so the very first navigation is the first thing to go. Read network before the page has had time to bury it. An empty list next to a non-zero dropped means "evicted", not "quiet".
Both topics work in companion current mode and in the Selenium-backed modes; the Selenium path reads Chrome's browser and performance logs instead of the extension's buffer, so field coverage is close but not identical. It caps in-flight requests at 500 and keeps no byte budget. Its performance log is switched on when the session starts, and its page console hook is registered with Page.addScriptToEvaluateOnNewDocument before the session navigates anywhere, so it runs ahead of each document's own first script and the first page's boot output is covered too. That hook used to be installed the first time the console topic was read, which lost exactly those lines and made a reload the price of seeing them.
Forms and multi-step flows
The short version: open the page, read it, fill everything in one call, attach
files through files or the upload action, reread every consequential live
choice, submit exactly once, then prove the outcome with fresh DOM/text.
{
"actions": [
{"action": "fill", "session_id": "intake",
"fields": {"#full-name": "Ada Lovelace", "#topic": "demo", "#subscribe": true}},
{"action": "submit", "session_id": "intake", "form_selector": "#intake-form"}
]
}fillapplies what it can and returnsfilled,field_values,files_uploadedand a per-selectorerrorsmap;successisfalsewhenevererrorsis non-empty. Every write is read back off the control, so amaxlengthtruncation and anoninputhandler that rewrote the value are errors naming both values instead of successes.field_valuesanswers for every selector you sent, failures included — a refused control shows what it still holds, andnullmeans nothing could be read back, because the selector matched nothing or the control has left the page.Sanitisation is not refusal. Whitespace trimmed off a
type=emailvalue,\r\nnormalised to\n, a handler that lower-cases the input — the control kept what you gave it, in its own terms, and all three read back as filled. A byte-for-byte comparison used to call them failures.A
<select multiple>is the one control that takes a list of values and reads back as one. A scalar sent to it replaces the whole selection rather than extending it, and a list sent to anything else is refused.Date, time, datetime-local, month, week, range and colour inputs are set rather than typed, since typing into them depends on the browser's locale. The value is rehearsed on a throwaway input first, so an unparseable one is refused without touching the control — no more valid-looking wrong date, no more slider dropped to its midpoint — and the error names the format the control wants.
A
contenteditableeditor — TipTap, ProseMirror, Slate, Quill, a webmail message body — is written as a real edit and keeps its paragraphs: the text goes in a line at a time with a soft break between the lines (Shift+Enter, never Enter, which in a chat composer would send what is written so far), because inserting it in one go puts\ninside a single text node where it renders as a space. The read-back isinnerText, nottextContent, which ran every paragraph together and reported a fill that had worked as a refusal. If an editor folds the breaks away regardless, the error says exactly that and names the way through: put the text on the clipboard withrun_script(user_gesture=true,navigator.clipboard.writeText) and paste it with a realCtrl+Vthroughinput. Some chat composers never pick up a DOM write at all — their send control stays inert — so paste there always.uploadno longer treats an empty file input as proof of failure. Any Dropzone-style widget takes the file off the input and uploads it itself, so reading the input back finds nothing a millisecond later while the file is already stored and its chip with the name is on the screen.upload_statesays how far the evidence goes:attached(the input holds the files, which is exact),taken_by_widget(the input was emptied, and either the page now names the file or a POST/PUT/PATCH followed the attach) orunconfirmed(nothing vouches for it either way — which is not a refusal;notesays to look for the name withpage_text/elementsand for the request in thenetworktopic before attaching again).successis false only when the attach itself failed.fillwithfilesreports the same thing inupload_statesandupload_notes.uploadsays how the file got in:attach_method: "set_file_input_files"(Chrome read the path itself) or"streamed"withstream_reason. In your own Chrome, if Chrome refuses access to the file (a bare "Not allowed"), the server reads the file and hands its bytes to the page and reportsattach_method: streamed; the page then sees syntheticinput/changeevents (isTrusted=false). Chrome gives the same "Not allowed" when the companion's "Allow access to file URLs" is off and, apparently, when an administrator forbids file access - the two cannot be told apart. A refusal whose text names a policy is returned as an error.fetch_linksreturns a list of URLs; when thelimitor the byte budget cut it, the last line starts with#(# truncated=true: …,# body_cut=true: …) and is never a URL.submitruns native validation first and reportsvalidation_passedwith the offending field ids, thensubmit_triggeredfrom the firedsubmitevent or from the document being replaced — including a reload of the same URL, which a check on url and title alone reported as a failed submit.submit_default_preventedmarks the SPA case where a handler cancelled the navigation on purpose.submit_evidencestates in a sentence what the verdict rests on, andnew_tab_openedwarns that atarget="_blank"result landed in a tab this session does not own — so theurlandtitlebeside it are still this page's, not the answer you are looking for.A checkbox takes
1/yes/y/on/check/checkedor0/no/n/off/uncheck/unchecked/""and refuses anything else, a<select>takes an optionvalueor its visible text, and a file input must go throughfiles/upload.Every control
fillwrites is blurred afterwards, because that is the only way the last field of a fill ever fires itschangeevent. Three consequences follow and the third bites: focus ends ondocument.body, an autocomplete list the fill opened is dismissed, and a followingpress_keys(["ENTER"])goes to the body rather than the field — passtarget_selector, orfocus_mode="click", when you mean to submit by keyboard. Thefilesentries are the exception: nothing is typed into them, so nothing is blurred.fill,click,submit,uploadandwaittakeframe_selector, so a form inside an iframe is addressable by name and not only through a ref.File inputs honor that frame context too: same-origin uploads are resolved inside the selected frame in the tab target, while cross-origin uploads use the child debugger target. A top-level matching input is never chosen as a decoy, and valid same-origin uploads no longer surface Chrome's opaque
Uncaughtevaluation error.waittakespresent,visible, orclickable, and honours itstimeout_secondsas passed (it defaults to 10). A timeout names the seconds actually waited in its message, so you can ask for as long as the target needs and read back what really happened.After any navigation or step change, read
page_outlineagain: old ref handles are stale by definition.dom_epochonly tests one direction of that. A different epoch proves your refs are dead; the same epoch proves nothing, because the epoch belongs to the document and a wizard step that swaps its own markup in place keeps it while every ref it issued goes with the old nodes. A ref read inside a frame carries that frame's epoch rather than the page's, so a mismatch there is normal and not staleness at all.
For an application, payment, message, deletion, or other consequential terminal
action, keep submit_attempted=false. From a fresh page_elements read match the
exact target (href where available) and verify each critical selected value in
the live control; remembered defaults are not evidence. Set the flag as the
terminal button is clicked once. After any response or timeout, never retry that
button before inspecting URL, text, elements, console, and network — the first
click may already have succeeded. Stop immediately on terminal success text.
→ Complex forms covers the rest with worked calls:
choosing between the three locator forms, finding a field by meaning when
selectors are generated, partial-failure recovery, multi-step wizards and SPAs,
fields inside Shadow DOM and same-origin iframes, what to do when
challenge_detected turns true, and how to trace a submit that silently failed
through network and network_body.
Reviewable macros
A macro is a JSON file of web_action steps that lives in the calling project. Write it
with an editor, check it with op=validate, and replay it by name. Before a consequential
replay, preview the exact resolved steps with the same variables that will be passed to
run:
{"actions":[{"action":"macro","op":"preview","name":"document-submit","variables":{"target_url":"https://forms.example/requests/42","resource_path":"C:/docs/request-42.pdf"},"project_root":"C:/work/my-project"}]}preview loads the saved macro, validates that every placeholder was supplied,
and returns the resolved steps with executed: false. It never dispatches an
action or changes browser state. Review the canonical URL, form values, upload
path, and whether a terminal consequential action is present before sending the corresponding
op: "run" call.
It also checks every resolved step against the schema its action publishes, and reports
steps_valid with a problems list naming the offending step index and error. Because macro
files are written and edited by hand, the defect in one is usually a mistyped parameter rather
than a wrong click path — and without this check it surfaces mid-replay, after the steps
before it have already dispatched.
Preview is a review boundary, not proof that replay will succeed. After run,
check the batch result and freshly inspect the page. For applications, payments,
messages, or other consequential actions, use the guarded two-phase path below so the exact
terminal Submit or safe Click is attempted only once after live-state validation.
Checking a macro before it runs
op=validate reads a macro file and dispatches nothing at all. It is the check that
belongs before the first live run of a hand-written or generated macro, and it reports
every finding with the step index, what is wrong, and how to fix it.
{"actions":[{"action":"macro","op":"validate","name":"document-submit","project_root":"auto"}]}Errors are things that cannot work: an action name the server does not have, a required
parameter missing, a parameter that is not part of that action and would be refused at
dispatch, a placeholder used but not declared in variables — and, above all, a
{{placeholder}} inside a run_script script. That last one is the expensive mistake: the
value is pasted into the JavaScript as raw text, so anything containing a newline, a quote
or a backslash produces a syntactically broken program and the step fails with an opaque
Uncaught from inside the page. Pass the value through args and read it from
arguments[0] instead.
Warnings never make a macro invalid, because each of them can be deliberate: a variable
declared and never used, steps that drift between two different session_id values, and a
macro whose last meaningful step neither waits nor reads anything back — one that sends
something and then reports whatever the dispatcher happened to return.
A replay can be pointed somewhere else: session_id on run or preview rewrites the
session of every step, so a macro written for one tab runs in another without editing the
file. A macro that already drives two sessions is refused rather than collapsed into one,
because two tabs were the point; give those steps a {{placeholder}} session instead.
Where macros live
There are two stores, and every macro answer reports which one it used, as scope,
project_root, and storage. A macro that appears to have vanished is almost always a
macro saved into the other one.
| Store |
omitted | The per-user store, unless |
| The project found by walking up from the working directory: |
an absolute path | Exactly that project. |
op=list additionally reports other_store, so the question "is it in the other one?" is
answered by the call that raised it. The nearest marker wins, so a package that keeps its
own .web-search-neo directory inside a larger repository is its own project. Macro names
are traversal-safe, and the resolved store is additionally required to remain beneath
project_root; a symlink or junction that escapes is refused.
Macro files
A project's macros are one .json file each under
<project_root>/.web-search-neo/macros/, intended to be committed with the project and
edited by hand. The shortest valid file is the step list itself:
[{"action": "open", "url": "{{url}}", "session_id": "docs"}]The full form adds the description and the placeholder defaults:
{
"name": "open-docs",
"description": "Open one documentation page",
"steps": [{"action": "open", "url": "{{url}}", "session_id": "docs"}],
"variables": {"url": "https://example.com"}
}Every step is exactly one web_action action object, so
web_info(topic="action_schema", params={"action": "<name>"}) is the reference for what a
step may contain. The variables map is optional: placeholders are declared from the steps
themselves, so a hand-written file reports what it wants through op=list and op=show
without one. Name a variable only to give it a default. Each store gets a README.md stating this next to the files, because the
next reader of a committed macro directory is often a model that has never seen the schema.
A file that cannot be read as a macro is reported by op=list as broken rather than
skipped, so one bad hand edit never hides the macros beside it. The guarded-operation
ledger lives in the same directory and is never listed as a macro.
There is no write API
A macro is a file, and that is the whole of it. Creating, editing, renaming, copying,
committing and deleting macros are ordinary file operations, done with the tools that
already do them well; the macro action only reads (list, show, validate), resolves
(preview) and replays (run, guarded_stage, guarded_commit). Moving a set between
machines is copying a directory.
This used to include a recorder, and the recorder is gone. It was never self-sufficient:
its own contract told the caller to save the recording and then hand-edit the JSON to turn
the changing parts into {{placeholders}}, so the path ended in an editor either way. Over
a full day of live use, every macro that actually worked was written directly as JSON and
the recorder was not used once, while it carried a class of defects of its own — races
between concurrent batches, steps attributed to the wrong open recording, a name quietly
borrowed by an explicit save. op=validate replaces it with the part that was actually
missing: a check that runs before the page does.
Project-local and guarded macros
The engine is universal and domain-neutral. Domain rules belong in a saved macro or its
calling project, never in Web Search Neo core. Pass project_root — an absolute path, or
"auto" — to any macro operation to use that project's independent macro set and its own
guarded-operation ledger.
Consequential flows use a generic two-phase path. A guarded macro must end in exactly one
consequential action: explicit submit, or a safe click. A terminal click must be either
a plain CSS selector (required to match exactly one live element immediately before commit)
or exact rendered text plus an explicit role (the semantic dispatcher refuses zero or
multiple matches). Coordinates, trusted=true, substring text, ref handles, piercing paths,
and non-terminal submits fail closed. guarded_stage resolves the macro, then fails closed unless
guard provides:
equal
target_urlandcanonical_url, plus an optional domain-definedidentity_key; query parameters are preserved because they may carry the target/requisition identity, while fragments are ignored;an explicit
allowed_hostspolicy and optionaldenied_hostspolicy;an existing absolute
resource_paththat the resolved macro uploads exactly;the exact 64-hex
resource_sha256for the current file bytes (case-insensitive input, normalized to lowercase after verification);a stable 16-128 character
idempotency_tokenunique to this target;one or more live
assertionsover staged action results.
Host policy is data. Core contains no built-in site categories or denylist. A project can deny platforms, production hosts, vendors, or any other domain-specific routes in its own macro/configuration. The resolved macro must open the exact canonical target URL.
This neutral document-request example assumes result 2 is a fresh page_text checkpoint:
{
"actions": [{
"action": "macro",
"op": "guarded_stage",
"name": "document-request",
"project_root": "C:/work/my-project",
"variables": {
"target_url": "https://forms.example.com/requests/42",
"resource_path": "C:/work/my-project/artifacts/request-42.pdf"
},
"guard": {
"target_url": "https://forms.example.com/requests/42",
"canonical_url": "https://forms.example.com/requests/42",
"identity_key": "request-42",
"allowed_hosts": ["example.com"],
"denied_hosts": ["staging.example.com"],
"resource_path": "C:/work/my-project/artifacts/request-42.pdf",
"resource_sha256": "0123456789abcdef0123456789abcdef0123456789abcdef0123456789abcdef",
"idempotency_token": "request-42-20260820",
"assertions": [
{"result_index": 2, "path": "data.text", "contains": "Request 42"}
]
}
}]
}guarded_stage dispatches every step except the terminal consequential action. Only after all staged actions
succeed, the supplied SHA-256 matches the current resource bytes, and every semantic assertion
passes does it reserve a checkpoint. The normalized digest is returned by staging and stored
with the checkpoint. Review the
result and commit once from the same project store:
{"actions":[{"action":"macro","op":"guarded_commit","checkpoint":"guard-request-42-20260820","project_root":"C:/work/my-project"}]}The project-local ledger is marked <action>_attempted before terminal dispatch. A timeout,
lost response, second call, new token for the same target identity, or resource reused for
another target refuses replay. Every one of those refusals is final by design, so each names
the ledger file it is refusing from — otherwise a stage whose commit never ran leaves a target
that can never be staged again and nothing to look at. The guard proves one guarded attempt, not server acceptance;
inspect durable confirmation separately. Concrete domain macros should live with their
projects and are not bundled in this repository.
guarded_commit never retries the terminal action and does not infer remote acceptance. After
the single attempt, use ordinary read-only inspection actions (for example fresh page text or a
screenshot) to collect destination-specific proof.
Check your own site
Seven actions for a developer reviewing a site they own before a release. The full reference, the scoring table and a worked example are in docs/site-checks.md.
Action | What it answers |
| What is configured unsafely and how to fix it: CSP directives, HSTS, framing, |
| The map of the page's calls to its backend (method, path template, frequency, origins, XHR/fetch/WebSocket/SSE/beacon) and what every own-site API response says: CORS against credentials (wildcard+credentials and |
| Secrets and endpoints in the page's own scripts (cloud keys, tokens, JWT, private keys, high-entropy literals - values masked), routes from fetch/axios literals, a referenced openapi.json, sourcemaps and agent files (llms.txt, skills, MCP cards); sign-in forms; third-party scripts named, never fetched. |
| Opt-in active checks: CORS preflight with a foreign Origin, OPTIONS/TRACE methods, open redirects on own links, one inert reflected token. Only own scope, budget and requests_made; never POST/PUT/DELETE, payloads, auth or fuzzing. |
| TTFB, FCP, LCP and CLS with Web Vitals ratings, resources by initiator, the largest files, render-blocking files and what to change, from the page's own Performance API after a cold isolated load. |
| The session's network journal as a HAR 1.2 file; |
| A regression scenario: every step is an ordinary action plus |
security_report is passive and stays inside the scope you give it: exact hosts (with
optional scheme and port) only — wildcards and public suffixes are refused, and a redirect,
link or sitemap entry leaving the scope is named, never followed. It sends ordinary GETs of
the pages, of http://host/ (no path or query), of security.txt, robots.txt and the
sitemap, and one ordinary TLS handshake per https host — at most 200 requests, all listed,
redacted, in requests_made. With scope: "page" the page also loads once in a fresh
isolated browser, as any visit does; browser_requests summarises that traffic. It guesses
no paths, scans no ports, tries no origins, fuzzes nothing and never gets past a login or a
CAPTCHA; point it only at sites you own or may test. Cookie values never appear in the
answer.
{"actions":[{"action":"test_run","url":"http://127.0.0.1:8000/signup","session_id":"reg","steps":[
{"action":"fill","fields":{"#email":"qa@example.com"}},
{"step_name":"submit","action":"click","selector":"#send",
"expect":{"text":"Check your inbox","no_console_errors":true}}]}]}Dialogs, downloads, navigation
Action | What it does |
| Reads the |
| Lists the files the session downloaded. Browsers the server launches download into their own folder under the download directory, never into |
| Goes to a URL in an open session, keeping its browser, profile and viewport. |
Canvas and WebGL games
Browser automation is not limited to DOM forms. The compact contract covers common HTML5 game controls:
Call | What it does |
| Reports canvases, 2D/WebGL context, iframe surfaces, document focus, sampled FPS under |
| Mixes per-key |
| Keyboard-only shortcut: 1-8 |
|
|
| Acquires, releases, or reports pointer lock for first-person controls; while locked, |
| Control the animation gate and safely reset held input. |
| Turns a pointer-locked camera by exactly |
| Lets N animation frames render - between the click that focuses a canvas and the first keys, which would otherwise be lost. |
| When |
| Every capture reports image size, CSS box, |
| Targets a cross-origin game iframe such as the one used by Yandex Games. |
The verb of each action is namespaced because the dispatcher already owns
action: key_action, pointer_action, touch_action, and pointer_lock's
operation.
Key names also accept the DOM spellings (ArrowLeft, KeyW, Digit1, ShiftLeft) and
"Control+Shift+K" chords; several keys in one press_keys are pressed together as a
chord. ENTER is the main Enter key (NUMPAD_ENTER the keypad one) and produces the
same keydown/keypress/insertLineBreak sequence as a real keyboard in both input paths.
Keyboard coverage includes F1-F12, NUMPAD0-NUMPAD9 with the numeric keypad location, MULTIPLY/ADD/SUBTRACT/DECIMAL/DIVIDE, META (also WIN, CMD, COMMAND), the arrow, navigation, and editing keys, and any single printable character. A key held as W is released by w as well, and the release dispatches exactly the character that was pressed. A modifier held with hold is carried into subsequent mouse and touch events, so Shift-click and Ctrl-click behave as a user's would. Giving a canvas keyboard focus no longer costs a synthetic click, which used to reach the game as a shot or a jump.
Example for a game hosted inside an iframe:
{
"actions": [
{"action": "open", "url": "https://yandex.ru/games/app/geometry-dash-ufo-2d-371298", "session_id": "ufo", "headless": false},
{"action": "render", "mode": "step", "session_id": "ufo", "frame_selector": "#game-frame"},
{
"action": "input",
"session_id": "ufo",
"frame_selector": "#game-frame",
"key_actions": [
{"key": "W", "action": "hold"},
{"key": "S", "action": "release"},
{"key": "SPACE", "action": "tap"},
{"key": "E", "action": "tap"}
],
"pointer_actions": [
{"action": "hover", "x": 640, "y": 360},
{"action": "wheel", "x": 640, "y": 360, "delta_y": -240},
{"action": "move", "x": 15, "y": -5, "coordinate_mode": "delta"}
]
},
{"action": "step", "frames": 5, "session_id": "ufo"},
{"action": "release_inputs", "session_id": "ufo"},
{"action": "render", "mode": "normal", "session_id": "ufo"}
]
}Send the batch to web_action; use web_info(topic="game_probe", params={"session_id": "ufo", "frame_selector": "#game-frame"}) or the screenshot topic to observe it.
A batch like that one ends with release_inputs and render mode=normal for a reason, and web_action stops at the first action that fails unless continue_on_error=true. So the failure of a single input in the middle is the case where the cleanup at the end never runs and the page is left frozen with keys held. Either send the cleanup as its own call, or pass continue_on_error=true on any batch whose tail is cleanup.
For a top-level canvas application such as https://redoschool.ru/demo/?auto=true, omit frame_selector. Pointer coordinates are relative to the selected top-level viewport or iframe.
→ Playing games is the full walkthrough: a complete
run from open to the win condition, why a tapped key is held across the frame,
the auto-repeat that keeps a held key alive after a respawn, iframe coordinate
handling, pointer lock, touch games, and the cleanup that must happen at the end.
Render modes
Mode | Behavior |
| Removes the render gate and returns to the page's normal |
| Continuously releases animation callbacks at no more than |
| Holds queued animation callbacks. The |
While the gate is engaged — in throttled as well as in step — page time is frozen:
performance.now()andDate.now()advance by exactly one frame delta of 16.667 ms per released frame;setTimeout,setInterval, andrequestIdleCallbackare queued against that same virtual clock and run immediately before the frame's animation callbacks. Without this a game reads the agent's thinking time as itsdeltaTime, and a batch of frames released back to back arrives with a delta near zero;promises and
queueMicrotaskare not gated, andnew Date()still reports wall-clock time;queued timers are handed back to the real scheduler, with their remaining delay intact, when the mode returns to
normal;page time never runs backwards. Stepping carries the clock ahead of wall time — sixty frames released back to back advance it by a second in barely a moment — and returning to
normalkeeps that gap instead of dropping it. Handing the native clock back used to lose the whole gain at once, measured at −5.9 s, which a game reads as a negative frame delta. A session that has stepped keeps the clock wrapper for the rest of its life, since only the wrapper can go on applying the offset; one that never stepped gets the untouched native clock back.
render also accepts frame_delta_ms, freeze_time, and gate_timers when an
engine needs something other than those defaults. One knob stays Python-only:
browser_tools.set_render_control(..., key_repeat=False) disables the auto-repeat
that keeps a held key alive for games which latch input on keydown.
One input action can hold one key, release another, tap two more, turn the wheel, and move the pointer by a delta before releasing exactly one frame. A tapped key is pressed together with the rest of the batch, stays down for the whole released frame, and is lifted afterwards, so an engine that polls key state once per frame — Phaser, Godot, a hand-written canvas loop — actually observes the press. In step mode the game never observes a partially applied intermediate input state. Outside step mode the actions are still serialized, but the page continues rendering normally.
The gate is installed into every new document of a session, so it survives a page reload or a game iframe that reloads itself; a step that lands on a fresh document re-applies step mode once and reports gate_reinstalled. render reports frame_delta_ms, time_frozen, timers_gated, pending_callbacks, and input_advances_frame; step reports frames, callbacks, pending_timers, and virtual_now.
The render controller gates JavaScript requestAnimationFrame, which covers typical canvas/WebGL and Unity WebGL loops. It does not change video decoding, CSS compositor animations, the monitor refresh rate, or guarantee an exact GPU hardware frame rate on every engine.
While a gate is engaged, game_probe does not try to sample FPS. Frames are released by
hand, so the measurement could only expire against its own script timeout and then report
a fabricated zero; the probe returns immediately with an animation object holding
fps: null, animation_suspended: true, and a reason naming the active render mode and
the gated frame. Every frame-rate field lives in that nested object — animation.fps,
animation.frames, animation.elapsed_ms, animation.available — never at the top level
of the result.
console_messages holds the warnings and errors that appeared since the previous
game_probe call, which is what the probe's console_scope field says in the result. A
polling loop therefore reports a problem when it happens instead of re-reading everything
the session has ever logged, and its output does not grow with the length of the run. The
flip side is that each entry is delivered exactly once: a probe result that is thrown away
takes its console entries with it. Read console_messages on every probe you make, or use
the console topic, which keeps its own place in the same buffers — polling the probe never
hides entries from console, and reading console never hides them from the probe.
The probe reads the in-page hook as well as Chrome's browser log, so console.warn,
console.error, and uncaught exceptions are covered on both backends. In companion
current mode that browser log is not available at all, and earlier builds, which read only
it, reported nothing there.
Input latency
Game control is only useful if a round trip is cheap, so each action reads the page once
and never sleeps by default. Measured through web_action against the bundled platformer
fixture in step mode on a headless Chrome session, 30 iterations each, median and p95:
Action | Median | p95 | Same call at the former |
| 34 ms | 47 ms | 233 ms |
| 25 ms | 26 ms | 230 ms |
| 27 ms | 27 ms | 220 ms |
| 11 ms | 12 ms | — |
The published wait_seconds default of every input action dropped from 0.2 to 0.0, which
is the whole of the last column: the action itself was never the cost. include_summary
is now part of the MCP contract as well as the Python API, and setting it to false skips
the post-action page read: input 34 ms → 29 ms, press_keys 25 ms → 21 ms, step
11 ms → 8 ms. It saves nothing measurable on pointer, whose summary is already the
cheapest of the four. Absolute numbers depend on the machine; the ratios do not.
These are step-mode numbers: the page does no work between actions. In normal
render mode the same call waits on the page's own main thread, so on a heavy page (a
144 FPS canvas, a busy SPA) or a throttled window it takes 0.3-2 s. game_probe
reports frame_health (measured raf_fps, throttled, throttle_reason) so a slow
round trip can be told apart from a slow page.
Two-tool MCP contract
Tool | Responsibility |
| Return the whole contract, the built-in automation skill, or one action schema on demand; read search, current Chrome tabs, browser, page outline/text/find, console, network, game, screenshot, or time state. |
| Execute one or up to 32 ordered setup, search, fetch, tab attach/open, form, input, render, audit, and close actions. Supports fail-fast or |
Start with web_info(). With no arguments it returns actions with each action's summary and its required parameter names, action_groups, info_topics, recipes, pitfalls, limits, and worked examples. Optional names, types, and defaults are deliberately left out of it. Request only the needed, generated JSON Schema with web_info(topic="action_schema", params={"action": "input"}), then invoke it through web_action. Two narrower entry points exist for a model that does not want the whole contract at once: web_info(topic="actions") is the action index alone, with params={"group": "macro"} to narrow it, and web_info(topic="skill") is the runtime playbook. The same call describes an observation topic — params={"action": "find"} returns find's parameters — which matters because a topic refuses any argument it does not list, and that list appears nowhere else. This follows the on-demand Tool Search principle used by official Unreal MCP: keep the eager tool list small, disclose schemas only when needed, and dispatch actions through a meta-tool. Web Search Neo combines Unreal's list/describe discovery tools into one web_info, so only two tools are advertised.
Measured on the current build, summing each advertised tool's name, description, and serialized inputSchema: the compact surface is 1,605 characters across two tools, against the tens of thousands of the legacy tool list. The self-describing contract behind web_info() is about 13,900 characters (a test keeps it under 14,000), and it is fetched only when an agent asks for it. Each action in it lists its required parameter names; optional names, types, and defaults stay in action_schema, where they cost nothing until needed — and where an observation topic's parameters live too, since a topic accepts exactly the list it publishes and nothing else.
Every action is declared once in a single registry that also generates its published schema, and arguments are validated against that same model before the handler runs. An unknown or malformed field returns the offending names and the list of allowed parameters instead of an internal TypeError:
ValueError: action 'input': unknown parameter(s) ['frames']. Allowed: ['key_actions', 'pointer_actions',
'session_id', 'target_selector', 'frame_selector', 'wait_seconds', 'include_summary']. Call
web_info(topic='action_schema', params={'action': '<name>'}) for the full schema.Existing direct Python imports remain available. For temporary MCP-client migration only, set WEB_SEARCH_NEO_LEGACY_TOOLS=1 before starting the server to advertise the former narrow tool list instead of the compact default.
Browser state is keyed by session_id, and that key is also the boundary between agents working at the same time: one session_id is one tab, so parallel agents must not share one. The open_many action can create as many independent sessions concurrently as the session cap allows (eight by default), and the companion popup or WEB_SEARCH_NEO_MAX_SESSIONS moves that ceiling. A viewport screenshot with no dimensions preserves the current viewport in every mode. An explicit viewport width/height pair is exact in isolated Selenium modes and is refused in current mode, where the MCP never resizes the user's Chrome; use mode="region" there for an exact-size image. Full-page captures above 3840x10000 fail explicitly instead of returning an unlabelled partial image.
Image-guided clicking is available through pointer with
pointer_action="click" and viewport CSS x/y. Take a fresh viewport screenshot;
if its PNG dimensions differ from the reported viewport dimensions, scale both
axes proportionally. Full-page and region pixels do not map directly to the
viewport: scroll the target into view and recapture. Any scroll, zoom, resize,
navigation, animation, or rerender invalidates the old image coordinates.
Optional agent skill
web_info(topic="skill") returns a compact built-in playbook intended even for
small local models: inspect → act → verify, schema discovery before guessing
optional parameters, element pagination and lazy-page scrolling, the three
screenshot modes, safe visual-coordinate clicks, current-Chrome CSS-only action
locators, exact href/value matching, and a one-shot final-submit guard. It is
about 6.5 KB and can be fetched once at the start of an automation task.
It also names its detailed sections, which are the part a small model needs when
a specific operation starts rather than at the beginning of the session. Each is opened
on demand with web_info(topic="skill", params={"section": "<name>"}) and carries
what a JSON Schema cannot: when the section applies, the calls in order, the rules
that are not guessable, and the mistakes to avoid.
Section | Covers |
| Choosing the surface, the first three calls, session naming. |
| Inspect, act once, verify — and what counts as proof. |
| CSS, refs, piercing paths, and which action accepts which. |
| Filling, choice widgets, uploads, the one-shot terminal submit. |
| Project files, placeholders, validate, preview, replay. |
| The two-phase path for an action that cannot be taken twice. |
| Session discipline for subagents sharing one server. |
| Search and fetch without a browser. |
| Console, network, response bodies, page scripts. |
| The render gate, per-frame input, probes. |
| Symptom, cause, and the call that fixes it. |
The troubleshooting section maps the server's own error text — an unknown
parameter, a stale ref:, a macro that "does not exist" because it was saved into
the other store — onto the call that resolves it, which is what turns a repeated
failure into one corrected retry.
web_info() called with no arguments still returns the entire agent-facing contract: every action with its summary and required parameter names, the observation topics, ready-made recipes, the common mistakes, the hard limits, and runnable examples. Optional parameters — for an action or for a topic — come from action_schema, one at a time. An agent that reads either contract needs no external instructions, so the bundled filesystem skill is a convenience, not a requirement.
The repository still includes a short Web Search Neo skill for clients that prefer a resident description of when to reach for the server at all.
Install it locally by copying skills/web-search-neo into your Codex skills directory, then restart Codex:
Copy-Item -Recurse -Force skills\web-search-neo "$env:USERPROFILE\.codex\skills\web-search-neo"Invoke it explicitly as $web-search-neo, or let its task description trigger it for web search, visible Chrome automation, authorized attach sessions, form work, and browser-game testing.
Command-line client
scripts/mcp_cli.py drives the server over stdio for scripts and agents without an MCP
connector. call and run start a server per invocation (a cold start is about 12 s and
sessions end with it). For iterative work keep one server alive:
python scripts/mcp_cli.py serve --port 47811 # prints READY; stops after 30 idle minutes
python scripts/mcp_cli.py send web_action '{"actions":[{"action":"open","url":"https://example.com","session_id":"s","profile_mode":"isolated"}]}'
python scripts/mcp_cli.py send web_info '{"topic":"screenshot","params":{"session_id":"s"}}'
python scripts/mcp_cli.py repl # one "tool json" per line
python scripts/mcp_cli.py stopThe daemon listens on 127.0.0.1 only (a busy port is an error, and on Windows the port is
bound exclusively) and answers only callers holding the random token it writes to
~/.web-search-neo/cli-<port>.token, made readable by your user only (chmod 600; on
Windows an icacls ACL for the current user alone - the daemon refuses to start if that fails). Images are
written to --out-dir (default downloads/cli) and named in the answer; JSON text parts are
decoded.
Optional configuration
$env:WEB_SEARCH_NEO_REGION = "us-en"
$env:WEB_SEARCH_NEO_PROXY = "socks5://127.0.0.1:9050"
$env:WEB_SEARCH_NEO_BROWSER_USER_AGENT = "..."
$env:WEB_SEARCH_NEO_PROFILE_ROOT = "D:\BrowserProfiles"
$env:WEB_SEARCH_NEO_DEBUGGER_ADDRESS = "127.0.0.1:9222"
$env:WEB_SEARCH_NEO_ALLOW_PLAIN_HTTP = "1"
$env:WEB_SEARCH_NEO_PROJECT_ROOT = "C:\work\my-project"
python main.pyWEB_SEARCH_NEO_PROJECT_ROOT names the project whose .web-search-neo/macros/ set every
macro call uses when it passes no project_root of its own. Setting it per project in the
MCP client's configuration is the way to give a model a project's macros without it having
to spell a path on every call. WEB_SEARCH_NEO_MACRO_ROOT still points the per-user store
somewhere specific, and a named project wins over it.
The bridge daemon has five switches of its own. Set them for the MCP server process: a daemon inherits the environment of whichever server spawns it.
Variable | Effect |
| Loopback port shared by the daemon and the servers, default |
| How long a daemon with neither a companion nor a client stays up, default |
|
|
| How long a server keeps trying to reach a daemon, including one it started itself, default |
| How long server startup waits for that first attempt before continuing in the background, default |
One more is about the server rather than the bridge:
Variable | Effect |
| Browser sessions one server holds at once, default |
Only use a proxy you are authorized to use. HTTP sessions use desktop browser headers, connection pooling, bounded response sizes, and conservative retry/backoff. Rendered pages use the installed Chrome's native matching User-Agent unless explicitly overridden.
Transport policy
Unencrypted http:// to a public host is refused, for plain fetches and for browser open alike, with an error that names the host and the override. Loopback, private, link-local, and unspecified addresses stay reachable over plain HTTP, as do localhost and any host ending in .local, .localhost, .internal, .home.arpa, .lan, .home, .intranet, .private, or .corp — a local ComfyUI, Ollama, or dev server keeps working unchanged. So does any single-label host, http://nas/ or http://raspberrypi:8080/: a name with no dot cannot be published on the public DNS, and deciding it by resolution would cost a lookup on every URL and every redirect hop. Set WEB_SEARCH_NEO_ALLOW_PLAIN_HTTP=1 to accept plain HTTP everywhere.
Redirects are followed one hop at a time, at most five, and every hop is validated again, so a public HTTPS URL cannot quietly land on plain HTTP.
Tests
python -m pip install -r requirements-dev.txt
python -m pytest
python -m pytest --cov=. --cov-report=term-missingThe deterministic suite is 576 tests, grouped by what they protect:
Area | Covered |
MCP contract | Exactly two advertised tools, on-demand action discovery, ordered multi-action calls. |
Search | Routing, fallback, cooldown and cache, live status probes that leave the real cooldown untouched, accurate Bing challenge detection, manual challenges. |
HTTP | Fetches, the plain-HTTP policy, and per-hop redirect checks. |
Perception | The accessibility outline including open shadow roots and same-origin frames, refs minted in the document that owns them and resolved back into it, selectors verified unique before they are handed over, page text extraction with its non-empty |
Locators | The three forms and their escaping, refs that refuse to resolve in another document or after their element was removed, valid CSS that merely contains |
Diagnostics | The companion's console and network buffers with their ring-buffer eviction, capture armed at open so the very first navigation is reported, a claimed tab recorded only from the claim, and a subscription that failed at open repaired on the next read. |
Companion bridge | The handshake — a client without the token, a hello without a nonce, a newer companion evicting an older one, and WebCrypto HMAC agreeing with the Python signature — origin checks, tab grouping, and confirmation-gated Chrome setup that publishes the token before touching Chrome. |
Bridge daemon | Two MCP servers with commands in flight at once, two servers starting together converging on one daemon, a client never mistaken for the browser, an in-flight command failing instead of hanging when the daemon dies, a daemon of the same version left alone while an outdated one is replaced, a newer daemon used rather than evicted by the server it outranks, versions ranked component by component so |
Forms | Multipart upload, form filling, native validation, values read back off the control so a refused write is not a success, exact PNG viewport size, a guarded terminal click that counts its matches in the page and so works in the user's own Chrome as well as a Selenium one, and guarded refusals that name the ledger they refuse from. |
Challenges | A captcha that blocks told apart from one merely present on the page, and the providers a top-level query walked past — DataDome, AWS WAF, a challenge one frame down, and one in a shadow root. |
Games and input | Canvas probing, normal/throttled/step rendering, atomic mixed input, held modifiers across a batch, gate and held-input reset on navigation, coordinates that follow a frame's CSS transforms, key spellings that release each other, and a virtual clock whose intervals keep their period. |
Site checks | CSP parsing and the browser's nonce/hash/ |
Sessions | Sessions that close the tabs they own, concurrent sessions, persistent storage, a session dropped when the browser it was opened in is gone, a borrowed tab handed back instead of navigated, a real managed-Chrome attach/detach that leaves Chrome running, a second caller inside one session reported rather than swallowed (and a reentrant one not mistaken for it), and a full session cap that names the setting which lifts it. |
Public search engines may rate-limit an IP or region, so live internet smoke checks are kept separate from deterministic tests. scripts/live_smoke.py runs a search, opens two pages concurrently, fills and submits a public Selenium test form, uploads a file, and verifies exact screenshot dimensions; it does not cover games. scripts/companion_live_smoke.py reaches the network too, though its subject is the extension: it starts a disposable Chromium with the companion loaded and drives whatever --url names, which defaults to https://example.com. Public game sites change without notice, so the frame gate, input atomicity, and held-input recovery are verified by the deterministic local suite instead.
End-to-end check with a local model
scripts/live_agent_game.py exercises the whole stack the way a real client does: it starts main.py as an MCP stdio subprocess, hands the two advertised tools to a model served by LM Studio on 127.0.0.1:1234, and lets that model play the bundled platformer fixture while every tool call is timed.
lms load qwen3.5-4b-mtp --context-length 16384 --parallel 1
python scripts/live_agent_game.py --model qwen3.5-4b-mtpIt reports the eager tool-schema size and the median and p95 latency of both MCP tool calls and model turns, so a regression in either is visible immediately. Thinking is disabled through reasoning_effort: "none", because a mechanical control loop pays for it without gaining anything. The script needs Chrome and a running LM Studio server and is not part of pytest. Whether the model finishes the level is a property of the model, not of the server: a 4B does it, but not on every attempt.
Safety notes
Visible or attached sessions may contain authenticated accounts. The MCP client can act with the permissions of those accounts.
The companion declares five permissions:
alarms,debugger,storage,tabs, andtabGroups. It ships no content scripts and asks for nohost_permissions, butdebuggeris the broad one: it lets the extension attach the Chrome DevTools Protocol to a tab and from there read and modify that page, its console, and its network traffic. Chrome shows a "started debugging this browser" banner whenever it is attached.alarmsis the narrow one — it only wakes a suspended service worker to retry the bridge, and Chrome shows no extra warning for it. Install the companion only from this repository.The loopback bridge is authenticated in both directions with an HMAC challenge-response over fresh nonces: the daemon proves it holds the machine-local token first, and the extension or MCP server answers only after checking that proof, so the token itself never crosses the socket (since 1.15.0). Until 1.3.0 the port accepted any local client that spoke the protocol, and the extension trusted whatever answered on it.
That secret is a file readable by the user account that owns it, so it does not defend against a malicious process already running as you: such a process can read the token and impersonate either side. It removes the race in which any local program that binds
127.0.0.1:8765before the server inherits DevTools access to every signed-in tab. Chrome Native Messaging, which needs no listening port at all, is the actual fix and is tracked in TODO.md.The port is now held by a daemon that outlives each agent call, so it is reachable for as long as Chrome keeps the companion attached rather than only while an agent runs. The authentication is unchanged and same-user processes were never excluded by it, but the window in which one could use the token is wider. Stop the daemon with
python main.py --bridge --stop, or let its idle exit close the port fifteen minutes after the last companion and client are gone.Authentication is not authorization. An authenticated peer may call
cdp.sendon any tab it holds, limited only to the companion's DevTools method allowlist; the daemon refuses commands for a tab another connected client has claimed.setup_current_chromeopens no page, navigates nothing, and reads no browsing data. It publishes the shared secret and returns the steps. The single exception, since 1.3.1, is that it may tell an already-installed companion older than the bundled build to reload itself; installing the companion stays a deliberate user action in Chrome's own UI.Plain
http://to public hosts is refused unlessWEB_SEARCH_NEO_ALLOW_PLAIN_HTTP=1is set; loopback and private-network addresses are reachable when addressed directly, but a public host redirecting to one is refused, and link-local/cloud-metadata addresses are always blocked. Cross-origin redirects drop credential headers, andsave_towrites only underWEB_SEARCH_NEO_DOWNLOAD_DIR(default./downloads), never over an existing file withoutoverwrite=true.File upload tools can upload local paths supplied to the tool. Review agent actions and scope filesystem access appropriately.
Browser automation may be restricted by a site's terms of service. Use it only where you are authorized.
Manual challenge mode hands control to you. A paid CAPTCHA solving service is contacted only on
captchamode='solve'(ormode='auto'withWEB_SEARCH_NEO_CAPTCHA_AUTO_SOLVE=1) and only whenWEB_SEARCH_NEO_CAPTCHA_KEYis set.
Contributing
Issues and focused pull requests are welcome. A new search engine only needs a SearchProvider implementation plus register_search_provider(provider); status, cooldown, cache, and fallback routing update automatically.
See TODO.md for the current roadmap.
Focused interactive controls
On a noisy page, ask for one deduplicated list of controls instead of all categories:
{"topic":"page_elements","params":{"session_id":"my-task","category":"interactive","visible_only":true,"enabled_only":true,"limit":30,"max_chars":8000}}Narrow further with role: "button", text_pattern: "Save" (case-insensitive
accessible name, rendered text, label or placeholder), or href_pattern. Other
categories are links, buttons, fields, forms, iframes; all preserves
the legacy default output. Interactive includes native controls, editable areas,
ARIA widgets and focusable elements, without duplicate rows for the same element.
Filters run before pagination and the response budget. Continue with
range.interactive.next_offset; found.interactive counts matching controls
within the collector limit. Visible means rendered, not necessarily in the viewport;
scroll and verify before acting on a control outside the viewport.
Available Tools
2 toolsweb_actionA
Execute 1-32 ordered actions. Read web_info action_schema before unfamiliar actions or after validation errors. fill uses fields={CSS_selector: value}. Use a unique session_id per task/agent; never close another agent's tabs.
| Name | Required | Description | Default |
|---|---|---|---|
| actions | Yes | ||
| continue_on_error | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden of behavioral disclosure. It covers ordering, count limits, session isolation, tab-safety, and the fill field syntax. It doesn't explain error propagation, atomicity, or side effects of continue_on_error, but the guidance provided is far beyond a generic 'execute actions' statement.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tightly packed sentences, with the core behavior first and operational guidance after. Every sentence earns its place; no filler or restatement of the name.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a flexible tool that accepts 1-32 heterogeneous actions, delegating full action details to web_info action_schema is a sensible division of labor. The description includes the essential constraints and a concrete example. It is not fully self-contained, but it gives the agent a clear path to the missing details while the output schema covers return values.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It does explain the fill action format (fields={CSS_selector: value}), which adds real parameter meaning. However, it does not explain the structure of the actions array items or the continue_on_error parameter, leaving substantial gaps.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Execute 1-32 ordered actions.' It adds critical scope details (count limit, ordering) that differentiate it from sibling web_info, which is clearly the informational/schema tool. An agent can tell what web_action does without inspecting the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly directs the agent to 'Read web_info action_schema before unfamiliar actions or after validation errors,' which effectively explains when to consult the sibling vs. act. It also gives operational constraints: use a unique session_id per task/agent and never close another agent's tabs. This is clear context, though it stops short of explicit 'use X instead of Y' phrasing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
web_infoB
Read the contract, the page, the console, the network, or search/browser state.
Called with no arguments it returns the whole contract, including recipes and
common mistakes, so no external skill file is needed. Every result (dict
payloads) also carries the current local date/time and UTC-offset region
under the top-level now key.
| Name | Required | Description | Default |
|---|---|---|---|
| topic | No | capabilities | |
| params | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, and it does disclose the no-argument behavior and the guaranteed top-level 'now' key in results. However, it does not explain the behavior of all topic values, notably 'execute_js', which may not be purely read-only, leaving some behavioral ambiguity.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is brief, front-loaded with the core purpose, and every sentence adds useful information. It wastes no words and is easy to scan.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description is reasonably complete for a meta/introspection tool because it tells the agent to call with no arguments to pull the full contract, which effectively self-documents the rest. It still leaves topic-specific output and parameter expectations unspecified, but the contract mechanism compensates.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description needed to compensate, but it only mentions the no-argument default behavior. It does not explain what values of 'topic' mean or how the free-form 'params' object should be structured for each topic.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a clear verb ('Read') and a set of resources ('the contract, the page, the console, the network, or search/browser state'), which maps well to the tool's intent. It does not explicitly differentiate from the sibling 'web_action', but 'read' versus action is reasonably implied.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives a concrete usage rule for the no-argument case: calling with no arguments returns the whole contract and removes the need for an external skill file. It does not explicitly say when to prefer this tool over 'web_action' or when a mutation-style action should be used instead.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
2 tool updates
v1.15.0- First observed
web_action - First observed
web_info
TDQS
Scored across 2 tools
The two tools have completely distinct roles: web_action executes operations, while web_info retrieves state and context. There is no functional overlap, so an agent should never confuse them.
Both tools follow the same 'web_' prefix with a clear noun suffix ('action' vs 'info'). The naming is consistent, short, and predictable, even though it uses nouns rather than verb_noun.
Two tools is on the low end and feels thin for a web automation domain. While every tool is broad and earns its place, the server likely relies on complex schemas to compensate for the minimal count.
The action/info split covers the fundamental read-operate cycle of browser automation, and web_action's ordered actions suggest broad capability. Minor gaps may exist (e.g., no obvious dedicated session-management tool), but the design appears to cover core needs.
Maintenance
Related MCP Connectors
Web and URL utilities over MCP: shorten URLs, screenshot pages, read page metadata, encode URLs.
Web MCP: scrape/crawl sites, web search, brand assets, app stores, YouTube, Reddit, Hacker News.
Scrape, crawl and search the web for AI agents via MCP.
Web search, URL content extraction to Markdown, site mapping, and recursive web crawler.
Related MCP Servers
- AlicenseNot gradedqualityCmaintenanceEnables web browsing capabilities for locally served LLMs through URL text fetching, link extraction, and web search using Brave and DuckDuckGo engines. Designed to enhance LLMs with real-time web access through the MCP protocol.MIT
- FlicenseDqualityDmaintenanceProvides web connectivity tools for searching the web via DuckDuckGo or SerpAPI, fetching URL content, and extracting readable text from web pages.32-
- AlicenseBqualityBmaintenanceEnables AI agents to perform web searches and extract full-text content from web pages via standard MCP tools, with fallback search and semantic reranking.22MIT
- AlicenseNot gradedqualityCmaintenanceEnables web search through one or two SearxNG instances and retrieval of web pages, GitHub READMEs, and article text via MCP tools.149 npmGPL 2.0