Skip to main content
Glama

WebSense MCP

Release License: MIT Chrome MV3

Non-vision, AI-native web automation via a Chrome extension. No screenshots, no debug port, no bot detection — the model reads structured JSON and acts through the browser's own input pipeline.

WebSense gives an AI agent hands on a real, logged-in Chrome profile. The agent gets a lossless, addressable map of the page, acts on it by reference, and is told — with a grouped diff — whether the page actually changed. No vision model, no headless browser, no remote-debugging port.


Requirements

  • Node.js ≥ 18

  • Google Chrome / Chromium (Manifest V3, offscreen WebSocket bridge). Chrome-only: there is no Firefox code in this repo.

  • OS-level input is Windows-only. Everything else (browse / find / act, trusted input, frames, diffs) is cross-platform. act{how:"os"}, real_click, real_paste, real_activate_tab and dialog{keystroke} use Windows SendInput / PowerShell.

  • main_world (the CSP-proof MAIN-world read path) needs Chrome 138+ with the per-extension "Allow User Scripts" toggle enabled.

Related MCP server: live-mcp

Install

git clone https://github.com/spliffspliff70-wq/websense-mcp
cd websense-mcp
npm install

1. Load the extension. Open chrome://extensions, turn on Developer mode, click Load unpacked, and select the extension/ folder. It connects to the WebSocket hub on ws://127.0.0.1:38401 automatically — there is no launcher page. (For main_world, open the extension's Details and enable Allow User Scripts.)

2. Register the MCP server with your client.

stdio — one client per server process:

{ "mcpServers": { "websense": { "command": "node", "args": ["src/server.js"] } } }

streamable HTTP — one server, many clients:

node src/server.js --http --http-port 9222
# then point each client at http://localhost:9222/mcp

3. Call websense_guide (or just browse). The guide is the runtime source of truth and documents every tool.

Bridge port. Default 38401, plain ws:// on 127.0.0.1 (loopback is exempt from mixed-content blocking, so it works from HTTPS pages). Override with --port <n> on the server and the matching PORT constant in extension/offscreen.js. If the port is already taken the hub logs a warning and the server keeps running — MCP still works, the bridge just isn't claimed. Run isolated servers with different --port values.

The model's surface: 8 listed tools

A model sees exactly eight tools. The rest stay callable by name but are not listed, so a model never has to choose between a wall of one-verb tools.

tool

what it does

browse

TOOL 1 — open (or bind) a tab and map it in one call: navigate, seed the diff baseline, collect a lossless inventory, and return only the INDEX + the page's vocabulary + a region outline.

find

TOOL 2 — search the stored inventory. Every hit answers WHERE (region + branch chain, resolved from parent pointers) and WHAT (role / name / attrs / state).

act

Do something: click · hover · rightclick · drag · type · key · form · upload · scroll · dialog. Add how:"trusted" for the browser's own input pipeline.

page_slice

Full-fidelity records for one slice of the inventory (tag / role / region / vp / interactive / query), or any cached part of a diff.

tabs

Tab/window ops: list · switch · close · bind · frames · windows · focus · move · transfer · switchread.

debug

WebSense itself + raw reads: status · session · logs · cookies · clipboard · screenshot · ax · evaluate · main_world · explore_page · reload · respawn · guide.

preflight

Call this first when anything is refused. Walks launch → server → extension → binding → page-ready, stops at the first broken link, and returns the command that fixes it. Read-only by default.

websense_guide

Start here — returns the full in-tool guide.

38 tools are registered; 30 of them are unlisted but callable by name. The listable extras are summarised under Other registered tools.

Quick start — the loop

  1. browse {url} — one call that navigates (or binds), seeds the diff baseline, stores a lossless inventory of every element (nothing filtered, capped or truncated) and returns the small INDEX + a region outline. The full records stay server-side, addressable by index.

  2. find {query} — locate the control. Each hit tells you WHERE it is and WHAT it is, so five controls called "New" are distinguishable.

  3. act {action, ref, ...} — do it. Reach for how:"trusted" when the page checks isTrusted, reads coordinates, or a default action must run. Works in a background tab — no focus steal.

  4. Read the diff the reply carries.

The diff (did it land?)

Every mutating action returns a second block: a grouped DIFF against your browse baseline.

  • structure — the page's shape changed (elements added/removed; tag/role/name/attrs changed). Page truth.

  • content — the same element's value/text changed and its shape did not. The page answered you.

  • viewport — only vp/x/y differ. That is scroll/layout churn and is not a mutation.

mutated is true only when structure or content moved — viewport churn can never make an action look landed. The reply ends with a FULL DIFF: <handle> line; fetch any part with page_slice{diff:"<handle>", part:"structure|content|visual|viewport"}. The full delta is cached because it can be hundreds of KB for one action — the summary is what changed, the handle is the rest. Pass verify:false to skip the diff on a call you don't need checked.

A navigation is the strongest confirmation and is not in the groups: when a click or press_key (Enter/Space) replaces the document, the result carries effect:"confirmed" plus a navigation {from,to}, and the diff line says so explicitly — a diff across a navigation compares two different documents, so its groups are meaningless.

Verdicts, and when they are wrong

effect is derived from page_state — url, title, readyState, scroll. That is a deliberately weak signal, and an action that only changes the DOM does not move any of them. Two rules make that honest:

  • The diff can upgrade the verdict. When the page-side differ measures a real structure/content move (mutated:true), an unverifiable or suspected_noop verdict is upgraded to confirmed and carries effectSource:"page_diff", and its stale escalation advice is removed with it. Measured on x.com: every trusted type into the thread composer used to answer unverifiable while the same reply said mutated:true — a verdict contradicting its own evidence, which reads as "no proof" and makes a caller re-run an action that already worked. It only ever upgrades: a failed verdict (the action layer refused — disabled, read-only) stays failed. The measurement happened; classifyEffect simply could not see it.

  • An ambiguous selector is refused, not guessed. querySelector returns the first match with no word, so a selector matching two elements silently drove the wrong one. Measured on x.com's /compose/post: [data-testid="tweetTextarea_0"] matches twice — the dialog's real composer and the empty page-level inline composer. You now get a refusal naming the count and the tag (ambiguous selector … matches 2 elements — refusing to pick one silently). Scope it and retry: [role="dialog"] [data-testid="tweetTextarea_0"], or any selector that is unique on that page. This is also why find returns all matches with no cap — a cap would hide the second element and make the ambiguity invisible.

effect remains weak evidence for everything else: confirmed means "the page measurably moved", never "the app accepted and persisted it". Re-read the field or the page when the outcome matters.

Trusted input — act{how:"trusted"}

how:"trusted" drives chrome.debugger + Input.dispatchMouseEvent / Input.dispatchKeyEvent, so the page receives isTrusted events and the browser itself runs the default action (a link navigates, an Enter submits, a checkbox toggles, an arrow key moves a slider) instead of the tool guessing. It works in a background tab — no focus steal, no window activation. Chrome shows its "debugging this browser" infobar while attached.

Off-screen targets are scrolled into view first. Measured: a trusted click aimed at y=1727 below the fold hit nothing; after the fix it re-aimed at y=938 and the page recorded the click.

how:"os" is the OS-level rung (Windows SendInput) — reach for it only when a page rejects programmatic input outright, or for a raw-input / canvas surface. It lands on the frontmost window and so steals focus.

Dialog handling

  • JS dialogs, by type — this is measured behavior, not a blanket rule. The MAIN-world hook shadows alert only (its return is undefined, so nothing can branch on it) — it is captured into pendingDialogs and auto-dismisses after 30s, matching what Chrome itself does in a background tab. confirm/prompt stay NATIVE: a hooked confirm returns a Promise (always truthy), so every if(confirm(...)) would take the TRUE branch regardless of the answer — the auto-answer-after-30s behavior was a bug, not a safety net, and it is gone. A native confirm/prompt blocks the page; answer it on request with dialog{native:true, action:"accept"|"dismiss", value:promptText} — that goes through Page.handleJavaScriptDialog (chrome.debugger), so the page's own branch follows the agent's real choice, on a background tab, with no focus steal. Nothing ever auto-answers a decision the agent did not make.

  • DOM modals ([role=dialog], most in-app modals) are closed by ref; status.hasModal / dialogCount come from a visibility-blind scan (a hidden modal still counts).

  • OS-level dialogs (HTTP basic-auth, proxy-auth, print) cannot be intercepted by JS. dialog{keystroke:true, key:"enter"|"escape"} injects a global keystroke through Windows control (PowerShell SendKeys) — Windows-only.

  • File picker: handled by form{action:"upload"} (DataTransfer API) — no OS dialog.

Iframes / frames

Same-origin iframes are walked and clickable — verified: the frame echoes the click. List them with tabs{action:"frames"} (pass tabId; omit it for your bound tab) and pass frameId to any element tool — this unlocks Gmail compose, Notion, Figma and any site that renders key UI inside child frames. Cross-origin frames are skipped: nothing in the page can read them and no click can be aimed inside them.

Other registered tools

These are registered and callable by name — the standalone forms that act / debug absorbed are noted inline:

  • explore_page — quick look at a page's actions (SAG). compact:true, intent:"submit", goal:"log in", preload:true, incremental:true. For a full page map use browse + find instead.

  • read — page text: text · content · markdown · diff · scrollextract · preload.

  • click — click a ref (default) · mode:"hover" · "rightclick" · "drag" (fromRef/toRef) · x,y for canvas.

  • trusted_click · trusted_key — the standalone trusted mouse / keyboard paths.

  • press_key — synthetic key events only; runs no default action. Use trusted_key / act{how:"trusted"} when the default matters.

  • type_text — fill one input (React-safe native setter) or fields:[{ref,text},…] for a verified batch.

  • form — state · select · toggle · special · upload.

  • reveal — pre-extract hidden content: dropdown · tabs · accordion.

  • scroll — direction+amount (ticks) · y absolute · intoView.

  • status — page · bridge · doctor · downloads.

  • wait — poll conditions (ANDed) until met, or wait for an event.

  • evaluate — run JS and return its value (auto-reroutes through the MAIN world on a CSP block) or a no-eval query DOM read.

  • main_world — run a compiled function in the page's MAIN world (CSP-proof; needs "Allow User Scripts").

  • ax — native accessibility tree via chrome.debugger (for canvas SPAs and chrome:// pages).

  • screenshot — captureVisibleTab → PNG/JPEG dataUrl, for a vision model.

  • dialog — accept | dismiss (+ value); captures the page's own JS alert/confirm/prompt.

  • session — reset · map · mermaid.

  • network_log · console_log — captured page fetch/XHR · console + JS errors.

  • cookies — list · get · clear (values are masked on other surfaces).

  • clipboard — copy · read.

  • inspect — element · geometry · relation.

  • navigate — navigate a tab (reuses your bound tab; no tab spam).

  • page_snapshot — collect / return the lossless inventory index directly.

  • respawn_offscreen · extension_reload — extension maintenance (MV3 traps).

  • real_activate_tab · real_click · real_paste — Windows OS-level input; the last rung.

Architecture

MCP Client (Claude / Cline / Cursor / Hermes)
    ↔ stdio or streamable HTTP
WebSense MCP Server (src/server.js)
    ↔ WebSocket  ws://127.0.0.1:38401
Chrome Extension (extension/)
    ├── background.js    service worker, tab management, binding
    ├── offscreen.js     WebSocket client, auto-reconnect
    └── websense-cs.js   SAG extraction + native DOM interaction (CSP-safe)
    ↔ chrome.runtime.sendMessage / chrome.debugger (trusted input)
Live DOM

What's new in 2.0

  • A 7-tool listed surface — browse · find · act · page_slice · tabs · debug · websense_guide. act and debug are facades that dispatch to the real handlers, so the 30 unlisted tools remain callable by name with identical behaviour.

  • browse → find → act → read the diff replaces explore_page → click/type as the primary loop. browse returns an INDEX over a lossless inventory; find returns WHERE + WHAT per hit; page_slice loads one branch at full fidelity.

  • A grouped auto-diff after every mutating action, with mutated derived from structure + content only (viewport churn cannot fake a landing), a summary in the reply, and the full delta cached behind FULL DIFF: <handle>.

  • Trusted input: act{how:"trusted"} runs through chrome.debugger + Input.dispatchMouseEvent / Input.dispatchKeyEvent, so the page sees isTrusted events and the browser performs the default action — in a background tab. Off-screen targets are scrolled into view first.

  • Same-origin iframes are walked and clickable.

Known limitations (honest)

  • Trusted drag works end-to-end (verified on the fixture): act{action:"drag", how:"trusted"} produces trusted dragstart/dragenter/dragover and a real drop, from a background tab, with both ends scrolled into view automatically. The drop only fires if the target accepts drops (preventDefault() on dragover) — a non-zone target correctly gets dragleave, exactly as a real mouse would, so check the target before blaming the driver. The plain drag mode still fires the whole sequence but its events are not trusted — use how:"trusted" when the page checks isTrusted.

  • OS-click equivalence is unproven. act{how:"trusted"} is the browser's own input pipeline; it is not proven byte-identical to a real OS click. real_click (Windows SendInput) is the only genuinely OS-level path.

  • Canvas / WebGL: coordinate clicks are TRUSTED (fixed 2026-10-02 — this line said "not trusted" and was stale). act{action:"click", how:"trusted", x, y} drives Input.dispatchMouseEvent at the raw viewport point, so a canvas that inspects isTrusted now sees a real one. The untrusted form still exists as how:"auto" and lands within 1px; prefer how:"trusted" on canvas/WebGL, which is exactly the surface class that checks the flag.

  • Chrome-only. MV3 + offscreen WebSocket bridge; no Firefox code.

  • OS-level input is Windows-only (real_*, dialog{keystroke}).

  • OS-level input additionally needs Python 3 with pyautogui + pywinauto (Windows only). Everything else — the whole browse/find/act loop, trusted input, frames, diffs — needs neither Python nor Windows. The server discovers an interpreter (py → python → python3); override with WEBSENSE_PYTHON=/path/to/python. If none is found the error says so instead of failing as a mysterious page problem (fixed 2026-10-02 — the interpreter path was previously hard-coded to the maintainer's machine, so OS input could not work for anyone else).

  • main_world requires the "Allow User Scripts" toggle (Chrome 138+). Without it the CSP-proof MAIN-world path is unavailable.

  • evaluate{script} runs your JS and returns its value; the isolated-world new Function path is blocked by the extension's own MV3 CSP, so it transparently re-routes through the MAIN world (chrome.userScripts, no eval) and reports via:"main_world". evaluate{query:{…}} remains the lighter path for plain DOM reads.

  • A JS dialog raised while the tab is hidden can be auto-dismissed by Chrome before anyone sees it — act fast: dialog{native:true, action, value} answers a still-pending native confirm/prompt via Page.handleJavaScriptDialog on the background tab (no activation needed). A hooked alert is still recorded (status.recentDialogs); native confirm/prompt are NOT in the hook's queue (they are never shadowed), so check status and answer immediately if the branch matters.

  • ax attaches chrome.debugger and shows Chrome's warning banner while attached; it requires an explicit tabId.

  • One profile, per-tab isolation. Concurrent jobs share one Chrome profile — there is no cookie/storage isolation between them. Scope work with tabs{action:"bind", tabId} + an explicit tabId. Session state (map/history) is per-session: session{action:"reset"} clears only your own history.

  • Logged-in sites (LinkedIn, etc.) must already be authenticated in that Chrome profile; navigate opens a fresh tab that needs an existing session cookie.

  • Refs are stable — E# refs are assigned in viewport order on the first scan, then held by element identity (a per-element cache + a data-websense-ref attribute), so they survive re-explores, scrolls and framework re-renders. A ref dies only when its element leaves the DOM with nothing to heal from. CSS-selector refs (#id) remain the safest choice for anything long-lived or across navigations.

Testing

# Regression suite — no Chrome needed (hub, diff, snapshot, trusted-input,
# guide-truth and doc-drift guards). Currently 170 tests.
npm test

# Print the LISTED surface from a running server (tools/list is filtered to it)
node tools/tools-list.mjs          # names-only preflight
node tools/tools-list.mjs --full   # name + one-line description

# End-to-end MCP client test (needs Chrome + the extension loaded)
node test/mcp-client-test.js

# Keep MODEL_PROMPT.md in sync with the in-tool guide (also enforced inside npm test)
node tools/export-guide.mjs --check

File structure

websense-mcp/
├── src/
│   ├── server.js          # MCP server: tool registration, the 7-tool listed surface, facades
│   ├── hub.js             # WebSocket hub on ws://127.0.0.1:38401
│   ├── session.js         # exploration map + per-session state
│   ├── snapshot.js        # lossless page inventory (collector + slicer)
│   ├── diff-cache.js      # cached grouped diffs behind FULL DIFF handles
│   ├── diff-collector.js  # in-page diff collection
│   └── climb.js, incr.js, upload.js, summarize.js, mermaid.js
├── extension/
│   ├── manifest.json      # Chrome MV3
│   ├── background.js      # service worker (tab mgmt, offscreen lifecycle)
│   ├── offscreen.js       # WebSocket client (auto-reconnect)
│   ├── websense-cs.js     # built content script (generated — do not hand-edit)
│   └── cs-src/            # content-script sources (edit here; build with tools/build-cs.mjs)
├── test/
│   └── mcp-client-test.js # end-to-end MCP client test
└── tools/                 # build + measurement scripts (build-cs, export-guide, tools-list…)

License

MIT — see LICENSE.

Available Tools

7 tools
actC

DO something: click, hover, rightclick, drag, type, key, form, upload, scroll, dialog.…

ParametersJSON Schema
NameRequiredDescriptionDefault
xNo
yNo
howNoauto = normal path, trusted = browser in…
keyNo
refNo
textNo
tabIdNo
toRefNo
valueNo
actionYes
amountNo
frameIdNoframeId (omit=top)
fromRefNo
filePathNo
selectorNo
directionNo
modifiersNo

TDQS

C2/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden, and it discloses nothing: no mention of destructive effects, focus requirements, wait/retry behavior, auth needs, or what happens on failure. For a 17-parameter interaction tool spanning drag, upload, and form submission, this is a serious omission.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The single line is short, but nearly all of its content duplicates the action enum already present in the schema, and the trailing ellipsis signals an unfinished thought. It is under-specified rather than efficient, so size is not the problem—value per word is.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With 17 parameters, 12% schema coverage, no annotations, and no output schema, the description is the only place an agent could learn behavioral expectations—and it provides none. Nothing about per-action parameter requirements, side effects, or result handling is covered, leaving the definition wholly inadequate for the tool's complexity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 12% across 17 parameters, so the description must compensate, and it does not. It restates the action enum but says nothing about how x/y pair with coordinates, how ref/selector/fromRef/toRef relate, or which params are required per action (e.g., type needs text/value, upload needs filePath).

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names the tool's domain by enumerating the actions it supports (click, hover, drag, type, form, upload, scroll), which hints at browser/page interaction. However, 'DO something' is an empty verb and it never states the resource being acted on, nor does it distinguish act from siblings like browse or find. Purpose is inferable from the enum list, not from a clear statement.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use act versus browse, find, page_slice, or tabs, and no exclusions or prerequisites. An agent must infer that this is the interaction tool purely from the action names. No when-to-use context is supplied at all.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

browseB

TOOL 1. Go to a page and map it in ONE call: navigate (or bind an existing tab), seed the auto-diff BASELINE f…

ParametersJSON Schema
NameRequiredDescriptionDefault
urlNo
freshNoRe-collect even if a snapshot for this t…
tabIdNo
newTabNoForce a fresh tab rather than reusing th…
frameIdNoframeId (omit=top)

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It does disclose a meaningful behavioral trait that annotations would not cover: that it seeds the auto-diff BASELINE, which tells the agent state is being persisted for later diffing. It does not explain what 'map' produces, whether prior baselines are overwritten, or any side effects, and the text is cut off mid-sentence.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The leading 'TOOL 1.' is wasted tokens and the whole entry appears truncated mid-word ('BASELINE f…'), which harms structure. The core action is front-loaded reasonably well, so density is decent, but the artifact prefix and cutoff keep this from scoring higher.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With 5 parameters, no annotations, and no output schema, the description should explain what 'map' returns and how the baseline behaves. Instead it is truncated and omits all of that, leaving an agent unable to predict results or side effects for a tool with no other structured guidance.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 60%, so baseline is 3. The description adds marginal meaning by framing url as a navigation target and tabId as an existing tab to bind, which the schema leaves undocumented. It says nothing about fresh, newTab, or frameId beyond the truncated schema hints.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: 'Go to a page and map it in ONE call: navigate (or bind an existing tab)'. This distinguishes it from siblings like tabs and page_slice by emphasizing the combined navigate-and-map behavior in a single call. However, the 'TOOL 1.' prefix is noise and the description is truncated, leaving the mapping output undefined.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Implies usage context by noting it can navigate or bind an existing tab, and that it seeds an auto-diff baseline, so an agent can infer this is the entry point for page analysis. But it never states when to prefer this over siblings such as page_slice, find, or tabs, nor any preconditions, so guidance stays implicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

debugD

WebSense itself, and raw reads: op=status, session, logs, cookies, clipboard, screenshot, ax, evaluate, main_w…

ParametersJSON Schema
NameRequiredDescriptionDefault
opYeswhat to inspect or maintain
urlNo
funcNo
kindNofor logs: network | console
queryNo
tabIdNo
actionNo
frameIdNoframeId (omit=top)

TDQS

D1.5/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description must carry the full behavioral burden, but it only says 'raw reads' and lists ops without explaining side effects, permissions, rate limits, or safety for mutating ops like reload, respawn, or evaluate. It also truncates before finishing the list.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single truncated fragment ending in an ellipsis, lacking a front-loaded purpose. It is under-specified rather than concise, wasting the opportunity to state what the tool does and how to use it.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given 13 ops, 8 parameters, no annotations, no output schema, and incomplete parameter documentation, the description is far too incomplete to guide correct invocation. It leaves critical gaps in purpose, usage, and parameter mapping.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 38%, so the description should compensate for undocumented parameters. It merely enumerates some op values and leaves url, func, query, tabId, action, and frameId without added semantic meaning, adding little beyond the enum itself.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description is a truncated fragment that lists op values rather than stating a clear verb+resource. 'raw reads' gives a hint, but the overall purpose is vague and does not distinguish this tool from siblings like act, browse, or websense_guide.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance, no alternatives, and no prerequisites are provided. The description offers no routing information among the sibling tools, leaving the agent to guess when this tool is appropriate.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

findC

TOOL 2. Search the stored page inventory and get WHERE and WHAT the hit is: its own role/name/attributes/state…

ParametersJSON Schema
NameRequiredDescriptionDefault
vpNotrue = in viewport only
tagNoElement tag
attrNoMatch any attribute the page wrote
roleNo
fieldNotrue = only form controls (platform-repo…
limitNo
queryNo
tabIdNo
regionNoRegion substring (region is derived from…
frameIdNoframeId (omit=top)
indicesNoFetch exact records by inventory index —…
focusableNotrue = only focusable elements (el.…
branchDepthNoHow many ancestors to include in the bra…
interactiveNotrue = only controls (derived: focusable…

TDQS

C2.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations exist, so the description carries the full burden. It conveys one useful behavioral fact — that searching operates over a stored/cached page inventory rather than the live page — but says nothing about read-only nature, staleness, pagination, or the shape of the result, and the sentence is cut off.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single truncated fragment ending in an ellipsis, prefixed with an unexplained 'TOOL 2' label. It is neither complete nor front-loaded with actionable content.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 14-parameter, zero-required search tool with no annotations and no output schema, the description is far too thin: it omits result shape, empty/no-match behavior, and the meaning of the untagged filter parameters.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 14 parameters at 71% schema coverage, several parameters (role, limit, query, tabId) have no schema description at all, and the description adds zero parameter meaning. Since coverage is below the 80% threshold, the description should compensate and does not.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a verb (search) and resource (stored page inventory) and hints at the return payload ('WHERE and WHAT the hit is: its own role/name/attributes/state'). However it is truncated mid-sentence and gives no differentiation from siblings like page_slice or act, so an agent cannot confidently tell it apart.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'TOOL 2' implies a pipeline position, but the description never states when to prefer this over page_slice, browse, or act, nor any preconditions (e.g. that an inventory must already be captured). No when-to-use or when-not guidance is present.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

page_sliceC

Full-fidelity records from the stored page snapshot, filtered to ONE slice: indices / tag / role / region / vp…

ParametersJSON Schema
NameRequiredDescriptionDefault
vpNotrue = in viewport only, false = off-vie…
tagNoFilter by tag, e.g. input
attrNoFilter by ANY attribute the page wrote, …
diffNoA FULL DIFF handle from a DIFF block (e.…
partNoWhich part of a cached diff (default all…
roleNo
fieldNotrue = form controls only (platform-repo…
limitNo
queryNo
tabIdNo
regionNoFilter by region substring — region is d…
frameIdNoframeId (omit=top)
indicesNoFetch exact records by inventory index —…
focusableNotrue = focusable only (el.…
interactiveNotrue = actionable only, DERIVED (focusab…

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are supplied, so the description carries the full burden. It does hint that results come from a stored snapshot rather than a live page, which is useful context, but it says nothing about whether the operation is read-only, whether snapshots must pre-exist, what happens on empty slices, pagination, or response shape.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is a single front-loaded sentence with no padding, which is good, but the sentence is cut off mid-enumeration ('vp…'), so its brevity stems partly from incompleteness rather than discipline. Structure is acceptable, content is not fully delivered.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 15-parameter, no-required, no-annotation, no-output-schema tool, the description is far too thin. Nothing explains how the many filters combine, whether 'ONE slice' is enforced, or what the records look like, leaving major gaps an agent must guess at.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 73%, below the high-coverage baseline, so the description should help disambiguate the 15 filters. Instead it only echoes a few slice names (indices/tag/role/region/vp) with no syntax, interaction, or precedence guidance, and truncates before finishing. It adds essentially no meaning beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description identifies a retrieval verb and resource (full-fidelity records from a stored page snapshot), which is more than a tautology. However, it is truncated mid-sentence ('vp…') and the claim of 'filtered to ONE slice' sits uneasily against 15 independent filter parameters, leaving the actual selection model vague. It does not clearly distinguish itself from siblings like find or browse.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit when-to-use, when-not-to-use, or alternative named. An agent cannot infer from the text whether this is preferred over find or browse for locating elements. Usage is only weakly implied by the phrase 'from the stored page snapshot.'

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

tabsC

Tab/window ops. bind routes page ops to a tab WITHOUT focus — page ops NEVER need activation.…

ParametersJSON Schema
NameRequiredDescriptionDefault
tabIdNo
toTabNotransfer: destination tab
actionYes
fromTabNotransfer: source tab
activateNobind: ALSO make this the OS-active tab.…
selectorNo
useValueNotransfer: copy input VALUE instead of vi…
windowIdNo
toSelectorNotransfer: destination selector
fromSelectorNotransfer: source selector

TDQS

C2.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden of behavioral disclosure. It reveals one behavior (bind operates without focus/activation) but remains silent on the side effects, prerequisites, or outcomes of other actions (close, switch, move, transfer, etc.). It does not contradict any annotations since none exist.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is very brief and front-loads the broad category 'Tab/window ops' before a specific detail about 'bind'. It is concise, but the truncation with an ellipsis suggests the full description may include more content. As presented, it is efficient but structurally minimal.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a multi-action tool with 10 parameters, 1 required, no output schema, and no annotations, this description is severely incomplete. It does not explain the different actions, parameter usage, return values, or error conditions. An agent would need to inspect the schema thoroughly and still lack context for the undocumented parameters. This is far from a usable definition.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 60%, leaving four parameters (tabId, action, selector, windowId) without any schema-level descriptions. The tool description does not mention any parameters or add meaning beyond the schema. For the uncovered parameters, nothing is provided, so the description fails to compensate for the gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states the tool handles 'Tab/window ops' and highlights the 'bind' action's specific behavior (routing page ops without focus). This gives a general sense of purpose but does not enumerate the full range of actions (list, switch, close, frames, windows, focus, move, transfer, switchread). It differentiates itself from some siblings by mentioning 'bind' but not clearly against all overlapping tools like real_activate_tab.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies that page ops never need activation, which hints that for such operations you may not need to use tab activation tools, but it does not explicitly state when to use this tool vs alternatives like real_activate_tab or session. No exclusions or explicit conditions for each action are provided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

websense_guideB

START HERE. The listed tools: browse, find, act, page_slice, tabs, debug, guide. Call once before you act.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

B3.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It implies a one-time, presumably side-effect-free read via 'call once', but never states that it is read-only, what it returns, or whether it has any cost or auth requirement.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Very brief and front-loaded with 'START HERE', which is exactly right. The enumeration of sibling tool names is somewhat redundant given they are already exposed, but it costs little.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a parameterless, output-schema-less orientation tool, the description is thin: it tells the agent to call it but not what the returned guidance contains, so the agent cannot anticipate its value beyond 'read this first'.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so there is no parameter semantics to document; the baseline of 4 applies and no schema gap exists to compensate for.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description positions the tool as an entry point ('START HERE', 'guide'), which does distinguish it from the action-oriented siblings. However, it never states what the tool actually returns or what kind of guidance it provides, and the enumerated tool list largely restates the sibling names already visible to the agent.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'Call once before you act' gives an explicit, actionable trigger for invocation and implies it should not be repeated. It stops short of any when-not guidance or explanation of what happens if you skip it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 33 tool updatesv2.0.0
    • Addedact
    • Removedax
    • Addedbrowse
    • Removedclick
    • Removedclipboard
    • Removedconsole_log
    • Removedcookies
    • Addeddebug
    • Removeddialog
    • Removedevaluate
    • Removedexplore_page
    • Removedextension_reload
    • Addedfind
    • Removedform
    • Removedinspect
    • Removedmain_world
    • Removednavigate
    • Removednetwork_log
    • Changedpage_slice8 fields changed
      • addedInput schema / properties / attr
        Added value: +{
        +  "anyOf": [
        +    {
        +      "type": "string"
        +    },
        +    {
        +      "properties": {
        +        "name": {
        +          "type": "string"
        +        },
        +        "value": {
        +          "type": "string"
        +        }
        +      },
        +      "required": [
        +        "name"
        +      ],
        +      "type": "object"
        +    }
        +  ],
        +  "description": "Filter by ANY attribute the page wrote, …"
        +}
      • addedInput schema / properties / diff
        Added value: +{
        +  "description": "A FULL DIFF handle from a DIFF block (e.…",
        +  "type": "string"
        +}
      • addedInput schema / properties / field
        Added value: +{
        +  "description": "true = form controls only (platform-repo…",
        +  "type": "boolean"
        +}
      • addedInput schema / properties / focusable
        Added value: +{
        +  "description": "true = focusable only (el.…",
        +  "type": "boolean"
        +}
      • addedInput schema / properties / indices
        Added value: +{
        +  "description": "Fetch exact records by inventory index —…",
        +  "items": {
        +    "type": "number"
        +  },
        +  "type": "array"
        +}
      • changedInput schema / properties / interactive / description
        Previous value: -"true = actionable elements only"New value: +"true = actionable only, DERIVED (focusab…"
      • addedInput schema / properties / part
        Added value: +{
        +  "description": "Which part of a cached diff (default all…",
        +  "enum": [
        +    "structure",
        +    "content",
        +    "visual",
        +    "viewport",
        +    "all"
        +  ],
        +  "type": "string"
        +}
      • changedInput schema / properties / region / description
        Previous value: -"Filter by region substring, e.g.…"New value: +"Filter by region substring — region is d…"
    • Removedpage_snapshot
    • Removedpress_key
    • Removedread
    • Removedreal_activate_tab
    • Removedreal_click
    • Removedreal_paste
    • Removedrespawn_offscreen
    • Removedreveal
    • Removedscreenshot
    • Removedscroll
    • Removedsession
    • Removedstatus
    • Removedtype_text
    • Removedwait
  2. 1 tool updatev1.4.7
    • Changedevaluate2 fields changed
      • changedInput schema / properties / script / description
        Previous value: -"JS to execute (eval mode)"New value: +"JS to execute. A bare expression or stat…"
      • addedInput schema / properties / tabId
        Added value: +{
        +  "type": "number"
        +}
  3. 14 tool updatesv1.4.5
    • Changedclick1 field changed
      • addedInput schema / properties / verify
        Added value: +{
        +  "description": "Set false to SKIP the automatic post-act…",
        +  "type": "boolean"
        +}
    • Changeddialog1 field changed
      • addedInput schema / properties / verify
        Added value: +{
        +  "description": "Set false to SKIP the automatic post-act…",
        +  "type": "boolean"
        +}
    • Changedevaluate1 field changed
      • addedInput schema / properties / verify
        Added value: +{
        +  "description": "Set false to SKIP the automatic post-act…",
        +  "type": "boolean"
        +}
    • Changedexplore_page4 fields changed
      • addedInput schema / properties / contentMaxLen
        Added value: +{
        +  "description": "Cap on extracted body text (default 8000…",
        +  "type": "number"
        +}
      • addedInput schema / properties / fresh
        Added value: +{
        +  "description": "Force a real re-scan. Without it, a repe…",
        +  "type": "boolean"
        +}
      • changedInput schema / properties / maxActions / description
        Previous value: -"Cap for compact mode (default 250)"New value: +"Cap on RETURNED actions (default 200).…"
      • addedInput schema / properties / settle
        Added value: +{
        +  "description": "false = skip the SPA hydration settle wa…",
        +  "type": "boolean"
        +}
    • Changedform1 field changed
      • addedInput schema / properties / verify
        Added value: +{
        +  "description": "Set false to SKIP the automatic post-act…",
        +  "type": "boolean"
        +}
    • Addedmain_world
    • Changednavigate1 field changed
      • addedInput schema / properties / tabId
        Added value: +{
        +  "type": "number"
        +}
    • Addedpage_slice
    • Addedpage_snapshot
    • Changedpress_key1 field changed
      • addedInput schema / properties / verify
        Added value: +{
        +  "description": "Set false to SKIP the automatic post-act…",
        +  "type": "boolean"
        +}
    • Changedreal_click1 field changed
      • addedInput schema / properties / verify
        Added value: +{
        +  "description": "Set false to SKIP the automatic post-act…",
        +  "type": "boolean"
        +}
    • Changedreal_paste1 field changed
      • addedInput schema / properties / verify
        Added value: +{
        +  "description": "Set false to SKIP the automatic post-act…",
        +  "type": "boolean"
        +}
    • Changedtabs1 field changed
      • changedInput schema / properties / activate / description
        Previous value: -"bind: also activate the tab"New value: +"bind: ALSO make this the OS-active tab.…"
    • Changedtype_text1 field changed
      • addedInput schema / properties / verify
        Added value: +{
        +  "description": "Set false to SKIP the automatic post-act…",
        +  "type": "boolean"
        +}
  4. 5 tool updatesv1.1.1
    • Addedextension_reload
    • Changedform2 fields changed
      • changedInput schema / properties / action / enum
        Previous value: -[
        -  "state",
        -  "select",
        -  "toggle",
        -  "upload"
        -]New value: +[
        +  "state",
        +  "select",
        +  "toggle",
        +  "special",
        +  "upload"
        +]
      • addedInput schema / properties / clearAll
        Added value: +{
        +  "description": "select on multi-select: deselect non-mat…",
        +  "type": "boolean"
        +}
    • Addedreal_activate_tab
    • Addedreal_click
    • Addedreal_paste
  5. 24 tool updatesv1.0.0
    • First observedax
    • First observedclick
    • First observedclipboard
    • First observedconsole_log
    • First observedcookies
    • First observeddialog
    • First observedevaluate
    • First observedexplore_page
    • First observedform
    • First observedinspect
    • First observednavigate
    • First observednetwork_log
    • First observedpress_key
    • First observedread
    • First observedrespawn_offscreen
    • First observedreveal
    • First observedscreenshot
    • First observedscroll
    • First observedsession
    • First observedstatus
    • First observedtabs
    • First observedtype_text
    • First observedwait
    • First observedwebsense_guide

TDQS

C2.5/5.0

Scored across 7 tools

Disambiguation4/5

Each tool has a distinct primary purpose (browse for navigation/mapping, find for search, act for interactions, page_slice for snapshot extraction, debug for raw reads, tabs for management, websense_guide for orientation). However, debug and page_slice both read page data, and find and page_slice both retrieve page inventory entries, so minor confusion is possible.

Naming Consistency2/5

Tool names mix verbs (browse, find, act) with nouns (tabs, page_slice) and a prefixed noun phrase (websense_guide). There is no consistent verb_noun or other predictable pattern, making the naming style inconsistent.

Tool Count5/5

Seven tools is well within the 3–15 range for a browser automation server. Each tool covers a broad area (e.g., act handles many actions, debug handles many raw reads), so the count feels well-scoped without redundancy.

Completeness4/5

The set covers navigation, search, interaction, tab management, raw reads, and page data extraction, which are core browser automation needs. Minor gaps include no explicit wait/synchronization tool and no download handling, though these can likely be worked around via debug.evaluate or act.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    B
    quality
    F
    maintenance
    Enables AI agents to directly control your real Chrome browser with full context including login sessions, cookies, and open tabs. It provides tools for page scanning, JavaScript execution, CDP control, screenshots, and physical mouse/keyboard input for authentic browser automation.
    20
    245
    MIT
  • F
    license
    Not graded
    quality
    B
    maintenance
    Enables AI agents to declaratively control web pages using real mouse and keyboard events via Chrome DevTools Protocol, without executing page JavaScript.
    10 npm
    1
    -