websense
WebSense is a non-vision, AI-native browser automation MCP server that drives a real logged-in Chrome profile via a Chrome extension, letting an agent map pages, act by reference, and verify changes with grouped diffs.
browse— open/bind a tab and map it in one call: navigate, seed a diff baseline, collect a lossless element inventory, and return an INDEX, page vocabulary, and region outline.find— search the stored inventory by query, tag, role, region, attr, index, viewport, focusable/interactive/field flags; each hit reports WHERE (region + ancestor branch chain) and WHAT (role/name/attrs/state), returning all matches with no cap.act— perform actions:click · hover · rightclick · drag · type · key · form · upload · scroll · dialog, targeting aref,selector, or rawx/y; supportshow:"auto" | "trusted" | "os"(trusted = browser's own input pipeline, background tab, no focus steal),modifiers,fromRef/toRef,filePath, andframeId.page_slice— pull full-fidelity records for one slice of the snapshot (indices, tag, role, region, vp, interactive, focusable, field, attr, query) or any cached part of a diff (structure | content | visual | viewport | all) via a FULL DIFF handle.tabs— tab/window management:list · switch · close · bind · frames · windows · focus · move · transfer · switchread, including moving a selector/value between tabs and binding page ops to a tab without activation.debug— introspect WebSense and read raw page data:status · session · logs · cookies · clipboard · screenshot · ax · evaluate · main_world · explore_page · reload · respawn · guide, with optionaltabId/frameId.websense_guide— returns the full in-tool guide (the runtime source of truth); call it first before acting.Verification built in — every mutating action returns a grouped DIFF (structure/content/viewport), an
effectverdict (upgraded toconfirmedon real DOM movement), and navigation confirmation; viewport churn never fakes a landing.Cross-cutting support — same-origin iframes walked and clickable, trusted input on canvas/WebGL, JS/DOM/OS dialog handling, Windows-only OS-level input via
real_*/how:"os", and stdio or streamable-HTTP transport (30 further unlisted-but-callable tools).
Enables browser automation of GitHub, allowing AI agents to navigate repositories, click, fill forms, and read state changes through the Semantic Action Graph.
Enables automation of the Gmail web interface for reading, composing, and managing mail via the browser extension.
Supports automation of Google web properties, including search and other Google pages, with structured actions and before/after diffs.
Supports automation of Telegram Web through the accessibility-tree ax tool, covering canvas/WebGL-heavy interfaces.
Supports automation of TradingView charts and the web app via the accessibility-tree ax tool, especially for canvas/WebGL content.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@websenseOpen Hacker News and show me the top 3 stories"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
WebSense MCP
Non-vision, AI-native web automation via a Chrome extension. No screenshots, no debug port, no bot detection — the model reads structured JSON and acts through the browser's own input pipeline.
WebSense gives an AI agent hands on a real, logged-in Chrome profile. The agent gets a lossless, addressable map of the page, acts on it by reference, and is told — with a grouped diff — whether the page actually changed. No vision model, no headless browser, no remote-debugging port.
Requirements
Node.js ≥ 18
Google Chrome / Chromium (Manifest V3, offscreen WebSocket bridge). Chrome-only: there is no Firefox code in this repo.
OS-level input is Windows-only. Everything else (browse / find / act, trusted input, frames, diffs) is cross-platform.
act{how:"os"},real_click,real_paste,real_activate_tabanddialog{keystroke}use WindowsSendInput/ PowerShell.main_world(the CSP-proof MAIN-world read path) needs Chrome 138+ with the per-extension "Allow User Scripts" toggle enabled.
Related MCP server: live-mcp
Install
git clone https://github.com/spliffspliff70-wq/websense-mcp
cd websense-mcp
npm install1. Load the extension. Open chrome://extensions, turn on Developer mode, click Load unpacked, and select the extension/ folder. It connects to the WebSocket hub on ws://127.0.0.1:38401 automatically — there is no launcher page. (For main_world, open the extension's Details and enable Allow User Scripts.)
2. Register the MCP server with your client.
stdio — one client per server process:
{ "mcpServers": { "websense": { "command": "node", "args": ["src/server.js"] } } }streamable HTTP — one server, many clients:
node src/server.js --http --http-port 9222
# then point each client at http://localhost:9222/mcp3. Call websense_guide (or just browse). The guide is the runtime source of truth and documents every tool.
Bridge port. Default 38401, plain ws:// on 127.0.0.1 (loopback is exempt from mixed-content blocking, so it works from HTTPS pages). Override with --port <n> on the server and the matching PORT constant in extension/offscreen.js. If the port is already taken the hub logs a warning and the server keeps running — MCP still works, the bridge just isn't claimed. Run isolated servers with different --port values.
The model's surface: 8 listed tools
A model sees exactly eight tools. The rest stay callable by name but are not listed, so a model never has to choose between a wall of one-verb tools.
tool | what it does |
| TOOL 1 — open (or bind) a tab and map it in one call: navigate, seed the diff baseline, collect a lossless inventory, and return only the INDEX + the page's vocabulary + a region outline. |
| TOOL 2 — search the stored inventory. Every hit answers WHERE (region + branch chain, resolved from parent pointers) and WHAT (role / name / attrs / state). |
| Do something: |
| Full-fidelity records for one slice of the inventory ( |
| Tab/window ops: |
| WebSense itself + raw reads: |
| Call this first when anything is refused. Walks launch → server → extension → binding → page-ready, stops at the first broken link, and returns the command that fixes it. Read-only by default. |
| Start here — returns the full in-tool guide. |
38 tools are registered; 30 of them are unlisted but callable by name. The listable extras are summarised under Other registered tools.
Quick start — the loop
browse {url}— one call that navigates (or binds), seeds the diff baseline, stores a lossless inventory of every element (nothing filtered, capped or truncated) and returns the small INDEX + a region outline. The full records stay server-side, addressable by index.find {query}— locate the control. Each hit tells you WHERE it is and WHAT it is, so five controls called "New" are distinguishable.act {action, ref, ...}— do it. Reach forhow:"trusted"when the page checksisTrusted, reads coordinates, or a default action must run. Works in a background tab — no focus steal.Read the diff the reply carries.
The diff (did it land?)
Every mutating action returns a second block: a grouped DIFF against your browse baseline.
structure — the page's shape changed (elements added/removed; tag/role/name/attrs changed). Page truth.
content — the same element's value/text changed and its shape did not. The page answered you.
viewport — only
vp/x/ydiffer. That is scroll/layout churn and is not a mutation.
mutated is true only when structure or content moved — viewport churn can never make an action look landed. The reply ends with a FULL DIFF: <handle> line; fetch any part with page_slice{diff:"<handle>", part:"structure|content|visual|viewport"}. The full delta is cached because it can be hundreds of KB for one action — the summary is what changed, the handle is the rest. Pass verify:false to skip the diff on a call you don't need checked.
A navigation is the strongest confirmation and is not in the groups: when a click or press_key (Enter/Space) replaces the document, the result carries effect:"confirmed" plus a navigation {from,to}, and the diff line says so explicitly — a diff across a navigation compares two different documents, so its groups are meaningless.
Verdicts, and when they are wrong
effect is derived from page_state — url, title, readyState, scroll. That is a deliberately weak signal, and an action that only changes the DOM does not move any of them. Two rules make that honest:
The diff can upgrade the verdict. When the page-side differ measures a real structure/content move (
mutated:true), anunverifiableorsuspected_noopverdict is upgraded toconfirmedand carrieseffectSource:"page_diff", and its stale escalation advice is removed with it. Measured on x.com: every trustedtypeinto the thread composer used to answerunverifiablewhile the same reply saidmutated:true— a verdict contradicting its own evidence, which reads as "no proof" and makes a caller re-run an action that already worked. It only ever upgrades: afailedverdict (the action layer refused — disabled, read-only) staysfailed. The measurement happened;classifyEffectsimply could not see it.An ambiguous selector is refused, not guessed.
querySelectorreturns the first match with no word, so a selector matching two elements silently drove the wrong one. Measured on x.com's/compose/post:[data-testid="tweetTextarea_0"]matches twice — the dialog's real composer and the empty page-level inline composer. You now get a refusal naming the count and the tag (ambiguous selector … matches 2 elements — refusing to pick one silently). Scope it and retry:[role="dialog"] [data-testid="tweetTextarea_0"], or any selector that is unique on that page. This is also whyfindreturns all matches with no cap — a cap would hide the second element and make the ambiguity invisible.
effect remains weak evidence for everything else: confirmed means "the page measurably moved", never "the app accepted and persisted it". Re-read the field or the page when the outcome matters.
Trusted input — act{how:"trusted"}
how:"trusted" drives chrome.debugger + Input.dispatchMouseEvent / Input.dispatchKeyEvent, so the page receives isTrusted events and the browser itself runs the default action (a link navigates, an Enter submits, a checkbox toggles, an arrow key moves a slider) instead of the tool guessing. It works in a background tab — no focus steal, no window activation. Chrome shows its "debugging this browser" infobar while attached.
Off-screen targets are scrolled into view first. Measured: a trusted click aimed at y=1727 below the fold hit nothing; after the fix it re-aimed at y=938 and the page recorded the click.
how:"os" is the OS-level rung (Windows SendInput) — reach for it only when a page rejects programmatic input outright, or for a raw-input / canvas surface. It lands on the frontmost window and so steals focus.
Dialog handling
JS dialogs, by type — this is measured behavior, not a blanket rule. The MAIN-world hook shadows
alertonly (its return is undefined, so nothing can branch on it) — it is captured intopendingDialogsand auto-dismisses after 30s, matching what Chrome itself does in a background tab.confirm/promptstay NATIVE: a hooked confirm returns a Promise (always truthy), so everyif(confirm(...))would take the TRUE branch regardless of the answer — the auto-answer-after-30s behavior was a bug, not a safety net, and it is gone. A nativeconfirm/promptblocks the page; answer it on request withdialog{native:true, action:"accept"|"dismiss", value:promptText}— that goes throughPage.handleJavaScriptDialog(chrome.debugger), so the page's own branch follows the agent's real choice, on a background tab, with no focus steal. Nothing ever auto-answers a decision the agent did not make.DOM modals (
[role=dialog], most in-app modals) are closed by ref;status.hasModal/dialogCountcome from a visibility-blind scan (a hidden modal still counts).OS-level dialogs (HTTP basic-auth, proxy-auth, print) cannot be intercepted by JS.
dialog{keystroke:true, key:"enter"|"escape"}injects a global keystroke through Windows control (PowerShellSendKeys) — Windows-only.File picker: handled by
form{action:"upload"}(DataTransfer API) — no OS dialog.
Iframes / frames
Same-origin iframes are walked and clickable — verified: the frame echoes the click. List them with tabs{action:"frames"} (pass tabId; omit it for your bound tab) and pass frameId to any element tool — this unlocks Gmail compose, Notion, Figma and any site that renders key UI inside child frames. Cross-origin frames are skipped: nothing in the page can read them and no click can be aimed inside them.
Other registered tools
These are registered and callable by name — the standalone forms that act / debug absorbed are noted inline:
explore_page— quick look at a page's actions (SAG).compact:true,intent:"submit",goal:"log in",preload:true,incremental:true. For a full page map usebrowse+findinstead.read— page text:text · content · markdown · diff · scrollextract · preload.click— click a ref (default)· mode:"hover" · "rightclick" · "drag" (fromRef/toRef) · x,yfor canvas.trusted_click·trusted_key— the standalone trusted mouse / keyboard paths.press_key— synthetic key events only; runs no default action. Usetrusted_key/act{how:"trusted"}when the default matters.type_text— fill one input (React-safe native setter) orfields:[{ref,text},…]for a verified batch.form—state · select · toggle · special · upload.reveal— pre-extract hidden content:dropdown · tabs · accordion.scroll—direction+amount(ticks) ·yabsolute ·intoView.status—page · bridge · doctor · downloads.wait— poll conditions (ANDed) until met, or wait for an event.evaluate— run JS and return its value (auto-reroutes through the MAIN world on a CSP block) or a no-evalqueryDOM read.main_world— run a compiled function in the page's MAIN world (CSP-proof; needs "Allow User Scripts").ax— native accessibility tree viachrome.debugger(for canvas SPAs andchrome://pages).screenshot—captureVisibleTab→ PNG/JPEG dataUrl, for a vision model.dialog—accept | dismiss(+value); captures the page's own JSalert/confirm/prompt.session—reset · map · mermaid.network_log·console_log— captured page fetch/XHR · console + JS errors.cookies—list · get · clear(values are masked on other surfaces).clipboard—copy · read.inspect—element · geometry · relation.navigate— navigate a tab (reuses your bound tab; no tab spam).page_snapshot— collect / return the lossless inventory index directly.respawn_offscreen·extension_reload— extension maintenance (MV3 traps).real_activate_tab·real_click·real_paste— Windows OS-level input; the last rung.
Architecture
MCP Client (Claude / Cline / Cursor / Hermes)
↔ stdio or streamable HTTP
WebSense MCP Server (src/server.js)
↔ WebSocket ws://127.0.0.1:38401
Chrome Extension (extension/)
├── background.js service worker, tab management, binding
├── offscreen.js WebSocket client, auto-reconnect
└── websense-cs.js SAG extraction + native DOM interaction (CSP-safe)
↔ chrome.runtime.sendMessage / chrome.debugger (trusted input)
Live DOMWhat's new in 2.0
A 7-tool listed surface —
browse · find · act · page_slice · tabs · debug · websense_guide.actanddebugare facades that dispatch to the real handlers, so the 30 unlisted tools remain callable by name with identical behaviour.browse→find→act→ read the diff replacesexplore_page → click/typeas the primary loop.browsereturns an INDEX over a lossless inventory;findreturns WHERE + WHAT per hit;page_sliceloads one branch at full fidelity.A grouped auto-diff after every mutating action, with
mutatedderived from structure + content only (viewport churn cannot fake a landing), a summary in the reply, and the full delta cached behindFULL DIFF: <handle>.Trusted input:
act{how:"trusted"}runs throughchrome.debugger+Input.dispatchMouseEvent/Input.dispatchKeyEvent, so the page seesisTrustedevents and the browser performs the default action — in a background tab. Off-screen targets are scrolled into view first.Same-origin iframes are walked and clickable.
Known limitations (honest)
Trusted drag works end-to-end (verified on the fixture):
act{action:"drag", how:"trusted"}produces trusteddragstart/dragenter/dragoverand a realdrop, from a background tab, with both ends scrolled into view automatically. The drop only fires if the target accepts drops (preventDefault()ondragover) — a non-zone target correctly getsdragleave, exactly as a real mouse would, so check the target before blaming the driver. The plaindragmode still fires the whole sequence but its events are not trusted — usehow:"trusted"when the page checksisTrusted.OS-click equivalence is unproven.
act{how:"trusted"}is the browser's own input pipeline; it is not proven byte-identical to a real OS click.real_click(WindowsSendInput) is the only genuinely OS-level path.Canvas / WebGL: coordinate clicks are TRUSTED (fixed 2026-10-02 — this line said "not trusted" and was stale).
act{action:"click", how:"trusted", x, y}drivesInput.dispatchMouseEventat the raw viewport point, so a canvas that inspectsisTrustednow sees a real one. The untrusted form still exists ashow:"auto"and lands within 1px; preferhow:"trusted"on canvas/WebGL, which is exactly the surface class that checks the flag.Chrome-only. MV3 + offscreen WebSocket bridge; no Firefox code.
OS-level input is Windows-only (
real_*,dialog{keystroke}).OS-level input additionally needs Python 3 with
pyautogui+pywinauto(Windows only). Everything else — the wholebrowse/find/actloop, trusted input, frames, diffs — needs neither Python nor Windows. The server discovers an interpreter (py→python→python3); override withWEBSENSE_PYTHON=/path/to/python. If none is found the error says so instead of failing as a mysterious page problem (fixed 2026-10-02 — the interpreter path was previously hard-coded to the maintainer's machine, so OS input could not work for anyone else).main_worldrequires the "Allow User Scripts" toggle (Chrome 138+). Without it the CSP-proof MAIN-world path is unavailable.evaluate{script}runs your JS and returns its value; the isolated-worldnew Functionpath is blocked by the extension's own MV3 CSP, so it transparently re-routes through the MAIN world (chrome.userScripts, no eval) and reportsvia:"main_world".evaluate{query:{…}}remains the lighter path for plain DOM reads.A JS dialog raised while the tab is hidden can be auto-dismissed by Chrome before anyone sees it — act fast:
dialog{native:true, action, value}answers a still-pending native confirm/prompt viaPage.handleJavaScriptDialogon the background tab (no activation needed). A hookedalertis still recorded (status.recentDialogs); nativeconfirm/promptare NOT in the hook's queue (they are never shadowed), so checkstatusand answer immediately if the branch matters.axattacheschrome.debuggerand shows Chrome's warning banner while attached; it requires an explicittabId.One profile, per-tab isolation. Concurrent jobs share one Chrome profile — there is no cookie/storage isolation between them. Scope work with
tabs{action:"bind", tabId}+ an explicittabId. Session state (map/history) is per-session:session{action:"reset"}clears only your own history.Logged-in sites (LinkedIn, etc.) must already be authenticated in that Chrome profile;
navigateopens a fresh tab that needs an existing session cookie.Refs are stable —
E#refs are assigned in viewport order on the first scan, then held by element identity (a per-element cache + adata-websense-refattribute), so they survive re-explores, scrolls and framework re-renders. A ref dies only when its element leaves the DOM with nothing to heal from. CSS-selector refs (#id) remain the safest choice for anything long-lived or across navigations.
Testing
# Regression suite — no Chrome needed (hub, diff, snapshot, trusted-input,
# guide-truth and doc-drift guards). Currently 170 tests.
npm test
# Print the LISTED surface from a running server (tools/list is filtered to it)
node tools/tools-list.mjs # names-only preflight
node tools/tools-list.mjs --full # name + one-line description
# End-to-end MCP client test (needs Chrome + the extension loaded)
node test/mcp-client-test.js
# Keep MODEL_PROMPT.md in sync with the in-tool guide (also enforced inside npm test)
node tools/export-guide.mjs --checkFile structure
websense-mcp/
├── src/
│ ├── server.js # MCP server: tool registration, the 7-tool listed surface, facades
│ ├── hub.js # WebSocket hub on ws://127.0.0.1:38401
│ ├── session.js # exploration map + per-session state
│ ├── snapshot.js # lossless page inventory (collector + slicer)
│ ├── diff-cache.js # cached grouped diffs behind FULL DIFF handles
│ ├── diff-collector.js # in-page diff collection
│ └── climb.js, incr.js, upload.js, summarize.js, mermaid.js
├── extension/
│ ├── manifest.json # Chrome MV3
│ ├── background.js # service worker (tab mgmt, offscreen lifecycle)
│ ├── offscreen.js # WebSocket client (auto-reconnect)
│ ├── websense-cs.js # built content script (generated — do not hand-edit)
│ └── cs-src/ # content-script sources (edit here; build with tools/build-cs.mjs)
├── test/
│ └── mcp-client-test.js # end-to-end MCP client test
└── tools/ # build + measurement scripts (build-cs, export-guide, tools-list…)License
MIT — see LICENSE.
Available Tools
7 toolsactC
DO something: click, hover, rightclick, drag, type, key, form, upload, scroll, dialog.…
| Name | Required | Description | Default |
|---|---|---|---|
| x | No | ||
| y | No | ||
| how | No | auto = normal path, trusted = browser in… | |
| key | No | ||
| ref | No | ||
| text | No | ||
| tabId | No | ||
| toRef | No | ||
| value | No | ||
| action | Yes | ||
| amount | No | ||
| frameId | No | frameId (omit=top) | |
| fromRef | No | ||
| filePath | No | ||
| selector | No | ||
| direction | No | ||
| modifiers | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden, and it discloses nothing: no mention of destructive effects, focus requirements, wait/retry behavior, auth needs, or what happens on failure. For a 17-parameter interaction tool spanning drag, upload, and form submission, this is a serious omission.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The single line is short, but nearly all of its content duplicates the action enum already present in the schema, and the trailing ellipsis signals an unfinished thought. It is under-specified rather than efficient, so size is not the problem—value per word is.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With 17 parameters, 12% schema coverage, no annotations, and no output schema, the description is the only place an agent could learn behavioral expectations—and it provides none. Nothing about per-action parameter requirements, side effects, or result handling is covered, leaving the definition wholly inadequate for the tool's complexity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 12% across 17 parameters, so the description must compensate, and it does not. It restates the action enum but says nothing about how x/y pair with coordinates, how ref/selector/fromRef/toRef relate, or which params are required per action (e.g., type needs text/value, upload needs filePath).
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names the tool's domain by enumerating the actions it supports (click, hover, drag, type, form, upload, scroll), which hints at browser/page interaction. However, 'DO something' is an empty verb and it never states the resource being acted on, nor does it distinguish act from siblings like browse or find. Purpose is inferable from the enum list, not from a clear statement.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use act versus browse, find, page_slice, or tabs, and no exclusions or prerequisites. An agent must infer that this is the interaction tool purely from the action names. No when-to-use context is supplied at all.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browseB
TOOL 1. Go to a page and map it in ONE call: navigate (or bind an existing tab), seed the auto-diff BASELINE f…
| Name | Required | Description | Default |
|---|---|---|---|
| url | No | ||
| fresh | No | Re-collect even if a snapshot for this t… | |
| tabId | No | ||
| newTab | No | Force a fresh tab rather than reusing th… | |
| frameId | No | frameId (omit=top) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It does disclose a meaningful behavioral trait that annotations would not cover: that it seeds the auto-diff BASELINE, which tells the agent state is being persisted for later diffing. It does not explain what 'map' produces, whether prior baselines are overwritten, or any side effects, and the text is cut off mid-sentence.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The leading 'TOOL 1.' is wasted tokens and the whole entry appears truncated mid-word ('BASELINE f…'), which harms structure. The core action is front-loaded reasonably well, so density is decent, but the artifact prefix and cutoff keep this from scoring higher.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With 5 parameters, no annotations, and no output schema, the description should explain what 'map' returns and how the baseline behaves. Instead it is truncated and omits all of that, leaving an agent unable to predict results or side effects for a tool with no other structured guidance.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 60%, so baseline is 3. The description adds marginal meaning by framing url as a navigation target and tabId as an existing tab to bind, which the schema leaves undocumented. It says nothing about fresh, newTab, or frameId beyond the truncated schema hints.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'Go to a page and map it in ONE call: navigate (or bind an existing tab)'. This distinguishes it from siblings like tabs and page_slice by emphasizing the combined navigate-and-map behavior in a single call. However, the 'TOOL 1.' prefix is noise and the description is truncated, leaving the mapping output undefined.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Implies usage context by noting it can navigate or bind an existing tab, and that it seeds an auto-diff baseline, so an agent can infer this is the entry point for page analysis. But it never states when to prefer this over siblings such as page_slice, find, or tabs, nor any preconditions, so guidance stays implicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
debugD
WebSense itself, and raw reads: op=status, session, logs, cookies, clipboard, screenshot, ax, evaluate, main_w…
| Name | Required | Description | Default |
|---|---|---|---|
| op | Yes | what to inspect or maintain | |
| url | No | ||
| func | No | ||
| kind | No | for logs: network | console | |
| query | No | ||
| tabId | No | ||
| action | No | ||
| frameId | No | frameId (omit=top) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description must carry the full behavioral burden, but it only says 'raw reads' and lists ops without explaining side effects, permissions, rate limits, or safety for mutating ops like reload, respawn, or evaluate. It also truncates before finishing the list.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single truncated fragment ending in an ellipsis, lacking a front-loaded purpose. It is under-specified rather than concise, wasting the opportunity to state what the tool does and how to use it.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given 13 ops, 8 parameters, no annotations, no output schema, and incomplete parameter documentation, the description is far too incomplete to guide correct invocation. It leaves critical gaps in purpose, usage, and parameter mapping.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 38%, so the description should compensate for undocumented parameters. It merely enumerates some op values and leaves url, func, query, tabId, action, and frameId without added semantic meaning, adding little beyond the enum itself.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description is a truncated fragment that lists op values rather than stating a clear verb+resource. 'raw reads' gives a hint, but the overall purpose is vague and does not distinguish this tool from siblings like act, browse, or websense_guide.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use guidance, no alternatives, and no prerequisites are provided. The description offers no routing information among the sibling tools, leaving the agent to guess when this tool is appropriate.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
findC
TOOL 2. Search the stored page inventory and get WHERE and WHAT the hit is: its own role/name/attributes/state…
| Name | Required | Description | Default |
|---|---|---|---|
| vp | No | true = in viewport only | |
| tag | No | Element tag | |
| attr | No | Match any attribute the page wrote | |
| role | No | ||
| field | No | true = only form controls (platform-repo… | |
| limit | No | ||
| query | No | ||
| tabId | No | ||
| region | No | Region substring (region is derived from… | |
| frameId | No | frameId (omit=top) | |
| indices | No | Fetch exact records by inventory index —… | |
| focusable | No | true = only focusable elements (el.… | |
| branchDepth | No | How many ancestors to include in the bra… | |
| interactive | No | true = only controls (derived: focusable… |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations exist, so the description carries the full burden. It conveys one useful behavioral fact — that searching operates over a stored/cached page inventory rather than the live page — but says nothing about read-only nature, staleness, pagination, or the shape of the result, and the sentence is cut off.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single truncated fragment ending in an ellipsis, prefixed with an unexplained 'TOOL 2' label. It is neither complete nor front-loaded with actionable content.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 14-parameter, zero-required search tool with no annotations and no output schema, the description is far too thin: it omits result shape, empty/no-match behavior, and the meaning of the untagged filter parameters.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 14 parameters at 71% schema coverage, several parameters (role, limit, query, tabId) have no schema description at all, and the description adds zero parameter meaning. Since coverage is below the 80% threshold, the description should compensate and does not.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a verb (search) and resource (stored page inventory) and hints at the return payload ('WHERE and WHAT the hit is: its own role/name/attributes/state'). However it is truncated mid-sentence and gives no differentiation from siblings like page_slice or act, so an agent cannot confidently tell it apart.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'TOOL 2' implies a pipeline position, but the description never states when to prefer this over page_slice, browse, or act, nor any preconditions (e.g. that an inventory must already be captured). No when-to-use or when-not guidance is present.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
page_sliceC
Full-fidelity records from the stored page snapshot, filtered to ONE slice: indices / tag / role / region / vp…
| Name | Required | Description | Default |
|---|---|---|---|
| vp | No | true = in viewport only, false = off-vie… | |
| tag | No | Filter by tag, e.g. input | |
| attr | No | Filter by ANY attribute the page wrote, … | |
| diff | No | A FULL DIFF handle from a DIFF block (e.… | |
| part | No | Which part of a cached diff (default all… | |
| role | No | ||
| field | No | true = form controls only (platform-repo… | |
| limit | No | ||
| query | No | ||
| tabId | No | ||
| region | No | Filter by region substring — region is d… | |
| frameId | No | frameId (omit=top) | |
| indices | No | Fetch exact records by inventory index —… | |
| focusable | No | true = focusable only (el.… | |
| interactive | No | true = actionable only, DERIVED (focusab… |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are supplied, so the description carries the full burden. It does hint that results come from a stored snapshot rather than a live page, which is useful context, but it says nothing about whether the operation is read-only, whether snapshots must pre-exist, what happens on empty slices, pagination, or response shape.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
It is a single front-loaded sentence with no padding, which is good, but the sentence is cut off mid-enumeration ('vp…'), so its brevity stems partly from incompleteness rather than discipline. Structure is acceptable, content is not fully delivered.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 15-parameter, no-required, no-annotation, no-output-schema tool, the description is far too thin. Nothing explains how the many filters combine, whether 'ONE slice' is enforced, or what the records look like, leaving major gaps an agent must guess at.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 73%, below the high-coverage baseline, so the description should help disambiguate the 15 filters. Instead it only echoes a few slice names (indices/tag/role/region/vp) with no syntax, interaction, or precedence guidance, and truncates before finishing. It adds essentially no meaning beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description identifies a retrieval verb and resource (full-fidelity records from a stored page snapshot), which is more than a tautology. However, it is truncated mid-sentence ('vp…') and the claim of 'filtered to ONE slice' sits uneasily against 15 independent filter parameters, leaving the actual selection model vague. It does not clearly distinguish itself from siblings like find or browse.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit when-to-use, when-not-to-use, or alternative named. An agent cannot infer from the text whether this is preferred over find or browse for locating elements. Usage is only weakly implied by the phrase 'from the stored page snapshot.'
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
tabsC
Tab/window ops. bind routes page ops to a tab WITHOUT focus — page ops NEVER need activation.…
| Name | Required | Description | Default |
|---|---|---|---|
| tabId | No | ||
| toTab | No | transfer: destination tab | |
| action | Yes | ||
| fromTab | No | transfer: source tab | |
| activate | No | bind: ALSO make this the OS-active tab.… | |
| selector | No | ||
| useValue | No | transfer: copy input VALUE instead of vi… | |
| windowId | No | ||
| toSelector | No | transfer: destination selector | |
| fromSelector | No | transfer: source selector |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden of behavioral disclosure. It reveals one behavior (bind operates without focus/activation) but remains silent on the side effects, prerequisites, or outcomes of other actions (close, switch, move, transfer, etc.). It does not contradict any annotations since none exist.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is very brief and front-loads the broad category 'Tab/window ops' before a specific detail about 'bind'. It is concise, but the truncation with an ellipsis suggests the full description may include more content. As presented, it is efficient but structurally minimal.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a multi-action tool with 10 parameters, 1 required, no output schema, and no annotations, this description is severely incomplete. It does not explain the different actions, parameter usage, return values, or error conditions. An agent would need to inspect the schema thoroughly and still lack context for the undocumented parameters. This is far from a usable definition.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 60%, leaving four parameters (tabId, action, selector, windowId) without any schema-level descriptions. The tool description does not mention any parameters or add meaning beyond the schema. For the uncovered parameters, nothing is provided, so the description fails to compensate for the gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states the tool handles 'Tab/window ops' and highlights the 'bind' action's specific behavior (routing page ops without focus). This gives a general sense of purpose but does not enumerate the full range of actions (list, switch, close, frames, windows, focus, move, transfer, switchread). It differentiates itself from some siblings by mentioning 'bind' but not clearly against all overlapping tools like real_activate_tab.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies that page ops never need activation, which hints that for such operations you may not need to use tab activation tools, but it does not explicitly state when to use this tool vs alternatives like real_activate_tab or session. No exclusions or explicit conditions for each action are provided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
websense_guideB
START HERE. The listed tools: browse, find, act, page_slice, tabs, debug, guide. Call once before you act.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It implies a one-time, presumably side-effect-free read via 'call once', but never states that it is read-only, what it returns, or whether it has any cost or auth requirement.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Very brief and front-loaded with 'START HERE', which is exactly right. The enumeration of sibling tool names is somewhat redundant given they are already exposed, but it costs little.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a parameterless, output-schema-less orientation tool, the description is thin: it tells the agent to call it but not what the returned guidance contains, so the agent cannot anticipate its value beyond 'read this first'.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so there is no parameter semantics to document; the baseline of 4 applies and no schema gap exists to compensate for.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description positions the tool as an entry point ('START HERE', 'guide'), which does distinguish it from the action-oriented siblings. However, it never states what the tool actually returns or what kind of guidance it provides, and the enumerated tool list largely restates the sibling names already visible to the agent.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Call once before you act' gives an explicit, actionable trigger for invocation and implies it should not be repeated. It stops short of any when-not guidance or explanation of what happens if you skip it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
33 tool updates
v2.0.0- Added
act - Removed
ax - Added
browse - Removed
click - Removed
clipboard - Removed
console_log - Removed
cookies - Added
debug - Removed
dialog - Removed
evaluate - Removed
explore_page - Removed
extension_reload - Added
find - Removed
form - Removed
inspect - Removed
main_world - Removed
navigate - Removed
network_log - Changed
page_slice8 fields changed- added
Input schema / properties / attrAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "properties": { + "name": { + "type": "string" + }, + "value": { + "type": "string" + } + }, + "required": [ + "name" + ], + "type": "object" + } + ], + "description": "Filter by ANY attribute the page wrote, …" +} - added
Input schema / properties / diffAdded value: +{ + "description": "A FULL DIFF handle from a DIFF block (e.…", + "type": "string" +} - added
Input schema / properties / fieldAdded value: +{ + "description": "true = form controls only (platform-repo…", + "type": "boolean" +} - added
Input schema / properties / focusableAdded value: +{ + "description": "true = focusable only (el.…", + "type": "boolean" +} - added
Input schema / properties / indicesAdded value: +{ + "description": "Fetch exact records by inventory index —…", + "items": { + "type": "number" + }, + "type": "array" +} - changed
Input schema / properties / interactive / descriptionPrevious value: -"true = actionable elements only"New value: +"true = actionable only, DERIVED (focusab…" - added
Input schema / properties / partAdded value: +{ + "description": "Which part of a cached diff (default all…", + "enum": [ + "structure", + "content", + "visual", + "viewport", + "all" + ], + "type": "string" +} - changed
Input schema / properties / region / descriptionPrevious value: -"Filter by region substring, e.g.…"New value: +"Filter by region substring — region is d…"
- Removed
page_snapshot - Removed
press_key - Removed
read - Removed
real_activate_tab - Removed
real_click - Removed
real_paste - Removed
respawn_offscreen - Removed
reveal - Removed
screenshot - Removed
scroll - Removed
session - Removed
status - Removed
type_text - Removed
wait
1 tool update
v1.4.7- Changed
evaluate2 fields changed- changed
Input schema / properties / script / descriptionPrevious value: -"JS to execute (eval mode)"New value: +"JS to execute. A bare expression or stat…" - added
Input schema / properties / tabIdAdded value: +{ + "type": "number" +}
14 tool updates
v1.4.5- Changed
click1 field changed- added
Input schema / properties / verifyAdded value: +{ + "description": "Set false to SKIP the automatic post-act…", + "type": "boolean" +}
- Changed
dialog1 field changed- added
Input schema / properties / verifyAdded value: +{ + "description": "Set false to SKIP the automatic post-act…", + "type": "boolean" +}
- Changed
evaluate1 field changed- added
Input schema / properties / verifyAdded value: +{ + "description": "Set false to SKIP the automatic post-act…", + "type": "boolean" +}
- Changed
explore_page4 fields changed- added
Input schema / properties / contentMaxLenAdded value: +{ + "description": "Cap on extracted body text (default 8000…", + "type": "number" +} - added
Input schema / properties / freshAdded value: +{ + "description": "Force a real re-scan. Without it, a repe…", + "type": "boolean" +} - changed
Input schema / properties / maxActions / descriptionPrevious value: -"Cap for compact mode (default 250)"New value: +"Cap on RETURNED actions (default 200).…" - added
Input schema / properties / settleAdded value: +{ + "description": "false = skip the SPA hydration settle wa…", + "type": "boolean" +}
- Changed
form1 field changed- added
Input schema / properties / verifyAdded value: +{ + "description": "Set false to SKIP the automatic post-act…", + "type": "boolean" +}
- Added
main_world - Changed
navigate1 field changed- added
Input schema / properties / tabIdAdded value: +{ + "type": "number" +}
- Added
page_slice - Added
page_snapshot - Changed
press_key1 field changed- added
Input schema / properties / verifyAdded value: +{ + "description": "Set false to SKIP the automatic post-act…", + "type": "boolean" +}
- Changed
real_click1 field changed- added
Input schema / properties / verifyAdded value: +{ + "description": "Set false to SKIP the automatic post-act…", + "type": "boolean" +}
- Changed
real_paste1 field changed- added
Input schema / properties / verifyAdded value: +{ + "description": "Set false to SKIP the automatic post-act…", + "type": "boolean" +}
- Changed
tabs1 field changed- changed
Input schema / properties / activate / descriptionPrevious value: -"bind: also activate the tab"New value: +"bind: ALSO make this the OS-active tab.…"
- Changed
type_text1 field changed- added
Input schema / properties / verifyAdded value: +{ + "description": "Set false to SKIP the automatic post-act…", + "type": "boolean" +}
5 tool updates
v1.1.1- Added
extension_reload - Changed
form2 fields changed- changed
Input schema / properties / action / enumPrevious value: -[ - "state", - "select", - "toggle", - "upload" -]New value: +[ + "state", + "select", + "toggle", + "special", + "upload" +] - added
Input schema / properties / clearAllAdded value: +{ + "description": "select on multi-select: deselect non-mat…", + "type": "boolean" +}
- Added
real_activate_tab - Added
real_click - Added
real_paste
24 tool updates
v1.0.0- First observed
ax - First observed
click - First observed
clipboard - First observed
console_log - First observed
cookies - First observed
dialog - First observed
evaluate - First observed
explore_page - First observed
form - First observed
inspect - First observed
navigate - First observed
network_log - First observed
press_key - First observed
read - First observed
respawn_offscreen - First observed
reveal - First observed
screenshot - First observed
scroll - First observed
session - First observed
status - First observed
tabs - First observed
type_text - First observed
wait - First observed
websense_guide
TDQS
Scored across 7 tools
Each tool has a distinct primary purpose (browse for navigation/mapping, find for search, act for interactions, page_slice for snapshot extraction, debug for raw reads, tabs for management, websense_guide for orientation). However, debug and page_slice both read page data, and find and page_slice both retrieve page inventory entries, so minor confusion is possible.
Tool names mix verbs (browse, find, act) with nouns (tabs, page_slice) and a prefixed noun phrase (websense_guide). There is no consistent verb_noun or other predictable pattern, making the naming style inconsistent.
Seven tools is well within the 3–15 range for a browser automation server. Each tool covers a broad area (e.g., act handles many actions, debug handles many raw reads), so the count feels well-scoped without redundancy.
The set covers navigation, search, interaction, tab management, raw reads, and page data extraction, which are core browser automation needs. Minor gaps include no explicit wait/synchronization tool and no download handling, though these can likely be worked around via debug.evaluate or act.
Maintenance
Related MCP Connectors
AI-powered browser automation — navigate, click, fill forms, and extract data from any website.
AI-powered web automation. Navigate websites using AI agents for one page or a thousand
AI-powered web automation. Navigate websites using AI agents for one page or a thousand
Stealth web automation for AI agents. Login, signup, navigate, screenshot.
Related MCP Servers
- AlicenseBqualityFmaintenanceEnables AI agents to directly control your real Chrome browser with full context including login sessions, cookies, and open tabs. It provides tools for page scanning, JavaScript execution, CDP control, screenshots, and physical mouse/keyboard input for authentic browser automation.20245MIT
- FlicenseNot gradedqualityBmaintenanceEnables AI agents to declaratively control web pages using real mouse and keyboard events via Chrome DevTools Protocol, without executing page JavaScript.10 npm1-
- FlicenseNot gradedqualityBmaintenanceEnables AI agents to autonomously control browsers with zero-hardcoding semantic DOM interaction, real-time network telemetry, Cloudflare/bot self-healing, and human-in-the-loop reasoning for web QA and automation.-
- AlicenseNot gradedqualityBmaintenanceEnables LLM agents to automate a live Chrome browser with compact DOM serialization, authenticated session continuity, and human-like input trajectories.4 npmMIT