Skip to main content
Glama

OctoLimb

MCP for any browser.

Give any MCP agent — Claude Code, Claude Desktop, OpenAI Codex, or your own — a real Chrome browser to drive.

License Node MCP Chrome

What it is · The Octopus family · Install · Use it with · Tools · Security


What it is

OctoLimb turns a normal Chrome browser into a tool any MCP agent can call. Click, type, scroll, read the page, run JavaScript, and replay saved step-by-step "recipes" — all through your everyday, logged-in browser, not a headless copy.

flowchart LR
    A["Any MCP agent<br/>Claude · Codex · Studio"] -->|stdio or HTTP| B["OctoLimb host<br/>the MCP server"]
    B -->|WebSocket<br/>ws://127.0.0.1:32528| C["Chrome extension<br/>the browser actuator"]
    C -->|CDP + DOM indexer| D["Your real Chrome"]

It ships as two pieces, one package:

Piece

What it does

Where

octolimb host

The MCP server. Speaks MCP to your agent and relays tool calls to the browser.

src/ → builds to dist/

Chrome extension

The browser actuator. Performs tool calls in Chrome (CDP input events + nanobrowser's DOM indexer).

extension/

Why two processes? A Chrome extension can't listen on a socket or serve stdio — that's a browser security boundary. So the MCP server is a separate host process, and the extension connects out to it.


Related MCP server: OpenCode in Chrome

The Octopus family

OctoLimb is one arm of a bigger octopus, all on the same account:

Repo

What it is

How OctoLimb fits

🧠 Octopus Studio

Local-first AI workspace — research, docs, data, media, automation, MCP, and development

The brain. Studio ships OctoLimb's bridge in-process, so its agent gets a real browser with one toggle. OctoLimb is also a standard MCP server, so it plugs straight into Studio's Connections catalog.

✂️ Octo Cut

Self-driving, non-destructive desktop video editor

The cut arm. OctoLimb browses and gathers clips, images, and audio from the web; Octo Cut assembles them on the timeline.

🐙 OctoLimb

MCP for any browser

The browser arm — drives any website for any agent.

The workflow: Octopus Studio is the hub. Its agent calls OctoLimb to go out and do things in the browser — fetch a stock clip, pull a reference, fill a form, post an update — and feeds the results back into the project, or into Octo Cut for editing. OctoLimb is how the octopus reaches beyond your machine into the open web.


Install

Load the extension

  1. Open chrome://extensions.

  2. Enable Developer mode (top-right toggle).

  3. Click Load unpacked and select the extension/ folder.

The extension dials out to ws://127.0.0.1:32528 and reconnects automatically whenever the host comes up.

Install the host

npm install -g octolimb      # from npm, or:
# git clone https://github.com/kaleemibnanwar/octolimb.git && cd octolimb && npm install && npm run build && npm link

Confirm it's on your path:

octolimb --help

Use it with

The host runs as a stdio server (default — the most universal) or a Streamable HTTP server. Both always start the WebSocket relay on :32528 for the extension.

Octopus Studio

Octopus Studio already ships an in-process OctoLimb bridge — no standalone host needed. Load the extension, then enable Settings → "Enable OctoLimb MCP Bridge". Studio's agent can then drive a browser for tasks like "post this on Reddit" or "pull the latest pricing from that page."

Don't run the standalone host at the same time — Studio's bridge and this host both bind :32527/:32528. Pick one. (To run both, start the host with --http-port 32529 --ws-port 32530.)

Claude Code

claude mcp add octolimb -- octolimb

Or via .mcp.json:

{
  "mcpServers": {
    "octolimb": { "command": "octolimb", "args": [] }
  }
}

Claude Desktop

Settings → Developer → Edit Config, then add to claude_desktop_config.json:

{
  "mcpServers": {
    "octolimb": { "command": "octolimb", "args": [] }
  }
}

Restart Claude Desktop. (HTTP is also possible — see below.)

Codex

Add to ~/.codex/config.toml:

[mcp_servers.octolimb]
command = "octolimb"
args = []

Any other MCP client

  • stdio (most common): configure the client to spawn octolimb as a command.

  • Streamable HTTP: run octolimb --http, then point the client at http://127.0.0.1:32527/mcp.

octolimb --http                                   # HTTP MCP on :32527 + WS relay on :32528
octolimb --http-port 32529 --ws-port 32530        # override ports

Tools

The host exposes 24 tools:

Category

Tools

🧭 Navigation

go_to_url, go_back, open_tab, close_tab, switch_tab, search_google

📖 Reading

read_page (URL/title + indexed elements), execute_js, get_dropdown_options

🖱️ Interaction

click_element, input_text, select_dropdown_option, send_keys, wait

📜 Scrolling

scroll_to_percent, scroll_to_top, scroll_to_bottom, previous_page, next_page, scroll_to_text

🔁 Recipes

find_recipe, run_recipe, done (auto-saves a task on success), cache_content

See extension/README.md for how the extension implements these (nanobrowser's DOM indexer + CDP input dispatch).

Saved recipes

done, find_recipe, and run_recipe give the agent a lightweight memory: a successfully completed task is saved and can be replayed later on the same domain in one run_recipe call instead of step-by-step. Recipes live as JSON at ~/.octolimb/recipes.json (override the directory with OCTOLIMB_HOME).


Security

  • The host binds to 127.0.0.1 only — never a network interface.

  • The WebSocket relay accepts only chrome-extension:// origins, so other local processes can't impersonate the browser.

  • Driving Chrome shows its unavoidable "OctoLimb Bridge is debugging this browser" banner on attached tabs.

  • execute_js runs arbitrary JavaScript in the page — grant the extension only to tasks you actually want automated.


License

Apache-2.0 — see LICENSE. The DOM-indexing script (extension/buildDomTree.js) is copied from nanobrowser and carries its own Apache-2.0 license files in extension/.

Available Tools

24 tools
cache_contentC

[Real Chrome browser via the OctoLimb extension — drive the user's browser; use only when the task requires acting on a real website.] Cache what you have found so far from the current page for future use

ParametersJSON Schema
NameRequiredDescriptionDefault
intentNoPurpose of this action
contentNocontent to cache (default '')

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are supplied, so the description carries the full behavioral burden. It hints at persistence via 'for future use' but never says where the cached content goes, whether it survives the session, how it is later retrieved, or whether it overwrites prior cache entries.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The functional sentence is short and front-loaded after the bracketed environment note. The bracket is shared boilerplate rather than this tool's content, but it does not obscure the action.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no annotations and no output schema, the description should explain the caching lifecycle but stops at 'for future use'. For a low-risk utility that is minimally viable, though an agent still cannot tell how cached content is consumed downstream.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both parameters (intent, content) are already documented; the baseline is 3. The phrase 'what you have found so far from the current page' adds modest context for what 'content' should hold, but no format or size expectations.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and object — caching content found on the current page for later use — which is clear enough to act on. It does not, however, distinguish this tool from nearby siblings like read_page or run_recipe, so the agent must infer the boundary itself.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The only guidance given is the shared bracketed boilerplate ('use only when the task requires acting on a real website'), which applies to the whole browser toolset, not to caching. There is no statement of when to cache versus when to read, act, or finish, and no prerequisites or exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

click_elementB

[Real Chrome browser via the OctoLimb extension — drive the user's browser; use only when the task requires acting on a real website.] Click element by index

ParametersJSON Schema
NameRequiredDescriptionDefault
indexYesindex of the element
xpathNoxpath of the element (optional fallback if the index is stale)
intentNoPurpose of this action

TDQS

B3.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It discloses the execution environment (real Chrome via extension, user's browser) which is useful, but says nothing about what clicking causes — navigation, page waits, permissions, or whether the action is reversible.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two terse components: a bracketed environment/usage caveat followed by the action statement. It is front-loaded and wastes no words, though the bracketed note is somewhat dense compared to the bare action phrase.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 3-parameter mutation tool with no annotations and no output schema, the description covers the environment and a usage boundary but leaves behavioral effects and expected outcomes unstated. It is minimally adequate rather than complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and all three parameters are documented in the schema, including the xpath fallback rationale and intent. The description adds no parameter detail beyond the index name already in the schema, so baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb (click) and resource (element), and the bracketed preamble identifies the environment as a real Chrome browser via the OctoLimb extension. It distinguishes the tool from recipe-based siblings by contrast, though it does not name a specific alternative like run_recipe.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

"use only when the task requires acting on a real website" gives an explicit when-to-use constraint that implicitly steers the agent away from the recipe tools. It stops short of naming the alternative or stating prerequisites, so it's clear but not exhaustive.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

close_tabB

[Real Chrome browser via the OctoLimb extension — drive the user's browser; use only when the task requires acting on a real website.] Close tab by tab id

ParametersJSON Schema
NameRequiredDescriptionDefault
intentNoPurpose of this action
tab_idYesid of the tab

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It does disclose a meaningful behavioral trait the schema cannot: that this drives the user's actual Chrome browser via an extension, implying real, externally visible side effects. It omits, though, whether closing is reversible, what happens to page state, or what occurs with an invalid tab id.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short clauses, front-loaded with the environmental context and then the action. Every element earns its place, though the bracketed environment note is likely repeated boilerplate across sibling tools rather than specific to this action.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple two-parameter, no-output-schema tool the description is adequate, but given zero annotation coverage a mutation-style action like closing a tab should say more about permanence and side effects than it does.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both parameters (intent, tab_id) are already documented structurally. The description only echoes 'tab id' without adding format, range, or sourcing detail, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Close tab') and identifies the key selector ('by tab id'), which an agent can trivially distinguish from siblings like open_tab and switch_tab. It stops short of differentiating itself explicitly, but the intent is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The bracketed prefix gives a real when-to-use constraint ('use only when the task requires acting on a real website'), which is genuine scoping guidance. However, it says nothing about when to close versus switch away, nor any prerequisites such as the tab existing or being non-active.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

doneA

[Real Chrome browser via the OctoLimb extension — drive the user's browser; use only when the task requires acting on a real website.] Complete task. Always call this once you're finished with a browser task, whether it succeeded or not — besides signaling completion, on success this automatically saves everything you did since the last find_recipe/done as a recipe (using text as its instruction), so a future task on this site can replay it via find_recipe instead of redoing every step. This is the only thing that triggers auto-saving — individual actions (go_to_url, click_element, etc.) are silently recorded as you go, but nothing is persisted until done() is called with success: true.

ParametersJSON Schema
NameRequiredDescriptionDefault
textYes
successYes

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and mostly discharges it: it discloses that actions are silently recorded but not persisted until done is called, that success:true gates the auto-save, and that a recipe is created from `text`. It does not describe the response or what is returned on failure, leaving a modest gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The bracketed browser-scope preamble is front-loaded and the core instruction ('Always call this once you're finished…') follows immediately. The closing sentence about silent recording is dense but earns its place by explaining the persistence model; overall slightly long but without filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a two-parameter tool with no annotations and no output schema, the description covers the essential behavior — completion signaling, auto-save trigger, recipe derivation from `text`. The main omission is any statement of what a caller receives back, but with no output schema the agent mainly needs invocation guidance, which is present.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate; it explains that `text` becomes the recipe's replay instruction and that `success` gating determines whether the sequence is persisted. This meaningfully clarifies both parameters, though it does not state the expected format or length of `text`.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource — completing the browser task and signaling the end of it — and the bracketed preamble scopes it to the real-browser workflow. An agent can distinguish it from action siblings like click_element or go_to_url, though it never explicitly names which sibling would be used instead for other purposes.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly says to always call it once finished with a browser task, succeeded or not, and explains that it is the only trigger for auto-saving a recipe. It also ties itself to find_recipe as the replay alternative, giving the agent a clear when-to-use and what-competing-tool-to-use story.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

execute_jsA

[Real Chrome browser via the OctoLimb extension — drive the user's browser; use only when the task requires acting on a real website.] Execute arbitrary JavaScript in this task's tab (page context) and return the result.

ParametersJSON Schema
NameRequiredDescriptionDefault
codeYesJavaScript to execute. Returns the value of the last expression, like a REPL (e.g. `document.title` returns the title with no explicit return needed).

TDQS

A3.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the burden. It usefully discloses the execution realm (page context, not an isolated sandbox), the real-browser/extension dependency, and that the last expression value is returned. It omits any warning that this is arbitrary code execution capable of mutating or destroying page state, and gives no permissions, timing, or error behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Compact and front-loaded: the bracketed environment/scope note comes first, then the action sentence. The bracket is slightly dense but every element (extension, real browser, page context, return) earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With one required parameter, full schema coverage, no output schema, and no annotations, the description covers context, execution realm, and return semantics adequately. A note on mutation risk or limits would round it out, but nothing needed to invoke it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the single parameter 'code' is fully documented in the schema, including the REPL-style last-expression return convention. The description adds no syntax, size, or format guidance beyond that, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Execute) and resource (arbitrary JavaScript) with the exact execution context (this task's tab / page context) and return behavior. It also implicitly separates itself from the recipe and page-reading siblings by being the raw code-execution path.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives a clear when-to-use gate: 'use only when the task requires acting on a real website,' plus the qualifier that it drives the user's real Chrome via the OctoLimb extension. It does not name the concrete alternative (e.g. run_recipe) for when a real browser is not required, so it stops short of full routing guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

find_recipeA

[Real Chrome browser via the OctoLimb extension — drive the user's browser; use only when the task requires acting on a real website.] List every previously saved step-by-step trajectory for this domain, before doing the task manually. Returns each saved recipe's own instruction text plus its fingerprint and success/failure counts — no steps yet. Read the instructions like retrieved documents and judge for yourself whether any of them is the same task you're about to do, even if worded differently ('post a tweet' vs 'tweet something'); ignore recipes that don't actually match. If one fits, pass its fingerprint straight into run_recipe to replay the whole task in one call instead of one tool call per step. Always try this first for any task that sounds like something you (or a prior session) may have already done on this site.

ParametersJSON Schema
NameRequiredDescriptionDefault
domainYesThe site's domain the task runs on (e.g. 'x.com')

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does so: it states the environment (real Chrome via OctoLimb extension), what is returned ('instruction text plus its fingerprint and success/failure counts - no steps yet'), and that matching is fuzzy so the agent must judge similarity. This is richer behavioral context than most definitions provide.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The browser-requirement caveat and the when-to-use instruction are front-loaded, which is good structure. It is a dense five-sentence block; the parenthetical example ('post a tweet' vs 'tweet something') earns its place, but the closing sentence partially restates the opening instruction.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, so the description must explain the return shape, and it does (per-recipe instructions, fingerprint, success/failure counts, no steps yet). Combined with the routing to run_recipe, an agent has everything needed to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Only one parameter, documented at 100% schema coverage with its own example ('x.com'), so the schema already does the work. The description adds only the implicit notion that recipes are scoped per-domain; no additional syntax or format meaning is supplied, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

It names a specific verb+resource ('list every previously saved step-by-step trajectory for this domain') and clearly distinguishes itself from the sibling run_recipe, which is where matches get replayed. An agent can tell this apart from read_page or search_google without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit when-to-use ('always try this first for any task that sounds like something you may have already done'), what to do with results ('read the instructions like retrieved documents and judge for yourself'), when to ignore ('ignore recipes that don't actually match'), and the named alternative ('pass its fingerprint straight into run_recipe').

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_dropdown_optionsB

[Real Chrome browser via the OctoLimb extension — drive the user's browser; use only when the task requires acting on a real website.] Get all options from a native dropdown

ParametersJSON Schema
NameRequiredDescriptionDefault
indexYesindex of the dropdown element
intentNoPurpose of this action

TDQS

B3.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. 'Get all options' implies a read, but there is no statement of side effects, whether the element must be visible/expanded, or what happens on an invalid index. For a zero-annotation tool this is thin.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One compact sentence after the shared browser-context prefix, with the action front-loaded. The prefix is boilerplate but justified for routing decisions among browser-vs-other tools.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no annotations and no output schema, the description should at least hint at the return shape (a list of option labels/values) and indexing semantics, neither of which appears. It is minimally adequate for a simple read tool but leaves real gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% ('index of the dropdown element', 'Purpose of this action'), so the schema already documents both parameters. The description adds no syntax, indexing base, or format detail beyond that, making 3 the correct baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (get all options) and resource (native dropdown), which clearly separates it from the write-side sibling select_dropdown_option. It stops short of naming that sibling, so differentiation is implied by verb rather than explicit.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The bracketed note 'use only when the task requires acting on a real website' gives a when-to-use condition, but it is generic boilerplate shared across this whole tool family rather than guidance specific to dropdown inspection. Nothing says whether the dropdown must be opened first or how it relates to select_dropdown_option.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

go_backA

[Real Chrome browser via the OctoLimb extension — drive the user's browser; use only when the task requires acting on a real website.] Go back to the previous page

ParametersJSON Schema
NameRequiredDescriptionDefault
intentNoPurpose of this action

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the burden. It usefully discloses the execution environment (real Chrome via OctoLimb extension), which matters for an agent reasoning about side effects. It says nothing about what happens if there is no history to go back to, or whether page state is lost.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Very short and front-loaded: the environment caveat precedes the action. Every clause earns its place, though the bracketed preamble is heavier than the one-line action it describes.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple, no-argument, no-output navigation tool with a fully documented schema, the description covers the essentials — environment and action. Only the history-exhausted/failure behavior and sibling disambiguation are missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the single 'intent' parameter is already documented in the schema, so the baseline of 3 applies. The description adds no extra meaning about what intent should contain for this action.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource: 'Go back to the previous page.' It is immediately understandable on its own. However, it does not differentiate itself from the sibling 'previous_page' (or 'next_page'), leaving genuine ambiguity about which navigation tool to pick.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The bracketed prefix gives a real constraint — 'use only when the task requires acting on a real website' — which is useful scoping. But there is no guidance on when to use this versus 'previous_page' vs history-independent alternatives, so the routing decision is left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

go_to_urlA

[Real Chrome browser via the OctoLimb extension — drive the user's browser; use only when the task requires acting on a real website.] Navigate to URL. Opens a new tab for this task if none is active yet, otherwise navigates the tab already in use for this task — never the user's own currently-focused tab, and every later action (click_element, read_page, etc.) keeps targeting that same tab even if the user clicks into a different one meanwhile.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYes
intentNoPurpose of this action

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does unusually well: it explains tab lifecycle (new tab if none active, otherwise the task's tab), guarantees it never touches the user's focused tab, and notes that later actions keep targeting the same tab regardless of user focus changes. It still omits error behavior, auth expectations, and return values.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Well front-loaded: the scope caveat comes first, then the core action, then the behavioral detail. Slightly long in the trailing tab-targeting clause, but every sentence adds actionable information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a navigation tool with no output schema and no annotations, the description covers the essential behavioral contract (which tab is targeted and how that persists). Only error/return semantics are absent, which is minor for this tool type.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 50%: 'intent' is documented in the schema but 'url' is not. The description's 'Navigate to URL' only marginally clarifies the url parameter, adding no format or resolution details beyond the parameter name. Baseline 3 is appropriate for this partial coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (navigate) and resource (URL), and additionally scopes it to the real Chrome browser via the OctoLimb extension, distinguishing it from hypothetical in-app navigation. An agent can tell what it does without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The bracketed clause 'use only when the task requires acting on a real website' gives a clear positive/negative usage condition. It does not name specific alternatives among siblings like open_tab or search_google, so routing between those still requires inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

input_textA

[Real Chrome browser via the OctoLimb extension — drive the user's browser; use only when the task requires acting on a real website.] Input text into an interactive input element

ParametersJSON Schema
NameRequiredDescriptionDefault
textYestext to input
indexYesindex of the element
xpathNoxpath of the element (optional fallback if the index is stale)
intentNoPurpose of this action

TDQS

A3.5/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description must carry the full behavioral burden. It says the tool drives the user's browser and inputs text, but does not explain whether existing text is replaced, whether events are triggered, whether the element must already be focused, or what errors can occur. Only minimal behavioral context is conveyed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences: a bracketed environment/usage note followed by a precise purpose statement. It is front-loaded, compact, and contains no redundant or wasted text.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers purpose and usage context, and the schema fully covers parameters. However, with no annotations and no output schema, it omits important behavioral details for a browser input action, such as replacement behavior, event triggering, and failure modes.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all four parameters, including the optional xpath fallback and the intent field. The description adds no additional meaning about parameters. This meets the baseline for a fully described schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'Input text into an interactive input element.' It also clarifies the execution environment as a real Chrome browser via the OctoLimb extension. It does not explicitly distinguish itself from the related sibling send_keys, but the core purpose is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The bracketed note gives a clear usage condition: 'use only when the task requires acting on a real website.' This tells the agent when this browser-based tool is appropriate. It does not name an alternative tool or elaborate on exclusions, but the context is strong enough for correct selection.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

next_pageA

[Real Chrome browser via the OctoLimb extension — drive the user's browser; use only when the task requires acting on a real website.] Scroll the document in the window or an element to the next page. If no index is specified, scroll the whole document.

ParametersJSON Schema
NameRequiredDescriptionDefault
indexNoindex of the element
intentNoPurpose of this action

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden: it discloses the default behavior when index is omitted ('scroll the whole document') but says nothing about what happens at the end of the document, whether the scroll is animated or waits for lazy-loaded content, or whether any permission/auth is required. That is adequate for a low-risk scroll action but leaves real behavioral gaps.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with the core action and the caveat folded in efficiently; nothing is redundant within the body. The bracketed preamble is boilerplate-ish framing rather than tool-specific information, which costs it the top score.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a two-parameter, no-output-schema, no-annotation tool, the description covers the action, the browser context, and the index default. Nothing critical is missing for correct invocation, though the page-end/termination behavior would be useful given there is no output schema to clarify results.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3, but the description goes beyond the schema by explaining the semantics of omitting index ('If no index is specified, scroll the whole document'), which the schema's 'index of the element' does not convey. It adds no guidance on the intent parameter, so it is not a 5.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Scroll the document in the window or an element to the next page'), which is distinct from the sibling scrolling tools scroll_to_top, scroll_to_bottom, scroll_to_percent and previous_page. The bracketed preamble further differentiates it from non-browser siblings by declaring it drives a real Chrome browser. It does not, however, explicitly name the sibling it is not (e.g. a one-page-at-a-time vs. arbitrary-offset distinction).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The preamble gives a clear when-to-use condition: 'use only when the task requires acting on a real website.' That is a genuine scoping rule rather than implied guidance. It stops short of routing within the scroll family (no mention of when previous_page or scroll_to_bottom would be preferable), so it is not a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

open_tabA

[Real Chrome browser via the OctoLimb extension — drive the user's browser; use only when the task requires acting on a real website.] Open URL in a new tab, and use it as this task's tab from now on.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYesurl to open
intentNoPurpose of this action

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It usefully discloses that this drives the user's real Chrome browser via an extension and that the tab becomes the task's active tab going forward — significant side effects — but says nothing about permissions, tab limits, or failure behavior when the extension is unavailable.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single sentence with the environment caveat bracketed up front and the action plus its persistent effect packed in afterward. No filler text.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple two-parameter navigation tool with no output schema and no nested objects, the description covers what the agent needs: it acts on a real browser, opens a tab, and makes that tab the task's working tab. The missing piece is differentiation from the many navigation siblings.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% with only two parameters ('url' and 'intent'), so the schema already documents both. The description adds no syntax, format, or URL requirements beyond what the schema states; the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Open URL in a new tab') and adds a meaningful side effect: the new tab becomes this task's tab. It does not explicitly distinguish itself from the adjacent sibling go_to_url, which an agent must infer covers in-place navigation rather than a new tab.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides clear gating context — 'use only when the task requires acting on a real website' — which tells the agent when not to use it. However it never names the alternative (go_to_url, switch_tab, run_recipe) or contrast conditions, so the routing decision remains partly inferred.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

previous_pageA

[Real Chrome browser via the OctoLimb extension — drive the user's browser; use only when the task requires acting on a real website.] Scroll the document in the window or an element to the previous page. If no index is specified, scroll the whole document.

ParametersJSON Schema
NameRequiredDescriptionDefault
indexNoindex of the element
intentNoPurpose of this action

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It discloses the key behavioral context that this drives the user's real Chrome browser via an extension, and that scrolling can target the document or a single element, which is useful. However it says nothing about scroll magnitude (viewport vs. full page), whether it blocks, or any auth/permission requirements.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded guidance in brackets followed by two tight sentences describing the action and default. The preamble is somewhat verbose but earns its place by stating the real-browser constraint; no filler beyond that.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a no-output, low-risk scroll tool with both parameters documented, the description supplies what an agent needs to invoke it correctly. The only residual gap is the ambiguity of 'previous page' (viewport scroll vs. paginated content), which slightly weakens completeness against the next_page sibling.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline would be 3, but the description adds real meaning by explaining the default behavior of the index parameter (whole document when omitted) that the terse schema does not convey. It leaves the 'intent' parameter meaning entirely to the schema, hence not a 5.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Scroll the document in the window or an element to the previous page') and scopes the fallback ('if no index is specified, scroll the whole document'). It is distinguishable from siblings like next_page or scroll_to_top by its direction, though it never explicitly contrasts itself with them.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The bracketed preamble gives a clear use condition: 'use only when the task requires acting on a real website,' which implicitly excludes non-browser or mocked contexts. It provides context for selection but names no specific alternative sibling (e.g., next_page or scroll_to_percent) the agent should pick instead.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

read_pageA

[Real Chrome browser via the OctoLimb extension — drive the user's browser; use only when the task requires acting on a real website.] Read the current page in this task's tab: URL, title, and an indexed list of interactive elements. Call this before click_element, input_text, or the dropdown tools to get valid indices. Body text is omitted by default — pass include_text when you actually need to read page content, not just interact with it.

ParametersJSON Schema
NameRequiredDescriptionDefault
include_textNoInclude the page's visible body text (default false — most steps only need element indices)
max_text_lengthNoMax characters of body text to return when include_text is true (default 4000)
viewport_expansionNoExpand the viewport by this many pixels when indexing elements (default 0 — only visible elements are indexed). Set to -1 to index all interactive elements on the page regardless of visibility.

TDQS

A4.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden and mostly succeeds: it discloses the environment (real Chrome via OctoLimb extension), the default omission of body text, and that indices must be refreshed before acting. It does not cover failure modes such as stale indices or empty/detached tabs, which leaves a small gap for a browser-driving tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The bracketed context note is front-loaded, followed by what is returned, when to call it, and the include_text caveat. Every sentence adds actionable information with no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema and no annotations, the description must explain return shape and it does (URL, title, indexed elements) as well as the default text behavior. An agent has everything needed to call it correctly and interpret the result.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is already 100%, so the baseline is 3, but the description adds intent beyond the schema: include_text should be passed only when page content is genuinely needed, not for interaction. It also reinforces that omitted text is the default, though max_text_length and viewport_expansion are left entirely to the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Read the current page in this task's tab') and enumerates exactly what comes back: URL, title, and an indexed list of interactive elements. This cleanly separates it from siblings like click_element or input_text, which act rather than read.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives both when-to-use ('use only when the task requires acting on a real website') and an explicit ordering rule: call this before click_element, input_text, or the dropdown tools to get valid indices. It also names the condition for the text-reading variant versus mere interaction.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

run_recipeA

[Real Chrome browser via the OctoLimb extension — drive the user's browser; use only when the task requires acting on a real website.] Execute a sequence of OctoLimb tool steps in one call, without round-tripping to you between each one — use this to batch obvious action chains (e.g. go_to_url, then click_element, then input_text) or to replay steps returned by find_recipe. Runs steps exactly as given, in order, via the real browser — nothing is added or skipped, so include every step you'd need if doing this by hand one call at a time (e.g. a read_page before an index-based step on a page you haven't indexed yet in this session). Stops at the first step that errors, returning every result so far plus which step failed — pick up manually from there. Pass fingerprint (from find_recipe) to replay and auto-update its success/failure count instead of passing steps directly. Pass save_as alongside steps to save this task once it fully succeeds — this saves every tool call made in the session so far (including ones made outside this run_recipe call, e.g. an earlier read_page), not just the steps in this call, so replay later needs nothing extra.

ParametersJSON Schema
NameRequiredDescriptionDefault
stepsNoAd-hoc steps to run now, each a real OctoLimb tool call (mutually exclusive with fingerprint)
save_asNoIf provided with `steps` and every step succeeds, save this as a recipe for future find_recipe lookups
fingerprintNoReplay a saved recipe by the fingerprint returned from find_recipe (mutually exclusive with steps)

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and meets it: it discloses that steps run exactly as given in order, that it stops at the first error and returns all prior results plus the failing step, and that saving captures every session tool call, not just this call's steps. These are non-obvious behaviors an agent could not infer from the schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core action, then the operational caveats in descending importance. It is dense and long, but nearly every clause adds a distinct operational fact; the parenthetical examples are slightly verbose but still load-bearing.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a nested, no-annotation, no-output-schema tool, the description covers execution order, error semantics, return-on-failure behavior, replay, and save scope. Nothing an agent needs to invoke it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds real meaning beyond the schema: fingerprint replays and auto-updates a success/failure counter, and save_as persists the whole session's calls rather than just steps. The mutual-exclusivity and the need to include a read_page before an index-based step are also clarified.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Execute a sequence of OctoLimb tool steps in one call') and immediately scopes it as a batching primitive, distinguishing it from the single-step siblings like click_element and go_to_url. It also carves out its relationship to find_recipe (replay) explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit when-to-use conditions ('batch obvious action chains', 'replay steps returned by find_recipe') and a when-not ('use only when the task requires acting on a real website'). It also states the selection rule between the steps and fingerprint parameters.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

scroll_to_bottomB

[Real Chrome browser via the OctoLimb extension — drive the user's browser; use only when the task requires acting on a real website.] Scroll the document in the window or an element to the bottom

ParametersJSON Schema
NameRequiredDescriptionDefault
indexNoindex of the element
intentNoPurpose of this action

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It does disclose meaningful context: this drives a real Chrome browser via the OctoLimb extension and acts on the user's browser. It does not say whether the scroll is instant or animated, whether it waits for load, or what happens if no index is supplied.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One bracketed environment caveat plus one crisp action sentence; front-loaded and free of filler. Slightly longer than strictly necessary because the environment note duplicates context likely available elsewhere.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple two-optional-parameter scroll tool with no output schema, the description covers the action and environment adequately. It leaves gaps around element targeting (what 'index' refers to), scroll completion behavior, and no guidance on the result, but nothing critical is missing for a low-risk action.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both parameters are documented in the schema, and the description adds nothing about index or intent. Baseline 3 applies when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (scroll) and target (document in window or element) with a clear endpoint (bottom). It differentiates by outcome from scroll_to_top and scroll_to_percent, though it never names them explicitly, so the agent must infer the distinction from the sibling list.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The bracketed preface tells the agent to use this only when the task needs a real website, which is useful gating. However, it gives no guidance on when to choose this over scroll_to_percent, scroll_to_top, or scroll_to_text, so selection among scroll variants is left implicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

scroll_to_percentA

[Real Chrome browser via the OctoLimb extension — drive the user's browser; use only when the task requires acting on a real website.] Scrolls to a particular vertical percentage of the document or an element. If no index of element is specified, scroll the whole document.

ParametersJSON Schema
NameRequiredDescriptionDefault
indexNoindex of the element
intentNoPurpose of this action
yPercentYespercentage to scroll to - min 0, max 100; 0 is top, 100 is bottom

TDQS

A4.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It does disclose meaningful behavioral context: this drives the user's real browser through the OctoLimb extension, and that omitting index scrolls the whole document rather than an element. It does not describe return values, side effects on viewport state, or failure modes.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core action, followed by the conditional element-vs-document distinction. The bracketed environment note is boilerplate but earns its place by setting the real-browser context; no wasted sentences.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a three-parameter, no-output-schema scroll tool with no annotations, the description covers the action, the optional-index behavior, and the operating environment adequately. Only minor gaps (no mention of what is returned or error conditions) remain.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds genuinely useful semantics beyond the schema: the conditional behavior when no element index is specified (scroll whole document). The yPercent bounds and meaning are left to the schema, which documents them well.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (scroll) and resource (vertical percentage of document or element), and distinguishes itself from nearby siblings like scroll_to_top, scroll_to_bottom, and scroll_to_text by defining percentage-based scrolling. An agent can select it without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The bracketed preamble gives a clear usage condition — use only when the task requires acting on a real website via the real Chrome browser. It does not, however, explicitly route between this and sibling scroll tools (e.g., when to prefer scroll_to_percent vs scroll_to_top), leaving that to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

scroll_to_textA

[Real Chrome browser via the OctoLimb extension — drive the user's browser; use only when the task requires acting on a real website.] If you dont find something which you want to interact with in current viewport, try to scroll to it

ParametersJSON Schema
NameRequiredDescriptionDefault
nthNowhich occurrence of the text to scroll to (1-indexed, default: 1)
textYestext to scroll to
intentNoPurpose of this action

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description must carry behavioral disclosure. It usefully states that this drives a real Chrome browser via the OctoLimb extension and should only be used for real website tasks. However, it does not describe what happens when the text is not found, whether scrolling is smooth or instant, or any side effects, leaving some behavioral gaps.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is short and front-loads the environmental constraint before the usage condition. It contains a minor grammatical typo ('dont') but no wasted sentences, and every part contributes useful context.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple scroll-to-text action with a fully described input schema and no output schema, the description provides adequate completeness: it names the browser environment, the real-website restriction, and the viewport condition. Nothing essential for correct invocation is missing, though failure behavior remains unstated.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema fully documents all three parameters including nth and intent. The description adds no additional parameter-level meaning beyond what the schema already provides, which sets the baseline at 3.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states a specific action (scroll to text) and its triggering condition (when the target is not visible in the current viewport). It does not explicitly distinguish this tool from sibling scroll tools like scroll_to_percent or scroll_to_top, but the name and description together make the purpose unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives a clear when-to-use condition: when the desired element is outside the current viewport. The bracketed context also specifies that it should only be used when acting on a real website via the OctoLimb extension. It lacks explicit when-not or alternative scroll-tool guidance, but the context is sufficient.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

scroll_to_topB

[Real Chrome browser via the OctoLimb extension — drive the user's browser; use only when the task requires acting on a real website.] Scroll the document in the window or an element to the top

ParametersJSON Schema
NameRequiredDescriptionDefault
indexNoindex of the element
intentNoPurpose of this action

TDQS

B3.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description must carry the full behavioral burden. It only states the action and mentions the real-browser context; it does not disclose side effects, required permissions, whether scrolling is instant or smooth, or what happens if the page/element is not scrollable. This is a significant gap for a tool with zero annotation coverage.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is short: one main sentence with a bracketed context prefix. It is mostly front-loaded, though the context note precedes the actual action, slightly delaying the specific purpose. Every part earns its place, but the ordering could be improved.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple scroll tool with no output schema and a straightforward two-parameter input, the description covers the basic purpose. However, without annotations, it should disclose more behavioral details (e.g., side effects, return behavior) to fully equip the agent. The real-browser context note is helpful but generic.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description adds that the scroll target can be 'the window or an element,' which clarifies the role of the optional index parameter (element selection) beyond the schema's brief 'index of the element.' This is a small addition, but the schema already documents both parameters adequately.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific action: 'Scroll the document in the window or an element to the top.' It clearly covers scrolling to the top and targets both window and element. However, it does not distinguish itself from sibling scroll tools like scroll_to_bottom or scroll_to_percent, leaving the agent to infer based on the name alone.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It provides a contextual condition: '[Real Chrome browser... use only when the task requires acting on a real website.]' This is useful guidance for when to use real-browser tools generally. But it offers no direction on when to choose this tool over the numerous other scroll variants (scroll_to_percent, scroll_to_bottom, scroll_to_text, etc.).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_googleA

[Real Chrome browser via the OctoLimb extension — drive the user's browser; use only when the task requires acting on a real website.] Search the query in Google. Opens a new tab for this task if none is active yet, otherwise reuses the tab already in use for this task — never the user's own currently-focused tab. The query should be a search query like humans search in Google, concrete and not vague or super long. More the single most important items.

ParametersJSON Schema
NameRequiredDescriptionDefault
queryYes
intentNoPurpose of this action

TDQS

A4.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden. It usefully discloses tab lifecycle behavior: it opens a task tab or reuses the task's tab and never the user's focused tab. However, it does not describe return behavior, wait conditions, or failure modes.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The definition is front-loaded and compact, with the core purpose stated before behavioral details. The phrase 'More the single most important items' is awkward and slightly reduces polish, but otherwise the structure is efficient.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no annotations and no output schema, the description does a good job covering scope, browser context, tab handling, and query quality. It leaves some gaps around what the tool returns after the search, but for a browser-action tool in this sibling set, it is largely sufficient.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 50%, so the description must compensate for the required query parameter. It does so by explaining that the query should be human-like, concrete, not vague or overly long, and focused on the single most important item. The optional intent parameter is left to its schema description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'Search the query in Google.' It also scopes the tool to a real Chrome browser via the OctoLimb extension, distinguishing it from generic URL navigation or page-reading siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives an explicit condition: 'use only when the task requires acting on a real website.' That is clear usage guidance, but it does not name or compare against alternative sibling tools such as go_to_url or run_recipe.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

select_dropdown_optionA

[Real Chrome browser via the OctoLimb extension — drive the user's browser; use only when the task requires acting on a real website.] Select dropdown option for interactive element index by the text of the option you want to select

ParametersJSON Schema
NameRequiredDescriptionDefault
textYestext of the option
indexYesindex of the dropdown element
intentNoPurpose of this action

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the burden. It usefully discloses that this drives the user's actual Chrome browser through the OctoLimb extension (a real side-effecting action), but says nothing about failure modes (option not found), whether the dropdown must be open first, or timing/auto-wait behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Compact — one bracketed environment note plus one core sentence. Front-loaded with the selection action, with no redundant padding.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Adequate for a simple interactive action with no output schema, but with no annotations and no sibling routing it leaves the agent to infer how to obtain valid option texts and whether the dropdown must be opened beforehand. Complete enough to call, incomplete on the workflow around it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description restates the schema's own semantics (select by option text, target element index) without adding format details, and the optional 'intent' parameter is not mentioned at all.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (select) and resource (dropdown option) plus the selection key (option text, element index). It is clearly distinct from get_dropdown_options and click_element in intent, though it never names those siblings to sharpen the distinction.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The bracketed clause gives real context on when to use it: only when the task requires acting on a real website via the user's Chrome browser. It does not, however, explain the relationship to sibling get_dropdown_options (e.g., call it first to learn valid option texts), so alternatives are left implicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

send_keysB

[Real Chrome browser via the OctoLimb extension — drive the user's browser; use only when the task requires acting on a real website.] Send strings of special keys like Backspace, Insert, PageDown, Delete, Enter. Shortcuts such as Control+o, Control+Shift+T are supported as well. This gets used in keyboard press. Be aware of different operating systems and their shortcuts

ParametersJSON Schema
NameRequiredDescriptionDefault
keysYeskeys to send
intentNoPurpose of this action

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden. It usefully discloses that this drives a real Chrome browser via the OctoLimb extension and warns about OS-specific shortcut differences, but it omits where the keys land (focused element? page body?), what happens on failure, and any side effects such as page navigation from Enter.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The purpose and examples are front-loaded and earn their place, but 'This gets used in keyboard press' is filler restating the name, and the closing OS warning is vague rather than actionable. The bracketed preamble is boilerplate shared across a tool family.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a two-parameter real-browser action tool with no annotations and no output schema, the description covers purpose and key format but leaves target-focus behavior, failure modes, and the intent parameter unexplained. Adequate to call it, thin for calling it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so baseline is 3, but the description goes beyond the terse 'keys to send' schema text by giving concrete syntax examples ('Control+o', 'Control+Shift+T') that define the expected key-string format. The second parameter (intent) is never mentioned, which keeps this from a 5.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('send strings of special keys') and enumerates examples (Backspace, Insert, Control+o), so the agent knows this is keyboard-event injection rather than text entry. It only implicitly distinguishes itself from the sibling input_text, which types characters into fields.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The bracketed preamble gives a real routing condition ('use only when the task requires acting on a real website'), which is genuinely useful for choosing this over a non-browser path. However, it never says when to prefer send_keys over input_text or execute_js for keyboard work, leaving the sibling choice to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

switch_tabA

[Real Chrome browser via the OctoLimb extension — drive the user's browser; use only when the task requires acting on a real website.] Switch to tab by tab id, and make it the tab this task uses from now on (e.g. after a link opened a new tab).

ParametersJSON Schema
NameRequiredDescriptionDefault
intentNoPurpose of this action
tab_idYesid of the tab to switch to

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the burden and does disclose a meaningful stateful side effect: the switched-to tab becomes the tab this task uses from now on. It does not mention failure behavior for an invalid tab_id or any permission requirements, leaving a small gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The bracketed environment/safety preamble is front-loaded, followed by a single tight sentence carrying the operation and its persistent effect. No filler and nothing redundant.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple one-required-parameter tool with no output schema or annotations, the description covers environment constraint, action, and the persistent tab-selection effect. Only error/edge-case behavior is unstated, which is a minor omission.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both tab_id and intent are already documented in the schema; the description only restates that switching is by tab id. Baseline 3 applies since the schema does the heavy lifting and the description adds no format or range detail.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (switch) plus resource (tab) and adds the crucial scope detail that the tab becomes the task's active tab going forward. This distinguishes it cleanly from siblings like open_tab, close_tab, and go_to_url without needing the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives a clear condition of use ('use only when the task requires acting on a real website') and a concrete trigger example ('after a link opened a new tab'). It does not explicitly name the alternative tools (open_tab/close_tab) or state when not to use it, so it stops short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

waitA

[Real Chrome browser via the OctoLimb extension — drive the user's browser; use only when the task requires acting on a real website.] Wait for x seconds default 3, do NOT use this action unless user asks to wait explicitly

ParametersJSON Schema
NameRequiredDescriptionDefault
intentNoPurpose of this action
secondsNoamount of seconds (default 3)

TDQS

A4.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations exist, so the description must carry behavioral disclosure. It does disclose that this drives the user's real browser via the OctoLimb extension, which is meaningful context, but says nothing about blocking vs non-blocking behavior, side effects, or what a call returns.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Short and mostly front-loaded, with the critical 'do NOT use' warning present. The bracketed browser preamble comes first, which delays the actual action, and the 'default 3' claim duplicates the schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-required-parameter wait tool with full schema coverage and no output schema, the description supplies what is needed: what it does, when it is permitted, and the environment it operates in. Only the blocking/return behavior is left unspecified.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both 'intent' and 'seconds' are already documented. The description's 'default 3' simply restates the schema default and adds no new syntax or constraint information.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific action ('Wait for x seconds, default 3') and frames it in the real-Chrome/OctoLimb browser context. No sibling tool in the list performs waiting, so the operation is unambiguously distinguishable.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives an explicit exclusion ('do NOT use this action unless user asks to wait explicitly') plus a context condition ('use only when the task requires acting on a real website'). That is precisely the when/when-not guidance agents need, and it is rare to see it stated so directly.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 24 tool updatesv1.0.0
    • First observedcache_content
    • First observedclick_element
    • First observedclose_tab
    • First observeddone
    • First observedexecute_js
    • First observedfind_recipe
    • First observedget_dropdown_options
    • First observedgo_back
    • First observedgo_to_url
    • First observedinput_text
    • First observednext_page
    • First observedopen_tab
    • First observedprevious_page
    • First observedread_page
    • First observedrun_recipe
    • First observedscroll_to_bottom
    • First observedscroll_to_percent
    • First observedscroll_to_text
    • First observedscroll_to_top
    • First observedsearch_google
    • First observedselect_dropdown_option
    • First observedsend_keys
    • First observedswitch_tab
    • First observedwait

TDQS

A3.5/5.0

Scored across 24 tools

Disambiguation4/5

Most tools target clearly distinct browser actions, and the descriptions explicitly distinguish tab-scoped navigation from browser history (go_back) and scrolling. There is some potential overlap between scroll variants and pagination tools, and execute_js could be seen as a catch-all alternative, but the intended boundaries are mostly clear.

Naming Consistency4/5

Nearly all names use readable snake_case, and action tools generally follow verb_noun or verb patterns. Minor deviations include one-word controls like done and wait, and noun-only names like previous_page and next_page, but the set remains predictable overall.

Tool Count3/5

At 24 tools, the set is on the heavy side for a browser automation server and includes several granular scroll and tab primitives that could be consolidated. However, each tool does correspond to a real browser operation, so the count is not entirely unjustified.

Completeness4/5

The surface covers core browser workflows: navigation, inspection, clicking, text input, scrolling, tabs, dropdowns, JavaScript execution, and recipe replay. Some gaps remain, such as explicit hover, drag-and-drop, file upload, screenshots, or forward navigation, but these are relatively minor for many tasks.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    B
    maintenance
    Hosted Chrome as an MCP skill. Enables any MCP-compatible agent to drive a real Chromium browser for tasks like navigation, clicking, typing, and taking screenshots.
    19 npm
    MIT
  • A
    license
    Not graded
    quality
    A
    maintenance
    Enables local opencode agents to control a live Chrome browser via MCP tools, including tab management, JavaScript execution, clicking, form filling, page reading, screenshots, and console log retrieval.
    MIT
  • F
    license
    A
    quality
    C
    maintenance
    Enables coding agents to drive a real Playwright-backed browser session, navigating, observing, interacting, and capturing proof on web pages via MCP tools across stdio and Streamable HTTP transports.
    15
    2
    -