Skip to main content
Glama

Sleeper

English · 简体中文 · 日本語 · Português · Español · Deutsch

Sleeper browser control for agents

Let your agent control the browser you already use, whether it’s on your desktop or Android phone. Sleeper connects Firefox or Chromium through MCP or CLI to navigate websites, read pages, fill forms, take screenshots, and extract structured data without copying session credentials into agent configuration.

Features · Install · Commands · Agent setup · Privacy · Security · Benchmarks · Contribute

Features

  • 📱 Firefox for Android (beta): Run Sleeper on Android through a private Tailscale Serve connection to your desktop daemon. No Caddy, router port, or public listener is required.

  • 🖥️ Firefox + Chromium desktop: Control pages in an existing browser session through the extension. Linux and macOS are supported; Windows is best-effort.

  • 🗂️ Profiles and tabs: Each browser installation has its own persistent ID. Agents discover connected browsers and target the intended tab.

  • 🎯 Element targeting: Find controls by CSS selector, accessible role and name, or a reference returned by snapshot.

  • 📝 Page interaction: Fill inputs, press keys, click controls, and wait for selectors or text before the next action.

  • 📋 Structured extraction: Read one element, collect matching elements, or extract a JSON map. Save repeatable work as recipes and schemas.

  • 📸 Screenshots: Capture the viewport or full page, with optional annotations. PNG output has a size limit; capture options explain it.

  • 🌐 Network and APIs: Inspect captured requests and call allowed HTTPS APIs with credentials kept in the browser and bound to their source host.

  • 🔒 Secret masking: Structured results receive best-effort redaction before reaching the CLI or MCP client. Screenshots can still contain private information.

  • 🔌 CLI and MCP: Run shell commands or call MCP tools through the local daemon.

  • 📖 Included skill: Give agents instructions for session selection, action verification, and connection recovery.

Related MCP server: open-browser-control

Install

Install uv first, then run from a checkout on Linux or macOS (the supported desktop platforms):

./install.sh

Windows works on a best-effort basis and is not a launch-supported platform; run in PowerShell:

py scripts/install.py

Choose Everything or Customize to select MCP, native agent plugins, and standalone instructions. The installer starts the daemon, installs the CLI, registers Firefox's extension-scoped native host for automatic XPI pairing, keeps Chromium's extension in a permanent folder, and copies browser packages to Downloads.

Firefox for Android connects through Tailscale Serve (beta). Run sleeper mobile setup, then open the QR code in Firefox; see the Android setup guide. The daemon remains loopback-only.

Troubleshooting

Toolbar icon opens a thin full-width strip instead of the popup — fixed by using a fixed popup width; rebuild and reinstall the extension package (see docs/commands.md for daemon environment configuration).

"no browser connected" / popup shows "Daemon disconnected" — check curl http://127.0.0.1:8790/health (daemon must be running), then run sleeper sessions (your browser should appear under profiles). For Firefox, rerun ./install.sh update if the native host was not registered; the fields under Advanced connection are manual recovery only. Extension origins are pinned to the first browser add-on that pairs with the daemon; restrict them further with SLEEPER_ALLOW_ANY_EXTENSION=false and SLEEPER_ALLOWED_EXTENSION_IDS=<uuid>,... in the daemon's startup service (systemd user unit on Linux, LaunchAgent on macOS — see docs/commands.md).

Run sleeper --version for the installed CLI version. sleeper sessions reports daemon, protocol, and connected add-on versions; an incompatible or older add-on is marked in its browser record.

Codex and Claude Code

The main installer can register Sleeper's native plugin when you choose Everything or Customize. For manual setup, install the Sleeper runtime first, then run the commands for your agent from the repository root.

Codex

codex plugin marketplace add .
codex plugin add sleeper --marketplace sleeper-local

Claude Code

claude plugin marketplace add . --scope user
claude plugin install sleeper@sleeper --scope user
# Codex
codex plugin marketplace add https://github.com/shy-tangerine/Sleeper.git
codex plugin add sleeper --marketplace sleeper-local

# Claude Code
claude plugin marketplace add https://github.com/shy-tangerine/Sleeper.git --scope user
claude plugin install sleeper@sleeper --scope user

Restart the agent after installation so it loads the bundled skill and MCP server. See agent setup for public-GitHub installation, verification, updates, and removal commands.

Browser

Finish installation

Chromium

Open the extensions page, enable Developer mode, choose Load unpacked, and select the folder printed by the installer.

Firefox

Install the signed sleeper-firefox.xpi release asset through Add-ons → Install Add-on From File. Signed releases are pending Mozilla signing credentials.

Browser permission approval is required once. Installation, updates, and removal.

Prefer manual client setup? Use the verified Codex and Claude Code plugin commands.

curl -fsSL https://raw.githubusercontent.com/shy-tangerine/Sleeper/main/install.sh | bash

Inspect the open tabs and current page:

sleeper tabs
sleeper snapshot
sleeper type --role textbox --name Search 'your query' --clear
sleeper press Enter

While an agent is working, Sleeper can show a small on-page activity message such as “Opening a page…”, “Clicked an element”, or “Timed out waiting for the page.” These messages use plain language rather than CLI/MCP command names. Typed values, uploaded file contents, and other sensitive action values are not shown. Visible activity can be turned off in the extension settings; the toolbar still exposes connection and access state.

The included agent skill covers session selection, action verification, and connection recovery. Agents can use MCP tools directly; the command guide covers tab targets, extraction, recipes, and API calls.

Waiting

Working

Benchmarks

Five verified runs per interface, on one machine. Each sequence navigates, reads a heading, types, clicks, waits for text, and reads the result.

Tested on Linux with Helium 0.17.0.1 (Chromium 153.0.8010.36). Every interface used the same browser build and a fresh profile.

Percentages use the rounded medians below.

Interface

Median sequence

Median browser RSS

Task-text tokens

Protocol-text tokens

Sleeper CLI

934 ms

1,139 MiB

441

409 (HTTP JSON)

Sleeper MCP

281 ms

1,127 MiB

502

695 (JSON-RPC)

OpenCLI

4,011 ms

1,218 MiB

364

29,152 (HTTP JSON)

Playwright MCP

1,635 ms

1,027 MiB

536

734 (JSON-RPC)

Direct CDP

291 ms

1,038 MiB

Not applicable

956 (CDP)

Task-text tokens estimate agent-facing input and output with o200k_base. Protocol tokens count internal HTTP JSON, JSON-RPC, or CDP traffic; they are not model usage. Direct CDP has no agent-facing task-text interface.

Browser RSS sums browser process memory, can double-count shared pages, and excludes daemons. Timing and protocol captures use separate runs. These results come from one machine, one browser build, and one run date (2026-09-11) measuring deterministic browser primitives, not agent task completion or general speed; see the canonical methodology caveats, which must accompany these figures wherever they are quoted. Raw samples, methodology, discovery costs, and feature comparison.

Browser

Version

Result

Firefox

155.0.1

Passed

Chrome

151.0.7922.47

Passed

Zen

1.22b

Passed

Helium

0.17.0.1 (Chromium 153.0.8010.36)

Passed

The Chrome check used Google Chrome for Testing, the automation distribution of Chrome. Its exact version is shown above.

Contribute

Build and test · Changelog · Third-party notices · Privacy · Security · MIT license · Sponsorship

Directory

Contents

extension/

Browser manifests, page handlers, popup, icons

daemon/

HTTP/WebSocket relay and MCP server

cli/

CLI, recipes, adapter support

skills/

Agent instructions

examples/

Recipes and extraction schemas

test/

Behavioral, transport, and package checks

Star history

View the star-history chart

Python package build (PyPI)

The browser extension release workflow is separate from Python packaging. Python package metadata is in pyproject.toml; uv.lock governs development. Run uv lock --check, uv sync --locked, and uv build to produce an sdist and wheel in dist/. After the owner approves a release and PyPI credentials/trusted publishing are configured, uv publish can upload those artifacts. Do not treat building as release authorization.

Available Tools

49 tools
sleeper_apiC

Call page backend API with captured auth token (background fetch)

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYes
bodyNo
methodNo
profileNoBrowser ID from sleeper_sessions; optional with one connected browser.

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden and only partially does so. It discloses that auth is handled via a captured token and runs in the background, but says nothing about whether requests can mutate server state (method/body are accepted), permission prerequisites, rate limits, or failure/return behavior — significant gaps for a request-issuing tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single short sentence with the verb front-loaded and no filler. The parenthetical is slightly cryptic but the definition is tight and well-organized, if perhaps too sparse.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 4-parameter tool with no annotations, no output schema, and 25% schema coverage, the description is too thin. It omits method semantics, mutation risk, and any indication of what the call returns, leaving the agent under-informed for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 25% — only 'profile' is documented. The description mentions the captured auth token but adds no meaning for the undocumented url, body, or method parameters, so it fails to compensate for the coverage gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Call page backend API') and adds a useful qualifier about the captured auth token and background execution mode. It is clear what the tool does, though it does not explicitly distinguish itself from likely siblings such as sleeper_exec or sleeper_network that could also issue requests.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use, when-not-to-use, or alternative-tool guidance. The '(background fetch)' parenthetical hints at a usage mode but never states when an agent should pick this over sleeper_exec or the network/wait tools, leaving the choice to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

sleeper_backD

History back

ParametersJSON Schema
NameRequiredDescriptionDefault
tabNo
profileNoBrowser ID from sleeper_sessions; optional with one connected browser.

TDQS

D1.7/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden, yet it discloses nothing: not whether navigation waits for load, what happens at the start of history, or whether it affects the active tab only. A one-line restatement of the name leaves the mutation/read profile entirely opaque.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two words is not conciseness but under-specification; there is no front-loaded purpose statement, scope, or outcome. Nothing is wasted, but nothing useful is conveyed either.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a navigation tool with no annotations, no output schema, and an undocumented tab parameter, the description should at minimum define the action, target tab semantics, and failure behavior. None of that is present, leaving the agent unable to invoke it confidently.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 50%: the profile parameter is documented as a 'Browser ID from sleeper_sessions,' but the tab parameter has no description in either schema or description. The description adds zero parameter meaning, so it fails to compensate for the uncovered half.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

"History back" essentially restates the tool name sleeper_back rather than stating a verb+resource, e.g. "navigate the current tab back one entry in browser history." An agent can infer browser-back navigation, but the description adds no distinguishing detail against siblings like sleeper_goto or sleeper_newtab.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance, no mention of when to prefer sleeper_goto (explicit URL) over back navigation, and no preconditions (e.g. does it fail on the first page?) are given. The agent gets no routing help at all.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

sleeper_batchB

Run ordered browser actions in one request; stops at the first failure by default

ParametersJSON Schema
NameRequiredDescriptionDefault
actionsYes
profileNoBrowser ID from sleeper_sessions; optional with one connected browser.
stop_on_errorNo

TDQS

B3.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations, so the description carries the full burden. It usefully discloses that execution is ordered and halts on the first failure by default, but omits what a failure returns, whether partial results are reported, the 50-action cap, and any tab/profile requirements.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single tight sentence with the core behavior front-loaded and zero filler; appropriately sized, though perhaps too terse given the tool's complexity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex batch tool with a nested action array, no annotations, and no output schema, the description is far too sparse. It never explains what cmd values are valid or how results/errors are surfaced, leaving essential invocation knowledge missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 33%: profile is documented in the schema and stop_on_error's default is echoed by the description, but the critical 'cmd' field and the nested 'tab'/'args' semantics are explained nowhere, so the description fails to compensate for the coverage gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: 'Run ordered browser actions in one request.' This clearly frames it as a batch executor over the many single-action siblings, though it never explicitly says it batches those sibling commands.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Implies usage via 'in one request' (batch multiple actions for efficiency), but gives no explicit when-to-use vs the individual action tools, no prerequisites, and no exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

sleeper_checkC

Check checkbox

ParametersJSON Schema
NameRequiredDescriptionDefault
tabNo
profileNoBrowser ID from sleeper_sessions; optional with one connected browser.
selectorYes

TDQS

C2.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden and supplies almost none of it. It does not say whether the call is idempotent when the box is already checked, whether it waits for the element, or what errors occur on a missing selector.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

At two words this is not concise so much as under-specified; there is no structure to speak of and no front-loaded context beyond the bare action. It earns its place only as a label, not as usable documentation.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a browser-automation mutation tool with three parameters, no annotations, and no output schema, 'Check checkbox' omits everything an agent would need: selector format, tab/frame context, and the effect on page state. It is wholly inadequate for the tool's complexity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 33% (only 'profile' is documented, referencing sleeper_sessions), and the description adds nothing about any parameter. The 'selector' and 'tab' parameters are left entirely undocumented in both schema and description, so the description fails to compensate for the coverage gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a verb ('Check') and a resource ('checkbox'), so the basic action is inferable, but it is only two words and provides no detail on scope (which checkbox, in which frame/tab). It implicitly distinguishes itself from the sibling sleeper_uncheck, but does not do so explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this versus sleeper_click, sleeper_click_text, or sleeper_uncheck. The usage is only weakly implied by the word 'checkbox', with no conditions, prerequisites, or alternatives named.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

sleeper_clickC

Click element by CSS selector (supports :has-text)

ParametersJSON Schema
NameRequiredDescriptionDefault
tabNo
profileNoBrowser ID from sleeper_sessions; optional with one connected browser.
selectorYes

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It does not disclose what happens if the selector matches nothing, whether it waits/auto-retries, whether the click scrolls the element into view, or any error/permission behavior for a mutation-style action.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with zero filler, correctly leading with the action and target. It is efficiently sized, though minimally informative given the room available.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema, no annotations, and one of three parameters undocumented. For a browser mutation action in a large sibling set, the description does not supply enough context about behavior, targeting alternatives, or results to be complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 33% (only 'profile' is documented). The description does add meaningful value by noting ':has-text' selector support, but the 'tab' parameter is undocumented in both schema and description, so the gap is not fully compensated.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (click) and resource (element by CSS selector), and the ':has-text' note signals selector-based targeting, distinguishing it from text-based siblings like sleeper_click_text. However, it does not explicitly name or contrast with those siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use or when-not-to-use guidance is given. With close siblings such as sleeper_click_text, sleeper_click_all, and sleeper_dblclick available, the description leaves the agent to infer that this tool is the selector-based variant without stating it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

sleeper_click_allC

Click every matching element

ParametersJSON Schema
NameRequiredDescriptionDefault
tabNo
profileNoBrowser ID from sleeper_sessions; optional with one connected browser.
selectorYes

TDQS

C2.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure, and it discloses almost nothing. It does not say whether clicks are sequential, whether navigation/awaits occur between them, whether the tool fails or no-ops on zero matches, or that this is a mutating action on a live page.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The single sentence is front-loaded and free of filler, which is good, but at this length it reads as under-specification rather than genuine conciseness for a three-parameter mutating tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no annotations, no output schema, 33% parameter coverage, and a dense sibling cluster, the definition leaves too much for the agent to infer. A bulk-click tool needs at least match-count behavior, failure semantics, and its relationship to sleeper_click.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 33% — only 'profile' is documented, while 'selector' and 'tab' have no schema descriptions. The description adds nothing about selector syntax, matching semantics, or the meaning of 'tab', so it fails to compensate for the coverage gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('click') plus a scope qualifier ('every matching element'), which does distinguish it from the single-element sibling sleeper_click. It stops short of naming the alternative or the browser-automation context explicitly, so the agent must infer the rest from the sibling list.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance at all. With siblings sleeper_click, sleeper_click_text, and sleeper_dblclick present, the description never states when bulk-clicking is preferable to a single click or what happens if the selector matches zero elements.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

sleeper_click_textC

Click element by text

ParametersJSON Schema
NameRequiredDescriptionDefault
tabNo
textYes
exactNo
indexNo
verifyNo
profileNoBrowser ID from sleeper_sessions; optional with one connected browser.

TDQS

C2.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden, and it discloses almost nothing: no mention of whether the element is scrolled into view, whether the click auto-waits, whether ambiguous text raises an error or clicks the first match, or what a verify failure looks like. A mutation-style action with zero annotation coverage deserves far more disclosure than one terse line.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The single three-word sentence is not wasting words, but it is under-specification rather than conciseness. Front-loading is irrelevant when the entire payload is three words with no supporting detail.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 6-parameter interaction tool with no annotations and no output schema, the description is wholly inadequate: no matching semantics, no failure behavior, no parameter meaning. An agent cannot call this correctly without opening the schema and guessing at runtime behavior.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 17% (just the profile param), and 6 parameters exist, so the description is expected to compensate but does not. The words 'by text' only loosely gesture at the text parameter and say nothing about tab, exact, index, or verify, leaving five parameters undocumented in both places.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The phrase states a verb (Click) and a resource (element) with a locating mechanism (by text), which is enough to tell it apart from a generic selector-based click. However, it never names the sibling it contrasts with (sleeper_click, sleeper_click_all), so an agent must infer the distinction. Purpose is clear but minimally articulated.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus sleeper_click, sleeper_click_all, sleeper_find_text, or sleeper_focus. No prerequisites, no note about what to do when multiple elements match or when the target is off-screen. Usage must be entirely inferred from the name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

sleeper_consoleD

Page console capture

ParametersJSON Schema
NameRequiredDescriptionDefault
tabNo
limitNo
profileNoBrowser ID from sleeper_sessions; optional with one connected browser.

TDQS

D1.8/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure, yet it says nothing about whether the tool is read-only or destructive, whether it clears the console, how much data is returned, or any auth/rate-limit requirements. This is a complete omission for a browser-console tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely short, but it is under-specified rather than concise. There is no structure or front-loaded information to help an agent understand what the tool does before invoking it.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no output schema, no annotations, and three parameters with mostly missing descriptions, the definition is incomplete for correct invocation. The description should explain scope, return behavior, and parameter roles, but provides none of that.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 33% (only the profile parameter is described). The description adds no meaning for the tab or limit parameters and does not compensate for the low coverage, leaving parameter usage largely undocumented.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Page console capture' names a resource and a vague action, but it does not specify what exactly is captured (console logs, messages, errors) or distinguish the tool from siblings like sleeper_snapshot, sleeper_read, or sleeper_network. It is a vague purpose rather than a clear verb+resource statement.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no when-to-use guidance, no alternatives, and no exclusions. With many sibling tools that also inspect page state, the agent is left to infer when sleeper_console is appropriate.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

sleeper_dblclickC

Double-click element

ParametersJSON Schema
NameRequiredDescriptionDefault
tabNo
profileNoBrowser ID from sleeper_sessions; optional with one connected browser.
selectorYes

TDQS

C2/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden and delivers nothing: no word on whether the selector is auto-waited/scrolled into view, whether the double-click dispatches full event sequences, what errors occur on a missing element, or whether the tab/profile context matters. This is effectively zero behavioral disclosure for an interactive tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two words is not conciseness but under-specification; the terse phrasing leaves every sentence-sized gap an agent would need filled. What is present is front-loaded, but there is essentially nothing to front-load.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 3-parameter interactive browser tool with no annotations and no output schema, the description omits essentially all context: selector semantics, tab/profile scoping, wait behavior, and failure modes. It is not sufficient for reliable invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 33% — only 'profile' is documented in the schema, while 'selector' and 'tab' are bare strings. The description adds no parameter meaning at all (no selector syntax guidance, no explanation of tab vs. profile scoping), so it fails to compensate for the coverage gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb (double-click) and its target (element), so the basic purpose is clear. However, it does nothing to distinguish this tool from the adjacent sleeper_click or sleeper_click_text siblings, leaving the agent to guess when a double-click is wanted over a single click.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no mention of the single-click alternative, and no prerequisites or ordering constraints. The agent gets no help deciding between this and the many other click-family tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

sleeper_dialogC

Install auto accept/dismiss dialog hook

ParametersJSON Schema
NameRequiredDescriptionDefault
profileNoBrowser ID from sleeper_sessions; optional with one connected browser.

TDQS

C2.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It discloses that a hook is installed and dialogs are auto accepted/dismissed, but omits whether the hook persists, how it interacts with browser/profile state, whether it can override user intent, and whether it requires specific permissions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single short phrase with no filler and is front-loaded. It is efficient, though the slash construction 'accept/dismiss' is slightly ambiguous and not expanded for clarity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with no annotations and no output schema, the description is too thin. It does not explain persistence, scope, interaction with the profile, side effects of auto-accepting dialogs, or how to reverse the hook, leaving important context missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There is one optional parameter, and schema description coverage is 100%, so the schema already explains the profile parameter fully. The description adds no additional parameter meaning or constraints, matching the baseline of 3 for high schema coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states an action and object: 'Install auto accept/dismiss dialog hook.' However, 'hook' and 'auto accept/dismiss' are ambiguous, and it does not distinguish this tool from the related sleeper_wait_dialog sibling. An agent can infer it installs automatic dialog handling, but the exact effect is not fully clear.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives no when-to-use guidance, prerequisites, or alternatives. It does not tell the agent whether to use this instead of sleeper_wait_dialog or when a dialog hook is appropriate. Only a vague implied context is present.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

sleeper_dragC

Drag source to target

ParametersJSON Schema
NameRequiredDescriptionDefault
tabNo
sourceYes
targetYes
profileNoBrowser ID from sleeper_sessions; optional with one connected browser.

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations present, the description carries the full behavioral burden but discloses almost nothing: no indication of what kind of drag is performed (HTML5 drag-and-drop vs. mouse move), whether intermediate hover steps occur, or what happens on failure. Only the bare mutation intent is conveyed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The four-word phrasing is front-loaded and free of filler, but its extreme brevity on a 4-parameter browser-automation tool reads as under-specification rather than disciplined conciseness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a parameterized browser-automation mutation with no annotations, no output schema, and under-documented params, the description is not complete enough for an agent to invoke it confidently. It omits selector semantics, tab/profile targeting context, and any note on success behavior.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 25% (just 'profile' is documented), so the description must compensate for 'source', 'target', and 'tab'. It hints at the source→target relationship but gives no format or resolution rules for the selector-like values, leaving the key required params under-specified in both places.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a concrete verb ('drag') and a two-endpoint resource ('source' to 'target'), so the core action is identifiable. However, it never clarifies what source/target denote (CSS selectors, text, element refs) and does not differentiate itself from click/hover/dblclick siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use drag versus alternatives such as sleeper_dblclick, sleeper_click, or sleeper_hover, and no preconditions (element visibility, awaiting load) are stated. Usage can only be inferred from the tool name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

sleeper_execC

Execute arbitrary JS in Chromium; unavailable in signed Firefox builds

ParametersJSON Schema
NameRequiredDescriptionDefault
tabNo
codeYes
profileNoBrowser ID from sleeper_sessions; optional with one connected browser.

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden. It usefully discloses the Firefox limitation, but for an arbitrary-code-execution tool it says nothing about execution context (page vs. isolated world), side effects on page state, return value, or error/timeout behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler. It is efficient, though the extreme brevity borders on under-specification rather than disciplined conciseness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No annotations, no output schema, and two of three parameters undocumented. For a tool that executes arbitrary code, the description omits return format, execution semantics, and failure modes that an agent needs before invoking it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 33% — 'tab' and the required 'code' have no descriptions. The description adds no parameter meaning at all, so it fails to compensate for the coverage gap, leaving the agent to guess what 'tab' accepts.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a specific verb and resource ("Execute arbitrary JS") plus the execution environment (Chromium), which no sibling tool duplicates. It stops short of explicitly contrasting with the closest relative, sleeper_api, so the agent must infer which to reach for.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The only conditional guidance is a platform limitation ("unavailable in signed Firefox builds"). There is no statement of when to prefer this over sleeper_api or sleeper_batch, no prerequisites, and no exclusion cases.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

sleeper_extractD

Extract selector->text map

ParametersJSON Schema
NameRequiredDescriptionDefault
mapYes
tabNo
profileNoBrowser ID from sleeper_sessions; optional with one connected browser.

TDQS

D1.7/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are supplied, so the description carries the full burden of behavioral disclosure, and it says nothing about side effects, required session/auth state, or failure modes. A terse 'extract' fragment leaves the agent blind to whether this mutates state, requires an active tab, or blocks.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is a single short fragment; brevity here reflects under-specification rather than efficient front-loading. There is no wasted sentence, but there is also not enough substance to constitute a usable definition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with a required nested-object parameter, no annotations, no output schema, and low schema coverage, the description leaves everything an agent needs to invoke it correctly unstated. It is not complete at any useful level.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Three parameters with only 33% schema coverage: only 'profile' is documented, while the required nested object 'map' and 'tab' are bare. The phrase 'selector->text map' hints at the 'map' object's purpose but is ambiguous about whether it is an input specification or the return shape, so it does not compensate for the coverage gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The fragment 'Extract selector->text map' gestures at a verb (extract) and a result (selector-to-text mapping) but is grammatically incomplete and doesn't state what is being extracted, from where, or how it differs from siblings like sleeper_get, sleeper_read, or sleeper_find_text. An agent cannot reliably distinguish this from related extraction/read tools on the text alone.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no mention of prerequisites (e.g., needing a page loaded or a browser session), and no reference to any alternative tool among the ~50 siblings. The description provides no routing signal at all.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

sleeper_fill_formC

Fill form by name/id/placeholder (JSON map)

ParametersJSON Schema
NameRequiredDescriptionDefault
tabNo
fieldsYes
profileNoBrowser ID from sleeper_sessions; optional with one connected browser.

TDQS

C2.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries full behavioral burden. It does not disclose side effects (e.g., whether filling triggers events or submits), error handling, or any post-conditions. Only the targeting mechanism is mentioned, leaving most behavioral traits opaque.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded phrase with no wasted words. Yet it is arguably too terse for a tool with a nested object parameter, risking under-specification that could be mistaken for conciseness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the absence of annotations, output schema, and full parameter documentation, the description is insufficient. It does not explain return values, side effects, or how to handle partial failures, leaving significant gaps for safe invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 33% (only 'profile' documented). The description adds meaning for the required 'fields' parameter by indicating it is a JSON map keyed by name/id/placeholder, which partially compensates. However, 'tab' remains undocumented and no further syntax or value format is provided.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Fill') and resource ('form') and specifies how fields are targeted ('by name/id/placeholder'), making the purpose clear. However, it does not distinguish itself from siblings like sleeper_type or sleeper_submit, so an agent must infer that this tool fills multiple fields at once rather than typing into a single element.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance, prerequisites, or alternatives are provided. The description does not mention when to prefer this over sleeper_type or how it relates to sleeper_submit. Usage is only implied by the verb 'Fill'.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

sleeper_findC

Inspect elements by CSS selector

ParametersJSON Schema
NameRequiredDescriptionDefault
tabNo
profileNoBrowser ID from sleeper_sessions; optional with one connected browser.
selectorYes

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure, and it discloses almost nothing. "Inspect" weakly implies a read, but there is no statement about what is returned, behavior on zero matches, whether elements must be visible/attached, or pagination/limits for large result sets.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single short sentence with no filler, which is structurally clean, but the brevity is under-specification rather than earned conciseness given the tool's 3 parameters and dense sibling set. Nothing is front-loaded beyond the bare purpose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 3-parameter tool with no annotations and no output schema, the description should compensate by describing the return shape or result semantics, and it does not. It also omits how this differs from the many other selector/text-reading siblings, so an agent lacks the context needed to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 33%: profile is documented in the schema, tab is bare, and selector is bare. The description's phrase "by CSS selector" only restates the selector parameter's obvious role, adding no syntax guidance (e.g., match-all vs first match, shadow DOM, frames) and nothing at all for tab.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

"Inspect elements by CSS selector" gives a verb (inspect) and a resource (elements), but "inspect" is vague about the actual outcome — does it return matches, counts, or handles? It also never distinguishes itself from close siblings like sleeper_find_text, sleeper_get, or sleeper_extract, leaving the agent to guess which selector-based reader to use.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use, when-not-to-use, or alternative routing guidance. With siblings such as sleeper_find_text (text matching), sleeper_get, and sleeper_extract, the absence of any disambiguation is a real gap — the agent must infer that CSS selectors are the differentiator.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

sleeper_find_textC

Find element by innerText

ParametersJSON Schema
NameRequiredDescriptionDefault
tabNo
textYes
exactNo
profileNoBrowser ID from sleeper_sessions; optional with one connected browser.

TDQS

C2.4/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden of behavioral disclosure, and it discloses essentially nothing. It does not say what is returned, what happens when no match is found, whether the tab is scrolled into view, or how case/matching behaves.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The single sentence is tight and front-loaded with the verb and mechanism, but its brevity reflects under-specification rather than deliberate economy for a four-parameter tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with four parameters, no annotations, no output schema, and 25% schema coverage, the description is far too thin. An agent lacks the information needed to predict return values, failure behavior, or matching semantics.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 25% (just 'profile'), and the description adds nothing about 'tab', 'text', or 'exact'. Notably the boolean 'exact' semantics (exact vs substring, case sensitivity) are undocumented anywhere, so the description fails to compensate for the coverage gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a clear verb (find) and resource (element), plus the locating mechanism (innerText), which helps separate it from selector-based sleeper_find. However, it never explicitly contrasts itself with siblings like sleeper_find or sleeper_click_text, so it stops short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this versus the many sibling find/click/wait tools (sleeper_find, sleeper_click_text, sleeper_wait_text). No prerequisites or conditions are stated, leaving routing entirely to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

sleeper_focusD

Focus element

ParametersJSON Schema
NameRequiredDescriptionDefault
tabNo
profileNoBrowser ID from sleeper_sessions; optional with one connected browser.
selectorYes

TDQS

D1.5/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the entire burden of behavioral disclosure, and it discloses nothing. It does not say whether focusing scrolls the element into view, whether it mutates page state, whether it requires the element to already exist or be visible, or what the call returns.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two words is terse but not concise in the useful sense: the brevity comes from under-specification rather than efficiency. There is no front-loaded explanation because there is no explanation.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no annotations, no output schema, three parameters at 33% coverage, and a required selector that is undocumented, the description is fully inadequate for an agent to call this tool correctly or predict its effect.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 33% (only profile is documented; tab and the required selector have none). The description adds no meaning for any of the three parameters, leaving the required selector and tab semantically undefined in both places.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

"Focus element" is essentially a two-word restatement of the tool name sleeper_focus with its object tacked on. It gives no scope (single element? page focus? tab focus?) and never distinguishes itself from the many sibling interaction tools (click, hover, dblclick, check). An agent cannot tell from this what focusing actually accomplishes.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use, when-not-to-use, or alternative guidance of any kind. Nothing tells the agent whether to reach for sleeper_focus instead of sleeper_click or sleeper_hover, which are adjacent sibling actions on elements.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

sleeper_formsC

Enumerate form inputs

ParametersJSON Schema
NameRequiredDescriptionDefault
tabNo
profileNoBrowser ID from sleeper_sessions; optional with one connected browser.
selectorNo

TDQS

C2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. 'Enumerate' weakly implies a read-only, non-destructive operation, but nothing is said about return format, whether it waits for the page, whether it covers shadow DOM/iframes, or what authorization/state is assumed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three words contain no filler, but this is under-specification rather than conciseness; the description is too small to front-load any useful information for a three-parameter browser automation tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no annotations, no output schema, low parameter coverage, and a dense sibling set, the definition leaves virtually everything an agent needs unanswered. It is not adequate for invoking this tool correctly on the first try.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 33% (two of three parameters undocumented), and the description adds no parameter information at all. It does not explain what 'tab' or 'selector' mean or how they scope the enumeration, so it fails to compensate for the low coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('enumerate form inputs'), so the basic operation is graspable. However, it gives no scope (current page? frame? whole session?) and does not differentiate from siblings like sleeper_find, sleeper_extract, sleeper_snapshot, or sleeper_fill_form, which an agent could easily confuse with this one.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no preconditions, and no mention of alternative tools. With 50+ siblings including several read/enumerate-style tools, the absence of any routing signal is a real gap.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

sleeper_framesC

List cross-origin iframes

ParametersJSON Schema
NameRequiredDescriptionDefault
profileNoBrowser ID from sleeper_sessions; optional with one connected browser.

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It does not state that this is a read-only operation, nor does it clarify scope (e.g., current page only, all tabs, or all frames including same-origin). For a browser automation tool, this leaves key behavioral traits unspecified.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely concise and front-loaded, consisting of only four words that directly convey the tool's action and target. Every word earns its place with no wasted phrasing.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a browser automation tool with no annotations and no output schema, the description is too sparse. It omits essential context such as whether the list is scoped to the current page or all pages, whether it requires an active tab, and what information is returned. The schema covers the single parameter, but the description does not compensate for the lack of behavioral detail.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the 'profile' parameter is already fully documented in the schema, including its optional nature when one browser is connected. The description adds no additional meaning or context about the parameter, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('List') and a specific resource ('cross-origin iframes'), making the tool's purpose immediately clear. It does not explicitly differentiate itself from sibling tools like sleeper_tabs or sleeper_find, but the resource is distinctive enough that an agent can identify its function.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives no guidance on when to use this tool versus alternatives. There is no mention of prerequisites (e.g., needing an active page), no indication of whether it should be called before interacting with iframes, and no exclusions. It merely states what the tool does.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

sleeper_getC

Get page/element properties

ParametersJSON Schema
NameRequiredDescriptionDefault
tabNo
profileNoBrowser ID from sleeper_sessions; optional with one connected browser.
selectorNo

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden, and it discloses almost nothing: no statement that it is read-only, no return format, no error behavior for invalid selectors, and no note on what happens if selector is omitted. The phrase "page/element" weakly implies the selector toggles the target, but this is minimal.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single short phrase and is front-loaded, with no filler or repetition. However, it is under-specified rather than genuinely concise — the brevity reflects missing content, not efficient communication.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with 3 parameters, zero annotations, no output schema, and a large sibling family, the definition is inadequate: it omits usage routing, parameter format, and return expectations. Only the schema's partial 'profile' description provides any support, which is not enough for an agent to invoke this correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 33% — only 'profile' is documented. The description's "page/element" phrasing adds partial meaning for 'selector' (present = element, absent = page), but 'tab' is completely undocumented in both places and no format, default, or matching semantics are given. With low coverage the description needed to compensate and largely does not.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

"Get page/element properties" states a verb (Get) and a resource (page/element properties), so the basic action is inferable. However, it does not distinguish this tool from closely related siblings like sleeper_state, sleeper_read, sleeper_extract, or sleeper_snapshot, all of which could plausibly return property-like data. The word "Get" is generic enough that an agent cannot confidently route to this tool over its neighbors.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no mention of alternatives, and no conditions for choosing this over sleeper_state/sleeper_read/sleeper_extract. The only implicit hint is that a selector yields element properties rather than page properties, but this is left for the agent to infer.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

sleeper_gotoC

Navigate to URL

ParametersJSON Schema
NameRequiredDescriptionDefault
tabNo
urlYes
profileNoBrowser ID from sleeper_sessions; optional with one connected browser.

TDQS

C2.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden, yet it discloses nothing about behavior: whether it waits for page load, what happens on a failed navigation, whether it navigates in the current tab or a new one, or how it interacts with the session/profile state.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The four-word description is certainly concise and front-loaded, but this is under-specification rather than effective brevity — there is no wasted sentence because there is essentially only one.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a state-mutating browser-navigation tool with no annotations, no output schema, and two of three params undocumented, the description is far too thin. An agent cannot tell whether navigation waits, which tab context is affected, or what a failure looks like.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 33% (just the optional 'profile' param), and the description adds no meaning for 'url' or 'tab'. With low coverage the description should compensate by explaining what 'tab' selects and how navigation interacts with profiles, but it says nothing.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Navigate') and resource ('URL'), which is enough to separate it from siblings like sleeper_back, sleeper_newtab or sleeper_click. It does not, however, contrast itself with the similarly-navigation-oriented siblings, so it stops short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no prerequisites, and no mention of alternatives such as sleeper_newtab for opening a fresh tab. The agent must infer usage entirely from the tool name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

sleeper_hoverD

Mouseover element

ParametersJSON Schema
NameRequiredDescriptionDefault
tabNo
profileNoBrowser ID from sleeper_sessions; optional with one connected browser.
selectorYes

TDQS

D1.7/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden and discloses nothing: not whether the hover persists, whether it waits for the element, whether it scrolls into view, whether it triggers dependent UI (menus/tooltips), or whether it fails on invisible elements. A single two-word fragment leaves a mutation-ish interaction completely opaque.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two words with no waste, but this is under-specification rather than conciseness. There is no front-loaded action context, no scope statement, and nothing an agent can act on beyond the name.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No annotations, no output schema, 33% schema coverage, and a dense sibling ecosystem of ~50 browser tools. For an interaction tool in that context, the description is far too thin to let an agent pick it correctly over sleeper_click, sleeper_focus, or sleeper_drag.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 33% (only 'profile' is documented, and that documentation lives in the schema, not the description). The description adds no meaning for 'selector' or 'tab', so two of three parameters remain undocumented in both places.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

"Mouseover element" restates the tool name rather than describing the operation in a distinguishing way. It does convey hover semantics, but there is no differentiation from siblings like sleeper_click, sleeper_focus, or sleeper_drag, and 'element' is left undefined (selector? text? coordinates?).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance, no mention of alternatives such as sleeper_click or sleeper_focus, and no stated prerequisites (e.g., element must exist or be scrolled into view). The agent must infer usage entirely from the name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

sleeper_keysC

Dispatch key events on active element

ParametersJSON Schema
NameRequiredDescriptionDefault
tabNo
keysYes
profileNoBrowser ID from sleeper_sessions; optional with one connected browser.

TDQS

C2.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. 'Dispatch key events on active element' hints that focus matters, but it does not say whether keys are held, whether modifier combos are supported, what happens if no element is focused, or what the response looks like. One clause is far too thin for a no-annotation tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single sentence with no padding and the action front-loaded, which is good. But the brevity crosses into under-specification rather than economy — nothing here earns extra credit for tightness when critical detail is absent.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Three-parameter input tool, no annotations, no output schema, and only 33% schema coverage. The description should fill the gaps on key syntax, tab/profile targeting, and focus behavior; it supplies none of that, leaving the definition materially incomplete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 33% — only 'profile' is documented in the schema. The description says nothing about 'keys' (array-of-strings format, key names vs. combinations, e.g. 'Control+C') nor about 'tab' (tab targeting). With two of three parameters undocumented anywhere, the description fails to compensate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and target: 'Dispatch key events on active element.' That is clearly a keyboard-input operation on the current focus. It does not, however, differentiate itself from the sibling sleeper_press, which appears to be a very similar keyboard tool — an agent cannot tell which to pick from the description alone.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives no condition for when to use this versus sleeper_press, sleeper_type, or sleeper_exec, all of which can emit keystrokes. No prerequisites, no mention of focus state requirements, no exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

sleeper_mediaC

Per-tab media request log

ParametersJSON Schema
NameRequiredDescriptionDefault
tabNo
clearNo
sinceNo
profileNoBrowser ID from sleeper_sessions; optional with one connected browser.

TDQS

C2.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure and largely fails it. It never states whether the log is captured by default, whether it is bounded or overwritten, what 'clear' destroys, or whether the log is scoped to the current tab only; 'per-tab' is the sole behavioral hint.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is a four-word noun fragment with no waste, but it is under-specified rather than concise: there is no sentence structure, no front-loaded verb, and no information for an agent to act on. Brevity here costs comprehension.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a four-parameter tool with no annotations, no output schema, and only 25% schema description coverage, the description supplies essentially nothing — no return shape, no default, no side effects. An agent cannot invoke it confidently.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 25% — just the optional 'profile' parameter is documented. The description says nothing about 'tab', 'clear', or 'since', leaving half the tool's behavior (scoping, destructive clearing, time filtering) undocumented in both the schema and the description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The fragment 'Per-tab media request log' names the resource (media request log) and its scope (per-tab), which distinguishes it from sleeper_network and sleeper_console. However, it contains no verb, so an agent cannot tell whether the tool lists, streams, or clears the log — only the presence of the 'clear' parameter hints at mutation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus sleeper_network, sleeper_console, or sleeper_state, all of which are plausible log-inspection siblings. No prerequisites, no exclusions, no conditions are stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

sleeper_networkD

Per-tab request log or fetch

ParametersJSON Schema
NameRequiredDescriptionDefault
tabNo
urlNo
bodyNo
clearNo
fetchNo
mediaNo
sinceNo
profileNoBrowser ID from sleeper_sessions; optional with one connected browser.
segmentNo

TDQS

D1.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden, yet it discloses nothing about read vs. write semantics, whether the log accumulates across calls, or what the 'clear' and 'fetch' flags do. With a 'clear' parameter present, the absence of any mutation/reversibility note is a real gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Five words is not conciseness here but under-specification: for a nine-parameter tool the description is too sparse to be front-loaded around anything useful. There is no wasted filler, but there is also almost no content.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a nine-parameter network-inspection tool with no annotations and no output schema, the description omits everything an agent would need: what is returned, how 'since'/'media'/'segment' filter results, and when a fetch vs. log read is intended. It is grossly incomplete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 11% (just 'profile'), so the description must compensate for nine largely undocumented parameters. It mentions no meaning for tab, url, body, clear, fetch, media, since, or segment, leaving most of the schema opaque.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The phrase 'Per-tab request log or fetch' gestures at the resource (network requests) but never states a clear verb or action. The dangling 'or fetch' leaves it ambiguous whether the tool logs, retrieves, or fetches a URL, and nothing distinguishes it from siblings like sleeper_media, sleeper_wait_xhr, or sleeper_api.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no mention of prerequisites, and no routing to alternatives such as sleeper_wait_xhr for waiting on a specific request. The agent is left to guess the intended scenario from the name alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

sleeper_newtabC

Open new tab via extension API

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYes
profileNoBrowser ID from sleeper_sessions; optional with one connected browser.

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It discloses only 'via extension API', which is implementation trivia, and says nothing about tab focus/activation, whether the tab must be cleaned up later, or what is returned. For a state-creating tool this is a significant gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single short, front-loaded sentence with no filler. Brevity is appropriate, though the terseness contributes to the missing guidance rather than serving it.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no annotations, no output schema, and one of two parameters undocumented, the description leaves key questions unanswered: does the new tab become active, does it return a tab handle, and how does it relate to sleeper_goto/sleeper_tabs. Inadequate for a tool that mutates browser state.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 50%: url is completely undocumented in both schema and description, and the description adds no format hints (absolute URL? about: pages allowed?). The profile parameter's meaning comes entirely from the schema's reference to sleeper_sessions, not from the description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a clear verb+resource: opening a new tab. It is distinguishable from most siblings in intent, but it never differentiates itself from sleeper_goto (navigate an existing tab) or sleeper_tabs (list tabs), which are the closest alternatives an agent would weigh.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance at all. The agent is not told whether to use this versus sleeper_goto when it wants a page loaded, nor whether the new tab becomes the active target for subsequent calls.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

sleeper_pressC

Keyboard events on element/active

ParametersJSON Schema
NameRequiredDescriptionDefault
keyYes
tabNo
profileNoBrowser ID from sleeper_sessions; optional with one connected browser.
selectorNo

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden, yet it discloses almost nothing: it doesn't say whether it dispatches keydown/keyup, whether it targets a selector or the active element by default, or whether focus/permission is required. The only context is the terse 'on element/active' hint.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

At six words it is certainly terse with no wasted filler, but the size reflects under-specification rather than disciplined conciseness. A slightly longer front-loaded sentence would serve the agent better.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With four parameters, no output schema, and no annotations, the description leaves too much unsaid for an agent to invoke it confidently. Key syntax and targeting behavior are the minimum an agent needs and both are missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 25% (only 'profile' has a description), so the description must compensate and does not. It never explains the 'key' string format (e.g. 'Enter' vs 'Control+A'), what 'selector' targets, or the role of 'tab'.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The phrase 'Keyboard events on element/active' vaguely indicates a key-pressing tool, but 'events' is imprecise and it never states it sends a single key press to a targeted element. It does not distinguish itself from close siblings like sleeper_keys (key combos) or sleeper_type (text entry).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this versus sleeper_keys, sleeper_type, or sleeper_submit. The fragment 'on element/active' hints at targeting, but no conditions, prerequisites, or alternatives are stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

sleeper_readC

Read single element (text/html/value/attr)

ParametersJSON Schema
NameRequiredDescriptionDefault
tabNo
attrNo
whatNo
profileNoBrowser ID from sleeper_sessions; optional with one connected browser.
selectorYes

TDQS

C2.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations and no output schema, the description carries the full behavioral burden, yet it says nothing about return format, error behavior when the selector matches nothing or multiple nodes, or how the mode changes the result. It only restates the operation the name already conveys.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The single parenthetical phrase is front-loaded and wastes no words, but its extreme brevity reflects under-specification rather than disciplined conciseness given the tool's five parameters.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a five-parameter tool with no annotations and no output schema, the description omits far too much: the output structure, the attr/what dependency, tab scoping, and selector behavior. An agent cannot call this confidently across all modes without opening the schema and guessing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 20% (just the profile param). The description lists the values for 'what' (text/html/value/attr), but those already appear as the schema enum, and it says nothing about tab, attr (which is clearly required when what=attr, an unstated dependency), or selector matching semantics.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource (read a single element) and enumerates the four extraction modes (text/html/value/attr). The word 'single' implicitly distinguishes it from the sibling sleeper_read_all, but it never names that alternative, so the differentiation is only implied rather than explicit.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance and no mention of alternatives. The word 'single' hints that sleeper_read_all handles bulk reads, and the several other read-oriented siblings (sleeper_get, sleeper_find, sleeper_extract) are unaddressed, leaving the agent to infer routing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

sleeper_read_allC

Read all matching elements

ParametersJSON Schema
NameRequiredDescriptionDefault
tabNo
attrNo
whatNo
profileNoBrowser ID from sleeper_sessions; optional with one connected browser.
selectorYes

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are supplied, so the description carries the full burden, yet it adds nothing beyond "read." It is silent on return format, whether attr/what alter the output, whether profile is required, and any limits on the number of matched elements returned.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is a single front-loaded phrase with no wasted words, but it is terse to the point of under-specification rather than genuinely concise. Brevity here costs clarity instead of buying it.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 5-parameter tool with no annotations, no output schema, and 20% schema coverage, the description is far too thin. An agent would not know what a call returns or how attr/what/selector interact.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 20% (just profile is documented), leaving tab, attr, what, and selector unexplained. The description's phrase "matching elements" loosely gestures at selector but adds no meaning for the other three parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ("read") and resource ("matching elements") with a plural scope, so the basic intent is inferable. However, it does not distinguish itself from close siblings like sleeper_read, sleeper_get, sleeper_find, or sleeper_extract, leaving the agent to guess how it differs.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no indication of when to use this tool versus sleeper_read (singular) or sleeper_extract. The word "all" hints at multi-element retrieval, but no when-to-use or exclusion guidance is provided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

sleeper_screenshotC

Alias for shot — viewport PNG screenshot

ParametersJSON Schema
NameRequiredDescriptionDefault
tabNo
profileNoBrowser ID from sleeper_sessions; optional with one connected browser.

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It does not disclose what happens to the captured image (returned inline, written to disk, path returned), whether the capture is read-only and side-effect free, or any failure modes on a missing/backgrounded tab.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single short clause, front-loaded with the alias relationship and the output type. Nothing is wasted, though the brevity is partly under-specification rather than true economy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

A tool with two parameters, no annotations, and no output schema needs the description to explain the return artifact and where it goes; it does not. An agent cannot tell whether it receives image bytes, a file path, or a URL.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 50%: the 'profile' parameter is documented in the schema, but the 'tab' parameter is undocumented in both schema and description. The description adds no parameter meaning at all, so it fails to compensate for the coverage gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a concrete output (viewport PNG screenshot) and frames the tool as an alias for the sibling sleeper_shot. That is enough to distinguish it from sleeper_snapshot (likely full-page) without opening either schema, though the reliance on 'Alias for shot' means the agent must already know the sibling naming scheme.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use, when-not-to-use, or alternative guidance is given. Saying it is an alias for sleeper_shot implicitly raises the question of which to call, but the description never resolves it, nor does it mention any prerequisite such as a connected browser.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

sleeper_scrollC

Scroll by pixels or to selector

ParametersJSON Schema
NameRequiredDescriptionDefault
yNo
tabNo
profileNoBrowser ID from sleeper_sessions; optional with one connected browser.
selectorNo

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It does not say whether the scroll waits for completion, how sticky/fixed elements behave, whether it is instant or smooth, or what happens if the selector is missing. For a browser action with zero annotation coverage this is a significant gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single tight phrase with no filler, and the core distinction is front-loaded. It is appropriately sized for its content, though brevity here leans toward under-specification rather than crisp conciseness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With 4 parameters, no annotations, no output schema and no defaults disclosed, the description omits too much: activation context, wait/settle behavior, and the tab vs profile browser-targeting semantics. An agent cannot call this reliably without guessing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 25% (only 'profile' documented), so the description must compensate. It does add meaning for two of four params ('by pixels' -> y, 'to selector' -> selector), but leaves tab and the y/selector interaction (mutually exclusive?) undefined.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (scroll) and the two modes of operation (by pixels or to selector), so an agent knows exactly what it does. It does not differentiate from the sibling sleeper_scroll_until, which is the closest alternative and is not mentioned.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this versus sleeper_scroll_until, nor any prerequisites such as needing a page/tab context first (sleeper_goto/sleeper_tabs). The agent receives no signal about which scroll variant to pick.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

sleeper_scroll_untilC

Scroll until text/selector appears

ParametersJSON Schema
NameRequiredDescriptionDefault
tabNo
textNo
profileNoBrowser ID from sleeper_sessions; optional with one connected browser.
selectorNo
containerNo
max_itersNo

TDQS

C2.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It conveys only the polling condition and omits critical details: timeout behavior, what happens if the target never appears, default max_iters, scrolling container semantics, and any side effects or auth requirements.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The single short phrase is front-loaded and waste-free, but it is drastically undersized for a tool with six parameters and no annotations. Appropriate brevity would still require more actionable detail.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given six parameters, zero required fields, no output schema, and no annotations, the description is far too sparse to be complete. It omits usage context, parameter meanings, and behavioral expectations that an agent needs to call the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 17% (only the 'profile' parameter is documented). The description mentions no parameters, so the agent receives no guidance for 'tab', 'text', 'selector', 'container', or 'max_iters'. It fails to compensate for the low coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Scroll') and condition ('until text/selector appears'), making the core action clear. It does not differentiate itself from siblings like sleeper_scroll, sleeper_wait_until, or sleeper_wait_text, so it falls short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description offers no guidance on when to use this tool versus alternatives such as sleeper_wait_text or sleeper_scroll. No prerequisites, no context, and no exclusions are provided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

sleeper_selectC

Set select value + change event

ParametersJSON Schema
NameRequiredDescriptionDefault
tabNo
valueYes
profileNoBrowser ID from sleeper_sessions; optional with one connected browser.
selectorYes

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden. It does disclose one meaningful trait — that it triggers a change event, which matters for sites reacting to input events — but says nothing about whether it waits for the element, what happens on failure, or what it returns.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is a single compact phrase with no waste and the action is front-loaded, but the brevity tips into under-specification rather than disciplined conciseness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No annotations, no output schema, and 25% parameter coverage leave the description doing very little for a 4-parameter browser-automation tool. An agent has no way to confirm selector format, session requirements, or error behavior from this definition alone.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 25% (only 'profile' is documented, as the browser ID from sleeper_sessions). The description adds no meaning for selector, value, or tab, leaving three of four parameters undocumented in both places for a tool whose correctness hinges on how the selector is expressed.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

"Set select value + change event" names a verb (set) and a target (select value), implying it assigns a value to a <select> element and dispatches a change event. That is more specific than a tautology, but it never clarifies what a "select" is (HTML dropdown vs. a selection of elements) or how it differs from the many sibling interaction tools like sleeper_type, sleeper_click, or sleeper_fill_form.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this instead of alternatives such as sleeper_fill_form, sleeper_type, sleeper_check, or a generic click-and-pick flow. An agent gets no signal about prerequisites (e.g., a connected browser session) or when this tool is the wrong choice.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

sleeper_sessionsB

List connected stable browser installation IDs for routing

ParametersJSON Schema
NameRequiredDescriptionDefault
profileNoBrowser ID from sleeper_sessions; optional with one connected browser.

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. 'List connected stable browser installation IDs' honestly conveys this is a read-only enumeration of available browsers, but it omits return format, whether a profile filter narrows results, and any environment requirements.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler, and the key resource is placed early. It is efficient, though slightly terse enough that 'for routing' reads as an afterthought rather than useful context.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple one-parameter discovery tool with no output schema, the description names what is returned (installation IDs) but does not explain the result shape an agent needs to route to subsequent calls. Adequate but with a clear gap given the absence of an output schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the single 'profile' parameter is already documented in the schema, including the condition that it is optional with one connected browser. The description adds nothing beyond what the schema provides, so baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (List) and resource (connected stable browser installation IDs), which clearly distinguishes it from the action-oriented siblings like sleeper_goto and sleeper_click. The trailing 'for routing' is slightly vague but the core purpose is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit guidance on when to call this tool versus alternatives. 'For routing' hints that the returned IDs feed other calls, but it never states the trigger condition or that this is the discovery step before tools requiring a profile. An agent must infer usage from the param note alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

sleeper_shotC

Viewport PNG screenshot (data URL)

ParametersJSON Schema
NameRequiredDescriptionDefault
tabNo
profileNoBrowser ID from sleeper_sessions; optional with one connected browser.

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It discloses the return format (PNG as a data URL), which is genuinely useful, but says nothing about permissions, viewport sizing, whether the result is base64-inline, or any side effects, which is thin for a no-annotation definition.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single fragment with no waste, but also no substance — the brevity here reflects under-specification rather than disciplined conciseness. Nothing is front-loaded beyond the artifact name.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool in a 50-sibling browser-automation set with no annotations and no output schema, the definition is markedly incomplete: it omits tab semantics, capture scope versus sibling capture tools, and any indication of how the data URL should be handled.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Only 50% schema description coverage: the 'profile' parameter is documented in the schema, but 'tab' has no description anywhere. The tool description adds nothing about either parameter, so the tab selector and its relationship to sleeper_tabs is undocumented.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

States the action and artifact clearly enough: 'Viewport PNG screenshot (data URL)'. However, it gives no differentiation from close siblings like sleeper_screenshot and sleeper_snapshot, so an agent cannot tell which capture tool to pick.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance at all, and no mention of the alternatives (sleeper_screenshot, sleeper_snapshot) that appear in the sibling list. The viewport-vs-full-page distinction that would drive selection is left entirely implicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

sleeper_snapshotC

Headings/inputs/buttons/links summary

ParametersJSON Schema
NameRequiredDescriptionDefault
tabNotab index or URL substring
profileNoBrowser ID from sleeper_sessions; optional with one connected browser.

TDQS

C2.5/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden and discloses almost nothing. It does not state whether the call is a pure read (implied but unconfirmed), whether it covers only the top frame or all frames, whether hidden elements are included, or what the returned structure looks like.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

At five words it is maximally terse and front-loaded with the element categories, with zero filler. However, the brevity here reads as under-specification rather than efficient economy, since key scoping information is simply absent.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, so the description is the only place the return shape could be explained, and it only hints at four element categories. Combined with a crowded sibling set and no annotations, an agent lacks enough context to pick this tool confidently or predict its output.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and both parameters (tab, profile) are documented in the schema, so the baseline of 3 applies. The description adds nothing about how tab selection or the optional profile interacts with the snapshot behavior.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The fragment 'Headings/inputs/buttons/links summary' names the resource categories it covers, so an agent can guess it returns a page-level digest of interactive and structural elements. But it is a noun phrase with no verb and no scope statement, and it does not differentiate itself from close siblings like sleeper_read, sleeper_read_all, or sleeper_extract.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance at all: no condition selecting it over sleeper_read_all, sleeper_extract, or sleeper_state, and no note on when a snapshot is preferable to a targeted query. The agent must infer the use case entirely from the name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

sleeper_stateB

Get page state for the selected Sleeper profile (url, title, focus, visibility)

ParametersJSON Schema
NameRequiredDescriptionDefault
profileNoBrowser ID from sleeper_sessions; optional with one connected browser.

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. 'Get' implies a non-mutating read, and the parenthetical discloses the shape of the returned state, but there is no statement about permissions, whether it waits for the page to settle, or what happens if no profile is connected.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One front-loaded sentence with the returned fields in parentheses; every token earns its place and there is no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description must convey what comes back, and it does at a coarse level (url, title, focus, visibility). For a one-optional-parameter read tool with no annotations this is nearly sufficient, though it says nothing about error or disconnected-profile behavior.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the single optional 'profile' parameter is fully documented in the schema (Browser ID from sleeper_sessions). The description's mention of the 'selected Sleeper profile' loosely echoes that parameter but adds no syntax or default-behavior detail beyond the schema, so baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Get') and resource ('page state') and enumerates the returned fields (url, title, focus, visibility), which is more informative than the bare name. It does not, however, distinguish itself from close siblings like sleeper_snapshot, sleeper_tabs, or sleeper_read, so an agent still has to guess which state-reading tool applies.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives no when-to-use guidance and no exclusions. With siblings such as sleeper_snapshot, sleeper_tabs, sleeper_read, and sleeper_find all plausibly overlapping, the agent receives no routing signal for picking this one.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

sleeper_submitC

Submit form

ParametersJSON Schema
NameRequiredDescriptionDefault
tabNo
profileNoBrowser ID from sleeper_sessions; optional with one connected browser.
selectorNo

TDQS

C2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. 'Submit form' implies a mutating action but does not disclose side effects, navigation behavior, required browser context, or what happens on success or failure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is only two words and front-loaded, but it is under-specified rather than concise. It does not earn its brevity by providing sufficient information for invocation.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a three-parameter browser automation tool with no annotations and no output schema, the description is completely inadequate. It leaves the agent without enough context to know what a selector or tab refers to or how to invoke the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There are three parameters and only the profile parameter has schema documentation, leaving tab and selector undocumented. The description adds no meaning for any parameter, so it fails to compensate for the low schema description coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a bare verb-resource pair ('Submit form'), but it does not distinguish this tool from closely related siblings such as sleeper_fill_form, sleeper_forms, or sleeper_click. It is more specific than a pure tautology but still too vague to route an agent confidently.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no prerequisites, and no named alternatives. The description gives no context for choosing this over the many other form-related tools in the sibling list.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

sleeper_tabsB

List tabs in the selected Sleeper browser profile

ParametersJSON Schema
NameRequiredDescriptionDefault
profileNoBrowser ID from sleeper_sessions; optional with one connected browser.

TDQS

B3.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It does not state that this is a read-only operation, what each tab entry contains (URL, title, index), ordering, or whether it targets only the active profile. For a navigation-inventory tool with zero annotation coverage, this leaves real gaps.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single, front-loaded sentence with no wasted words. It is tight, though its brevity borders on under-specification rather than optimal structure.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema and no annotations, the description should say more about what is returned and the scope of 'selected profile'. It is adequate for a simple list tool but leaves the return shape and multi-browser behavior to inference.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and the schema itself explains the single optional 'profile' parameter ('Browser ID from sleeper_sessions; optional with one connected browser'). The description adds nothing beyond the schema, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Specific verb+resource: 'List tabs' in the selected browser profile. An agent immediately understands this returns the set of tabs. It does not explicitly distinguish itself from the related sleeper_newtab sibling, so it stops short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The action name implies when it is useful (enumerating open tabs), but there is no explicit guidance such as when to call this versus sleeper_newtab or sleeper_state, and no stated prerequisites. Usage is only implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

sleeper_typeC

Type into input (React-compatible events)

ParametersJSON Schema
NameRequiredDescriptionDefault
tabNo
textYes
clearNo
profileNoBrowser ID from sleeper_sessions; optional with one connected browser.
stealthNo
selectorYes

TDQS

C2.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. The 'React-compatible events' note is a genuinely useful behavioral trait (it explains why it differs from a naive value assignment), but nothing is said about clear/stealth behavior, permissions, or what happens to existing input.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded line with zero waste, but the extreme brevity borders on under-specification for a six-parameter tool rather than being tight but complete.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With six parameters, 17% schema coverage, no annotations, and no output schema, the definition is too thin. An agent lacks enough to distinguish this from sibling input tools or to use the optional parameters correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 17% (just the profile param), so the description must compensate. It only implies the selector and text parameters and says nothing about the meaning of clear, stealth, tab, or profile, leaving four of six parameters semantically undocumented.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a clear verb+resource (type into input) and adds a useful qualifier ('React-compatible events'). However it offers no differentiation from siblings that also handle text entry or key input, such as sleeper_fill_form, sleeper_keys, or sleeper_press.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus the many siblings that overlap (sleeper_fill_form for multi-field entry, sleeper_keys/sleeper_press for keyboard input, sleeper_click_text). No prerequisites or exclusions are stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

sleeper_uncheckC

Uncheck checkbox

ParametersJSON Schema
NameRequiredDescriptionDefault
tabNo
profileNoBrowser ID from sleeper_sessions; optional with one connected browser.
selectorYes

TDQS

C2/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden, and it discloses nothing: not whether it waits for actionability, what happens if the checkbox is already unchecked, error behavior for a missing selector, or whether it dispatches events. A mutation tool with zero behavioral disclosure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three words is under-specification rather than conciseness; there is no structure, no front-loaded scope, and no sentence that informs invocation beyond the name itself.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no annotations, no output schema, a required undocumented parameter, and no mention of sibling tools for checking vs unchecking, the definition is not complete enough to call this tool confidently in a 49-tool browser-automation set.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 33% (only 'profile' is documented). The description adds no meaning for 'selector' (required) or 'tab' — it does not say whether selector is CSS, XPath, text, or an element handle, nor that the field must be a checkbox.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a concrete verb+resource ("Uncheck checkbox"), which is enough to know the operation, but it does not distinguish itself from the sibling sleeper_check or explain how it differs from a click on a checkbox. Minimal but not tautological.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance whatsoever on when to use this instead of sleeper_check, sleeper_click, or sleeper_select. The agent must infer usage entirely from the name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

sleeper_uploadB

Upload explicit base64 file contents to an input; native paths are unsupported. Total JSON request limit: 1 MiB.

ParametersJSON Schema
NameRequiredDescriptionDefault
tabNo
filesYes
profileNoBrowser ID from sleeper_sessions; optional with one connected browser.
selectorYes

TDQS

B3.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the behavioral burden. It usefully discloses the base64-only constraint and the 1 MiB JSON request limit, which are non-obvious operational traits. However it omits what happens to existing input contents and whether change/input events fire after upload.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences, front-loaded with the core action and immediately followed by the two hard constraints. No wasted words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 4-parameter tool with no annotations and no output schema, the description covers the format and size limit but leaves selector semantics and the nested file structure unexplained. Adequate but with clear gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 25% (only 'profile' is documented). The description clarifies that content must be base64 but says nothing about what 'selector' targets, the per-file name/mime_type fields, or 'tab' semantics, leaving most parameters undocumented in both schema and description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: uploads explicit base64 file contents to an input. It is clearly distinct from read/click/type siblings, but it does not explicitly name how it differs from related tools like sleeper_fill_form.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives a format constraint (base64 only, native paths unsupported) but no when-to-use guidance, no prerequisites, and no comparison to alternative tools for populating inputs. The agent must infer usage context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

sleeper_waitD

Wait for selector

ParametersJSON Schema
NameRequiredDescriptionDefault
tabNo
profileNoBrowser ID from sleeper_sessions; optional with one connected browser.
timeoutNo
selectorYes

TDQS

D1.7/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden, and it discloses nothing. It does not say what state the selector must reach (attached, visible, enabled), what happens on timeout, whether an error is thrown or a falsy result returned, or how tab/profile context affects the wait.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three words is terse but this is under-specification rather than conciseness; there is no front-loaded explanation of behavior. Nothing wasteful is present only because almost nothing is present.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 4-parameter tool with 25% schema coverage, no annotations and no output schema, the description leaves almost everything an agent needs unanswered. There is no indication of the wait criterion, timeout behavior, tab targeting, or return semantics.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 25% – just the 'profile' parameter is documented in the schema, while 'tab', 'timeout' and 'selector' have no descriptions. The description adds no unit, default, or semantics for 'timeout' and no meaning for 'tab', so it fails to compensate for the coverage gap beyond naming 'selector'.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Wait for selector' essentially restates the tool name plus the required parameter, giving no more than the name already implies. With siblings like sleeper_wait_text, sleeper_wait_until, sleeper_wait_url, sleeper_wait_dialog, sleeper_wait_xhr and sleeper_wait_download, it makes no attempt to distinguish what this variant waits on versus the others.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no condition for selecting this over sleeper_wait_text or sleeper_wait_until, and no mention of prerequisites. An agent must infer the routing from the tool name alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

sleeper_wait_dialogD

Wait for dialog

ParametersJSON Schema
NameRequiredDescriptionDefault
tabNo
textNo
profileNoBrowser ID from sleeper_sessions; optional with one connected browser.
timeout_msNo

TDQS

D1.5/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden, yet it discloses nothing: not whether the call blocks, what timeout means, what happens on timeout, which dialog types are caught, or whether the dialog must be dismissed with sleeper_dialog.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three words is not conciseness but under-specification; nothing is front-loaded because nothing of substance is present. Brevity here comes at the cost of all useful information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 4-parameter, zero-annotation, no-output-schema tool in a large browser-automation family, the definition is completely inadequate. An agent cannot determine the timeout semantics, dialog type coverage, or how this relates to sleeper_dialog.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Four parameters exist with only 25% schema description coverage (just the profile field). The description explains none of them — not the meaning of tab, text (matching text of the dialog?), or timeout_ms units/behavior — leaving three parameters undocumented in both places.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

"Wait for dialog" merely restates the tool name with no added specificity. It does not distinguish this from sibling wait tools (sleeper_wait, sleeper_wait_text, sleeper_wait_until, sleeper_wait_url) or explain that it is specifically for browser dialogs like alert/confirm/prompt.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this versus sleeper_wait or the other wait variants, nor when not to. The only hint is the implied context from the name itself, with no prerequisites, ordering, or alternatives stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

sleeper_wait_downloadC

Wait for a download matching a URL/filename pattern to finish

ParametersJSON Schema
NameRequiredDescriptionDefault
tabNo
patternYes
profileNoBrowser ID from sleeper_sessions; optional with one connected browser.
timeout_msNo

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden, yet it only says the wait ends when the download 'finishes'. It does not disclose what happens on timeout, whether partial downloads count, whether it blocks or errors, or any permission/session requirements.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One front-loaded sentence with no filler or redundancy, so every word earns its place. The terseness is arguably under-specification rather than padding, which is penalized under contextual completeness instead.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

A blocking wait tool with four parameters, no annotations, no output schema, and unspecified timeout/error semantics needs considerably more description. An agent cannot know how long it will block, how to set timeout_ms, or what a failed wait looks like.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 25% (just 'profile'), leaving tab and timeout_ms undocumented anywhere. The description adds meaning only for the required 'pattern' argument ('URL/filename pattern'), which the schema leaves bare, but does not compensate for the other undocumented parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb (wait), resource (download), and matching criterion (URL/filename pattern). It is clearly distinguishable from sibling wait tools like sleeper_wait_url, sleeper_wait_xhr, and sleeper_wait_dialog, though it doesn't name any of them explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Nothing states when this tool should be preferred over the many other wait siblings (sleeper_wait, sleeper_wait_url, sleeper_wait_xhr, etc.), nor what prerequisite (a triggered download) is needed. Usage can only be inferred from the tool name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

sleeper_wait_textD

Wait for text

ParametersJSON Schema
NameRequiredDescriptionDefault
tabNo
textYes
exactNo
profileNoBrowser ID from sleeper_sessions; optional with one connected browser.
timeout_msNo

TDQS

D1.3/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden, and it discloses nothing. It does not explain blocking behavior, what happens on timeout, default timeout semantics, or whether it polls versus waits for a single event.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

At three words the description is under-specified rather than concise. There is no front-loaded context, actionable detail, or structure to guide invocation.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 5-parameter tool with no annotations and no output schema, the description is completely inadequate. An agent lacks the timeout defaults, exact-vs-substring matching behavior, and tab/profile targeting needed to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 5 parameters and only 20% schema description coverage, the description must compensate, but it adds zero meaning. Key parameters like tab, exact, and timeout_ms are entirely unexplained, and the relationship between 'text' and 'exact' is left ambiguous.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Wait for text' essentially restates the tool name sleeper_wait_text, providing a verb and resource but no additional detail. It does not distinguish itself from closely related siblings like sleeper_wait_until, sleeper_wait, or sleeper_find_text, so an agent cannot tell what makes this specific wait tool unique.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus the many other wait/find siblings (sleeper_wait, sleeper_wait_until, sleeper_wait_url, sleeper_find_text). No conditions, exclusions, or alternatives are mentioned.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

sleeper_wait_untilC

Wait for a structured page condition; JavaScript predicates are Chromium-only

ParametersJSON Schema
NameRequiredDescriptionDefault
tabNo
profileNoBrowser ID from sleeper_sessions; optional with one connected browser.
timeoutNo
intervalNo
conditionNoStructured condition with selector or text, plus optional state, count, attribute, value, operator, and exact fields
predicateNoChromium-only JavaScript predicate

TDQS

C2.5/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden. It discloses one important constraint—Chromium-only JavaScript predicates—but says nothing about timeout behavior, default interval, whether existing satisfied conditions return immediately, or what happens on timeout.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence with no waste, but for a 6-parameter wait tool with no annotations and no output schema, it is severely under-specified. Brevity here is not appropriate sizing; it leaves critical details unstated.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity—six optional parameters, a nested condition object, no output schema, and no annotations—the description is incomplete. It lacks timeout/interval semantics, profile/tab context, and condition-field guidance, though it does note the Chromium-only predicate limitation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 50%, so the description should compensate. It only alludes to condition and predicate; it adds no meaning for tab, profile, timeout, or interval, and does not explain the nested condition fields beyond what the schema already partially covers.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'Wait for a structured page condition.' It distinguishes the tool from simple text/URL waits by focusing on structured conditions and predicates, but it does not explicitly name or differentiate from siblings like sleeper_wait_text or sleeper_wait_url.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The only usage guidance is the constraint that JavaScript predicates are Chromium-only. It does not say when to choose this tool over alternatives such as sleeper_wait, sleeper_wait_text, or sleeper_wait_url, nor does it describe prerequisites or exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

sleeper_wait_urlC

Wait until URL contains pattern

ParametersJSON Schema
NameRequiredDescriptionDefault
tabNo
patternYes
profileNoBrowser ID from sleeper_sessions; optional with one connected browser.
timeout_msNo

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden, yet it says nothing about how long it blocks, what happens on timeout (timeout_ms exists but is unexplained), whether it polls, or whether it throws vs returns a status. For a blocking wait tool this is a notable gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded clause with zero filler, which is appropriately terse for the action. Brevity here reflects under-specification rather than wasted words, so it is efficient but not maximally informative.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No annotations, no output schema, and three of four parameters undocumented, so the agent lacks timeout, return-value, and failure-mode context. A blocking wait tool needs more disclosure than one line.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 25% — only profile is described in the schema. The description adds nothing about pattern (substring vs regex), tab, or timeout_ms defaults, so it fails to compensate for the coverage gap on a 4-parameter tool.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Wait) plus resource (URL) and the matching condition (contains pattern), which parses cleanly. It does not explicitly differentiate itself from sibling wait tools like sleeper_wait_text or sleeper_wait_until, so the agent must infer the distinction.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance, no mention of when it fails or blocks, and no reference to alternatives such as sleeper_wait_until or sleeper_wait_text. Usage is only implied by the tool name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

sleeper_wait_xhrC

Wait for XHR matching URL substring

ParametersJSON Schema
NameRequiredDescriptionDefault
tabNo
methodNo
profileNoBrowser ID from sleeper_sessions; optional with one connected browser.
timeout_msNo
url_substringYes

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden and largely drops it. It does not say whether the call blocks until match, what happens on timeout (raise vs. return null), whether an already-completed XHR satisfies the wait, or what a timeout_ms of omitted means, despite exposing a timeout_ms parameter.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is a single front-loaded sentence with zero filler, which is structurally fine. But it is so sparse that conciseness tips into under-specification for a 5-parameter, behaviorally complex tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No annotations, no output schema, 5 parameters at 20% coverage, and a one-line description leave the agent without the blocking/timing/error semantics it needs to invoke this reliably. Much more disclosure is required for a wait primitive in a suite with five other wait tools.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 20% (only 'profile' is documented in the schema), so the description must compensate for tab, method, timeout_ms, and url_substring. It only implicitly gestures at url_substring via 'matching URL substring' and adds nothing about method matching semantics, tab targeting, or timeout units/defaults.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The fragment 'Wait for XHR matching URL substring' identifies the verb (wait) and the resource (XHR request whose URL contains a substring), so the core action is inferable. However, it is a terse fragment with no differentiation from the many similar wait siblings (sleeper_wait, sleeper_wait_url, sleeper_wait_text, sleeper_wait_until, sleeper_network), leaving the agent to guess what kind of wait this is.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance, no prerequisites, and no mention of alternatives. The description never says to prefer this over sleeper_wait_url (navigation) or sleeper_wait (generic), which is exactly the ambiguity an agent faces in this dense sibling set.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 49 tool updatesv2.0.1
    • First observedsleeper_api
    • First observedsleeper_back
    • First observedsleeper_batch
    • First observedsleeper_check
    • First observedsleeper_click
    • First observedsleeper_click_all
    • First observedsleeper_click_text
    • First observedsleeper_console
    • First observedsleeper_dblclick
    • First observedsleeper_dialog
    • First observedsleeper_drag
    • First observedsleeper_exec
    • First observedsleeper_extract
    • First observedsleeper_fill_form
    • First observedsleeper_find
    • First observedsleeper_find_text
    • First observedsleeper_focus
    • First observedsleeper_forms
    • First observedsleeper_frames
    • First observedsleeper_get
    • First observedsleeper_goto
    • First observedsleeper_hover
    • First observedsleeper_keys
    • First observedsleeper_media
    • First observedsleeper_network
    • First observedsleeper_newtab
    • First observedsleeper_press
    • First observedsleeper_read
    • First observedsleeper_read_all
    • First observedsleeper_screenshot
    • First observedsleeper_scroll
    • First observedsleeper_scroll_until
    • First observedsleeper_select
    • First observedsleeper_sessions
    • First observedsleeper_shot
    • First observedsleeper_snapshot
    • First observedsleeper_state
    • First observedsleeper_submit
    • First observedsleeper_tabs
    • First observedsleeper_type
    • First observedsleeper_uncheck
    • First observedsleeper_upload
    • First observedsleeper_wait
    • First observedsleeper_wait_dialog
    • First observedsleeper_wait_download
    • First observedsleeper_wait_text
    • First observedsleeper_wait_until
    • First observedsleeper_wait_url
    • First observedsleeper_wait_xhr

TDQS

C2.4/5.0

Scored across 49 tools

Disambiguation3/5

Descriptions generally clarify boundaries, but several clusters overlap heavily: element reading (get, find, find_text, read, read_all, extract) and clicking (click, click_text, click_all) require careful reading to pick correctly, and shot/screenshot are literal aliases. press vs keys vs type also border on each other.

Naming Consistency4/5

Nearly all tools follow a consistent sleeper_ prefix with lowercase snake_case verb or verb_noun names (goto, click_text, wait_url), which is highly predictable. Minor deviations like the shot/screenshot alias and single-word verbs slightly break the pattern.

Tool Count2/5

49 tools is well above a comfortable range for a single server and includes redundant entries (shot/screenshot alias, multiple near-duplicate click/read/find variants). The surface could be consolidated substantially without losing capability.

Completeness4/5

Coverage of browser automation is broad: navigation, interaction, waiting, forms, uploads, dialogs, tabs, frames, network/console capture, and screenshots. Only some gaps remain (e.g. cookie/storage management), but core workflows are fully supported.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    B
    quality
    F
    maintenance
    Enables AI agents to directly control your real Chrome browser with full context including login sessions, cookies, and open tabs. It provides tools for page scanning, JavaScript execution, CDP control, screenshots, and physical mouse/keyboard input for authentic browser automation.
    20
    243
    MIT
  • A
    license
    A
    quality
    D
    maintenance
    Enables AI agents to control the user's Chrome or Firefox browser, leveraging existing sessions for tasks requiring authentication and user handoff.
    18
    57 npm
    17
    MIT
  • A
    license
    A
    quality
    A
    maintenance
    Enables AI agents to drive your real, logged-in Chrome browser with existing sessions and cookies, bypassing CAPTCHA and anti-bot measures, with support for multi-session and human-in-the-loop workflows.
    40
    46
    MIT
  • A
    license
    Not graded
    quality
    B
    maintenance
    Enables any MCP-compatible AI agent to drive your own Chrome browser with your existing login state, filling forms, clicking elements, fetching data, and handling captchas without API keys or re-authentication.
    609 npm
    255
    MIT