Skip to main content
Glama

๐Ÿ”ญ SceneScout

Your coding agent, turned into an exploratory QA tester for any web app.

Works with Claude Code ยท Cursor ยท VS Code (Copilot) ยท Codex CLI ยท Gemini CLI ยท Copilot CLI ยท Windsurf ยท any MCP client

test npm license: MIT node >= 20

๐Ÿ“– Guide ยท ๐Ÿ› What it catches ยท ๐ŸŽฌ QA with evidence ยท ๐Ÿš€ Get started ยท ๐Ÿšฆ CI ยท ๐Ÿ”’ Safety ยท ๐Ÿ“š Docs

Scripted end-to-end tests answer one question: does this exact flow still work? They say nothing about the rest of the app. SceneScout lets the agent you already use explore a running web app like a curious, thorough tester. It clicks, fills forms, switches roles and calls the API behind a hidden button, then writes a report of what is broken and what could be better, with a picture and evidence for every line.

Try it on any app you are allowed to test. No account, no API key, no setup:

npx -y scenescout http://localhost:3000
flowchart LR
  A["๐Ÿง  Your coding agent<br/>Claude Code, Cursor, Copilotโ€ฆ"] -- MCP --> S["๐Ÿ”ญ SceneScout<br/>browser ยท checks ยท memory"]
  S --> App["๐ŸŒ Your running app"]
  S --> R["๐Ÿ“‹ Report<br/>plain words + pictures"]
  S --> E["๐ŸŽฌ Evidence<br/>replay, videos, test report"]

๐Ÿ‘ฅ Who it's for

๐Ÿ› ๏ธ Developers

๐Ÿง‘โ€๐Ÿ’ผ QA, product and non-technical teams

Findings with the request that failed (GET /api/orders?status=archived โ†’ 500), the steps, and a Playwright regression-test skeleton

Each problem in plain words first: what was done, what was expected, what happened, and a picture of the page

Next to the source, the file behind the bug and a likely fix

Visible problems shown, not described: a covered button, a broken image, text too faint to read

A deterministic gate for pull requests, with SARIF for code scanning

Journeys recorded step by step, as frames and video, so a pass is something you can watch

Runs from your editor's agent, or unattended in CI

A test report laid out by your template, with expected and actual results, deviations and blank sign-off rows

Related MCP server: PixelCheck

๐Ÿ› What it catches

A real run against the small demo app in this repository, which has bugs planted on purpose. The two you can see are outlined:

It filed twelve findings. A few from the report:

๐Ÿ”ด [HIGH] A double-click on Create order creates two orders 2ร— click fired the same state-changing request 2ร— (POST /api/orders)

๐Ÿ”ด [HIGH] A clerk can approve an order by calling the endpoint the page hides from them POST /api/orders/1037/approve 200 as clerk: the button was hidden, the server did not agree.

๐Ÿ”ด [HIGH] Filtering orders by Archived fails, and the page shows an empty table instead of an error GET /api/orders?status=archived โ†’ HTTP 500

๐ŸŸ  [MEDIUM] The "New: bulk import" badge sits on top of the All orders button (callout 1) "All orders" overlaps "New: bulk import" (81%), measured from layout boxes, no screenshot needed.

What it looks for, on every page, after every action:

  • ๐Ÿ“ Broken layout, from geometry: controls that overlap, sit off-screen, hide under a sticky bar or can never be scrolled into view. The CSS bugs a person spots at a glance, found without one.

  • ๐Ÿงจ Real breakage: console errors, crashes, failed requests, 4xx and 5xx responses, broken images, dead-end pages.

  • ๐Ÿ”“ Permission leaks: it calls the app's own API as each role, so "the button is hidden" becomes "the server refuses it", or doesn't.

  • ๐Ÿคฅ Pages that lie: "Saved!" after the server refused the save, or an empty table after the request failed.

  • ๐Ÿ‘† Impatient users: a double-click that sends the same order twice.

  • โ™ฟ Accessibility and craft: contrast, focus, labels, target sizes, spacing and type, with a 0 to 100 score per page.

  • ๐Ÿงญ Friction: how many steps a task takes, and where a user had to go back.

  • ๐Ÿ’‰ Security smells: typed markup that comes back as an element, and tokens posted to any window.

Everything it checks.

๐ŸŽฌ Automated QA with evidence

Save the journeys that matter, the happy paths and the ones that must fail politely, and scenescout check replays them on every pull request with no model involved. Recorded, each run leaves proof you can watch and hand to someone who never opens a terminal.

flowchart LR
  PR["๐Ÿ”€ Pull request"] --> C["๐Ÿšฆ scenescout check<br/>--record --video --template"]
  F["๐Ÿ“ Saved journeys<br/>.scenescout/flows/*.json"] --> C
  C --> G["โœ… / โŒ A gate on the pull request"]
  C --> RP["๐Ÿ–ผ๏ธ replay.html<br/>every step, with frames"]
  C --> V["๐ŸŽž๏ธ A video per journey"]
  C --> TR["๐Ÿ“„ Test report<br/>your template, SHA-256 manifest,<br/>blank sign-off rows"]

A journey is a few lines of JSON, written by hand or kept from a flow your agent just walked. This is the demo's happy path:

{
  "name": "place an order",
  "id": "TC-01",
  "requirements": ["REQ-ORD-1", "REQ-ORD-2"],
  "steps": [
    { "action": "navigate", "target": "/orders-new.html", "expected": "The new-order form is shown" },
    { "action": "type", "target": "testid=new-order-customer", "value": "Harbour Bakery", "expected": "The customer is filled in" },
    { "action": "type", "target": "testid=new-order-items", "value": "4", "replace": true, "expected": "The item count is 4" },
    { "action": "click", "target": "testid=new-order-submit", "expected": "The order is created" },
    { "action": "expect-request", "request": "POST /api/orders", "status": "2xx" },
    { "action": "expect-element", "target": "testid=new-order-created-link", "state": "visible", "expected": "A link to the new order is shown" }
  ]
}

Against the demo app, with the three journeys in examples/flows (one creates an order, so the check is allowed to send it):

npx scenescout check http://127.0.0.1:4173 --flows examples/flows --mode read-only --flow-writes allow \
  --record --video --template examples/report-template.json

The journey, as it ran (one of the videos it filmed):

The replay page: each journey with a pass or fail badge. A journey that broke opens at the step that broke, with the page as it was after every step:

The test report, laid out by a template you write once: test IDs, the requirements each covers, expected and actual results, a screenshot per step, every deviation listed again for the reviewer, and the SHA-256 of each piece of evidence. SceneScout signs nothing; the sign-off rows are for your people.

Everything is one self-contained HTML page per report, with no scripts, nothing loaded from the network, and a layout that prints. Recording a check and a test report from a template have the details, including what to keep out of the pictures.

๐Ÿ“บ Watch it work

Each exploratory run opens a live view on your machine, with one card per agent: what it is doing, the page it is on, and a feed of every action. Three agents are testing the demo app in parallel here:

You can read the report while the agents are still working, and scrub back through any session's timeline. Beside report.md, every run writes report.html, the whole run as one self-contained page; ask for a recorded run and it also keeps a frame after every action.

โœจ Why it's different

  • ๐Ÿง  Your agent is the brain, so no extra API key. The engine contains no model. It gives the agent you already pay for deterministic tools and the testing method to use them.

  • ๐Ÿ“ It reads structure, not pixels. The agent sees every element with its role, state and layout box, so overlap and broken images are measured, not guessed from a screenshot.

  • ๐Ÿ›ก๏ธ Safety is enforced on the network, not requested in a prompt. Nothing existing is changed unless you allow it, and a blocked write never reaches your server.

  • โœ… "Done" is a contract. The report lists everything not tested, and at the extensive level refuses to finish while any known page is unvisited.

  • ๐Ÿง  It remembers. Each run starts from what the last one learned, and re-tests the bugs earlier runs left open.

  • ๐Ÿ“Š It is measured, not asserted. Every change to how it explores is scored against an app with planted bugs and a held-out app it is never tuned on (the log):

xychart-beta
  title "Planted defects found per run, demo app (13 planted)"
  x-axis "Run" ["0", "1", "2", "3", "4", "5", "6", "7", "8", "9", "10", "11", "12"]
  y-axis "Found" 0 --> 13
  bar [11, 12, 9, 10, 11, 11, 10, 10, 11, 13, 11, 11, 12]

Each bar is one run of eight parallel agents, re-scored against today's answer key. One run is noisy: read the trend across runs, not a single bar.

๐Ÿš€ Get started

1. Install for your agent (Node 20 or newer):

npx -y scenescout install                      # Claude Code: skill, MCP server and the test browser
npx -y scenescout install --client cursor      # or vscode, codex, gemini, copilot, windsurf
npx -y scenescout doctor                       # every line should be a โœ“
  • Claude Code plugin: /plugin marketplace add brunoboto96/SceneScout, then /plugin install scenescout@scenescout-marketplace. The command becomes /scenescout:scenescout.

  • Claude Desktop: download scenescout-X.Y.Z.mcpb from the latest release and open it. No terminal needed.

  • Any other MCP client: add a stdio server whose command is npx -y scenescout serve. Each client's config.

2. Start your app, then a fresh session of your agent, and ask:

Use SceneScout to test http://localhost:3000

In Claude Code there is a command too: /scenescout --url http://localhost:3000 --level medium. On its own, /scenescout asks you four plain questions instead: where the app is, how you sign in, what to check, and whether it holds real data.

3. Read the report in .scenescout/report.md. It opens in plain words (each problem, its steps, what was expected and what happened), with the technical detail one click away. A complete example.

No app handy? Clone this repository and run npm run demo:serve: the demo app starts on http://127.0.0.1:4173, and its README lists every planted bug.

Behind a sign-in? Sign in once yourself, SSO and MFA included, and every session reuses it:

npx -y scenescout login http://localhost:3000 --role admin    # then: /scenescout --role admin

Signing in covers roles, expiry and scripted sign-in for CI.

๐Ÿšฆ In CI

What it does

Needs a model?

scenescout check

A deterministic gate: measures every page, replays your saved journeys and visual baselines, fails only on what it can prove; recorded, it leaves the evidence above

No

scenescout ci

An unattended exploratory run, driven by the Anthropic or OpenAI API. It reports and never fails the build

An API key

/scenescout qa

A comment on a pull request that tests its preview deploy and replies with the results; /scenescout qa check runs your own check instead

An API key (qa check: none)

scenescout export

Files the findings as GitHub or Jira issues, each once

No

A gate on every pull request that keeps its evidence, as a GitHub Action:

- uses: brunoboto96/SceneScout@v3
  with:
    url: http://127.0.0.1:3000
    record: on
    video: on
    template: tests/report-template.json

docs/ci.md has complete workflows, every option, and the same check on GitLab CI, CircleCI or any shell.

๐Ÿ”’ Safe by default

Mode

What may leave the page

๐Ÿ”ต observe

Reads only. The default for a first look, for scenescout check, and for an agent's run on a remote site with no source

๐ŸŸข read-only

Reads and ordinary form posts; nothing existing is changed or deleted. The default for an agent's run on a local app

๐ŸŸก safe-write

Creates records, and edits or deletes only the ones it created

๐Ÿ”ด destructive

Everything. Only when you say the data is disposable; the agent never picks it

The policy sits on the network, so a blocked request never reaches your server. The page gets a refusal instead, which is how SceneScout catches a page that claims success anyway. A saved journey that creates something, like the order above, runs only when the check is told --flow-writes allow, against a test environment. Findings, memory and reports stay in a .scenescout/ folder that keeps itself out of git.

IMPORTANT

Only test sites you own or are allowed to test.Safety model.

๐Ÿ“š Documentation

Start here

A first look, install, a first run, reading the report, the live view

Ways to use it

Interactive runs, parallel agents, CI, pull-request QA, filing issues

What it checks

Every check and every scout_* tool

Signing in ยท Safety model

Roles and saved logins; what each mode refuses and why

Recipes

Setups for seven kinds of project

Configuration reference

Every option, environment variable and action input

Troubleshooting

Symptoms and fixes, upgrading and uninstalling

Running it in CI

Saved journeys, recording, test reports from templates, every workflow

How it works

Diagrams of a run, an action, the write policy, lanes

Benchmark ยท Validation

How runs are scored against answer keys, and runs on public apps

Design decisions

Why the rules are what they are

๐Ÿ”ง Contributing

Start with VISION.md (what is in scope) and CONTRIBUTING.md (local setup and how changes land). AGENTS.md holds the house rules for people and coding agents alike.

Found a way past the write policy, or another security problem? Report it privately: SECURITY.md.

MIT licensed.

Available Tools

37 tools
scout_attachA

Launch a browser and attach to a running web app. First attach in this conversation and you have read neither the SceneScout skill nor scout_playbook? Call scout_playbook before this. Write policy is enforced at the NETWORK layer: mode='observe' blocks EVERY request that is not a GET (login and token refresh excepted, and POSTs the user named in readPosts) โ€” choose it for a target that holds real data, where even an ordinary form submission would create a record; mode='read-only' (default) blocks destructive-labeled elements AND all PUT/PATCH/DELETE + destructive POSTs, but lets ordinary form POSTs through; mode='safe-write' allows creating data and permits updates/deletes ONLY on resources this session created (use when the user wants create/edit flows tested); mode='destructive' allows everything โ€” ONLY when the user explicitly confirmed a disposable/seeded environment. Pass role to sign in with a login the user saved by scenescout login <url> --role <name>, or a Playwright storage-state JSON as storageStatePath. Pass session to keep MULTIPLE roles alive at once (one browser each, genuinely concurrent) for collaboration testing โ€” target each directly with every tool's session param, or use scout_session to set which one is the default; coverage and findings merge into one project memory.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYesBase URL of the running app, e.g. http://localhost:3000
modeNoWrite policy (see tool description). Never choose 'destructive' yourself โ€” user opt-in only.read-only
openNoOpen the live view (on this attach) and/or report.html (when scout_report writes it) in the user's default browser: 'live', 'report', 'both' or 'none'. Default: the SCENESCOUT_OPEN environment variable, else 'both' on a local desktop session (headed or headless) and 'none' in CI, over SSH, or with no display. Pass it only when the user asked for something other than the default.
roleNoSign in as a role whose login was saved with `scenescout login <url> --role <name>` (kept in the project's .scenescout/auth/). Every session given the same role gets its own browser built from that one login. Not with storageStatePath. No saved login for the role: the attach is refused and names the command to run โ€” ask the user to run it, since it opens a window for them to sign in.
taskNoWhat this session is doing right now, shown under its objective from the moment it appears, e.g. Signing in and taking stock. Passed alone it is read as the 2.0 spelling of `objective`. Defaults to a placeholder so a fresh card never reads as idle.
dedupNoHow this run tells a filed finding from one already recorded. Default: the SCENESCOUT_DEDUP environment variable, else 'rule', the store's rule alone. 'judge': the rule first, then, for a filing the rule keeps apart from everything recorded, a model is asked whether it is one of the open findings on its page, and merges it when it says so. It needs ANTHROPIC_API_KEY or OPENAI_API_KEY in the server's environment, and sends each pair's titles, categories and evidence, and the page's path, to that provider โ€” ONLY when the user asked for it. Applies to every session of the project until the run ends.
headedNoShow the browser window
paceMsNoA floor between actions, in milliseconds, for when a person is watching and needs to keep up โ€” following a flow, taking notes, demonstrating. Default 0: as fast as the page allows, which is what a run wants otherwise. Changeable mid-run with scout_session {paceMs}.
recordNoKeep a frame of the page after every action, under .scenescout/recordings/, and show it beside that step in report.html. Default: the SCENESCOUT_RECORD environment variable (on or off), else off: a recording is pictures of the app under test sitting in the project folder. Turn it on for QA work, where the run is evidence and not only a report.
browserNoBrowser to drive. Default: the SCENESCOUT_BROWSER environment variable, else chromium. A build that is not on disk is downloaded on this attach, once, except in CI or with SCENESCOUT_BROWSER_DOWNLOAD=off, where the attach names the command to run. Use firefox or webkit for a cross-browser pass; stay on chromium otherwise.
sessionNoSession name for multi-role runs (e.g. 'admin', 'qa'). Creates/replaces that session's browser and makes it the default. Default: 'default'.
evidenceNoWhat happens to the picture each scout_finding takes of what it is about. 'inline': kept under .scenescout/recordings/, shown in report.html, and returned in the scout_finding result so the conversation shows it. 'file': kept and shown in the report only. 'off': none taken. Default: the SCENESCOUT_EVIDENCE environment variable, else 'inline', or 'file' in a CI job. Pictures are bounded in size and in how many one session returns (see the configuration reference).
objectiveNoThis session's objective: the whole remit you were given, in one sentence ("Admin lane: ยง2 registers, ยง7 plan gating", "Approve and reject orders as a manager"). It sits above the task, which is what the session is doing at any moment. Shown to whoever is watching the run; worth setting whenever more than one session is live.
readPostsNoPOST endpoints that only read, e.g. ["POST /api/search", "POST https://api.example.com/reports/query"], which observe mode then lets out โ€” ONLY when the user named them. Never add one yourself, even when the gap ledger lists a refused POST: ask the user. Exact paths; * stands for one path segment. Still refused when the path or body looks destructive or the body is a GraphQL mutation. Observe mode only. Default: the SCENESCOUT_READ_POSTS environment variable, else none.
projectPathNoAbsolute path to the project (memory + report live in .scenescout/ here). Pass it whenever you have a project or working folder. Omitted: the client's workspace folder, else a folder per tested site under the user's documents folder (Documents/SceneScout/<host>/, or SCENESCOUT_PROJECTS_DIR), which the result names โ€” tell the user where it is.
navTimeoutMsNoHow long a page may take to load, in ms. Default: the SCENESCOUT_NAV_TIMEOUT_MS environment variable, else 20000 (crawled pages 15000; a value set here applies to them too). Raise it only when timeouts come from a loaded machine rather than the app.
trustedEmbedsNoOrigins of embedded frames (e.g. "https://pay.example.com") whose writes out of the app may go out โ€” ONLY when the user named them, typically a provider in test mode, and only in safe-write mode. Never add one yourself. Hostile input, repeated-click probes and uploads stay refused in them.
viewportWidthNoViewport width (default 1280); use e.g. 390 for a mobile pass
viewportHeightNoViewport height (default 900)
actionTimeoutMsNoHow long one click, keystroke, hover or pick may take, in ms. Default: the SCENESCOUT_ACTION_TIMEOUT_MS environment variable, else 5000. Raise it only when timeouts come from a loaded machine rather than the app.
storageStatePathNoOptional Playwright storage-state JSON path for authenticated exploration. Not with `role`.

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does so richly: it discloses that write policy is enforced at the network layer, exactly what each mode blocks/allows, that readPosts is user-named only, that role attaches are refused without a saved login, and that dedup='judge' sends evidence to a third-party provider. This is behavioral context well beyond the schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with purpose and prerequisites, and no sentence is pure filler. However the write-policy sentence is an extremely long run-on cramming four modes and their exceptions into one breath, which hurts parseability even though the content is valuable.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 21-parameter tool with no annotations and no output schema, the description covers the critical routing decisions (mode, role, session, dedup) thoroughly and leaves the mechanical params to the 100%-covered schema. It does not describe the attach result or what a successful attach returns, which is the main residual gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents parameters and baseline is 3. The description adds meaning on the highest-stakes params: it explains the mode taxonomy (which the schema delegates to it), the role/session interaction and multi-role concurrency, and the readPosts restriction โ€” genuine added value over the enumerated values.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ("Launch a browser and attach to a running web app") and separates itself from siblings like scout_login, scout_navigate and scout_run_plan. An agent can immediately tell this is the entry/attach tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly directs the agent to call scout_playbook first when new to the conversation, and gives conditional selection rules for every mode: observe for real data, read-only as default, safe-write for create/edit flows, destructive only on user-confirmed disposable envs. Names alternatives (storageStatePath vs role) and preconditions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

scout_backB

Go back in browser history (tests back-button resilience).

ParametersJSON Schema
NameRequiredDescriptionDefault
taskNoWhat you are DOING right now, in a few words: the action, not the acceptance criteria. "Filtering the documents register by status", "Filling the deviation form with invalid dates", "Signing in as QA_Team" โ€” NOT "ยง2.4 filtering narrows the set and the filter is reflected in the URL", which is what you are CHECKING, not what you are doing. Naming the item you are on is fine ("ยง2.4: filtering the documents register"); keep the rest to what a colleague would see over your shoulder. It stays set until you pass a different one, so a batch costs a few words, not one per call. Required on the tools that act unless a journey or an earlier call already set one.
leaveNoHow to answer if the page asks to confirm leaving (a beforeunload prompt over unsent input): true leaves and discards that input, false stays. Omitted, observe and read-only stay and other modes leave. The result says when the page asked.
sessionNoTarget this session directly instead of the active one โ€” pass it explicitly when dispatching to MULTIPLE sessions in one turn (e.g. two scout_click calls with different `session`), which then run CONCURRENTLY rather than queueing. Omit for single-session sequential use.
objectiveNoOld name for `task` (2.0). Prefer `task`.

TDQS

B3.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It does not say what happens when there is no history to go back to, whether the page is reloaded/re-verified, whether in-page state survives, or whether the tool waits for load โ€” all material for a navigation action in a test harness.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single short sentence, action first, with the parenthetical earning its place by giving the testing rationale. No waste.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a no-annotation, no-output-schema navigation tool with four well-documented params, the definition is minimally adequate but silent on outcome (no return shape, no statement of what the agent should do if history is empty). It is callable but not fully self-explanatory.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents task, leave, session and objective in depth. The description adds nothing about any parameter; baseline 3 applies when the schema does all the work.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: 'Go back in browser history', with a parenthetical scoping it to back-button resilience testing. That is enough to tell it apart from scout_navigate and scout_click, though it never names the sibling it complements.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The parenthetical '(tests back-button resilience)' implies when this tool is relevant, but there is no explicit when-to-use statement, no exclusions, and no mention of scout_navigate as the forward-navigation counterpart.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

scout_captureA

Save a PNG of ONE element โ€” its bounds plus a margin, from a real screenshot โ€” under .scenescout/captures/, to show a person how it looks. Give the element's ref from the latest scout_snapshot. Not a way to judge design: scout_design_audit measures it.

ParametersJSON Schema
NameRequiredDescriptionDefault
keyNoInstead of a ref: the element's key from an earlier scout_capture result, to capture the same element on another deployment
refNoThe element's ref from the latest scout_snapshot
nameNoThe file's name: letters, digits and dashes. Default: capture
marginNoCSS pixels kept around the element. Default 8
sessionNoTarget this session directly instead of the active one โ€” pass it explicitly when dispatching to MULTIPLE sessions in one turn (e.g. two scout_click calls with different `session`), which then run CONCURRENTLY rather than queueing. Omit for single-session sequential use.

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It discloses the output format (PNG), content (element bounds plus margin), and write location (.scenescout/captures/), but does not address permissions, error cases, overwrite behavior, or what is returned.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, front-loaded with the core action and output path. The ref-sourcing instruction and the design-audit exclusion each earn their place, with no wasted wording.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a no-output-schema, no-annotation tool with five parameters, the description covers purpose, element selection, output location, and sibling routing well. It does not state what the tool returns (e.g., a capture key), though the schema's key parameter implies a return artifact.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all five parameters are already documented in the schema. The description reinforces that ref comes from the latest scout_snapshot and adds conceptual meaning to margin, but does not add syntax or constraints beyond what the schema provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: save a PNG of ONE element from a real screenshot, with scope (bounds plus margin) and output location (.scenescout/captures/). It distinguishes itself from the full-page sibling scout_screenshot by emphasizing 'ONE element' and from scout_design_audit by name.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Tells the agent to supply the element's ref from the latest scout_snapshot, and explicitly routes design judgment to scout_design_audit. It implies the contrast with full-page scout_screenshot via 'ONE element' but does not name that alternative directly.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

scout_clickA

Click an element by its ref from the latest scout_snapshot. Returns the outcome plus any oracle violations triggered. clicks=2 (or 3) probes IMPATIENT-USER behaviour: a rapid multi-click that fires the same state-changing request twice means the control is not guarded against double submission (button stays enabled, endpoint not idempotent) โ€” use it on every important submit/create button once; the result says explicitly whether duplicates fired.

ParametersJSON Schema
NameRequiredDescriptionDefault
refYesElement ref, e.g. e12
taskNoWhat you are DOING right now, in a few words: the action, not the acceptance criteria. "Filtering the documents register by status", "Filling the deviation form with invalid dates", "Signing in as QA_Team" โ€” NOT "ยง2.4 filtering narrows the set and the filter is reflected in the URL", which is what you are CHECKING, not what you are doing. Naming the item you are on is fine ("ยง2.4: filtering the documents register"); keep the rest to what a colleague would see over your shoulder. It stays set until you pass a different one, so a batch costs a few words, not one per call. Required on the tools that act unless a journey or an earlier call already set one.
leaveNoHow to answer if the page asks to confirm leaving (a beforeunload prompt over unsent input): true leaves and discards that input, false stays. Omitted, observe and read-only stay and other modes leave. The result says when the page asked.
clicksNo1 = normal; 2-3 = rapid repeated clicks (double-submit probe)
sessionNoTarget this session directly instead of the active one โ€” pass it explicitly when dispatching to MULTIPLE sessions in one turn (e.g. two scout_click calls with different `session`), which then run CONCURRENTLY rather than queueing. Omit for single-session sequential use.
objectiveNoOld name for `task` (2.0). Prefer `task`.

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does well: it discloses the double-submit probe semantics, that the result explicitly reports whether duplicates fired, and that oracle violations are returned. It omits behavior on a stale ref or a ref missing from the snapshot, which is the main residual gap for a mutation tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The core action is front-loaded in the first sentence, and the detail sentences earn their place by explaining the probe. The middle sentence is a dense run-on with stacked em-dashes, so it is slightly heavier than necessary.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with no annotations and no output schema, the description explains both the action and the return semantics (outcome, duplicates, oracle violations), which is enough to call it correctly. What is missing is prerequisite/staleness handling for the ref and any error behavior.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so baseline is 3, but the description adds genuine interpretive meaning for the `clicks` parameter โ€” explaining that two fired requests indicate an unguarded control โ€” beyond the schema's terse '1 = normal; 2-3 = rapid repeated clicks'.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a precise verb+resource: clicks an element identified by its ref from the latest scout_snapshot, and notes the return payload (outcome plus oracle violations). It fails to explicitly separate itself from keyboard-style siblings like scout_press or scout_select, so an agent must infer the distinction from the 'ref' target.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives a concrete operative rule for the probing mode: use clicks=2/3 on every important submit/create button once, and it explains what a duplicate response implies. It does not, however, say when to prefer this tool over siblings such as scout_press or scout_type, so tool-choice guidance is incomplete.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

scout_closeA

Close a session's browser (memory persists on disk). Default: the DEFAULT session. Pass session to close a specific one, or all=true to close every live session at the end of a multi-role run. A session that is a lane of a parallel run (named by scout_lane_brief or scout_lane_report) is not closed until its lane report has been accepted by scout_lane_report, since folding needs the session attached; the refusal names every such lane.

ParametersJSON Schema
NameRequiredDescriptionDefault
allNoClose every live session
forceNoClose even a lane whose report has not been folded yet. Its decisions are then lost: nothing is kept for calibration and nothing checks its defects were filed. Prefer folding it with scout_lane_report first.
sessionNoSession to close (default: the default session)

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden, and it delivers: it discloses persistence, default/all behavior, lane-close refusal, and the destructive consequence of force. It also mentions that refusal names every such lane, so the agent knows what to expect. No contradictions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, front-loaded with the core action and default behavior. The lane sentence is long but carries necessary exception logic; no filler or repeated schema content.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a state-changing tool with no annotations and no output schema, it covers the main behaviors, including edge cases (lanes, force). It doesn't describe return values or error handling beyond the refusal, but that's a minor gap for a close operation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema already documents all three parameters at 100% coverage, so the baseline is 3. The description adds operational context: when to use all, what force destroys, and the default session behavior. This is meaningful added value, though the schema itself already covers force well.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Close a session's browser.' It also states the key side effect (memory persists on disk) and distinguishes default, specific, and all sessions, which separates it from sibling tools like scout_session or scout_snapshot.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives explicit operational guidance: default session, session parameter for a specific one, and all=true for multi-role runs. It also explains when not to closeโ€”lane sessions awaiting acceptanceโ€”and directs the agent to scout_lane_report for folding. It doesn't enumerate all alternatives, but the key routing is present.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

scout_coverageA

Show exploration coverage: states visited, which elements remain unexercised, which options of a dropdown used this run no session has chosen yet, and which forms seen this run no session has submitted with every text field blank. Use to decide where to explore next and when the level's budget is satisfied. In a parallel run it shows this session's own work by default โ€” the routes it reached this run and the forms it saw โ€” so one lane is not handed another's gaps; scope:"project" shows every session's, each form and route tagged with the sessions that saw it.

ParametersJSON Schema
NameRequiredDescriptionDefault
scopeNo'session': only the routes this session reached this run and the forms it saw. 'project': every route in the memory, across runs and sessions. Default: 'session' when other sessions share this project, else 'project'.
sessionNoTarget this session directly instead of the active one โ€” pass it explicitly when dispatching to MULTIPLE sessions in one turn (e.g. two scout_click calls with different `session`), which then run CONCURRENTLY rather than queueing. Omit for single-session sequential use.

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does substantial work: it discloses the default scope in parallel runs, the session-vs-project scoping behavior, and that each form/route is tagged with the sessions that saw it. It does not state read-only safety or return format, but for a read-only coverage query the parallel-run semantics are the key behavioral fact and they are covered.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with what the tool shows, but the prose is long and dense, with nested clauses such as 'which options of a dropdown used this run no session has chosen yet' that are hard to parse. It also duplicates the schema's scope explanation, so not every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, so the description must characterize returns, and it does enumerate the coverage categories reported. Combined with the parallel-run scoping explanation and the documented defaults, an agent has enough to call it correctly; only the output shape (format, grouping) is left implicit.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already defines both parameters thoroughly, including the default rule for `scope`. The description largely restates that (session vs project) and adds only the rationale ('so one lane is not handed another's gaps'), which is marginal value over the structured field.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Show exploration coverage') and then enumerates exactly what is shown: visited states, unexercised elements, unselected dropdown options, unsubmitted blank-field forms. An agent can distinguish it from status/report siblings by content, though no sibling is named to sharpen the contrast.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear when-to-use context: 'Use to decide where to explore next and when the level's budget is satisfied.' That is a concrete decision context, but no alternative tool is named and no explicit when-not condition is given.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

scout_crawlA

Engine-side route sweep in ONE call: visits each path (default: all known routes not yet visited), records states into coverage memory, and returns a per-route health summary (HTTP status, element count, what the main area holds, oracle violations, dead-ends, auth-redirects, and ERROR-VIEW or STILL-LOADING for a main area showing only an alert or a loading placeholder). Navigation-only โ€” safe in read-only mode. Use this FIRST for broad coverage; explore interactively only where it flags problems or where journeys matter.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathsNoPaths to visit, e.g. ['/orders','/settings'], or full URLs on the attached origin. A path resolves from the origin's root, whatever page the session attached on. Omit to crawl all unvisited known routes.
sessionNoTarget this session directly instead of the active one โ€” pass it explicitly when dispatching to MULTIPLE sessions in one turn (e.g. two scout_click calls with different `session`), which then run CONCURRENTLY rather than queueing. Omit for single-session sequential use.

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations present, the description carries the full burden and delivers: it discloses side effects (records states into coverage memory), the safety profile ('Navigation-only โ€” safe in read-only mode'), and the exact per-route signals returned (status, element count, oracle violations, dead-ends, auth-redirects, ERROR-VIEW/STILL-LOADING). This is unusually complete behavioral disclosure for an unannotated tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core action and scope in the first clause, then the output shape, then the routing guidance. The long parenthetical enumerating return signals is dense but each item earns its place; slightly heavy but not wasteful.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, and the description compensates by enumerating what the health summary contains, plus safety and coverage-memory side effects. Nothing an agent needs to call this correctly against its 35 siblings is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the schema already documents both parameters, including the default-to-unvisited-routes behavior and the concurrent multi-session use of `session`. The description adds no parameter-level detail beyond the schema, so baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('route sweep... visits each path') plus the default scope (all known routes not yet visited) and the return artifact. An agent can distinguish it from interactive siblings like scout_click/scout_navigate without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly says 'Use this FIRST for broad coverage; explore interactively only where it flags problems or where journeys matter,' giving both the when-to-use and the when-to-use-alternatives condition. This directly routes the agent between this bulk crawler and the interactive tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

scout_criterionA

Record whether one acceptance criterion of a ticket read with scout_tickets passed, failed or was not tested, with how sure you are. The link from a criterion to the findings that show it is YOUR judgement, stated with a confidence โ€” never matched on words. A "fail" names the findings that show it (file them with scout_finding first); "not-tested" says why in untestedBecause. Recording the same criterion again from the same session replaces your earlier verdict. Touches no browser.

ParametersJSON Schema
NameRequiredDescriptionDefault
reasonYesWhat you saw, or why it could not be tried, in a sentence
ticketYesThe ticket's id as scout_tickets gave it, e.g. "PROJ-12" or "T1"
sessionNoTarget this session directly instead of the active one โ€” pass it explicitly when dispatching to MULTIPLE sessions in one turn (e.g. two scout_click calls with different `session`), which then run CONCURRENTLY rather than queueing. Omit for single-session sequential use.
verdictYes"pass", "fail" or "not-tested"
findingsNoIds of the findings that show this criterion failing (required for a fail; may be given for a pass, none for not-tested)
criterionYesThe criterion's id, e.g. "AC2" (or just "2")
confidenceYesHow sure you are of this verdict and of the findings linked to it, from 0 to 1. State it honestly
untestedBecauseNoOnly with verdict "not-tested": "no-access" (the role this run used could not reach it), "observe-blocked" (it needs a change sent and this session is in observe mode), "out-of-scope" (it is outside what this run could check, such as an email or another system)

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does well: it discloses idempotent replacement semantics ('Recording the same criterion again from the same session replaces your earlier verdict'), the fact that link judgement is never word-matched, and that it 'Touches no browser.' It stops short of stating auth/permission needs or any failure/error behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the core action and scope in the first clause, then layers the linkage rule, fail/not-tested requirements, and idempotency. Dense but all sentences earn their place; the parenthetical 'file them with scout_finding first' is well placed.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For an 8-param, no-annotation, no-output-schema mutation of run state, the description covers prerequisites, verdict semantics, idempotency, and browser footprint. Only the return/confirmation behavior and any permission constraints are left unstated, which is a minor gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% so a baseline of 3 applies, but the description adds cross-parameter semantics the schema does not spell out: a 'fail' requires naming findings, 'not-tested' requires untestedBecause, and confidence must be an honest judgement rather than a match. That is genuine meaning beyond the field-level docs.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (record) and resource (one acceptance criterion of a ticket), names the exact scope, and ties the resource back to scout_tickets. It is immediately distinguishable from siblings like scout_finding or scout_note.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly tells the agent when to use it (recording a pass/fail/not-tested verdict), the prerequisite ordering ('file them with scout_finding first'), which sibling produced the ticket (scout_tickets), and what happens on re-record from the same session. This goes well beyond implied usage.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

scout_design_auditA

Computed-style design audit of the current page โ€” a design connoisseur's read WITHOUT screenshots. Measurable defects (โš ): WCAG contrast, tiny targets, clipped text, aspect-distorted images, horizontal overflow, missing keyboard-focus indicators (sampled with real Tab presses). Craft suggestions (โ†’): line measure and line-height rhythm, spacing-scale adherence, typography entropy, palette discipline (gray census, accent hue families, pure-#000 body text), elevation/control consistency, heading structure, indistinguishable links, and AI-slop tells (gradient text, glassmorphism, side-stripe borders, neon glows, violet gradients, identical card grids). Ends with a SYSTEM SUMMARY of design-system coherence. Run once per representative page; the โ†’ tier is improvement feedback โ€” file genuine opportunities as ux-polish findings with the concrete numbers, not just defects.

ParametersJSON Schema
NameRequiredDescriptionDefault
sessionNoTarget this session directly instead of the active one โ€” pass it explicitly when dispatching to MULTIPLE sessions in one turn (e.g. two scout_click calls with different `session`), which then run CONCURRENTLY rather than queueing. Omit for single-session sequential use.

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden and does a good job: it declares the audit is computed-style without screenshots, samples keyboard focus with 'real Tab presses', and lists both measurable defects and craft suggestions. It does not address side effects or permissions, but as a read-only audit the implied behavior is clear.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long but densely informative, with clear structure: purpose, defect types, suggestion types, summary, and usage note. It front-loads the core identity and every subsequent clause adds practical detail. A little trimming would be possible, but the length serves the complexity of the tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has no output schema and no annotations, the description compensates well by enumerating what the audit detects, the structure (defects vs. suggestions vs. system summary), and how to use the results. It does not specify the exact return format, but the content coverage is sufficient for an agent to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The only parameter, session, is already fully documented in the input schema with clear guidance on when to pass it explicitly. The description adds no additional parameter semantics, which is acceptable because schema coverage is 100%.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: a 'computed-style design audit of the current page', and explicitly distinguishes itself with 'WITHOUT screenshots'. It enumerates distinct audit categories (WCAG contrast, tiny targets, clipped text, etc.), making its purpose unmistakable compared to siblings like scout_screenshot or scout_scan.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives clear usage context: 'Run once per representative page' and instructs that improvement-tier findings should be filed as ux-polish findings with concrete numbers rather than just as defects. It does not explicitly name alternative tools or exclusion criteria, but the context is specific enough to guide correct invocation.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

scout_findingA

Record a structured finding (bug, UX issue, or improvement). Deduplicates across runs; automatically captures the recent action trace as the repro, and a picture of what it is about (the element ref names, else the viewport), kept for report.html and returned in this result as an image. Use for anything worth reporting: crashes, oracle violations you confirmed, dead ends, confusing UX, permission leaks, missing testids โ€” and design-audit improvement opportunities (ux-polish) with their concrete measurements.

ParametersJSON Schema
NameRequiredDescriptionDefault
refNoThe ref, from the latest scout_snapshot, of the element the finding is about. Its picture (the element plus a margin) is kept with the finding and shown in report.html. Omit it and the picture is the viewport.
titleYesOne-line summary of the defect
detailYesWhat happened, what was expected, and the evidence
sessionNoTarget this session directly instead of the active one โ€” pass it explicitly when dispatching to MULTIPLE sessions in one turn (e.g. two scout_click calls with different `session`), which then run CONCURRENTLY rather than queueing. Omit for single-session sequential use.
categoryYesPick the closest โ€” use 'other' only when nothing fits
evidenceNoCanonical machine signature for dedup, e.g. 'GET /api/reports/dashboard 403' or 'widget dashboard-summary-widget shows 0'. Same bug re-found later should produce the same string.
severityYes
conventionNoOnly for a WORTH-A-LOOK finding: the observation is real, and it is a defect only under a convention of this project you cannot see. Name that convention, e.g. 'a 4px spacing scale' or 'test ids on every control'. The report lists it under "Worth a look", apart from the defects, and does not count it as one. Not for "I could not tell": leave that unfiled or look closer. Omit for a defect.

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does well: it discloses automatic deduplication across runs, automatic capture of the recent action trace as the repro, and automatic capture of a picture (element `ref` or viewport) that is persisted for report.html and returned as an image. These are real side effects an agent should know about, though failure modes and any permission requirements are not covered.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is a single dense paragraph, but it is front-loaded with the purpose, then behavior, then usage scope, and every clause carries information (dedup, repro capture, image, category examples). The heavy em-dash listing makes it slightly harder to scan than an ideally structured definition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For an 8-parameter, no-annotation, no-output-schema tool, the description covers purpose, side effects, return of the captured image, and filing conventions. It is close to complete; only edge details like severity semantics and error behavior are left to the schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is already 88%, so the schema does most of the work, but the description adds meaning beyond it: it explains the `convention` field's special role (routed to a "Worth a look" section and not counted as a defect) and reinforces the `ref`/picture behavior. That is meaningful added value above the high-coverage baseline of 3.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a precise verb+resource ("Record a structured finding") and enumerates what qualifies (crashes, UX issues, improvements), so the agent knows exactly what this tool produces. It does not explicitly name or contrast a sibling tool (e.g. scout_note) to disambiguate where else one might file an observation, so it falls short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

"Use for anything worth reporting" followed by a concrete list of triggers (crashes, confirmed oracle violations, dead ends, confusing UX, permission leaks, missing testids, ux-polish) gives strong when-to-use guidance. There is also a when-not signal for the `convention` field ("Not for 'I could not tell'"), but no explicit alternative tool is named for cases this tool should not handle.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

scout_hoverA

Hover an element by ref like a user pausing the pointer on it, and report what it reveals: tooltips/popovers (diffed against pre-hover state), any other new page text that appeared (labelled as possibly unrelated on busy pages), the title attribute, and aria-describedby text โ€” each item truncated to 300 chars. Hovering does not count as exercising the element. Use on badges, icons, truncated text, and error indicators BEFORE concluding an element 'does nothing' โ€” hover-gated UI is invisible to snapshots and clicks.

ParametersJSON Schema
NameRequiredDescriptionDefault
refYesElement ref, e.g. e12
sessionNoTarget this session directly instead of the active one โ€” pass it explicitly when dispatching to MULTIPLE sessions in one turn (e.g. two scout_click calls with different `session`), which then run CONCURRENTLY rather than queueing. Omit for single-session sequential use.

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries full behavioral disclosure responsibility. It explains that results are diffed against pre-hover state, unrelated text may be labelled, items are truncated to 300 characters, and hover does not count as exercising the element. This gives a clear and honest model of the tool's behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three dense sentences with no filler: the first defines behavior and outputs, the second states a key caveat, and the third provides usage guidance. The most important information is front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Despite having no output schema and no annotations, the description explains the result composition, truncation, and possible noise. Combined with complete parameter schema and explicit usage conditions, an agent has everything it needs to invoke and interpret the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%: both ref and session already have full descriptions, including the concurrency behavior of session. The tool description adds little per-parameter detail, but that is acceptable because the schema already handles parameter semantics.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb and resource: hover an element by ref and report what it reveals. It enumerates concrete outputs (tooltips, new text, title attribute, aria-describedby) that clearly distinguish it from sibling tools like scout_click and scout_snapshot.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly says to use it on badges, icons, truncated text, and error indicators before concluding an element 'does nothing', which is strong when-to-use guidance. It also notes that hovering does not count as exercising the element, implying it is not a substitute for interaction, though it does not name scout_click as the explicit alternative.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

scout_intakeA

Ask the start-of-run questions (the site's address, whether and how to sign in, what to check, whether the site holds real data) when the person gave no settings. Where this client can show a form, the person answers there and this returns the settings to use: the scout_login and scout_attach calls, in order. Otherwise, or when they decline or close the form, it returns the questions for you to ask in chat. Never asks for a password. Touches no browser.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlNoThe site's address, when the person already said it: the form starts with it filled in

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does so well: it discloses that it returns scout_login and scout_attach calls in order, falls back to returning questions for chat, never asks for a password, and touches no browser. These are meaningful behavioral facts beyond a simple 'ask questions' framing, though the return payload shape is only sketched.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose is front-loaded, but the middle sentence ('Where this client can show a form, the person answers there and this returns the settings to use: the scout_login and scout_attach calls, in order') is dense and runs on. The content earns its place, yet the phrasing could be tightened without loss.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema and no annotations, the description must explain what comes back, and it does: either settings (scout_login/scout_attach calls in order) or a list of questions. It also covers the no-password and no-browser guarantees. Only the exact response structure remains slightly underspecified.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and there is a single optional parameter, so the schema already defines 'url' fully. The description adds only that the value pre-fills the form when the person has already stated the address, which is marginal beyond the schema text. Baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a specific verb and resource ('Ask the start-of-run questions') and enumerates what those questions cover (address, sign-in, what to check, real data). It also situates the tool relative to siblings by naming scout_login and scout_attach as its outputs, so an agent can tell it is the setup/intake step. It stops short of an explicit 'use this instead of X' contrast, which keeps it from a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It states a clear triggering condition: use it 'when the person gave no settings.' It also distinguishes the two execution paths (form vs. chat fallback) and when each applies ('otherwise, or when they decline or close the form'). No explicit when-not is given, but the gating condition is unambiguous.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

scout_journeyA

Measure how EASY a real task is, not just whether it works โ€” the question pass/fail e2e suites never answer. Wrap one user goal: scout_journey {action:'start', goal:'Create an order'}, perform it the way a first-time user would (navigate by CLICKING through the UI, not by jumping to a known deep URL โ€” a shortcut invalidates the measurement), then scout_journey {action:'end', completed:true|false}. Returns interaction cost (clicks, navigations, distinct screens, elapsed), the actual path taken, and friction signals: BACKTRACKS (returning to a screen already left โ€” the clearest sign the next step wasn't discoverable), screen count, and over-interaction. Run it on each module's primary journey; an abandoned journey is a high-severity finding.

ParametersJSON Schema
NameRequiredDescriptionDefault
goalNoFor start: the user-facing task, e.g. 'Create an order and assign it'
noteNoFor end: what made it hard or easy, in one line
actionYes'start' before attempting the task, 'end' when done or blocked
sessionNoTarget this session directly instead of the active one โ€” pass it explicitly when dispatching to MULTIPLE sessions in one turn (e.g. two scout_click calls with different `session`), which then run CONCURRENTLY rather than queueing. Omit for single-session sequential use.
completedNoFor end: did the user actually achieve the goal? false is a strong finding.

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries full behavioral burden and does so well: it discloses that the tool actually performs the journey by clicking, that shortcuts invalidate the measurement, what metrics are returned, and how to interpret abandonment. The live side-effect potential is implied by the 'Create an order' example, but the 'perform it' wording makes the behavior clear.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four dense sentences cover purpose, lifecycle, execution rule, return metrics, and usage guidance. There is no filler; the opening contrast with pass/fail e2e orients the agent, and the lifecycle example is front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Even without an output schema, the description enumerates the returned interaction metrics and friction signals, defines the invalidating shortcut, and states the severity of an abandoned journey. Session handling is covered by the schema's detailed session parameter, so an agent has enough to invoke and interpret the tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, and the description adds a concrete usage template pairing action:'start' with goal and action:'end' with completed, plus a real example goal. It doesn't add separate semantics for session/note, but those are already well documented in the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific measurement goal โ€” task ease rather than pass/fail โ€” and defines a clear start/end journey lifecycle with concrete signals (backtracks, interaction cost, over-interaction). This clearly distinguishes scout_journey from execution-oriented siblings like scout_click or scout_navigate.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly says when to use it: wrap a user goal with start/end and run it on each module's primary journey. It also emphasizes the correct execution method (clicking through UI, no deep-link shortcuts) and treats abandoned journeys as high-severity. However, it doesn't name alternative tools or state when not to use it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

scout_lane_briefA

For a run split across parallel agents (lanes). Divides the app's known routes between N lanes and returns each lane's session name, the objective to attach it with, and the routes it owns โ€” whole modules per lane, balanced by route count, so no two lanes audit the same area and none is left unopened. Call it after the first crawl, when route knowledge is complete. Touches no browser; pass the briefs to your lane agents, then use scout_lane_report for what they hand back.

ParametersJSON Schema
NameRequiredDescriptionDefault
goalNoWhat the whole run is for; each lane's objective is written against it
lanesYesHow many lanes to split across (1โ€“8)
routesNoRoutes to split. Omit to split every route this project knows about.
sessionNoTarget this session directly instead of the active one โ€” pass it explicitly when dispatching to MULTIPLE sessions in one turn (e.g. two scout_click calls with different `session`), which then run CONCURRENTLY rather than queueing. Omit for single-session sequential use.
runMinutesNoHow long the lanes will run, in minutes (default 60). When they attach by a saved role, the brief is refused if that login will not last this long plus the margin.
expiryMarginMinutesNoHow long past the run a saved role's login must still last, in minutes (default 10).

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does well: it discloses that the tool 'touches no browser' and that the caller must pass the briefs to lane agents itself. It also explains the balancing guarantee (no overlap, none left unopened) and the return shape, which is needed since there is no output schema. It doesn't cover failure modes or what happens if lanes exceed available routes.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four sentences, purpose front-loaded, zero filler. It is dense but every clause adds routing or behavioral information. Slightly long, and the return-value clause could be trimmed, but nothing is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 6-param tool with no annotations and no output schema, the description supplies the missing pieces: what it returns, that it has no browser side effects, and the ordering constraint relative to the crawl. Remaining gaps are minor โ€” no explicit mention of what a refused brief looks like or how to handle lane counts exceeding route counts.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. The description adds genuine meaning beyond the schema by explaining the splitting algorithm ('whole modules per lane, balanced by route count'), which tells the agent how the `lanes` and `routes` inputs actually partition work rather than merely restating their types.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource โ€” it divides an app's known routes across N parallel lanes and returns each lane's session name, objective, and owned routes. The scope ('whole modules per lane, balanced by route count, no two lanes audit the same area') is precise enough to separate it from scout_crawl or scout_lane_report without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit timing: 'Call it after the first crawl, when route knowledge is complete.' It also routes the agent forward, naming scout_lane_report as the counterpart for what lanes hand back. It stops short of stating when not to use it (e.g. a single-agent run with no parallel split).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

scout_lane_reportA

For a run split across parallel agents (lanes). Without reply: returns the paragraph to put in a lane's prompt, telling it to hand back ONE typed JSON object (verdicts, severities and categories from closed sets, a calibrated confidence per decision, routes covered, what blocked it). With reply: parses what the lane handed back and returns the one-line fold (defects, highs, unsure, mean confidence, routes) or the reason it was refused, to relay to the lane once. Touches no browser.

ParametersJSON Schema
NameRequiredDescriptionDefault
laneYesThe lane's name, as used in its session
replyNoThe text the lane handed back; omit to get the instruction instead

TDQS

A4.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It adds a helpful trait ('Touches no browser') and explains the two output behaviors, including the failure case ('reason it was refused'). However, it does not explicitly state whether it has any state-changing side effects or mention permissions/rate limits, leaving some ambiguity for a tool with no annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is information-dense and well-structured: it opens with the run context, then branches into the two modes with concrete outcomes, and closes with a safety note. Every sentence earns its place, though the density requires careful reading. It is slightly long but not wasteful.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, so the description must explain return values, and it does: it details the paragraph content for the no-reply mode and the one-line fold or refusal reason for the reply mode. It also explains the input contract for the lane. Missing are details like pagination or error semantics, but for a tool of this complexity it is reasonably complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description adds value beyond the schema by explicitly linking the two parameters: omitting `reply` yields the instruction, while providing it triggers parsing. This conditional relationship is not obvious from the schema alone, so the description enhances parameter understanding.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb+resource pair: it returns a lane prompt paragraph when `reply` is omitted and parses a lane's reply when `reply` is provided. The opening context 'For a run split across parallel agents (lanes)' clearly frames this as a lane-specific tool, distinguishing it from sibling tools like scout_report or scout_lane_brief.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit usage conditions: use it when a run is split across lanes, and it precisely states the two modes based on whether `reply` is omitted or provided. It does not name alternatives or exclusions, but the context is clear enough that an agent can decide when to invoke this tool over siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

scout_loginA

Open a visible browser window for the USER to sign in to the app as a role, and save that sign-in for scout_attach { role }. Use it when an attach is refused because no sign-in is saved for the role, or the saved one has expired. Tell the user first, in plain words: "A browser window is opening. Sign in there as you normally would; it closes by itself once you are in." The window saves once the user is back on the app with a new session (a round trip through a single sign-on provider is followed, not taken for the end) and closes. Never type credentials into it yourself. Returns once signed in and saved, or after waitSeconds with the window still open: then call scout_login again with the same role to keep waiting. Closing the window saves nothing. Needs a desktop: on a machine with no display, ask the user to run scenescout login <url> --role <name> where they can see the window.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYesWhere to sign in: the app's address or its sign-in page, e.g. http://localhost:3000/login
roleYesThe name to save the sign-in under, e.g. admin; scout_attach { role } signs in with it
browserNoBrowser to open. Default: the SCENESCOUT_BROWSER environment variable, else chromium
successUrlNoOnly when the user says how to tell: signed in once the URL's path contains this, or the URL starts with it (an absolute URL), instead of when a new session appears
projectPathNoAbsolute path to the project (the sign-in is saved in .scenescout/auth/ here), as for scout_attach. Omitted: the same folder an attach with no projectPath uses for this site, which the result names.
waitSecondsNoHow long this call waits for the user before returning with the window still open (default 120)

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden and does: window is visible and closes itself, credentials must never be typed by the agent, a round trip through SSO is followed, closing the window saves nothing, the call returns after waitSeconds with the window still open and must be re-called to keep waiting, and a display is required. This is unusually rich disclosure of side effects, failure modes, and interaction contract.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose and trigger are front-loaded, and every sentence carries operational content (user-facing message text, save semantics, retry loop, headless fallback). It is long for a tool description, but the length is justified by the complexity; only slight trimming would be possible.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no annotations and no output schema, the description still covers return conditions, timeout behavior, save semantics, prerequisites (desktop/display), and the fallback procedure. Nothing an agent needs to invoke this correctly appears to be missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3; the description goes beyond it by explaining the waitSeconds re-invocation loop (call scout_login again with the same role), the persistence semantics of role, and the projectPath default behavior. It adds real meaning rather than restating the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource (open a visible browser window for the user to sign in) plus the exact downstream artifact it produces (a saved sign-in for scout_attach { role }). An agent can distinguish this from scout_attach, scout_run_plan, and every other scout_* tool without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives an explicit trigger ('Use it when an attach is refused because no sign-in is saved for the role, or the saved one has expired'), names the related tool scout_attach, and provides an alternative path for headless machines (ask the user to run `scenescout login <url> --role <name>`).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

scout_navigateB

Navigate to a path on the attached origin (e.g. '/orders') or a full URL on it. A path resolves from the origin's root, whatever page the session attached on. Also supports 'back' via scout_back.

ParametersJSON Schema
NameRequiredDescriptionDefault
taskNoWhat you are DOING right now, in a few words: the action, not the acceptance criteria. "Filtering the documents register by status", "Filling the deviation form with invalid dates", "Signing in as QA_Team" โ€” NOT "ยง2.4 filtering narrows the set and the filter is reflected in the URL", which is what you are CHECKING, not what you are doing. Naming the item you are on is fine ("ยง2.4: filtering the documents register"); keep the rest to what a colleague would see over your shoulder. It stays set until you pass a different one, so a batch costs a few words, not one per call. Required on the tools that act unless a journey or an earlier call already set one.
leaveNoHow to answer if the page asks to confirm leaving (a beforeunload prompt over unsent input): true leaves and discards that input, false stays. Omitted, observe and read-only stay and other modes leave. The result says when the page asked.
targetYesAbsolute URL or path like /settings
sessionNoTarget this session directly instead of the active one โ€” pass it explicitly when dispatching to MULTIPLE sessions in one turn (e.g. two scout_click calls with different `session`), which then run CONCURRENTLY rather than queueing. Omit for single-session sequential use.
objectiveNoOld name for `task` (2.0). Prefer `task`.

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It does disclose one genuinely non-obvious behavioral trait: a path resolves from the origin's root regardless of which page the session attached on. Beyond that it says nothing about whether navigation waits for load, how failures/redirects are surfaced, or whether the session/target stays attached โ€” significant gaps for a tool whose whole job is a state-changing page transition.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, front-loaded with the primary target semantics before the path-resolution nuance. The trailing sentence about scout_back earns its place by routing the agent, though it reads slightly oddly as a feature of this tool rather than a redirect.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No annotations and no output schema, so the description must cover the behavioral profile โ€” it covers target/path resolution but omits load-waiting behavior, failure modes, and what the caller receives back. The remaining four parameters are fully documented in the schema, so the gap is behavioral rather than parametric.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description does add meaning for `target` by clarifying that a path is resolved relative to the origin root rather than the current page โ€” more than the schema's terse 'Absolute URL or path like /settings' โ€” but it says nothing about `task`, `leave`, or `session`, relying entirely on the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Navigate to') and resource ('a path on the attached origin or a full URL on it'), with a concrete example ('/orders'). It also routes the agent away from this tool for back-navigation ('Also supports \'back\' via scout_back'), which gives partial sibling differentiation, though it does not distinguish itself from scout_click/scout_attach in the same way.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied rather than stated: the agent can infer this is the tool for moving to a new page, and the description explicitly names scout_back as the alternative for back-navigation. However, there is no guidance on when to navigate vs. when the session is already on the target, nor any mention of prerequisites such as needing an attached session.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

scout_networkA

List the data requests (fetch and XHR) the current page made since its document loaded: method, path, status and time, oldest first, each marked with the route it was sent from when a client-side route change moved the page since. Use it when the page and the server seem to disagree โ€” an empty list where the API has data, a stale value after a save โ€” to tell a request that failed, one still pending and one that never ran apart. Read-only: it lists what the browser already saw and sends nothing. Credentials in query strings are redacted. A full page load starts a new list; scout_request's own calls are marked.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoHow many of the newest to list. Default 40.
sessionNoTarget this session directly instead of the active one โ€” pass it explicitly when dispatching to MULTIPLE sessions in one turn (e.g. two scout_click calls with different `session`), which then run CONCURRENTLY rather than queueing. Omit for single-session sequential use.
containsNoOnly requests whose path contains this text, e.g. /api/things

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does so: read-only, sends nothing, credentials in query strings redacted, list resets on full page load, and scout_request's own calls are marked. It also discloses ordering and the route-change annotation of entries โ€” unusually complete behavioral context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with what is listed and the ordering, then usage and safety notes. Dense but every sentence carries information; slightly long for a single paragraph, but no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema and no annotations, yet the description covers return content, ordering, the reset lifecycle, redaction, and integration with scout_request. An agent has everything needed to call and interpret it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all three parameters (limit, session, contains) are already documented in the schema, including the concurrency semantics of `session`. The description adds no parameter-specific detail, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a precise verb and resource ('List the data requests (fetch and XHR) the current page made since its document loaded') and enumerates the returned fields (method, path, status, time, ordering). It is clearly distinguishable from siblings like scout_request and scout_navigate.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives concrete trigger conditions ('when the page and the server seem to disagree โ€” an empty list where the API has data, a stale value after a save') and the diagnostic goal. It does not name an explicit alternative tool to use instead or state exclusion cases, so it stops short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

scout_noteA

Cumulative WRITTEN knowledge about the tested app โ€” .scenescout/ASSUMPTIONS.md, in prose a human can read and correct. memory.json stores coverage; this stores UNDERSTANDING, so every run starts smarter than the last. READ it at the start of every session ({action:'read'}). ADD durable learnings as you go ({action:'add', section, note}): what the app is for (app-model), who each role is and what they're FOR โ€” infer the persona from what the role can see and do, e.g. 'qa-role = reviewer: approves orders, cannot administer' (roles), UI patterns the app follows (conventions), rules discovered the hard way like 'an order can only ship once approved' (constraints), fragile areas worth re-testing every run (risks), domain terms (glossary), and how to get the app testable at all โ€” the command that regenerates an expired login state, what has to be running (setup), which the engine reads back to you the next time a storage state has expired. Notes are dated, attributed to the acting role, and deduplicated. Do NOT record session-specific facts (ids, counts) โ€” only durable knowledge.

ParametersJSON Schema
NameRequiredDescriptionDefault
noteNoFor add: the learning, one or two sentences, written for a future reader with no context
actionYes'read' the accumulated knowledge, or 'add' one durable learning
sectionNoFor add: which knowledge section this belongs to
sessionNoTarget this session directly instead of the active one โ€” pass it explicitly when dispatching to MULTIPLE sessions in one turn (e.g. two scout_click calls with different `session`), which then run CONCURRENTLY rather than queueing. Omit for single-session sequential use.

TDQS

A4.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With zero annotations, the description carries the full burden and discloses the cumulative/persistent nature ('every run starts smarter than the last'), the file location, and the side effects of adds: 'Notes are dated, attributed to the acting role, and deduplicated.' It stops short of exhaustive because edge cases like first-run behavior when ASSUMPTIONS.md does not yet exist, and the exact shape of the read response, are left implicit.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Every sentence earns its place and the purpose is front-loaded, but the whole description is one dense paragraph with a ~70-word parenthetical section list and a dangling clause about engine read-back at the end. The length is proportionate to the tool's conceptual richness (7 sections, 2 actions, content policy), but the lack of structural breaks makes it harder for an agent to scan.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no annotations and no output schema, the description alone must cover purpose, storage location, usage timing, section semantics, content policy, and behavioral traits โ€” and it covers all of these with concrete examples. Nothing an agent needs to invoke read/add correctly is missing; the only minor gaps (exact read return shape, first-run file creation) are self-evident for this kind of tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so baseline is 3; the description adds real value by giving a concrete example for every section enum โ€” e.g., roles: 'qa-role = reviewer: approves orders, cannot administer' and constraints: 'an order can only ship once approved.' It also reinforces note-format guidance and session concurrency semantics, meaningfully elevating what the schema states.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The opening sentence states the resource and function precisely: 'Cumulative WRITTEN knowledge about the tested app โ€” .scenescout/ASSUMPTIONS.md.' It further disambiguates from coverage tracking with 'memory.json stores coverage; this stores UNDERSTANDING,' and the twin actions read/add are made explicit. An agent can distinguish this from scout_coverage or scout_finding without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives direct operational directives: 'READ it at the start of every session' and 'ADD durable learnings as you go,' plus an explicit when-not rule: 'Do NOT record session-specific facts (ids, counts) โ€” only durable knowledge.' It also names the storage-state trigger ('reads back to you the next time a storage state has expired') and explains the alternative store for coverage, so an agent knows exactly when this tool is the right one.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

scout_playbookA

Return the SceneScout testing method: setup order, write modes, how to explore, what counts as done, how to report. Call this ONCE before the first scout_attach in a conversation, then follow it. If this client offers a SceneScout skill, load that instead โ€” it is the same text, so never read both. Takes no input and touches no browser.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden, and it states 'Takes no input and touches no browser,' making the read-only, side-effect-free nature clear. It also warns that calling it again is unnecessary and that it duplicates the skill text.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tightly packed sentences: content, usage timing/alternative, and input/side-effect constraints. No filler, and the most decision-relevant information appears first.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a no-input, no-output-schema tool, the description fully covers what the tool returns, when to invoke it, how it relates to the skill alternative, and that it has no browser side effects. Nothing an agent needs to call or follow it is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, so the schema already says nothing is needed. The description adds the explicit confirmation 'Takes no input,' which aligns with the empty schema, earning the zero-parameter baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb and resource: 'Return the SceneScout testing method' and enumerates its contents (setup order, write modes, exploration, done criteria, reporting). This clearly distinguishes it from the sibling execution tools like scout_attach and scout_report.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly says when to call it: 'ONCE before the first scout_attach in a conversation, then follow it.' It also names an alternative and provides a rule: if a SceneScout skill is offered, load that instead and 'never read both.'

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

scout_pressA

Press a keyboard key (e.g. Escape, Tab, Enter) โ€” useful for closing modals and testing keyboard navigation.

ParametersJSON Schema
NameRequiredDescriptionDefault
keyYes
taskNoWhat you are DOING right now, in a few words: the action, not the acceptance criteria. "Filtering the documents register by status", "Filling the deviation form with invalid dates", "Signing in as QA_Team" โ€” NOT "ยง2.4 filtering narrows the set and the filter is reflected in the URL", which is what you are CHECKING, not what you are doing. Naming the item you are on is fine ("ยง2.4: filtering the documents register"); keep the rest to what a colleague would see over your shoulder. It stays set until you pass a different one, so a batch costs a few words, not one per call. Required on the tools that act unless a journey or an earlier call already set one.
sessionNoTarget this session directly instead of the active one โ€” pass it explicitly when dispatching to MULTIPLE sessions in one turn (e.g. two scout_click calls with different `session`), which then run CONCURRENTLY rather than queueing. Omit for single-session sequential use.
objectiveNoOld name for `task` (2.0). Prefer `task`.

TDQS

A4.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden. It states the core behavior and even hints at consequences (modal dismissal, keyboard navigation), but it does not disclose details such as whether the press targets the active page, whether it dispatches keydown/keyup, or how unsupported keys are handled. This is acceptable for a simple action but not richly transparent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single front-loaded sentence with no filler. It immediately states the action, gives examples, and then explains when it is useful โ€” every word earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple key-press tool with a well-described parameter schema, the description is largely sufficient. It covers what the tool does and typical use cases. The only gap is an explicit relationship to sibling tools like scout_type, but this is minor given the overall clarity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 75%, and the tool description adds essential meaning to the undocumented `key` parameter by providing valid examples. The other parameters already have detailed schema descriptions, so the description's contribution is focused and useful for the one parameter that needs it.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action: 'Press a keyboard key' with concrete examples (Escape, Tab, Enter) and signals the intended resource (keyboard). It differentiates the tool from close/click/type siblings by focusing on key-based interaction rather than mouse or text entry, making intent unmistakable.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear use cases: 'closing modals and testing keyboard navigation.' It does not explicitly mention when to avoid this tool or name alternatives like scout_type, but the context is clear enough for an agent to choose it appropriately for key-press actions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

scout_reportA

Generate the final markdown report โ€” findings, page quality scores (worst first), role capability matrix, oracle rollup, and the GAP LEDGER (an explicit list of what was NOT tested). Writes the full document to .scenescout/report.md and returns a bounded SUMMARY (full reports exceed client token limits). Gates by level: 'minimal' needs all routes visited + โ‰ฅ1 design audit; 'medium' additionally needs several routes audited; 'extensive' REFUSES while the gap ledger is non-empty โ€” that refusal is the completeness guarantee: an extensive report only generates when nothing known is left untested. force=true overrides (only when the user capped the budget).

ParametersJSON Schema
NameRequiredDescriptionDefault
forceNoGenerate even though gates are unmet (only when the user capped the budget)
levelNoWhich completion contract to enforce โ€” match the level the run was asked formedium
reportNoWhich parts the report carries. 'both' (default): a plain-language section first โ€” a short summary, then each problem with numbered steps, what was expected, what happened, its picture and its impact (blocks users, annoying, cosmetic), each with its technical detail folded beneath โ€” followed by the technical report. 'qa': the plain section alone, for a tester or anyone not technical. 'dev': the technical report alone, as before the plain section existed. It applies to the files this call writes; the live view always shows both.both
historyNoHow much of the history to print. 'index' lists findings from earlier runs, and resolved ones, as a row each: id, severity, age, title. 'full' prints every one in full as before โ€” on one project that was 1.75 MB against 113 KB, nearly half of it findings already fixed. Use 'full' when handing the document to someone who has no access to the memory.index
sessionNoTarget this session directly instead of the active one โ€” pass it explicitly when dispatching to MULTIPLE sessions in one turn (e.g. two scout_click calls with different `session`), which then run CONCURRENTLY rather than queueing. Omit for single-session sequential use.

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden and does so well: it names the write target (.scenescout/report.md), explains that the return is a deliberately bounded SUMMARY because full reports exceed client token limits, and discloses the refusal semantics of 'extensive' as an intentional completeness guarantee rather than an error.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The artifact and its contents are front-loaded before the gating rules, so the key information arrives early. It is dense and long, and the level-gating sentence packs three contracts into one clause, but almost every sentence carries non-redundant information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, and the description compensates by explaining exactly what comes back (a bounded summary, not the full document) and where the full artifact lands. For a 5-parameter tool with gating behavior, an agent has everything needed to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% so the baseline is 3, but the description adds real semantics on top: it explains that level must match the contract the run was asked for, that force is a budget-cap escape hatch, and that session enables concurrency across multiple dispatches. The report/history enums are largely restated from the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and artifact (generate the final markdown report) and enumerates the exact contents (findings, page quality scores, role capability matrix, oracle rollup, gap ledger), which cleanly separates it from siblings like scout_coverage, scout_lane_report, or scout_status that produce different outputs.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit gating conditions per level and states when force=true is legitimate ('only when the user capped the budget'), which is a genuine when-not constraint. It does not, however, point at any sibling alternative for the partial-report use cases, so it stops short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

scout_requestA

Call the app's own API as this session, with the UI bypassed โ€” the check that turns a hidden or disabled control into a proven refusal. A button that is not shown proves nothing; the same action refused by the server does. The fetch runs IN the page, so it carries the session's cookies and replays the Authorization header the app itself last sent, and it passes through the same interception the write policy is enforced on: in safe-write a mutation on a record this session did not create is refused here exactly as it would be for a click, and that refusal is the engine's safety net, not a finding. Returns the status line, the timing, the headers that decide whether two responses are truly identical (content-type, location, www-authenticate, retry-after, cache-control), and the body โ€” its first 2000 characters, or the part named by select (one JSON value by path) or offset/limit (a window of characters). Unlike a shell call, every request is recorded in the run's trail and its signature is what a finding should quote. Paths are fenced to the attached origin: use another session to reach another host.

ParametersJSON Schema
NameRequiredDescriptionDefault
bodyNoRequest body, sent as application/json unless a content-type header is given
pathYesPath on the attached origin, e.g. /api/things/12, or a full URL on that same origin
taskNoWhat you are DOING right now, in a few words: the action, not the acceptance criteria. "Filtering the documents register by status", "Filling the deviation form with invalid dates", "Signing in as QA_Team" โ€” NOT "ยง2.4 filtering narrows the set and the filter is reflected in the URL", which is what you are CHECKING, not what you are doing. Naming the item you are on is fine ("ยง2.4: filtering the documents register"); keep the rest to what a colleague would see over your shoulder. It stays set until you pass a different one, so a batch costs a few words, not one per call. Required on the tools that act unless a journey or an earlier call already set one.
limitNoHow many characters to return with offset or select. Default 2000, or 8000 for a select with no offset.
methodNoDefault GET
offsetNoReturn the body (or the selected value) from this character on. The body is cut at 2000 characters by default; the result names the next offset.
selectNoReturn one value of a JSON response body by its dotted path, e.g. "stats.open" or "items.0.name", pretty-printed and up to 8000 characters. A path that is not there says which keys are.
headersNoExtra headers. One given here wins over the app's own, which is how a session tests a different or absent credential.
sessionNoTarget this session directly instead of the active one โ€” pass it explicitly when dispatching to MULTIPLE sessions in one turn (e.g. two scout_click calls with different `session`), which then run CONCURRENTLY rather than queueing. Omit for single-session sequential use.
objectiveNoOld name for `task` (2.0). Prefer `task`.

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations exist, so the description carries the full burden and does so: the fetch runs in-page, carries session cookies, replays the Authorization header, and passes through write-policy interception โ€” with the key caveat that a safe-write refusal here is the engine's safety net, not a finding. It also discloses the return shape (status line, timing, differentiating headers, body truncated at 2000 chars) and that every request lands in the run trail.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core purpose and every sentence carries behavioral information (safety-net semantics, trail recording, origin fencing), but the run-on clause stacking makes it heavier than it needs to be. Information-dense rather than padded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With 10 parameters, no annotations and no output schema, the description compensates by describing the return payload, truncation/windowing mechanics, and the write-policy interaction. An agent has enough to invoke it correctly and interpret the result.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents path, method, body, select, offset, limit, headers and session. The description restates body truncation and select/offset behavior but adds little syntax or format meaning beyond what the schema provides, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a precise verb+resource ('Call the app's own API as this session, with the UI bypassed') and immediately frames the scope against the obvious sibling behavior (clicking a control), so an agent can distinguish it from scout_click/scout_network without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly says when to reach for it ('the check that turns a hidden or disabled control into a proven refusal. A button that is not shown proves nothing; the same action refused by the server does') and when not to ('use another session to reach another host'), plus a contrast with a shell call. Selection criteria are unambiguous.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

scout_resolveA

Mark a finding as resolved (by its id, shown when recorded and in the report). Resolved findings move to the report's green โœ… Resolved section, and reopen automatically as flagged REGRESSIONS if re-found later. Use when the user says a bug is fixed, or when re-testing shows the evidence no longer reproduces.

ParametersJSON Schema
NameRequiredDescriptionDefault
idNoAlias for `findingId`.
sessionNoTarget this session directly instead of the active one โ€” pass it explicitly when dispatching to MULTIPLE sessions in one turn (e.g. two scout_click calls with different `session`), which then run CONCURRENTLY rather than queueing. Omit for single-session sequential use.
findingIdNoFinding id, e.g. a1b2c3d4e5

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and meets it by disclosing behavioral consequences: resolved findings move to the green Resolved section and reopen automatically as REGRESSIONS if re-found later. This adds meaningful context beyond the schema, though it does not address potential side effects like irreversibility or permissions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is three concise sentences with no filler. It front-loads the core action, then explains behavior and usage conditions. Every sentence contributes useful information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers purpose, usage triggers, and behavioral outcomes, which is sufficient for a simple action tool with no output schema and no required parameters. It does not explicitly explain the return value, but given the tool's simplicity and the strong schema coverage, the remaining gap is minor.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the parameters are already well-documented. The description adds minor value by hinting where to find the id ('shown when recorded and in the report'), but it does not elaborate on parameter formats beyond the schema, hence a baseline score.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb-resource pair ('Mark a finding as resolved') and clearly identifies the target (by id). It distinguishes itself from sibling tools by explaining the outcome (moving to the Resolved section and auto-reopening as REGRESSIONS), which is unique to this tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly states when to use it: when the user says a bug is fixed, or when re-testing shows the evidence no longer reproduces. It does not provide when-not-to-use guidance or name alternatives, but the usage context is clear and actionable.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

scout_run_planA

Execute up to 20 actions in ONE call โ€” use for mechanical sequences (fill a form, walk a wizard) so each step doesn't cost a round-trip. Targets resolve at execution time by semantic locator: 'testid=โ€ฆ', 'text=โ€ฆ', 'label=โ€ฆ' or 'role=button[name=Save]' (never snapshot refs). The same steps, saved to .scenescout/flows/.json with expect-text / expect-url / expect-request steps added, are replayed by scenescout check on every pull request. An upload step attaches a file as scout_upload does (target required โ€” the file input or the control that opens its chooser; value = a fixture kind or a project-relative path). The plan ABORTS at the first NEW oracle violation, policy refusal, or failed step, returning a transcript of how far it got; repeats of already-reported violations do not abort (they stay logged for the report). For a sweep of independent steps (tabs, filters, pages) pass onViolation "continue": a new error status or its console echo is listed on its step's line and the plan goes on; a failed step, a policy refusal or any other violation still stops it. A select step whose value names no option fails at once, listing the options.

ParametersJSON Schema
NameRequiredDescriptionDefault
taskNoWhat you are DOING right now, in a few words: the action, not the acceptance criteria. "Filtering the documents register by status", "Filling the deviation form with invalid dates", "Signing in as QA_Team" โ€” NOT "ยง2.4 filtering narrows the set and the filter is reflected in the URL", which is what you are CHECKING, not what you are doing. Naming the item you are on is fine ("ยง2.4: filtering the documents register"); keep the rest to what a colleague would see over your shoulder. It stays set until you pass a different one, so a batch costs a few words, not one per call. Required on the tools that act unless a journey or an earlier call already set one.
stepsYes
sessionNoTarget this session directly instead of the active one โ€” pass it explicitly when dispatching to MULTIPLE sessions in one turn (e.g. two scout_click calls with different `session`), which then run CONCURRENTLY rather than queueing. Omit for single-session sequential use.
objectiveNoOld name for `task` (2.0). Prefer `task`.
onViolationNostop (default): end the plan at the first new oracle violation, right for a form flow whose steps depend on each other. continue: list a new http_error or console_error on its step's line and run the next step, for a sweep of independent stepsstop

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does so richly: it discloses abort-on-first-new-violation semantics, that repeated violations do not abort, policy-refusal and failed-step stops, the returned transcript, and the exact continue-mode behavior. It also specifies execution-time target resolution and select-step failure, which an agent could not infer otherwise.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core purpose and batching rationale, and each sentence conveys a distinct behavioral rule. It is dense and long, but the length is largely justified by the failure-mode detail; only minor tightening is possible.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex batched-action tool with no annotations and no output schema, the description covers abort conditions, per-step failure, continue-mode behavior, and the returned transcript. Return-value detail beyond 'a transcript of how far it got' is thin, but overall it is complete enough to invoke correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is already 80%, but the description adds real meaning: targets resolve at execution time by semantic locator and 'never snapshot refs', plus the upload value being a fixture kind or project-relative path. Most other param detail (onViolation, task) mirrors the schema, so it earns above the baseline 3 without being exhaustive.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and scope ('Execute up to 20 actions in ONE call') and immediately names the intended scenarios ('fill a form, walk a wizard'), which clearly separates it from the single-action siblings like scout_click and scout_type. An agent can tell what this does and when it beats multiple round-trips without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear when-to-use context: mechanical sequences vs a 'sweep of independent steps' with onViolation 'continue'. It explains the stop-vs-continue decision well, but never explicitly names a sibling alternative (e.g. 'for a single action use scout_click'), so the routing is implied rather than stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

scout_scanA

Scan a project directory to discover the frontend workspace, framework, routes, dev command, Playwright auth storage states, and testid conventions. Run this first.

ParametersJSON Schema
NameRequiredDescriptionDefault
projectPathYesAbsolute path to the project root

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the burden of behavioral disclosure. It conveys a read-only reconnaissance behavior by saying 'scan a project directory to discover...' but does not mention whether state is persisted, what the return shape looks like, or if any environmental prerequisites exist. This leaves some behavioral ambiguity.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single efficient sentence plus the useful sequencing directive 'Run this first.' It frontloads the action and follows with a concrete list of discovery targets, containing no filler or redundant repetition of the tool name.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-parameter tool with no annotations and no output schema, the description provides the essential call context: what to pass, what will be discovered, and that it should be invoked first. It does not explicitly describe the return format, but the enumerated discovery targets partially compensate and make the tool callable.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The only parameter, projectPath, is already fully documented in the schema as 'Absolute path to the project root,' and the description adds no additional constraints, defaults, or format details. The discovered artifacts listed in the description concern tool output rather than parameter semantics, so the schema remains the primary source of parameter meaning.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('scan') and object ('project directory') and enumerates the concrete artifacts it uncovers: frontend workspace, framework, routes, dev command, Playwright auth storage states, and testid conventions. This makes its role as a project discovery tool unmistakable and distinguishes it from execution-oriented siblings. The directive 'Run this first' also makes its purpose immediately clear.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly instructs the agent to run this tool first, establishing clear sequencing relative to the many scout_* sibling tools. It does not name specific alternatives or conditions for choosing another tool, but the 'first' directive gives the agent enough contextual guidance for initial invocation.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

scout_screenshotA

Take a JPEG screenshot of the current viewport. LAST RESORT: geometry issues are in scout_snapshot and style/contrast/spacing issues are in scout_design_audit โ€” images that failed to load are listed in scout_snapshot under BROKEN IMAGES โ€” use a screenshot only for pixel-native content (a canvas, visual gestalt) that computed data cannot capture.

ParametersJSON Schema
NameRequiredDescriptionDefault
sessionNoTarget this session directly instead of the active one โ€” pass it explicitly when dispatching to MULTIPLE sessions in one turn (e.g. two scout_click calls with different `session`), which then run CONCURRENTLY rather than queueing. Omit for single-session sequential use.

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the burden, and it does well by disclosing the JPEG format, viewport scope, and the 'LAST RESORT' nature of the tool. It could be slightly more explicit about side effects or what the caller receives, but the core behavior is transparent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is one dense, front-loaded sentence with no filler. It packs the main action first and the routing caveats after, though the long em-dash chain is slightly less scannable than a short two-sentence structure would be.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description is complete enough for selecting and invoking the tool: it says what is captured, in what format, and when not to use it. With no output schema, it slightly underspecifies the result format or delivery mechanism, but that is a minor gap for a simple screenshot tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the sole optional `session` parameter is fully documented in the schema. The description adds no parameter-specific detail, but none is required beyond what the schema already provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Take a JPEG screenshot of the current viewport.' It also clearly distinguishes itself from siblings by naming what belongs in scout_snapshot and scout_design_audit instead.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

This is exemplary routing guidance. It tells the agent to use this tool only for pixel-native content that computed data cannot capture, and explicitly directs geometry issues to scout_snapshot, style/contrast/spacing issues to scout_design_audit, and broken-image checks to the BROKEN IMAGES section of scout_snapshot.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

scout_scrollA

Scroll like a user โ€” real apps hide their bugs below the fold. Reports the resulting position (px and %), and explicitly flags SCROLL LOCKED: scrollable content exists but the page will not move (the classic leaked modal scroll-lock that silently cuts users off from everything below the fold โ€” snapshots also detect this passively as an OVERLAY line). Without target it scrolls the page, falling back to the largest scrollable pane on app-shell layouts. Pass target to scroll ONE region instead (a sidebar nav, a dialog body, a table pane): the page-level pick is the LARGEST scroll port, so a smaller region beside it never moves and its content looks truncated when it is only scrolled away โ€” never call a nav item missing without scrolling its own container first. Use before judging a long page: the design audit measures at the current scroll position, so scroll + re-snapshot/re-audit deep sections; scroll also triggers lazy-loaded content whose failures then surface as oracle violations.

ParametersJSON Schema
NameRequiredDescriptionDefault
byNoScroll by px instead (positive = down). Default 600 when neither given.
toNoJump to an edge
targetNoScroll ONE region instead of the page: "testid=โ€ฆ", "text=โ€ฆ" or "label=โ€ฆ". Scrolls that element's nearest scrollable ancestor.
sessionNoTarget this session directly instead of the active one โ€” pass it explicitly when dispatching to MULTIPLE sessions in one turn (e.g. two scout_click calls with different `session`), which then run CONCURRENTLY rather than queueing. Omit for single-session sequential use.

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full behavioral burden, and it is unusually rich: it discloses the fallback to the largest scrollable pane on app-shell layouts, that smaller regions beside it never move, that it reports SCROLL LOCKED, and that scrolling triggers lazy-loaded content. This goes well beyond what the schema alone reveals.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long and dense, but nearly every clause carries behavioral or workflow value, and it is front-loaded with the core action. Some rhetorical framing ('real apps hide their bugs below the fold') could be trimmed, and the parenthetical caveats make it a wall of text rather than a scannable definition, but it is efficient for the complexity it documents.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given 4 optional parameters, no output schema, and no annotations, the description covers the essential missing context: result format (px and %), SCROLL LOCKED detection, fallback behavior, target-region semantics, and why scrolling matters for audits and lazy-loaded content. The schema handles the remaining parameter details, so nothing critical is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. The description adds real value on top by explaining the target semantics in depth โ€” single region, nearest scrollable ancestor, and the 'smaller region beside it' trap. It does not redundantly restate by/to/session, leaving the schema to cover those.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific behavior ('Scroll like a user'), the resource acted on (page or one target region), and the outcome (reports position in px and %, flags SCROLL LOCKED when scroll is blocked). It clearly differentiates itself from sibling tools like scout_snapshot and scout_design_audit by being the scroll-position-changing action that precedes them.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives explicit workflow guidance: use before judging a long page, then re-snapshot/re-audit deep sections, because the design audit measures at the current scroll position. It also warns against calling a nav item missing without scrolling its own container first, which is a concrete when-to-use versus when-to-target instruction.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

scout_selectA

Select an option in a by ref. The value is matched against the options before anything is picked: an exact value, an exact label, either ignoring case, then a label it starts with. A value matching no option, or several, is refused at once with the options listed.

ParametersJSON Schema
NameRequiredDescriptionDefault
refYes
taskNoWhat you are DOING right now, in a few words: the action, not the acceptance criteria. "Filtering the documents register by status", "Filling the deviation form with invalid dates", "Signing in as QA_Team" โ€” NOT "ยง2.4 filtering narrows the set and the filter is reflected in the URL", which is what you are CHECKING, not what you are doing. Naming the item you are on is fine ("ยง2.4: filtering the documents register"); keep the rest to what a colleague would see over your shoulder. It stays set until you pass a different one, so a batch costs a few words, not one per call. Required on the tools that act unless a journey or an earlier call already set one.
valueYesOption value or label (or the start of a label, when only one option has it)
sessionNoTarget this session directly instead of the active one โ€” pass it explicitly when dispatching to MULTIPLE sessions in one turn (e.g. two scout_click calls with different `session`), which then run CONCURRENTLY rather than queueing. Omit for single-session sequential use.
objectiveNoOld name for `task` (2.0). Prefer `task`.

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does well: it discloses the ordered matching algorithm (exact value, exact label, case-insensitive, prefix) and the failure mode (ambiguous or unmatched values are refused immediately with the option list). It omits success behavior and whether it dispatches change events or waits for options to populate.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, front-loaded with the action and then the resolution rules, with no filler. The middle sentence is dense but each clause carries distinct information about matching order.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a no-output-schema, no-annotation tool the description covers the critical decision logic and error behavior an agent needs to call it correctly. Minor gaps remain around success signaling and whether the element must be visible, but nothing essential is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 80%, so baseline is 3, and the description genuinely adds meaning for `value` by spelling out the match precedence and the ambiguity/prefix behavior that the schema's one-line note only hints at. The ref/task/session/objective params get no added explanation, but they are covered by the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Select an option in a <select> by ref'), which is inherently distinct from scout_click/scout_type because it targets a native select and uses option matching rather than coordinates or text entry. It never names a sibling to contrast against, so it stops short of full sibling differentiation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied by the matching rules (you reach for this when you have an option value or label), but the description never states when to prefer it over scout_click on an option, nor any preconditions such as the select needing to be present/visible. No explicit when/when-not guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

scout_sessionA

List live sessions, or set which one is the DEFAULT (used by any tool call that omits session). Prefer passing session directly on each tool call for multi-role work โ€” that's what lets concurrent dispatch happen; scout_session is for sequential convenience (skip repeating session on every call) and for checking what's live. Both browsers stay live and authenticated regardless of which is default โ€” re-snapshot a session after a break to see what changed while it was away.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameNoSession to make the default; omit to list sessions
paceMsNoChange how fast this session acts, mid-run: a floor between actions in milliseconds, for when a person is watching and needs to keep up. 0 restores full speed. With `name`, applies to that session; without, to every live session โ€” which is what 'slow everything down so I can follow' means.
sessionNoAlias for `name`.

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations supplied, the description carries full behavioral burden and meets it: it discloses default-session persistence, that switching default does not deauthenticate other sessions ('Both browsers stay live and authenticated'), and the effect of paceMs on live sessions. It also hints at re-snapshotting after a break, adding operational behavior beyond the schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three dense sentences, front-loaded with the core action, then usage guidance, then behavioral caveats. Every sentence earns its place and no tautology or padding.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 3-optional-parameter tool with no output schema, this is nearly complete: purpose, default behavior, persistence, and speed control are covered. It does not describe the shape of the listing result or edge cases like unknown session names, but the stated behavior is enough for an agent to decide confidently.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% so the schema already documents each parameter. The description still adds value by explaining that the default is what is used when a tool call omits session, and by contrasting direct session passing with default-based convenience. Minor: it could restate paceMs semantics, but schema already covers them.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description opens with a precise dual function: 'List live sessions, or set which one is the DEFAULT' and explains what DEFAULT means. This distinguishes scout_session from sibling action tools like scout_snapshot or scout_navigate while making its resource ('sessions') explicit.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit decision guidance: prefer passing session on each call for multi-role/concurrent work, and use scout_session for sequential convenience or checking what's live. It names the context that should steer an agent away from this tool, which is exactly what a usage guideline should do.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

scout_snapshotA

Capture the current page state: URL, state fingerprint, a one-line summary of the main area's heading and text, interactable elements with refs (e1, e2, โ€ฆ) and their state (pressed, selected, checked, expanded, current), what the page announces (alert and status regions, by their text), geometry issues, coverage, and oracle violations since the last action. Re-snapshotting a route returns a DIFF against its last snapshot, even after visiting elsewhere, and another tab of the same screen diffs against that screen's last tab; refs stay stable, including across a search or filter that rewrites only the query string. On a dense page, says what past the element cap was cut. Cheap โ€” prefer this over screenshots.

ParametersJSON Schema
NameRequiredDescriptionDefault
fullNoForce a full element list instead of a diff
sessionNoTarget this session directly instead of the active one โ€” pass it explicitly when dispatching to MULTIPLE sessions in one turn (e.g. two scout_click calls with different `session`), which then run CONCURRENTLY rather than queueing. Omit for single-session sequential use.

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden and does so richly: it discloses the diff semantics (re-snapshot returns a DIFF even after visiting elsewhere, per-tab diffing), ref stability across query-string rewrites, truncation behavior past the element cap, and relative cost. These are non-obvious behaviors an agent could not infer from the name or schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core purpose, and every clause carries information (return fields, diff behavior, ref stability, cost). It is dense and somewhat run-on in the middle, which slightly hurts scannability, but there is little wasted text.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, so the description must describe the return payload โ€” and it does, itemizing exactly what a snapshot contains (refs, states, announcements, geometry, coverage, oracle violations) plus diff/truncation behavior. An agent has everything needed to call and interpret it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both `full` and `session` are already documented in the schema, including the concurrency caveat. The description adds no format/semantic detail beyond that, so the baseline of 3 is correct when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Specific verb ("Capture") plus a fully enumerated resource: URL, fingerprint, summary, interactable elements with refs, announcements, geometry, coverage, oracle violations. It explicitly positions itself against a sibling ("prefer this over screenshots"), so an agent can distinguish it from scout_screenshot/scout_capture without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Clear usage context: it names the alternative (screenshots) and states the cost tradeoff, and it explains the multi-session concurrency case for the `session` param. It lacks explicit when-not guidance against the many other scout_* capture/verify siblings, but the routing signal is strong.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

scout_statusRun statusA

Show the person running you how the run stands: each session's objective and current task, open findings by severity, coverage, and the live view's address. Call it when the user wants to watch or asks how the run is going. A host that renders MCP Apps shows a pane that keeps itself up to date; every other host gets the same as text, and you pass the Live view: address on. Takes no input and touches no browser.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does well: it states it takes no input, touches no browser, and discloses host-dependent rendering (MCP Apps pane vs. plain text) plus the instruction to pass on the 'Live view:' address. It stops short of explicitly declaring the operation read-only/idempotent, which would fully close the gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, front-loaded with what the output contains before moving to when to call it and how it behaves across hosts. Every sentence carries information, though the phrasing ('the person running you how the run stands') is slightly indirect for a zero-param tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema and no annotations, so the description must supply the return shape itself, and it does: sessions, objectives, tasks, findings by severity, coverage, and the live-view address. An agent has everything it needs to call this correctly and describe the result.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Zero parameters, so the baseline is 4. The description reinforces this with 'Takes no input,' which matches the empty schema and prevents an agent from inventing arguments.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a concrete resource (the run's status) and enumerates exactly what it reports: per-session objective and current task, open findings by severity, coverage, and the live-view address. It implicitly separates itself from the browser-action siblings via 'touches no browser,' but never distinguishes itself from close cousins like scout_status_poll, scout_coverage, or scout_report.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives an explicit trigger: 'Call it when the user wants to watch or asks how the run is going.' That is clear usage context, but there is no when-not guidance and no named alternative for the overlapping status/coverage/report tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

scout_status_pollRun status (for the pane)A

Called by the scout_status pane every few seconds to refresh itself. Not for the agent: call scout_status instead.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations present, the description carries the full burden, and it does disclose real behavioral traits: the invocation cadence (every few seconds) and the intended consumer (a UI pane, not the agent). It stops short of describing the return payload or whether polling is cheap/idempotent, but the caller constraint is the behavior that matters most here.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, zero filler, and the audience constraint is front-loaded immediately after the mechanism. Nothing is padded or buried.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter, output-schema-less UI polling endpoint, the description supplies the one thing an agent needs: do not call this, call scout_status. Return-shape details are absent but legitimately irrelevant to the agent audience the description has already excluded.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so the schema already fully covers its surface and there is nothing for the description to disambiguate. Baseline for a parameterless tool.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names the specific caller (the scout_status pane) and the specific action (refreshing run status), so the resource and operation are identifiable. It is also clearly distinguishable from the sibling scout_status, though it never spells out what payload the poll actually returns.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives an explicit negative use case ('Not for the agent') plus a named alternative ('call scout_status instead'). An agent has everything it needs to route away from this tool with no inference required.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

scout_ticketsA

Read the tickets or acceptance criteria the person gave this run, pasted (text) or from a file (path), and keep them so the report answers each criterion: passed, failed or not tested. Recognises Given/When/Then scenarios, checklists, numbered and "AC1:" criteria, and lists under an "Acceptance criteria" heading; several tickets may be given at once. A ticket with no recognisable criteria is reported as such, never guessed at. Returns each criterion's id (AC1, AC2, โ€ฆ) to judge it by with scout_criterion. Touches no browser.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathNoA ticket file to read (.md, .markdown, .txt, .text, .feature), absolute or relative to the project folder. Pass this or `text`.
textNoThe tickets as pasted. Pass this or `path`.
sessionNoTarget this session directly instead of the active one โ€” pass it explicitly when dispatching to MULTIPLE sessions in one turn (e.g. two scout_click calls with different `session`), which then run CONCURRENTLY rather than queueing. Omit for single-session sequential use.

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does well: it lists the recognized formats (Given/When/Then, checklists, numbered, "AC1:", "Acceptance criteria" heading), notes multiple tickets may be given, states unrecognized tickets are reported rather than guessed, and discloses "Touches no browser." It omits auth/rate-limit details, but for a local parsing tool that is minor.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is a dense single paragraph that is front-loaded with the core action (read and retain criteria) before the parsing details. Slightly long, but every clause conveys a distinct behavior rather than padding.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, yet the description explains what is returned (criterion ids AC1, AC2, โ€ฆ) and the not-recognized case, so an agent knows what to expect. Combined with the named downstream tool, this is essentially complete for a parsing/ingest tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all three parameters, giving a baseline of 3. The description restates that criteria come from `text` or `path` but adds no format or size detail beyond what the schema provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Specific verb+resource: it reads tickets/acceptance criteria from `text` or `path` and stores them for the report. It clearly distinguishes itself from siblings by naming scout_criterion as the downstream tool used to judge each returned id, so an agent can place it in the workflow without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It states the context clearly ("the tickets or acceptance criteria the person gave this run") and routes the follow-up to scout_criterion. It does not, however, state when NOT to use it or contrast it against nearby capture tools like scout_playbook/scout_note, leaving that to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

scout_typeA

Type into a text input/textarea/composer by ref, the way a real user does: if the field already holds content (e.g. an @-mention chip a menu click inserted), the text is APPENDED at the end โ€” preserving that content โ€” and the result reports what was already there (a separating space is added only at a word-to-word boundary). Pass replace=true to clear the field first (correcting a previous entry); an empty textValue always clears. Appending fires input events but not keydown, so keydown-driven triggers (slash/mention menus) will not react to appended text. Use for both valid values and boundary/fuzz values (empty, very long, unicode, script tags).

ParametersJSON Schema
NameRequiredDescriptionDefault
refYesElement ref, e.g. e12
taskNoWhat you are DOING right now, in a few words: the action, not the acceptance criteria. "Filtering the documents register by status", "Filling the deviation form with invalid dates", "Signing in as QA_Team" โ€” NOT "ยง2.4 filtering narrows the set and the filter is reflected in the URL", which is what you are CHECKING, not what you are doing. Naming the item you are on is fine ("ยง2.4: filtering the documents register"); keep the rest to what a colleague would see over your shoulder. It stays set until you pass a different one, so a batch costs a few words, not one per call. Required on the tools that act unless a journey or an earlier call already set one.
valueNoAlias for `textValue`.
replaceNoClear the field before typing instead of appending to existing content
sessionNoTarget this session directly instead of the active one โ€” pass it explicitly when dispatching to MULTIPLE sessions in one turn (e.g. two scout_click calls with different `session`), which then run CONCURRENTLY rather than queueing. Omit for single-session sequential use.
objectiveNoOld name for `task` (2.0). Prefer `task`.
textValueNoText to type
pressEnterNoPress Enter after typing

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full disclosure burden and exceeds it. It details append semantics, the separating-space rule, empty-text clearing, replace behavior, the exact event side effects (input fires, keydown does not), and the fact that the result reports pre-existing content. This is exemplary behavioral documentation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but every clause earns its place: purpose, append behavior, replace semantics, clearing rule, event caveat, and usage scope. It is front-loaded with the core action and then systematically covers edge cases. Length is justified by the behavioral complexity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with no annotations and no output schema, the description covers the full operational surface: how content is entered, when it is cleared, what side effects occur, and what the result reports. Nothing an agent needs to invoke it correctly or anticipate its behavior is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. The description adds meaningful context beyond the schema by explaining the interaction between textValue, replace, and empty-string clearing, and by clarifying which event behaviors apply. It enriches the parameters without restating their schema descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb ('Type') and resource ('text input/textarea/composer by ref'), and immediately distinguishes the tool's core behavior โ€” appending vs. replacing โ€” which separates it from sibling input tools like scout_click or scout_press. The purpose is unambiguous and action-oriented.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly states a broad use case ('valid values and boundary/fuzz values') and gives a critical constraint (keydown-driven triggers won't react to appended text), which effectively tells the agent when behavior may not match expectations. It does not name a sibling alternative (e.g., 'use scout_press for keydown-sensitive actions'), but the guidance is sufficiently clear for selection.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

scout_uploadA

Attach a file to an upload control the way a user does. ref is either a visible (snapshots list these with role file) or the button/label/dropzone that opens the file chooser โ€” the chooser is intercepted and answered, which is how the hidden input behind a styled 'Choose file' control is reached. Omit ref to target the page's only file input, hidden or not (snapshots disclose hidden ones on a FILE INPUTS line). Nothing needs to exist on disk: a small VALID fixture (real PDF/PNG structure) is generated in memory, its kind inferred from the input's accept attribute or chosen with fixture; filePath uploads a real file but must live inside the attached project (fenced like navigation is fenced to the origin); name overrides the filename for boundary tests (wrong extension vs accept, very long, unicode). The result names the input, how the file reached it, flags a file that violates accept (a mismatch the app then accepts is a validation finding), warns if the app cleared the input after selection, and says whether a state-changing request fired on selection โ€” if none did, click the form's submit, or check the next snapshot for a client-side rejection.

ParametersJSON Schema
NameRequiredDescriptionDefault
refNoElement ref of the file input OR of the control that opens the file chooser; omit when the page has exactly one file input
nameNoFilename override (default scenescout-fixture.<kind>, or the disk file's own name)
taskNoWhat you are DOING right now, in a few words: the action, not the acceptance criteria. "Filtering the documents register by status", "Filling the deviation form with invalid dates", "Signing in as QA_Team" โ€” NOT "ยง2.4 filtering narrows the set and the filter is reflected in the URL", which is what you are CHECKING, not what you are doing. Naming the item you are on is fine ("ยง2.4: filtering the documents register"); keep the rest to what a colleague would see over your shoulder. It stays set until you pass a different one, so a batch costs a few words, not one per call. Required on the tools that act unless a journey or an earlier call already set one.
fixtureNoGenerated fixture kind; default: inferred from the input's accept attribute (pdf when there is none, or none we can generate)
sessionNoTarget this session directly instead of the active one โ€” pass it explicitly when dispatching to MULTIPLE sessions in one turn (e.g. two scout_click calls with different `session`), which then run CONCURRENTLY rather than queueing. Omit for single-session sequential use.
filePathNoA real file to upload โ€” absolute or relative to the project; must be inside the attached project. Exclusive with fixture.
objectiveNoOld name for `task` (2.0). Prefer `task`.

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries full behavioral burden and succeeds remarkably. It discloses the chooser interception, in-memory fixture generation, project-fencing of `filePath`, accept-violation flagging, warning if the input is cleared, and detection of state-changing requests. This goes far beyond a simple 'uploads a file' summary.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long but dense with operationally relevant detail; nearly every clause adds value. It is front-loaded with the core purpose and then proceeds from ref targeting to fixture/file selection to post-selection behavior. A few parentheticals could be trimmed, but the length is justified by the tool's complexity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity, zero annotations, and no output schema, the description is remarkably complete. It covers the input-targeting strategy, file source options, validation behavior, state-change detection, and necessary follow-up actions. An agent has enough context to invoke the tool correctly and interpret its result without additional structured metadata.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, but the description adds substantial meaning beyond the schema: `ref` resolution through styled controls, fixture kind inference from the accept attribute, project-fencing semantics for `filePath`, and `name` use for boundary tests. It also explains the result fields tied to these parameters, making parameter behavior concrete.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Attach a file to an upload control the way a user does.' It clearly distinguishes the tool's scope from general file operations by detailing `ref` targeting of file inputs, hidden inputs, and chooser-opening controls. The distinctive behavior of intercepting the chooser makes it unambiguous among sibling tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives extensive usage context: when to omit `ref`, how to distinguish fixture-based uploads from `filePath` uploads, and what to do after a selection that fires no state-changing request. It does not explicitly name alternative sibling tools or state when not to use this tool, but the parameter-level guidance is clear enough for correct invocation.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

scout_verifyA

Re-test findings earlier runs left open. With no arguments, returns the open findings in the order to re-test them โ€” worst route first, grouped so a route is walked once โ€” each with its evidence and repro steps. Pass ids to narrow it to specific findings. After re-testing one, call again with id and verdict to record what you saw: "gone" resolves it, "present" stamps it confirmed so the report stops calling it unverified, "changed" keeps it open and says the behaviour differs. Use after a fix wave, or at the start of a run against an app this project has tested before.

ParametersJSON Schema
NameRequiredDescriptionDefault
idNoThe finding being verified. Omit to get the worklist.
idsNoNarrow the worklist to these finding ids.
noteNoWhat you saw, in a sentence. Shown in the report beside the verdict.
sessionNoTarget this session directly instead of the active one โ€” pass it explicitly when dispatching to MULTIPLE sessions in one turn (e.g. two scout_click calls with different `session`), which then run CONCURRENTLY rather than queueing. Omit for single-session sequential use.
verdictNoWhat the re-test found: "gone", "present" or "changed". Requires id.

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden, and it delivers: it discloses ordering (worst route first), grouping, returned content (evidence and repro steps), and the state effects of each verdictโ€”'gone' resolves, 'present' confirms, 'changed' keeps open. This is unusually transparent about side effects.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Every sentence earns its place: the tool's action, the worklist behavior, narrowing, verdict recording, and usage timing. The information is front-loaded with the core purpose and flows logically without repetition or fluff.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with five parameters, no annotations, and no output schema, the description is remarkably complete. It covers what the worklist returns, how to narrow it, how to record verdicts, and what each verdict means, leaving no critical gap for an agent trying to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. The description adds meaning beyond the schema by explaining the interaction between id, ids, and verdictโ€”including that a verdict is recorded by calling again with id and verdictโ€”and clarifies what the verdict values do to report state.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: re-test findings left open by previous runs and record verdicts. It distinguishes itself from siblings like scout_scan or scout_resolve by describing a specific verification workflow with 'gone', 'present', and 'changed' outcomes.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives explicit when-to-use context: 'Use after a fix wave, or at the start of a run against an app this project has tested before.' It also explains the call pattern for listing, narrowing, and recording verdicts. It does not mention alternatives or when-not-to-use, so it stops short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 1 tool updatev3.18.0
    • Addedscout_intake
  2. 17 tool updatesv3.17.0
    • Changedscout_attach9 fields changed
      • changedInput schema / properties / browser / description
        Previous value: -"Browser to drive. Default: the SCENESCOUT_BROWSER environment variable, else chromium. firefox and webkit must be downloaded first (scenescout install --browser-only --browsers firefox). Use them for a cross-browser pass; stay on chromium otherwise."New value: +"Browser to drive. Default: the SCENESCOUT_BROWSER environment variable, else chromium. A build that is not on disk is downloaded on this attach, once, except in CI or with SCENESCOUT_BROWSER_DOWNLOAD=off, where the attach names the command to run. Use firefox or webkit for a cross-browser pass; stay on chromium otherwise."
      • addedInput schema / properties / dedup
        Added value: +{
        +  "description": "How this run tells a filed finding from one already recorded. Default: the SCENESCOUT_DEDUP environment variable, else 'rule', the store's rule alone. 'judge': the rule first, then, for a filing the rule keeps apart from everything recorded, a model is asked whether it is one of the open findings on its page, and merges it when it says so. It needs ANTHROPIC_API_KEY or OPENAI_API_KEY in the server's environment, and sends each pair's titles, categories and evidence, and the page's path, to that provider โ€” ONLY when the user asked for it. Applies to every session of the project until the run ends.",
        +  "enum": [
        +    "rule",
        +    "judge"
        +  ],
        +  "type": "string"
        +}
      • addedInput schema / properties / evidence
        Added value: +{
        +  "description": "What happens to the picture each scout_finding takes of what it is about. 'inline': kept under .scenescout/recordings/, shown in report.html, and returned in the scout_finding result so the conversation shows it. 'file': kept and shown in the report only. 'off': none taken. Default: the SCENESCOUT_EVIDENCE environment variable, else 'inline', or 'file' in a CI job. Pictures are bounded in size and in how many one session returns (see the configuration reference).",
        +  "enum": [
        +    "inline",
        +    "file",
        +    "off"
        +  ],
        +  "type": "string"
        +}
      • addedInput schema / properties / open
        Added value: +{
        +  "description": "Open the live view (on this attach) and/or report.html (when scout_report writes it) in the user's default browser: 'live', 'report', 'both' or 'none'. Default: the SCENESCOUT_OPEN environment variable, else 'both' on a local desktop session (headed or headless) and 'none' in CI, over SSH, or with no display. Pass it only when the user asked for something other than the default.",
        +  "enum": [
        +    "live",
        +    "report",
        +    "both",
        +    "none"
        +  ],
        +  "type": "string"
        +}
      • changedInput schema / properties / projectPath / description
        Previous value: -"Absolute path to the project (memory + report live in .scenescout/ here)"New value: +"Absolute path to the project (memory + report live in .scenescout/ here). Pass it whenever you have a project or working folder. Omitted: the client's workspace folder, else a folder per tested site under the user's documents folder (Documents/SceneScout/<host>/, or SCENESCOUT_PROJECTS_DIR), which the result names โ€” tell the user where it is."
      • addedInput schema / properties / readPosts
        Added value: +{
        +  "description": "POST endpoints that only read, e.g. [\"POST /api/search\", \"POST https://api.example.com/reports/query\"], which observe mode then lets out โ€” ONLY when the user named them. Never add one yourself, even when the gap ledger lists a refused POST: ask the user. Exact paths; * stands for one path segment. Still refused when the path or body looks destructive or the body is a GraphQL mutation. Observe mode only. Default: the SCENESCOUT_READ_POSTS environment variable, else none.",
        +  "items": {
        +    "maxLength": 300,
        +    "type": "string"
        +  },
        +  "maxItems": 20,
        +  "type": "array"
        +}
      • removedInput schema / properties / record / default
        Removed value: -false
      • changedInput schema / properties / record / description
        Previous value: -"Keep a frame of the page after every action, under .scenescout/recordings/, and show it beside that step in report.html. Off by default: a recording is pictures of the app under test sitting in the project folder. Turn it on for QA work, where the run is evidence and not only a report."New value: +"Keep a frame of the page after every action, under .scenescout/recordings/, and show it beside that step in report.html. Default: the SCENESCOUT_RECORD environment variable (on or off), else off: a recording is pictures of the app under test sitting in the project folder. Turn it on for QA work, where the run is evidence and not only a report."
      • changedInput schema / required
        Previous value: -[
        -  "url",
        -  "projectPath"
        -]New value: +[
        +  "url"
        +]
    • Changedscout_back1 field changed
      • addedInput schema / properties / leave
        Added value: +{
        +  "description": "How to answer if the page asks to confirm leaving (a beforeunload prompt over unsent input): true leaves and discards that input, false stays. Omitted, observe and read-only stay and other modes leave. The result says when the page asked.",
        +  "type": "boolean"
        +}
    • Changedscout_click1 field changed
      • addedInput schema / properties / leave
        Added value: +{
        +  "description": "How to answer if the page asks to confirm leaving (a beforeunload prompt over unsent input): true leaves and discards that input, false stays. Omitted, observe and read-only stay and other modes leave. The result says when the page asked.",
        +  "type": "boolean"
        +}
    • Changedscout_coverage1 field changed
      • addedInput schema / properties / scope
        Added value: +{
        +  "description": "'session': only the routes this session reached this run and the forms it saw. 'project': every route in the memory, across runs and sessions. Default: 'session' when other sessions share this project, else 'project'.",
        +  "enum": [
        +    "session",
        +    "project"
        +  ],
        +  "type": "string"
        +}
    • Changedscout_crawl1 field changed
      • changedInput schema / properties / paths / description
        Previous value: -"Paths to visit, e.g. ['/orders','/settings']. Omit to crawl all unvisited known routes."New value: +"Paths to visit, e.g. ['/orders','/settings'], or full URLs on the attached origin. A path resolves from the origin's root, whatever page the session attached on. Omit to crawl all unvisited known routes."
    • Addedscout_criterion
    • Changedscout_finding1 field changed
      • addedInput schema / properties / ref
        Added value: +{
        +  "description": "The ref, from the latest scout_snapshot, of the element the finding is about. Its picture (the element plus a margin) is kept with the finding and shown in report.html. Omit it and the picture is the viewport.",
        +  "maxLength": 20,
        +  "type": "string"
        +}
    • Addedscout_login
    • Changedscout_navigate1 field changed
      • addedInput schema / properties / leave
        Added value: +{
        +  "description": "How to answer if the page asks to confirm leaving (a beforeunload prompt over unsent input): true leaves and discards that input, false stays. Omitted, observe and read-only stay and other modes leave. The result says when the page asked.",
        +  "type": "boolean"
        +}
    • Addedscout_network
    • Changedscout_report1 field changed
      • addedInput schema / properties / report
        Added value: +{
        +  "default": "both",
        +  "description": "Which parts the report carries. 'both' (default): a plain-language section first โ€” a short summary, then each problem with numbered steps, what was expected, what happened, its picture and its impact (blocks users, annoying, cosmetic), each with its technical detail folded beneath โ€” followed by the technical report. 'qa': the plain section alone, for a tester or anyone not technical. 'dev': the technical report alone, as before the plain section existed. It applies to the files this call writes; the live view always shows both.",
        +  "enum": [
        +    "both",
        +    "qa",
        +    "dev"
        +  ],
        +  "type": "string"
        +}
    • Changedscout_request3 fields changed
      • addedInput schema / properties / limit
        Added value: +{
        +  "description": "How many characters to return with offset or select. Default 2000, or 8000 for a select with no offset.",
        +  "maximum": 8000,
        +  "minimum": 1,
        +  "type": "integer"
        +}
      • addedInput schema / properties / offset
        Added value: +{
        +  "description": "Return the body (or the selected value) from this character on. The body is cut at 2000 characters by default; the result names the next offset.",
        +  "maximum": 1000000,
        +  "minimum": 0,
        +  "type": "integer"
        +}
      • addedInput schema / properties / select
        Added value: +{
        +  "description": "Return one value of a JSON response body by its dotted path, e.g. \"stats.open\" or \"items.0.name\", pretty-printed and up to 8000 characters. A path that is not there says which keys are.",
        +  "maxLength": 500,
        +  "type": "string"
        +}
    • Changedscout_run_plan1 field changed
      • addedInput schema / properties / onViolation
        Added value: +{
        +  "default": "stop",
        +  "description": "stop (default): end the plan at the first new oracle violation, right for a form flow whose steps depend on each other. continue: list a new http_error or console_error on its step's line and run the next step, for a sweep of independent steps",
        +  "enum": [
        +    "stop",
        +    "continue"
        +  ],
        +  "type": "string"
        +}
    • Changedscout_select1 field changed
      • changedInput schema / properties / value / description
        Previous value: -"Option value or label"New value: +"Option value or label (or the start of a label, when only one option has it)"
    • Addedscout_status
    • Addedscout_status_poll
    • Addedscout_tickets
  3. 4 tool updatesv3.14.1
    • Changedscout_attach4 fields changed
      • addedInput schema / properties / actionTimeoutMs
        Added value: +{
        +  "description": "How long one click, keystroke, hover or pick may take, in ms. Default: the SCENESCOUT_ACTION_TIMEOUT_MS environment variable, else 5000. Raise it only when timeouts come from a loaded machine rather than the app.",
        +  "maximum": 120000,
        +  "minimum": 1000,
        +  "type": "integer"
        +}
      • addedInput schema / properties / navTimeoutMs
        Added value: +{
        +  "description": "How long a page may take to load, in ms. Default: the SCENESCOUT_NAV_TIMEOUT_MS environment variable, else 20000 (crawled pages 15000; a value set here applies to them too). Raise it only when timeouts come from a loaded machine rather than the app.",
        +  "maximum": 300000,
        +  "minimum": 1000,
        +  "type": "integer"
        +}
      • addedInput schema / properties / role
        Added value: +{
        +  "description": "Sign in as a role whose login was saved with `scenescout login <url> --role <name>` (kept in the project's .scenescout/auth/). Every session given the same role gets its own browser built from that one login. Not with storageStatePath. No saved login for the role: the attach is refused and names the command to run โ€” ask the user to run it, since it opens a window for them to sign in.",
        +  "maxLength": 40,
        +  "type": "string"
        +}
      • changedInput schema / properties / storageStatePath / description
        Previous value: -"Optional Playwright storage-state JSON path for authenticated exploration"New value: +"Optional Playwright storage-state JSON path for authenticated exploration. Not with `role`."
    • Addedscout_capture
    • Changedscout_finding1 field changed
      • addedInput schema / properties / convention
        Added value: +{
        +  "description": "Only for a WORTH-A-LOOK finding: the observation is real, and it is a defect only under a convention of this project you cannot see. Name that convention, e.g. 'a 4px spacing scale' or 'test ids on every control'. The report lists it under \"Worth a look\", apart from the defects, and does not count it as one. Not for \"I could not tell\": leave that unfiled or look closer. Omit for a defect.",
        +  "maxLength": 160,
        +  "minLength": 1,
        +  "type": "string"
        +}
    • Changedscout_lane_brief2 fields changed
      • addedInput schema / properties / expiryMarginMinutes
        Added value: +{
        +  "description": "How long past the run a saved role's login must still last, in minutes (default 10).",
        +  "maximum": 240,
        +  "minimum": 0,
        +  "type": "integer"
        +}
      • addedInput schema / properties / runMinutes
        Added value: +{
        +  "description": "How long the lanes will run, in minutes (default 60). When they attach by a saved role, the brief is refused if that login will not last this long plus the margin.",
        +  "maximum": 1440,
        +  "minimum": 1,
        +  "type": "integer"
        +}
  4. 1 tool updatev3.11.1
    • Changedscout_run_plan1 field changed
      • changedInput schema / properties / steps / items / properties / target / description
        Previous value: -"testid=โ€ฆ, text=โ€ฆ, label=โ€ฆ (or a path for navigate; 'top'/'bottom'/ยฑpx for scroll; for upload: the file input or the control that opens its chooser)"New value: +"testid=โ€ฆ, text=โ€ฆ, label=โ€ฆ, role=<role>[name=\"โ€ฆ\"] (or a path for navigate; 'top'/'bottom'/ยฑpx for scroll; for upload: the file input or the control that opens its chooser)"
  5. 1 tool updatev3.10.0
    • Changedscout_close1 field changed
      • addedInput schema / properties / force
        Added value: +{
        +  "default": false,
        +  "description": "Close even a lane whose report has not been folded yet. Its decisions are then lost: nothing is kept for calibration and nothing checks its defects were filed. Prefer folding it with scout_lane_report first.",
        +  "type": "boolean"
        +}
  6. 1 tool updatev3.9.0
    • Changedscout_attach1 field changed
      • addedInput schema / properties / trustedEmbeds
        Added value: +{
        +  "description": "Origins of embedded frames (e.g. \"https://pay.example.com\") whose writes out of the app may go out โ€” ONLY when the user named them, typically a provider in test mode, and only in safe-write mode. Never add one yourself. Hostile input, repeated-click probes and uploads stay refused in them.",
        +  "items": {
        +    "maxLength": 200,
        +    "type": "string"
        +  },
        +  "maxItems": 10,
        +  "type": "array"
        +}
  7. 16 tool updatesv3.4.0
    • Changedscout_attach4 fields changed
      • addedInput schema / properties / objective
        Added value: +{
        +  "description": "This session's objective: the whole remit you were given, in one sentence (\"Admin lane: ยง2 registers, ยง7 plan gating\", \"Approve and reject orders as a manager\"). It sits above the task, which is what the session is doing at any moment. Shown to whoever is watching the run; worth setting whenever more than one session is live.",
        +  "maxLength": 300,
        +  "type": "string"
        +}
      • addedInput schema / properties / paceMs
        Added value: +{
        +  "description": "A floor between actions, in milliseconds, for when a person is watching and needs to keep up โ€” following a flow, taking notes, demonstrating. Default 0: as fast as the page allows, which is what a run wants otherwise. Changeable mid-run with scout_session {paceMs}.",
        +  "maximum": 60000,
        +  "minimum": 0,
        +  "type": "integer"
        +}
      • addedInput schema / properties / record
        Added value: +{
        +  "default": false,
        +  "description": "Keep a frame of the page after every action, under .scenescout/recordings/, and show it beside that step in report.html. Off by default: a recording is pictures of the app under test sitting in the project folder. Turn it on for QA work, where the run is evidence and not only a report.",
        +  "type": "boolean"
        +}
      • addedInput schema / properties / task
        Added value: +{
        +  "description": "What this session is doing right now, shown under its objective from the moment it appears, e.g. Signing in and taking stock. Passed alone it is read as the 2.0 spelling of `objective`. Defaults to a placeholder so a fresh card never reads as idle.",
        +  "maxLength": 300,
        +  "type": "string"
        +}
    • Changedscout_back2 fields changed
      • addedInput schema / properties / objective
        Added value: +{
        +  "description": "Old name for `task` (2.0). Prefer `task`.",
        +  "maxLength": 120,
        +  "type": "string"
        +}
      • addedInput schema / properties / task
        Added value: +{
        +  "description": "What you are DOING right now, in a few words: the action, not the acceptance criteria. \"Filtering the documents register by status\", \"Filling the deviation form with invalid dates\", \"Signing in as QA_Team\" โ€” NOT \"ยง2.4 filtering narrows the set and the filter is reflected in the URL\", which is what you are CHECKING, not what you are doing. Naming the item you are on is fine (\"ยง2.4: filtering the documents register\"); keep the rest to what a colleague would see over your shoulder. It stays set until you pass a different one, so a batch costs a few words, not one per call. Required on the tools that act unless a journey or an earlier call already set one.",
        +  "maxLength": 120,
        +  "type": "string"
        +}
    • Changedscout_click2 fields changed
      • addedInput schema / properties / objective
        Added value: +{
        +  "description": "Old name for `task` (2.0). Prefer `task`.",
        +  "maxLength": 120,
        +  "type": "string"
        +}
      • addedInput schema / properties / task
        Added value: +{
        +  "description": "What you are DOING right now, in a few words: the action, not the acceptance criteria. \"Filtering the documents register by status\", \"Filling the deviation form with invalid dates\", \"Signing in as QA_Team\" โ€” NOT \"ยง2.4 filtering narrows the set and the filter is reflected in the URL\", which is what you are CHECKING, not what you are doing. Naming the item you are on is fine (\"ยง2.4: filtering the documents register\"); keep the rest to what a colleague would see over your shoulder. It stays set until you pass a different one, so a batch costs a few words, not one per call. Required on the tools that act unless a journey or an earlier call already set one.",
        +  "maxLength": 120,
        +  "type": "string"
        +}
    • Addedscout_lane_brief
    • Addedscout_lane_report
    • Changedscout_navigate2 fields changed
      • addedInput schema / properties / objective
        Added value: +{
        +  "description": "Old name for `task` (2.0). Prefer `task`.",
        +  "maxLength": 120,
        +  "type": "string"
        +}
      • addedInput schema / properties / task
        Added value: +{
        +  "description": "What you are DOING right now, in a few words: the action, not the acceptance criteria. \"Filtering the documents register by status\", \"Filling the deviation form with invalid dates\", \"Signing in as QA_Team\" โ€” NOT \"ยง2.4 filtering narrows the set and the filter is reflected in the URL\", which is what you are CHECKING, not what you are doing. Naming the item you are on is fine (\"ยง2.4: filtering the documents register\"); keep the rest to what a colleague would see over your shoulder. It stays set until you pass a different one, so a batch costs a few words, not one per call. Required on the tools that act unless a journey or an earlier call already set one.",
        +  "maxLength": 120,
        +  "type": "string"
        +}
    • Changedscout_note1 field changed
      • changedInput schema / properties / section / enum
        Previous value: -[
        -  "app-model",
        -  "roles",
        -  "conventions",
        -  "constraints",
        -  "risks",
        -  "glossary"
        -]New value: +[
        +  "app-model",
        +  "roles",
        +  "conventions",
        +  "constraints",
        +  "risks",
        +  "glossary",
        +  "setup"
        +]
    • Changedscout_press2 fields changed
      • addedInput schema / properties / objective
        Added value: +{
        +  "description": "Old name for `task` (2.0). Prefer `task`.",
        +  "maxLength": 120,
        +  "type": "string"
        +}
      • addedInput schema / properties / task
        Added value: +{
        +  "description": "What you are DOING right now, in a few words: the action, not the acceptance criteria. \"Filtering the documents register by status\", \"Filling the deviation form with invalid dates\", \"Signing in as QA_Team\" โ€” NOT \"ยง2.4 filtering narrows the set and the filter is reflected in the URL\", which is what you are CHECKING, not what you are doing. Naming the item you are on is fine (\"ยง2.4: filtering the documents register\"); keep the rest to what a colleague would see over your shoulder. It stays set until you pass a different one, so a batch costs a few words, not one per call. Required on the tools that act unless a journey or an earlier call already set one.",
        +  "maxLength": 120,
        +  "type": "string"
        +}
    • Changedscout_report1 field changed
      • addedInput schema / properties / history
        Added value: +{
        +  "default": "index",
        +  "description": "How much of the history to print. 'index' lists findings from earlier runs, and resolved ones, as a row each: id, severity, age, title. 'full' prints every one in full as before โ€” on one project that was 1.75 MB against 113 KB, nearly half of it findings already fixed. Use 'full' when handing the document to someone who has no access to the memory.",
        +  "enum": [
        +    "index",
        +    "full"
        +  ],
        +  "type": "string"
        +}
    • Addedscout_request
    • Changedscout_run_plan2 fields changed
      • addedInput schema / properties / objective
        Added value: +{
        +  "description": "Old name for `task` (2.0). Prefer `task`.",
        +  "maxLength": 120,
        +  "type": "string"
        +}
      • addedInput schema / properties / task
        Added value: +{
        +  "description": "What you are DOING right now, in a few words: the action, not the acceptance criteria. \"Filtering the documents register by status\", \"Filling the deviation form with invalid dates\", \"Signing in as QA_Team\" โ€” NOT \"ยง2.4 filtering narrows the set and the filter is reflected in the URL\", which is what you are CHECKING, not what you are doing. Naming the item you are on is fine (\"ยง2.4: filtering the documents register\"); keep the rest to what a colleague would see over your shoulder. It stays set until you pass a different one, so a batch costs a few words, not one per call. Required on the tools that act unless a journey or an earlier call already set one.",
        +  "maxLength": 120,
        +  "type": "string"
        +}
    • Changedscout_select2 fields changed
      • addedInput schema / properties / objective
        Added value: +{
        +  "description": "Old name for `task` (2.0). Prefer `task`.",
        +  "maxLength": 120,
        +  "type": "string"
        +}
      • addedInput schema / properties / task
        Added value: +{
        +  "description": "What you are DOING right now, in a few words: the action, not the acceptance criteria. \"Filtering the documents register by status\", \"Filling the deviation form with invalid dates\", \"Signing in as QA_Team\" โ€” NOT \"ยง2.4 filtering narrows the set and the filter is reflected in the URL\", which is what you are CHECKING, not what you are doing. Naming the item you are on is fine (\"ยง2.4: filtering the documents register\"); keep the rest to what a colleague would see over your shoulder. It stays set until you pass a different one, so a batch costs a few words, not one per call. Required on the tools that act unless a journey or an earlier call already set one.",
        +  "maxLength": 120,
        +  "type": "string"
        +}
    • Changedscout_session1 field changed
      • addedInput schema / properties / paceMs
        Added value: +{
        +  "description": "Change how fast this session acts, mid-run: a floor between actions in milliseconds, for when a person is watching and needs to keep up. 0 restores full speed. With `name`, applies to that session; without, to every live session โ€” which is what 'slow everything down so I can follow' means.",
        +  "maximum": 60000,
        +  "minimum": 0,
        +  "type": "integer"
        +}
    • Changedscout_type2 fields changed
      • addedInput schema / properties / objective
        Added value: +{
        +  "description": "Old name for `task` (2.0). Prefer `task`.",
        +  "maxLength": 120,
        +  "type": "string"
        +}
      • addedInput schema / properties / task
        Added value: +{
        +  "description": "What you are DOING right now, in a few words: the action, not the acceptance criteria. \"Filtering the documents register by status\", \"Filling the deviation form with invalid dates\", \"Signing in as QA_Team\" โ€” NOT \"ยง2.4 filtering narrows the set and the filter is reflected in the URL\", which is what you are CHECKING, not what you are doing. Naming the item you are on is fine (\"ยง2.4: filtering the documents register\"); keep the rest to what a colleague would see over your shoulder. It stays set until you pass a different one, so a batch costs a few words, not one per call. Required on the tools that act unless a journey or an earlier call already set one.",
        +  "maxLength": 120,
        +  "type": "string"
        +}
    • Changedscout_upload2 fields changed
      • addedInput schema / properties / objective
        Added value: +{
        +  "description": "Old name for `task` (2.0). Prefer `task`.",
        +  "maxLength": 120,
        +  "type": "string"
        +}
      • addedInput schema / properties / task
        Added value: +{
        +  "description": "What you are DOING right now, in a few words: the action, not the acceptance criteria. \"Filtering the documents register by status\", \"Filling the deviation form with invalid dates\", \"Signing in as QA_Team\" โ€” NOT \"ยง2.4 filtering narrows the set and the filter is reflected in the URL\", which is what you are CHECKING, not what you are doing. Naming the item you are on is fine (\"ยง2.4: filtering the documents register\"); keep the rest to what a colleague would see over your shoulder. It stays set until you pass a different one, so a batch costs a few words, not one per call. Required on the tools that act unless a journey or an earlier call already set one.",
        +  "maxLength": 120,
        +  "type": "string"
        +}
    • Addedscout_verify
  8. 2 tool updatesv1.2.0
    • Changedscout_attach1 field changed
      • addedInput schema / properties / browser
        Added value: +{
        +  "description": "Browser to drive. Default: the SCENESCOUT_BROWSER environment variable, else chromium. firefox and webkit must be downloaded first (scenescout install --browser-only --browsers firefox). Use them for a cross-browser pass; stay on chromium otherwise.",
        +  "enum": [
        +    "chromium",
        +    "firefox",
        +    "webkit"
        +  ],
        +  "type": "string"
        +}
    • Addedscout_playbook
  9. 48 tool updatesv1.1.0
    • Removedft_attach
    • Removedft_back
    • Removedft_click
    • Removedft_close
    • Removedft_coverage
    • Removedft_crawl
    • Removedft_design_audit
    • Removedft_finding
    • Removedft_hover
    • Removedft_journey
    • Removedft_navigate
    • Removedft_note
    • Removedft_press
    • Removedft_report
    • Removedft_resolve
    • Removedft_run_plan
    • Removedft_scan
    • Removedft_screenshot
    • Removedft_scroll
    • Removedft_select
    • Removedft_session
    • Removedft_snapshot
    • Removedft_type
    • Removedft_upload
    • Addedscout_attach
    • Addedscout_back
    • Addedscout_click
    • Addedscout_close
    • Addedscout_coverage
    • Addedscout_crawl
    • Addedscout_design_audit
    • Addedscout_finding
    • Addedscout_hover
    • Addedscout_journey
    • Addedscout_navigate
    • Addedscout_note
    • Addedscout_press
    • Addedscout_report
    • Addedscout_resolve
    • Addedscout_run_plan
    • Addedscout_scan
    • Addedscout_screenshot
    • Addedscout_scroll
    • Addedscout_select
    • Addedscout_session
    • Addedscout_snapshot
    • Addedscout_type
    • Addedscout_upload
  10. 24 tool updatesv0.23.3
    • First observedft_attach
    • First observedft_back
    • First observedft_click
    • First observedft_close
    • First observedft_coverage
    • First observedft_crawl
    • First observedft_design_audit
    • First observedft_finding
    • First observedft_hover
    • First observedft_journey
    • First observedft_navigate
    • First observedft_note
    • First observedft_press
    • First observedft_report
    • First observedft_resolve
    • First observedft_run_plan
    • First observedft_scan
    • First observedft_screenshot
    • First observedft_scroll
    • First observedft_select
    • First observedft_session
    • First observedft_snapshot
    • First observedft_type
    • First observedft_upload

TDQS

A3.9/5.0

Scored across 37 tools

Disambiguation4/5

Most tools have sharply distinct purposes, and the descriptions explicitly defuse the riskiest overlaps (scout_screenshot vs scout_capture vs scout_design_audit; scout_status vs scout_status_poll; scout_request vs scout_network; scout_crawl vs scout_run_plan). A few pairs remain close enough that an agent must read carefully to pick correctly, e.g. scout_intake/scout_playbook/scout_note all being 'read this at the start' guidance, and the two lane tools being mirror halves of one workflow.

Naming Consistency5/5

Every tool uses the same snake_case convention with a uniform `scout_` prefix and a noun/verb suffix describing the resource or action (scout_click, scout_type, scout_finding, scout_report). There are no camelCase or stylistic deviations, so the set is fully predictable.

Tool Count2/5

37 tools is well past the 25-tool threshold the rubric treats as excessive, and several are low-value or redundant for the agent (scout_status_poll exists only to refresh a pane, scout_intake/scout_playbook/scout_note overlap as upfront reading). The domain is genuinely broad, but the surface would benefit from consolidation into fewer composite tools.

Completeness5/5

The set covers the full lifecycle of the stated purpose: setup (scan, attach, login, session), exploration (crawl, navigate, snapshot, network, interaction verbs), analysis (design audit, journey, coverage), knowledge (note, tickets, criterion), and reporting (finding, verify, resolve, report). Nothing obvious is missing, and report gating plus a gap ledger shows dead-ends are explicitly handled.

Maintenance

ActivityMaintained
ResponsivenessResponsive

Related MCP Connectors

Related MCP Servers