Skip to main content
Glama

🔭 SceneScout

Exploratory UI testing, driven by the AI agent you already use.

Works with Claude Code · Cursor · VS Code (Copilot) · Codex CLI · Gemini CLI · Copilot CLI · Windsurf · any MCP client

test npm license: MIT node >= 20 MCP server

📖 Guide · 👀 See it work · ✨ Why · 🎯 Two ways to use it · 🚀 Quickstart · 🧰 Toolbox · 🔌 Other clients · 🔒 Safety · 🩺 Troubleshooting

SceneScout is an MCP server that hands an agent a structured view of a running web app — every element, its geometry, and a set of always-on correctness oracles — and lets the agent explore it like a curious user. Your coding agent is the brain; SceneScout is the hands, eyes, and memory. Any MCP client can drive it, and the testing method comes with the server, so the agent knows how to use the tools wherever it runs.

┌─────────────────────┐   MCP (stdio)   ┌───────────────────────────────┐
│ Your coding agent   │ ──────────────▶ │ SceneScout engine             │
│ (intent, judgment,  │ ◀────────────── │ Playwright · oracles · memory │
│  your subscription) │  tool results   │ findings · report — no LLM    │
└─────────────────────┘                 └───────────────────────────────┘

Scripted E2E suites answer one question — "does this exact flow still work?" — and say nothing about the 95% of the app they don't touch. SceneScout covers both gaps: it finds what's broken (crashes, dead ends, permission leaks) and reports how the product could be better (confusing flows, weak hierarchy, design-system drift), with concrete measurements.

👀 See it work

This is a real run against the small demo app bundled in this repository. The app has bugs planted in it on purpose, and two of them are visible on its dashboard:

The broken chart is the demo app's bug, not this page's — it is one of the twelve findings SceneScout filed, next to the badge sitting on a button. The red callouts were added for this README; the unmarked screenshots are the ones the engine took.

An excerpt of the report it wrote — read the whole thing:

🔴 [HIGH] A double-click on Create order creates two orders Evidence: 2× click fired the same state-changing request 2× (POST /api/orders) The submit button stays enabled while the request is in flight, and the endpoint accepts the repeat.

🔴 [HIGH] Filtering orders by Archived fails, and the page shows an empty table instead of an error Evidence: GET /api/orders?status=archived → HTTP 500

🔴 [HIGH] A clerk can approve an order by calling the endpoint the page hides from them Evidence: POST /api/orders/1037/approve 200 as clerk; POST /api/orders/1038/reject 403 as clerk — the button was hidden, the server did not agree.

🟠 [MEDIUM] The "New: bulk import" badge sits on top of the All orders button (callout 1) Evidence: "All orders" overlaps "New: bulk import" (81%) — measured from layout boxes, no screenshot needed.

🟡 [LOW] The dashboard chart image is missing (callout 2) Evidence: GET /img/weekly-chart.png → HTTP 404

Gap ledger — what was NOT tested: 9/12 visited routes never design-audited · single-role run, so permission boundaries are untested

Every finding comes with a repro trace and a Playwright regression-test skeleton. To try it yourself, clone this repository, run npm run demo:serve, then /scenescout --url http://127.0.0.1:4173 — see demo-app/. Its README lists every seeded defect and which oracle catches it.


Related MCP server: PixelCheck

✨ Why it's different

  • 🧠 Your agent is the brain — no API key. The engine contains no LLM. Exploration runs on the agent and subscription you already have (Claude Code, Cursor, Copilot, Codex, Gemini CLI and others); SceneScout just gives it deterministic tools and the method for using them.

  • 📐 Structured scene, not pixels. The agent reads element lists with layout geometry, not screenshots. Overlap and off-screen bugs are computed from boxes — deterministic, no vision guessing. Images that failed to load are read from the DOM too. (Screenshots exist only for pixel-native residue like a canvas or a rendering glitch.)

  • 🛡️ Read-only by default, enforced on the wire. Destructive actions are blocked at the network layer, not by asking the model nicely. Opt into writes only against disposable data.

  • ✅ Completion is a contract, not a vibe. The engine knows the app's routes and refuses to file an "extensive" report while any known route is unvisited, unexercised, or un-audited. "Explored a bit and stopped" is structurally impossible.

  • 🧭 It remembers. UI states are fingerprinted and stored in the project's .scenescout/. Run N+1 skips what run N already covered, and every run starts smarter than the last.


🎯 Two ways to use it

SceneScout needs only a URL. Give it the source code as well and it gets noticeably better.

🏠 Next to the codebase (recommended)

🌐 Against a remote URL

You run it from

the app's repository

any folder — an empty qa/ directory is fine

It plays the role of

a developer-tester who can read the code

a black-box QA tester, like a person with a browser

How it finds pages

📂 reads routes from the source and follows links: file-based routing (Next.js, SvelteKit, Nuxt) and router configuration written in code (React Router, Vue Router, Angular). Routes built at runtime are not seen

🔗 follows same-origin links only — pages nothing links to, or on another subdomain, stay unknown

"Did we cover everything?"

checked against the routes found in source plus discovered links — an unvisited one blocks the report

checked against the pages it managed to discover

Setup it figures out

framework, dev command, saved Playwright logins (playwright/.auth/), whether the app uses data-testid

none — you pass the URL, and the path to a login state if the app needs one

What a finding looks like

the symptom, plus the file behind it and a suggested fix

the symptom, a repro trace, and a regression-test skeleton

Typical target

localhost while you build

staging, a preview deploy, a client's site

Why the codebase helps. The agent driving SceneScout is a coding agent, which can already read your repository. With the source at hand it knows the app's static routes before opening the browser, so coverage is measured against the real app instead of whatever happened to be linked. It can also check a suspicion against the code before reporting it: "there is no way to export this table" is a much stronger finding once the agent has confirmed no export handler exists. And when something breaks it can open the component or handler responsible and tell you where and how to fix it — "the save button does nothing" becomes "OrderForm swallows the rejected promise in onSubmit; surface the error and re-enable the button".

Why it still works without it. Everything SceneScout observes comes from the running page — elements, layout geometry, console and network errors, design-audit scores, task-ease measurements — and none of that needs source code. Point it at a URL you are allowed to test and it behaves like a thorough QA tester: it explores, reproduces, and files findings with evidence.

# next to the code — run inside the app's repository
/scenescout --url http://localhost:3000

# remote — run from any folder; memory and the report are kept there
/scenescout --url https://staging.example.com --role ./auth/qa.json
IMPORTANT

Only test sites you own or are authorized to test. A remote environment is more likely to hold real data, so for a remote URL with no source the skill attaches inobserve mode: nothing but GET requests leaves the page. The default read-only mode blocks PUT/PATCH/DELETE and destructive-looking requests, but an ordinary form submission (a plain POST: contact form, comment, order, signup) still reaches the server and can create a record. Say so when that is acceptable on your target. See the safety model.


🚀 Quickstart

📦 Prerequisites

Node

≥ 20

An MCP client

Claude Code, Cursor, VS Code with Copilot, Codex CLI, Gemini CLI, GitHub Copilot CLI, Windsurf, or any other

A web app to test

SceneScout tests a live app: start yours locally first (e.g. npm run dev, make dev-up), or have the URL of a deployed one you're allowed to test

1️⃣ Install

It is on npm. Nothing to clone:

npx -y scenescout install                      # Claude Code: skill + server + Chromium (one-time download)
npx -y scenescout install --client cursor      # or: vscode, codex, gemini, copilot, windsurf (comma-separated for several)

Either way it downloads the browser and registers the server with the client you named. Claude Code also gets the method as a skill; every other client receives the same method from the server. What each client gets.

Prefer a Claude Code plugin? The skill and the server arrive together:

/plugin marketplace add brunoboto96/SceneScout
/plugin install scenescout@scenescout-marketplace

Then download the browser once with npx -y scenescout install --browser-only. The command becomes /scenescout:scenescout. A plugin's skill comes from this repository and its server from the latest npm release, so right after a release lands here the two can differ for a short while; /plugin marketplace update scenescout-marketplace brings the skill up to date.

A client that is not in that list? Run npx -y scenescout install --browser-only and add the server to its config by hand.

  1. puts the /scenescout skill into ~/.claude/skills/ (or $CLAUDE_CONFIG_DIR/skills/) — a scenescout folder it didn't create is moved aside to a .backup-… copy, never deleted,

  2. downloads the browser SceneScout drives (skipped if you already have it). By default that is Chromium, as two builds: the full browser for headed runs and the headless shell every other run uses. Choose something else with --browsers,

  3. registers the MCP server with Claude Code at user scope. Run through npx, the launcher is npx -y scenescout serve, with the absolute path of npx where one sits beside node, so it works under nvm/fnm. From a clone or a global install it is the absolute node path plus that install's dist/mcp-server.js,

  4. puts the scenescout command on your PATH, so scenescout status, scenescout watch and scenescout doctor work from any terminal. Run through npx, that is npm install -g of the version you just ran; from a clone it is npm link, so the command always runs what you last built. If npm refuses (a system-wide node usually needs sudo for this), the step prints the command to run by hand and the rest of the setup still counts as done: npx -y scenescout <command> works without it.

Re-run it any time: after moving the folder or switching node versions it refreshes the stored paths. It exits non-zero if a step the tool depends on failed, so it is safe to chain. Opt out of a step with --no-register, --skip-browser or --no-command.

If claude isn't on the PATH of the shell you ran it from, it prints the registration command instead of running it:

claude mcp add --scope user scenescout -- npx -y scenescout serve

2️⃣ Check it

npx -y scenescout doctor --engine   # any client: node + build + browser
npx -y scenescout doctor            # Claude Code: the above, plus the skill and the registration

Every line should be a ✓. Anything that isn't prints the exact command that fixes it. Then start a fresh session in your client so it picks up the new tools.

3️⃣ Run it

No app handy? Clone this repository and run npm run demo:serve: the demo app starts on http://127.0.0.1:4173.

Open your agent inside the project you want to test (or, for a remote URL, any folder) and ask:

Use SceneScout to test http://localhost:3000 at medium level

In Claude Code the skill gives you a command with flags for the same thing:

/scenescout --level medium --url http://localhost:3000 --role qa

The agent scans the project (if there is one), attaches read-only, explores, and writes findings to .scenescout/report.md. That's it.

Common flags — --level minimal|medium|extensive · --url <app> · --role <name\|path> (who to explore as: a login saved with scenescout login, a storage state found by the scan, or a path to a Playwright storage-state JSON) · --observe / --safe-write / --allow-destructive.

🔑 Signing in as a role

For an app behind SSO or MFA, sign in once yourself and let every session reuse it:

scenescout login http://localhost:3000 --role admin

A browser window opens at the URL. Sign in however the app asks, then come back to the terminal and press Enter: the session is saved as .scenescout/auth/admin.json in the project. Closing the window or pressing Ctrl+C saves nothing. The file is readable by your account only, .scenescout/ keeps itself out of git, and the command prints where it saved, how many cookies, origins and databases it holds, never what they are, and how long it will last: read from each cookie's expiry and the exp of any JWT in a cookie or in localStorage (the payload is decoded for that one claim, never verified, never printed). The profile keeps cookies, localStorage, IndexedDB and sessionStorage, so an app whose sign-in library keeps its token in sessionStorage or IndexedDB still comes back signed in; sessionStorage is put back only on the origin it came from, once per tab, so a lane that signs out stays signed out. A login saved by an earlier version has no sessionStorage or IndexedDB: record it again if the app keeps its token there. --project <dir> saves into another project; --browser firefox|webkit records in another browser.

Then /scenescout --role admin, or scout_attach { role: "admin" } from any agent. Every session attached with the same role gets its own browser built from that one login, so parallel lanes can all run as admin. A role with no saved login is refused with the command to run. role and storageStatePath are alternatives: pass one.

Sessions of one role share one saved login, so they share its refresh token too. An app that rotates refresh tokens and treats a second use of a spent one as theft would revoke the whole token family, and sign every session of that role out, the moment two of them refreshed with the same token. SceneScout stops that for a session attached by role. When the page is about to send a refresh token from the role's profile, the session first takes a lock beside the profile (.scenescout/auth/<role>.json.lock, owner-only, taken over if its holder has not touched it in 30 seconds). Holding the lock, it re-reads the profile: if another session has rotated the token in the meantime, it loads that profile into its own browser and sends the current token in place of the spent one. Once the page has stored the rotated token, the session writes its state back over the profile and releases the lock. Sessions in separate processes share the lock through the file. A refresh token is recognised by name (a cookie, a storage key, or a field inside a JSON storage value whose name contains refresh) and is never printed or logged. An app whose sign-in renews through the identity provider's own session cookie needs none of this, since no refresh token is shared. SCENESCOUT_REFRESH_BROKER=off turns the broker off.

In CI, where nobody can type, --script signs in headless as a test user from SCENESCOUT_LOGIN_USERNAME, SCENESCOUT_LOGIN_PASSWORD and, for a one-time code, SCENESCOUT_LOGIN_TOTP_SECRET, and saves the same profile. No credential value is ever printed. See signing in from CI for the options and the rules: a test tenant's user, never production or a real person's account.

Before a parallel run, scout_lane_brief checks that the planner's saved login will outlast it: runMinutes (default 60) plus expiryMarginMinutes (default 10). It refuses only when it is sure, meaning every credential in the profile has a date, none was set for another host, and the last of them ends before the run does, and then names the scenescout login command to run again. A profile holds cookies other than the sign-in (analytics, preferences), so the first one to expire is reported as a warning rather than a reason to refuse, and a profile with undated credentials in it (a session cookie, or a refresh token with no expiry) is a warning that its lifetime is unknown.


📺 Watching a run live

When a session attaches, the engine starts a small live view and hands the agent its address on a Live view: line, which the agent passes on to you. From a terminal, scenescout watch opens the same page. There is one card per session:

  • What it is doing: the tool it is running and for how long, the page it is on, and a thumbnail of that page. This works for headless runs too, which have no window to look at.

  • What it just did: a rolling feed of its actions, each with its target and how it turned out, with failures in red. It is the same trail a finding's repro trace uses. The engine never sees the agent's reasoning, so this is what the session did, not what it thought.

  • Stuck, not slow: a call still running past its own tool's watchdog budget turns the card red, so a wedged session is visible without asking. A crawl legitimately runs for minutes; it is judged against the crawl's budget, not a click's.

  • Live stream: switch it on for one card, or for all of them. Click a thumbnail for a close-up.

  • The report, as it stands: the Report button in the top bar shows the same document scout_report writes at the end, rendered from the run's current state, so findings can be read while the agents are still working.

  • What it is for: the close-up puts the feed beside the session's brief — the objective it was given when it attached (scout_attach {objective}), and underneath it the task it is on right now (scout_task), which the engine requires before any tool will act. Each task tints its own block of actions, so a change of task is a change of colour; point at a block and the brief names the task those actions served.

  • Scrub it back: under the page is a tick per action, coloured by task. Click one to see the frame from that moment, and Back to live to return. On a run that was not recorded the ticks still read the trail; they just have no picture behind them.

The view is served on 127.0.0.1 only, behind a token that changes every time the engine starts. It answers GET and nothing else, so a viewer can watch a run but not act in it, and no frame it shows is written to disk (ADR 7) unless the run was recorded, which is asked for and off by default (ADR 8). A stream runs only while someone is watching it. SCENESCOUT_LIVE=off keeps the port closed.

Try it with parallel agents. The demo app has three roles and several separate areas, so a run can be split between agents. Start it with npm run demo:serve, then ask your agent to explore it with several agents in parallel, one role and one area each. The pictures above come from a run of three. Two things keep a parallel run efficient:

  • Each agent opens its own session when it starts, and the planner closes it once it has folded that agent's report. An agent waiting for its turn then holds no browser. Opening every session up front leaves browsers idling while the machine runs out of memory for the agents that are working. Closing before the fold loses the lane's decisions, which have nowhere to be kept.

  • Slow it down to follow along. scout_attach {paceMs} (or scout_session {paceMs} mid-run) sets a floor between actions, for when you want to watch a flow rather than let it run as fast as the page allows.

  • Run about as many agents at once as your machine has cores, less two. Each one drives a real browser.


🎬 Recording a run, and reading it back

A report says what happened. For QA work that is not always enough — the point is often to show what was checked, not to assert it. Ask for a recorded run and the engine keeps a frame of the page after every action:

Use SceneScout to test http://localhost:3000, record the run

or, on the tool directly, scout_attach {record: true}.

Then scout_report writes two files side by side in .scenescout/: report.md as always, and report.html — the whole run as one self-contained page. It opens from the file system with nothing running, needs no network, and holds:

  • The report, rendered from the same Markdown.

  • The screenshots around each finding, in an accordion under it, from the session that filed it.

  • Every session's trail, in the blocks its tasks made, each step with the page as it was at that moment.

The live view serves the same document at run while the engine is still up, and sends you there when the run ends — so the address survives a refresh instead of a panel over a dead board.

What it costs. Frames are pictures of the app under test, inside the tested project's folder, and the secret redaction that protects everything else the engine writes cannot read a picture. That is why it is off unless asked for, capped per session, and written only under .scenescout/, which ignores itself so git add -A in the tested project cannot pick the frames up. The reasoning is in ADR 8.


🔄 How a run works

One curiosity loop, repeated — breadth first, then judgment where it matters:

scan ──▶ attach ──▶ crawl ──▶ investigate ──▶ measure ──▶ report
 │         │          │            │              │           │
routes   browser   every route  reproduce &   journeys +   gap-checked
& auth   (r/o)     in ONE call   file findings  design audit  markdown
  1. Scan the project — framework, routes, auth states.

  2. Attach a browser (read-only unless you said otherwise).

  3. Crawl every known route in a single call — per-route HTTP status, element counts, oracle violations, dead ends.

  4. Investigate what the crawl flagged: navigate, snapshot, reproduce, file a structured finding.

  5. Measure task ease (scout_journey) and design quality (scout_design_audit) on representative pages.

  6. Report — the engine checks the gap ledger and writes .scenescout/report.md.

Snapshots are cheap: re-snapshotting a route returns only what changed, with stable refs (measured on a 130-element page: 10.7 kB → 0.7 kB).


🧰 The toolbox

29 deterministic tools. The agent picks; you rarely call these by hand.

Phase

Tools

What they do

Set up

scout_playbook scout_scan scout_attach scout_session

Hand the testing method to an agent that has no skill loaded; discover routes; launch a browser in a write-mode; keep several authenticated roles alive at once

Explore

scout_crawl scout_coverage

Sweep every route in one call; ask what's still untested

Look

scout_snapshot scout_hover scout_screenshot scout_capture

Read the structured scene (diffed); reveal tooltips/hover cards; capture pixels only when needed; save one element as a PNG to show someone

Ask the server

scout_request

Call the app's own API as this session, with the UI bypassed — the check that turns a hidden button into a proven refusal

Act

scout_click scout_type scout_select scout_upload scout_press scout_scroll scout_navigate scout_back scout_run_plan

Drive the UI like a user; scout_run_plan batches a whole mechanical sequence into one call

Assess

scout_design_audit scout_journey

Score a page's craft/a11y/consistency; measure how hard a task is to complete

Record

scout_note scout_finding scout_resolve scout_report

Curate durable notes; file deduped findings; mark fixes; write the report, and on a recorded run the whole run as one page

Re-test

scout_verify

List the findings earlier runs left open, worst route first, and record whether each is gone, still present, or changed

Split the work

scout_lane_brief scout_lane_report

Divide the app between parallel agents by whole module, each with its own landing route and rules; fold what each hands back as one typed JSON object, and name any defect it judged but never filed

Close

scout_close

Tear down one session or all

A few that punch above their weight:

  • scout_crawl — the entire breadth pass in one tool call. No visiting routes one-by-one.

  • scout_run_plan — up to 20 actions (fill form → submit → check) with semantic targets (testid=…, text=…), aborting at the first anomaly.

  • scout_journey — wraps one goal and reports interaction count, screens seen, and backtracks; an abandoned journey is a finding no passing E2E suite can produce.

  • scout_upload — generates a valid in-memory fixture (real PDF/PNG, kind inferred from accept) so file-upload flows stop being a blind spot.

  • scout_click {clicks: 2} — the impatient-user probe: states whether a double-click fired the same state-changing request twice (the classic double-submit bug).

  • scout_request — calls the app's own API as the session, so "the button is hidden" becomes "the server refuses it" (or doesn't).

Beyond crashes and HTTP errors, two oracles catch a page contradicting the server: refused_empty (a list request was refused and the page shows its empty state with no error) and false_success (a save was refused and the page says it worked). A third, dom_injection, reports a typed markup value coming back as an element on any page any session opens. A fourth, postmessage_token, reports a page calling postMessage with targetOrigin "*" on a message that carries a token (a JWT, a Bearer value, or an opaque value under a key such as access_token): the report names where in the message it was, its shape and its first four characters, never the token.


📊 Test levels

Each level is a contract. scout_report enforces what the engine can see for itself — routes visited, pages audited, and at extensive an empty gap ledger (every route exercised and audited, every filled form submitted, a completed journey, two roles) — and the report's gap ledger discloses the rest.

Level

What it guarantees

Rough size

minimal

Every route visited, ≥1 design audit, key journeys as plans, crawl problems triaged. Remaining gaps disclosed.

~40 actions

medium (default)

minimal + design audits across several routes (enforced) + every element class exercised + every form submitted valid and invalid (asked of the agent)

~150 actions

extensive

medium + fuzzing, back/refresh/deep-link resilience, keyboard-only pass, a journey per module, ≥2 roles compared, anonymous auth-surface walk. Refuses to finalize while any gap in the ledger remains.

budget-capped

That refusal is the guarantee: an extensive report can only exist when nothing the engine can measure was left untested. Fuzzing, the keyboard pass and the auth-surface walk are the agent's to do; the engine cannot see whether they were done well.


🔒 Safety model

  • 🔵 observe (--observe) lets nothing but GET requests leave the page. The one exception is what a session needs in order to exist: logging in, logging out and refreshing a token. Signing up, changing or resetting a password and creating users are blocked like any other write. WebSocket frames are not inspected; the engine says so when the app opens a socket. It is what the skill picks for a remote URL with no source, where an ordinary form POST would create a real record. Forms that could not be submitted are listed in the gap ledger.

  • 🟢 read-only by default. Destructive-labeled elements (delete/revoke/archive/…) and all PUT/PATCH/DELETE + destructive POSTs are blocked at the network layer — see src/engine/policy.ts. Non-destructive POSTs are allowed, because submitting forms is how a tester finds validation bugs — so read-only means nothing existing is changed or removed, not nothing is ever created.

  • 🟡 safe-write (--safe-write) lets the agent create data and edit/delete only what it created this run — never pre-existing records.

  • 🔴 destructive (--allow-destructive) allows everything, and only ever when you confirm the environment is disposable. The skill will never choose this itself.

  • 📂 Findings, memory, and reports live in a .scenescout/ folder where you ran it. It ignores itself in git, so a stray git add -A never commits test data.

A 🛡 WRITE-POLICY blocked notice is the safety net doing its job, not an app bug. The server never sees a blocked request, but a page's own fetch or XHR is answered with a 403 in its place rather than dropped, so the page's handling of a refusal really runs: a page that then claims success is reported as a false_success (ADR 9).


📋 What you get

.scenescout/report.md — a deduplicated, worst-first report with:

  • 🐛 Findings with repro traces and generated Playwright regression-test skeletons.

  • 🔎 Worth a look — observations that are defects only under a convention of your project the run cannot see (a spacing scale, link styling in navigation, test ids on every control), each naming that convention. Listed below the findings and not counted as defects (ADR 13).

  • 💯 Page scores (0–100: a11y · craft · consistency · task-clarity), ranked worst-first, with stale scores from old runs marked as such.

  • 👥 A role capability matrix — what each role could and couldn't reach.

  • 🧾 A gap ledger — everything not done, so the report is honest about its own coverage.

  • ⏱️ How the run was paced — how closely each session kept working, and apart from that, how long finished lanes held their browsers waiting to be collected, so neither hides the other.

  • 🎯 How well the lanes judged — on a parallel run, whether the confidence each lane stated matched what the project went on to file, beside what later re-tests found (ADR 10).

.scenescout/report.html — the same report as one self-contained page, with every session's trail beside it, and on a recorded run the screenshots under each finding.

👀 Watch a run live: node dist/cli.js status <project-path>.


🚦 In CI: a deterministic check

An exploratory run is driven by a model, so two runs never find exactly the same things. That suits a report, but not a gate. scenescout check is the part that needs no model. It visits the start URL, the project's scanned routes and every same-origin link it finds, and measures each page:

  • HTTP and page errors

  • layout geometry (covered, clipped and overlapping controls, blocking overlays)

  • broken images

  • controls with no name, and fields whose only label is a placeholder

  • contrast and focus

  • pages with no way out

Some of what it measures is a defect only under a convention the check cannot see: paddings off a 4px grid, and links styled like body text. Those are listed under Worth a look, each with the convention that would make it a defect. They are never counted and never fail the gate, at any --fail-on; SARIF reports them at level note (ADR 13).

npx scenescout check http://127.0.0.1:3000 --fail-on high

Exit code

Meaning

0

Passed the gate

1

Failed it: something at the --fail-on severity or worse

2

Could not run, or not all of it: a bad argument, an app that never answered, only the sign-in page reached, a saved flow that is not valid, or a flow step the write policy refused

It writes report.md, check.sarif (for code-scanning dashboards) and check.json to .scenescout/check/, and on GitHub Actions it also puts the report on the job's summary page.

On GitHub Actions, this repository is also an action that installs everything and keeps the results:

- uses: brunoboto96/SceneScout@v3.10.0
  with:
    url: http://127.0.0.1:3000

docs/ci.md has a complete workflow (start the app, wait for it, check it), the action's inputs and outputs, code-scanning upload, and the same check on GitLab CI, CircleCI or any shell.

With the default settings its saved flows send no HTTP write (they replay under observe's rule), and its crawl runs under --mode observe or read-only; --flow-writes allow lets flows write as --mode allows. By default it fails only on facts that mean a page is broken: a page that did not load, an uncaught exception, a 5xx, a failure shown as success. Other options:

  • --fail-on medium or low makes the gate stricter.

  • --ignore <rule> drops a rule, worth-a-look rules included.

  • --paths /a,/b checks only those pages.

  • --storage-state <file> checks while signed in.

  • --flows <dir> or off chooses which saved flows to replay; --retest off skips re-testing open findings.

  • --flow-writes never|allow (default never): never replays flows under observe's rule whatever --mode says; allow replays them under --mode, so in read-only a flow's form submissions are sent to the target on every run.

  • --on-refused-step report|stop (default report): report marks a flow whose step was refused "could not run", keeps every other verdict and exits 2; stop exits 2 at that step with no results.

  • --gate-retests never|high|all (default high): which still-reproducing re-tested findings fail the gate.

The defaults are what an unconfigured check does, for a first try or an AI agent running it unattended: its flows send no HTTP write and it never silently hides a result. Each setting is a choice for the project; the report and check.json print the values a check ran with.

scenescout check --help lists every option. Why the defaults are what they are: ADR 11.

It also replays the flows saved in .scenescout/flows/*.json, with no model: the steps scout_run_plan takes (navigate, click, type, select, press) plus expect-text, expect-url and expect-request. A flow whose step breaks fails the gate, naming the flow and the step. And it re-tests the open findings earlier runs left in the project's memory that a page load can reproduce, reporting each as still reproducing or possibly fixed; by default a finding filed high that still reproduces fails the gate. docs/ci.md has the flow format; ADR 12 says why it works this way.

Beyond those flows it explores nothing and fills no forms. That is the exploratory run's job, and its findings belong in a report, not a gate.

🤖 In CI: an unattended exploratory run

scenescout ci runs the exploratory side in a CI job, with no person and no coding agent. A model reached through its API drives the same scout_* tools by the same method, and the run ends in the ordinary report:

export OPENAI_API_KEY=…          # or ANTHROPIC_API_KEY; read from the environment only
npx scenescout ci http://127.0.0.1:3000
  • It reports and never gates. Exit 0 when the run ran, whatever it found; exit 2 when it could not run (no key, a key the API refused, an app that never answered). Two runs of the same app find different things, so a finding is something to read, never a reason to fail a build. scenescout check is the gate.

  • Providers: the Anthropic Messages API (default model claude-sonnet-5) or the OpenAI Responses API (default gpt-6-luna), chosen by which key is set; with both set, --provider decides. --model and --effort (default low) override; --base-url points at another endpoint that implements the same API.

  • Caps: at most 40 model turns, 1,500,000 tokens and 20 minutes (--max-turns, --max-tokens, --max-minutes). The first cap reached ends the exploration; the report is still written, and says which cap ended it.

  • Mode: read-only by default; --mode observe sends no form at all, --mode safe-write lets the run create records and change only the ones it created. --mode destructive runs only with --allow-destructive as well.

  • Output, in .scenescout/ci/ (or --out): report.md and report.html (the report an agent's run writes), summary.md (also appended to the GitHub job summary), ci.json and ci.sarif, with a usage line: turns, tokens, time and an estimated cost where the model's price is known (--price-in, --price-out give one for any model).

There is a GitHub Action for it (uses: brunoboto96/SceneScout/ci@…). docs/ci.md has the workflow and every option; ADR 14 says why it works this way.

On a pull request, an allowed account can comment /scenescout qa to run it against that pull request's deployed preview and get the results as a reply. The job that holds the key checks out nothing and runs SceneScout from an exact release tag, so the pull request's code never runs beside the key. docs/ci.md has the workflow to copy and what a project configures; ADR 15 says why.


🩺 Troubleshooting

Run npx -y scenescout doctor first — it checks every setup item below (everything but the last row, which is about your app) and prints the fix.

Symptom

Cause and fix

/scenescout isn't a known command

The skill isn't linked, or the session predates it. npx -y scenescout install, then start a fresh Claude Code session.

The scout_* tools don't appear

The MCP server isn't registered, or points at an old path. npx -y scenescout install re-registers it; claude mcp list should show scenescout as connected.

"Executable not found in $PATH"

The server was registered with a bare node. npx -y scenescout install registers an absolute path.

Installed as a plugin, and the tools fail with "Executable not found in $PATH: npx"

A plugin starts the server with a bare npx, which Claude Code can only find if it was launched from an environment that has Node on its PATH. Under nvm or fnm that means starting Claude Code from a terminal, not from a dock or launcher. Or use npx -y scenescout install instead, which registers the absolute path of npx.

"… build has not been downloaded yet" on attach

The browser download was skipped or failed, or the run asked for a browser you did not install. Run the command the message names, for example npx -y scenescout install --browser-only --browsers firefox. On Linux, system libraries may be missing too: npx playwright install --with-deps chromium.

Tools broke after moving the folder or changing node version

The registration stores absolute paths. npx -y scenescout install refreshes them.

Attach fails or every route lands on the login page

Your app isn't running at --url, or the --role session has expired. For a saved login, run scenescout login <url> --role <name> again; for a storage-state file, regenerate it the way your project's Playwright setup does.

⬆️ Upgrading from an older version

  • Tools are now scout_*. Up to v0.23 they were prefixed ft_. The rename happened before the first npm release, with no aliases, so an agent's context carries one tool list rather than two. Re-run npx -y scenescout install so the installed skill matches the server.

  • Earlier names. This tool was previously called SceneCraft (and, before that, frontend-tester). scenescout install cleans up after both: it removes the old skill link and the old scenecraft MCP registration when they point at this install, and the first attach in a project moves its .scenecraft/ memory folder to .scenescout/ so earlier coverage and findings carry over.

🧹 Uninstall

# Claude Code
claude mcp remove --scope user scenescout
rm -rf ~/.claude/skills/scenescout
# Codex / Gemini / Copilot CLI
codex mcp remove scenescout        # likewise: gemini mcp remove …, copilot mcp remove …

For Cursor, Windsurf and VS Code, delete the scenescout entry from the client's MCP server list.

Nothing else is installed: npx runs the package from npm's cache. Per-project memory lives in each tested project's .scenescout/ folder; delete it there if you want it gone.


🌐 Choosing browsers

install downloads Chromium and nothing else unless you ask. --browsers takes one name, a comma-separated list, or all:

--browsers

What is downloaded

About, on disk

chromium (default)

the full browser and the headless shell

550 MB

chromium-headless-shell

the headless shell only: every run works except headed

200 MB

firefox

Firefox

270 MB

webkit

WebKit, the engine behind Safari

290 MB

all

Chromium, Firefox and WebKit

1.1 GB

npx -y scenescout install --browsers chromium-headless-shell   # the smallest working setup
npx -y scenescout install --browser-only --browsers firefox,webkit   # add two more later

Sizes vary by platform. The builds go to Playwright's shared cache, so a build another tool already fetched is not downloaded again.

To drive another browser, pass browser when attaching (scout_attach { browser: "firefox" }), or set SCENESCOUT_BROWSER=webkit in the server's environment to change the default. scenescout doctor checks the browser named by that variable in the shell it runs from, so check another one with SCENESCOUT_BROWSER=webkit scenescout doctor. Two things differ outside Chromium:

  • Service workers are not allowed to register in Firefox and WebKit. The write policy works by intercepting requests, and only Chromium lets a request issued by a service worker be intercepted. An app that depends on its worker may behave differently there.

  • A Firefox or WebKit left behind by a crash is not cleaned up on the next start the way a leftover Chromium is.

In every browser, pages are not given shared workers unless the mode is destructive: a request a shared worker sends cannot be intercepted anywhere, so the app is made to do that work on the page, where the policy sees it.

Time limits

An action on the page (a click, typing, a hover, a pick from a list) may take 5 s, and a page 20 s to load (15 s for a page the crawl opens). On a loaded machine these can run out while the app is fine; the timeout then says which limit ran out and how to raise it. Raise them per session with scout_attach { actionTimeoutMs: 15000, navTimeoutMs: 60000 }, for every session with SCENESCOUT_ACTION_TIMEOUT_MS and SCENESCOUT_NAV_TIMEOUT_MS in the server's environment, or on scenescout check and scenescout ci with --action-timeout-ms and --nav-timeout-ms. An option wins over the variable, and the variable over the default. The action limit takes 1000 to 120000 ms and the page-load limit 1000 to 300000 ms; anything else refuses the attach with a sentence naming the value to fix. Saving a login profile with scenescout login honours the two variables as well, and otherwise keeps its own longer waits (30 s for the page, 10 s for a field or the submit).

🔌 Other MCP clients

The engine is a plain MCP server over stdio, so any client can drive it, and the testing method reaches the agent through the server itself (see the end of this section). install can register it for you:

npx -y scenescout install --client cursor            # one client
npx -y scenescout install --client vscode,codex      # several; add claude-code to keep that one too

--client

How it is registered

claude-code (default)

claude mcp add, plus the skill

cursor

adds an entry to ~/.cursor/mcp.json, keeping the others

vscode

VS Code's own code --add-mcp. A code command that belongs to another editor is not used

codex

codex mcp add

gemini

gemini mcp add --scope user

copilot

copilot mcp add (GitHub Copilot CLI)

windsurf

adds an entry to ~/.codeium/windsurf/mcp_config.json, keeping the others

A config file that is not valid JSON is left untouched, and the entry to add by hand is printed instead; a config that is a link into a dotfiles repository is written through the link. When a client that is registered through its own command is not installed, install says so and prints the command to run later. Cursor and Windsurf are files, so their entry is written whether or not the editor is installed yet. On Windows, a client installed through npm is a .cmd shim that install cannot start; it prints the command for you to run instead. Then restart the client and ask its agent: "Use SceneScout to test http://localhost:3000".

What has been checked: registering through each command above was run against Codex CLI, Gemini CLI, GitHub Copilot CLI and VS Code, and Cursor's command line agent read the entry install wrote, connected and listed the tools. The Windsurf path follows its documentation. A full test session has been run in Claude Code, with and without the skill. If a client behaves differently for you, a correction is welcome (say which client version you checked).

To register by hand instead, the server entry is always the same command, npx -y scenescout serve:

{
  "mcpServers": {
    "scenescout": { "command": "npx", "args": ["-y", "scenescout", "serve"] }
  }
}
{
  "servers": {
    "scenescout": { "type": "stdio", "command": "npx", "args": ["-y", "scenescout", "serve"] }
  }
}
[mcp_servers.scenescout]
command = "npx"
args = ["-y", "scenescout", "serve"]
{
  "mcpServers": {
    "scenescout": { "command": "npx", "args": ["-y", "scenescout", "serve"] }
  }
}
{
  "mcpServers": {
    "scenescout": { "command": "npx", "args": ["-y", "scenescout", "serve"] }
  }
}
{
  "mcpServers": {
    "scenescout": { "command": "npx", "args": ["-y", "scenescout", "serve"], "disabled": false, "autoApprove": [] }
  }
}
{
  "context_servers": {
    "scenescout": { "command": "npx", "args": ["-y", "scenescout", "serve"], "env": {} }
  }
}

Most clients accept the same mcpServers JSON shape shown for Cursor.

The method travels with the server. The tools are only hands and eyes; skills/scenescout/SKILL.md is the method: what to look at first, when to stop, what counts as a finding. Claude Code loads it as a skill. Every other client gets the same text from the server, with nothing to copy:

  • the server's instructions tell the agent to call scout_playbook before its first attach, and that tool returns the method,

  • clients that list server prompts as commands also get an explore prompt, which loads the method and takes an optional URL, level and focus.

So in any client, a first message like "Use SceneScout to test http://localhost:3000" is enough. If an agent starts clicking without having called scout_playbook, tell it to call that first; how closely a model follows server instructions varies by client.

The CLI is also useful on its own:

npx -y scenescout scan <path>       # project discovery: framework, routes, saved logins
npx -y scenescout status <path>     # what every session of a running engine is doing right now
npx -y scenescout watch <path>      # the same, live in your browser, with each session's page
npx -y scenescout login <url> --role admin   # sign in once in a visible browser; sessions attach with role: "admin"

📁 Project layout

src/
  mcp-server.ts     the 29 tools + per-session dispatch
  scan.ts           project discovery (framework, routes, auth)
  cli.ts            scan · serve · install · doctor · check · ci · login · status · watch
  check-run.ts      drives a check: attach, crawl every route, collect what was measured
  ci-run.ts         drives a CI run: the MCP server as a child, the model's API, the agent loop
  login-run.ts      drives `scenescout login`: a visible browser, Enter to save the role's profile; or --script, headless from the environment
  installer.ts      setup logic (skill link, MCP registration, diagnostics)
  engine/
    browser.ts      the engine class: attach, snapshot, actions, crawl, plans
    probes.ts       in-page scroll + overlay + focus probes (needs a browser too)
    fingerprint.ts  route + element-set identity (state hashing)
    oracles.ts      console/page/network/HTTP error detection
    injection.ts    the DOM-injection oracle's rules (what to watch for, how to find it)
    claims.ts       when the page contradicts the server (refused_empty, false_success)
    request.ts      what a replayed API call may be and where it may go
    brief.ts        splitting the app between parallel lanes
    lane.ts         the typed report a lane hands back
    calibration.ts  whether a lane's confidence held up; what it judged and never filed
    pace.ts         how a run spent its time
    bench.ts        scoring a run against the demo app's answer key
    policy.ts       the write-policy safety net
    ownership.ts    safe-write: which records were created in this process?
    uploads.ts      disk uploads, fenced to the project by real path
    journey.ts      task-ease measurement from the action log
    design.ts       the design audit + page scoring
    memory.ts       cross-run storage + finding dedup
    profiles.ts     saved sign-ins: role names, where a profile lives, owner-only files, attach by role, sessionStorage restore
    refresh.ts      the refresh broker: which values are a role's refresh tokens, the lock beside the profile, swapping a spent token
    scripted-login.ts  a CI sign-in: env and flags, TOTP (RFC 6238), which field is which, redaction
    expiry.ts       how long a saved sign-in lasts: cookie dates and JWT exp, checked before lanes start
    report.ts       the gap ledger + report generation
    check.ts        the check's rules, gate, report and SARIF
    ci.ts           a CI run's options, provider choice, caps, key redaction, tools and files
    provider.ts     the Anthropic and OpenAI message shapes, and retries
    replay.ts       the run as one page: steps, tasks, frames under each finding
    …               collector · dispatch · fixtures · authloss · reaper
scripts/            the test suites (smoke/ holds the real-browser ones)
test-app/           fixtures for the real-browser smoke tests
skills/scenescout/   the testing method (SKILL.md): a skill in Claude Code, served by the server everywhere else
docs/how-it-works.md  what happens at each stage, in diagrams
docs/benchmark.md   measuring whether a change made runs better
docs/adr/           why it's built this way

Design principle: logic that doesn't need Playwright lives outside browser.ts, so it can be unit-tested without launching a browser. That's why fingerprint, policy, memory, report, etc. are their own modules.


🧠 Design decisions

How it works, stage by stage — diagrams of the run lifecycle, what happens inside one action, the write policy on the wire, how a violation becomes a finding, how a parallel run is split and folded, how roles hand work to each other, where a run's time goes, and how a lane's confidence is checked afterwards.

Measuring whether a change helped — the demo app's answer key, the scorecard (recall, precision, judged-not-filed, severity, calibration), and the results log of every run, including what did not help.

The load-bearing choices are recorded as ADRs — read the relevant one before changing a rule it covers:


🔧 Development

Working on SceneScout itself is the only reason to clone it:

git clone https://github.com/brunoboto96/SceneScout.git scenescout && cd scenescout
npm install        # installs dependencies and builds
npm run setup      # same as `scenescout install`, but registers THIS checkout (the skill is linked, so edits are live)
npm test           # build + 24 suites: 22 pure-logic suites (scan, oracle, policy, … bench, hygiene),
                   #                     then smoke and mcp-check (the server over stdio), both with real browsers
npm run bench -- --all   # re-score every archived benchmark run against the current answer key
npm run demo       # regenerate examples/ from the demo app

Contributing? Start with VISION.md (what is in scope) and CONTRIBUTING.md (how changes land), then see AGENTS.md for the house rules — chiefly: bug fixes need a regression test at the cheapest layer that can fail, keep the repo project-agnostic (ADR 6), and npm test must pass.

🔐 Security

Found a way past the write policy, or another security problem? Please report it privately — see SECURITY.md.

📄 License

MIT.

  • Structured render-state, not pixels. Element lists with geometry; screenshots reserved for pixel-native residue (canvas, rendering glitches). Images that failed to load are reported from the DOM, including ones whose URL answered 200 with something that is not an image.

  • Diff snapshots with stable refs. Re-snapshots return only what changed (10.7 kB → 0.7 kB on a 130-element page); old refs stay valid.

  • Geometry oracles. Overlap and off-screen defects computed from layout boxes.

  • Oracles after every action. Console errors, page errors, failed requests, HTTP 4xx/5xx drained into every tool result — and DOM injection: a markup-shaped value the agent typed that later renders as an element on any page (stored or reflected XSS).

  • Multi-role, genuinely concurrent. Commands to different sessions run in parallel; safe-write ownership is shared, so role A can create what role B approves. The report renders a role capability matrix.

  • Task ease, not just correctness. scout_journey measures interaction cost, distinct screens, path, and backtracks.

  • Design audit with page scores. Two tiers (⚠ measurable defects / → craft suggestions incl. AI-slop tells), per-page 0–100 score persisted per route, plus an automatic overlay/modal probe on every snapshot. Shared shell scored once, separately.

  • Scrolls like a user — and notices when it can't. Reports SCROLL LOCKED for a leaked modal scroll-lock, finds the real inner scroll pane on app-shell layouts, and flags UNREACHABLE controls clipped inside overflow:hidden.

  • Uploads like a user. Answers a styled file-chooser or sets a hidden input directly, with a valid in-memory fixture; filePath is fenced to the project under test; files violating accept are flagged at selection.

  • Auth via Playwright storage states. Expired tokens caught at attach; repeated login-bounces raise SESSION AUTH LOST (a session attached by role first re-attaches once from its role's latest saved profile and carries on); a bounced route is recorded as not covered — a dead session can't certify routes it never reached.

  • A trustworthy gap ledger. Entries must be actionable (a search box or wizard sub-step isn't "form filled but never submitted"); API/download URLs never enter the route contract.

  • Honest reporting. Shared chrome counted once, stale scores marked, role matrix compares only roles that actually attempted a route.

  • Cross-run written knowledge. scout_note curates .scenescout/ASSUMPTIONS.md — app model, personas, constraints, risks — in prose.

  • Daemon-grade robustness. Per-tool watchdogs, orphaned-browser reaping, bounded teardown, live status via scenescout status <project>, and a live view of every session's page: the agent gives you its address when it attaches, or run scenescout watch <project> (loopback only, read-only, nothing written to disk: ADR 7).

Available Tools

30 tools
scout_attachA

Launch a browser and attach to a running web app. First attach in this conversation and you have read neither the SceneScout skill nor scout_playbook? Call scout_playbook before this. Write policy is enforced at the NETWORK layer: mode='observe' blocks EVERY request that is not a GET (login and token refresh excepted) — choose it for a target that holds real data, where even an ordinary form submission would create a record; mode='read-only' (default) blocks destructive-labeled elements AND all PUT/PATCH/DELETE + destructive POSTs, but lets ordinary form POSTs through; mode='safe-write' allows creating data and permits updates/deletes ONLY on resources this session created (use when the user wants create/edit flows tested); mode='destructive' allows everything — ONLY when the user explicitly confirmed a disposable/seeded environment. Pass role to sign in with a login the user saved by scenescout login <url> --role <name>, or a Playwright storage-state JSON as storageStatePath. Pass session to keep MULTIPLE roles alive at once (one browser each, genuinely concurrent) for collaboration testing — target each directly with every tool's session param, or use scout_session to set which one is the default; coverage and findings merge into one project memory.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYesBase URL of the running app, e.g. http://localhost:3000
modeNoWrite policy (see tool description). Never choose 'destructive' yourself — user opt-in only.read-only
roleNoSign in as a role whose login was saved with `scenescout login <url> --role <name>` (kept in the project's .scenescout/auth/). Every session given the same role gets its own browser built from that one login. Not with storageStatePath. No saved login for the role: the attach is refused and names the command to run — ask the user to run it, since it opens a window for them to sign in.
taskNoWhat this session is doing right now, shown under its objective from the moment it appears, e.g. Signing in and taking stock. Passed alone it is read as the 2.0 spelling of `objective`. Defaults to a placeholder so a fresh card never reads as idle.
headedNoShow the browser window
paceMsNoA floor between actions, in milliseconds, for when a person is watching and needs to keep up — following a flow, taking notes, demonstrating. Default 0: as fast as the page allows, which is what a run wants otherwise. Changeable mid-run with scout_session {paceMs}.
recordNoKeep a frame of the page after every action, under .scenescout/recordings/, and show it beside that step in report.html. Off by default: a recording is pictures of the app under test sitting in the project folder. Turn it on for QA work, where the run is evidence and not only a report.
browserNoBrowser to drive. Default: the SCENESCOUT_BROWSER environment variable, else chromium. firefox and webkit must be downloaded first (scenescout install --browser-only --browsers firefox). Use them for a cross-browser pass; stay on chromium otherwise.
sessionNoSession name for multi-role runs (e.g. 'admin', 'qa'). Creates/replaces that session's browser and makes it the default. Default: 'default'.
objectiveNoThis session's objective: the whole remit you were given, in one sentence ("Admin lane: §2 registers, §7 plan gating", "Approve and reject orders as a manager"). It sits above the task, which is what the session is doing at any moment. Shown to whoever is watching the run; worth setting whenever more than one session is live.
projectPathYesAbsolute path to the project (memory + report live in .scenescout/ here)
navTimeoutMsNoHow long a page may take to load, in ms. Default: the SCENESCOUT_NAV_TIMEOUT_MS environment variable, else 20000 (crawled pages 15000; a value set here applies to them too). Raise it only when timeouts come from a loaded machine rather than the app.
trustedEmbedsNoOrigins of embedded frames (e.g. "https://pay.example.com") whose writes out of the app may go out — ONLY when the user named them, typically a provider in test mode, and only in safe-write mode. Never add one yourself. Hostile input, repeated-click probes and uploads stay refused in them.
viewportWidthNoViewport width (default 1280); use e.g. 390 for a mobile pass
viewportHeightNoViewport height (default 900)
actionTimeoutMsNoHow long one click, keystroke, hover or pick may take, in ms. Default: the SCENESCOUT_ACTION_TIMEOUT_MS environment variable, else 5000. Raise it only when timeouts come from a loaded machine rather than the app.
storageStatePathNoOptional Playwright storage-state JSON path for authenticated exploration. Not with `role`.

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden and does so unusually well: it discloses that write policy is enforced at the NETWORK layer, exactly what each mode blocks and permits, the auth failure behavior (attach refused, command named, user must sign in), that multiple sessions mean genuinely concurrent browsers, and that coverage/findings merge into one project memory.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose is front-loaded and the long paragraph is held together with dashes and semicolons, so a reader can scan to the mode they need. The opening sentence about the scout_playbook gate is contorted ('First attach in this conversation and you have read neither...'), which costs a little clarity in a description that is heavy by necessity for 17 parameters.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 17-parameter tool with no annotations and no output schema, the description covers the decisions that matter: write policy, authentication, multi-session orchestration, and prerequisites. It stops short of describing what a successful attach returns or how long it takes, though most remaining knobs (timeouts, pace, recording) are self-documented in the schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds real semantic depth beyond the schema — the four-way mode taxonomy and its consequences, the role/session interplay, and how session targeting works across every other tool. It meaningfully reduces the chance of a wrong mode or auth choice.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Launch a browser and attach to a running web app'), and immediately positions itself relative to siblings by gating on scout_playbook and explaining session/role handling versus scout_session. An agent can distinguish this from scout_navigate or scout_scan without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit when-to-use for every mode ('choose observe for a target that holds real data', 'safe-write when the user wants create/edit flows tested', 'destructive ONLY when the user explicitly confirmed a disposable/seeded environment'), plus a prerequisite rule (call scout_playbook first) and when to pass role vs storageStatePath vs session.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

scout_backA

Go back in browser history (tests back-button resilience).

ParametersJSON Schema
NameRequiredDescriptionDefault
taskNoWhat you are DOING right now, in a few words: the action, not the acceptance criteria. "Filtering the documents register by status", "Filling the deviation form with invalid dates", "Signing in as QA_Team" — NOT "§2.4 filtering narrows the set and the filter is reflected in the URL", which is what you are CHECKING, not what you are doing. Naming the item you are on is fine ("§2.4: filtering the documents register"); keep the rest to what a colleague would see over your shoulder. It stays set until you pass a different one, so a batch costs a few words, not one per call. Required on the tools that act unless a journey or an earlier call already set one.
sessionNoTarget this session directly instead of the active one — pass it explicitly when dispatching to MULTIPLE sessions in one turn (e.g. two scout_click calls with different `session`), which then run CONCURRENTLY rather than queueing. Omit for single-session sequential use.
objectiveNoOld name for `task` (2.0). Prefer `task`.

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations exist, so the description carries the full burden. It fully states the behavior (browser history navigation) and the testing intent (back-button resilience). For a simple navigation action, this is adequate disclosure; no hidden side effects are suggested.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The entire description is one short sentence with no filler. The core behavior comes first and the testing rationale follows in a parenthetical, preserving structure and brevity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-required-parameter navigation action with fully documented optional parameters and no output schema, the description is complete enough to invoke. It would benefit from explicit sibling routing, but that gap is minor for such a simple tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the parameter descriptions are rich, so the baseline is 3. The tool description itself adds no parameter-level meaning, but it does not need to because the schema already explains task, session, and objective.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb and resource ('Go back in browser history') and attaches a clear test purpose ('tests back-button resilience'). This makes it immediately distinguishable from siblings like scout_navigate or scout_crawl without needing their schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The parenthetical implies the tool should be used when exercising back-button behavior, but it does not explicitly state when not to use it or name alternatives. The usage context is inferable, not spelled out.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

scout_captureA

Save a PNG of ONE element — its bounds plus a margin, from a real screenshot — under .scenescout/captures/, to show a person how it looks. Give the element's ref from the latest scout_snapshot. Not a way to judge design: scout_design_audit measures it.

ParametersJSON Schema
NameRequiredDescriptionDefault
keyNoInstead of a ref: the element's key from an earlier scout_capture result, to capture the same element on another deployment
refNoThe element's ref from the latest scout_snapshot
nameNoThe file's name: letters, digits and dashes. Default: capture
marginNoCSS pixels kept around the element. Default 8
sessionNoTarget this session directly instead of the active one — pass it explicitly when dispatching to MULTIPLE sessions in one turn (e.g. two scout_click calls with different `session`), which then run CONCURRENTLY rather than queueing. Omit for single-session sequential use.

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It discloses the output format (PNG), content (element bounds plus margin), and write location (.scenescout/captures/), but does not address permissions, error cases, overwrite behavior, or what is returned.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, front-loaded with the core action and output path. The ref-sourcing instruction and the design-audit exclusion each earn their place, with no wasted wording.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a no-output-schema, no-annotation tool with five parameters, the description covers purpose, element selection, output location, and sibling routing well. It does not state what the tool returns (e.g., a capture key), though the schema's key parameter implies a return artifact.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all five parameters are already documented in the schema. The description reinforces that ref comes from the latest scout_snapshot and adds conceptual meaning to margin, but does not add syntax or constraints beyond what the schema provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: save a PNG of ONE element from a real screenshot, with scope (bounds plus margin) and output location (.scenescout/captures/). It distinguishes itself from the full-page sibling scout_screenshot by emphasizing 'ONE element' and from scout_design_audit by name.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Tells the agent to supply the element's ref from the latest scout_snapshot, and explicitly routes design judgment to scout_design_audit. It implies the contrast with full-page scout_screenshot via 'ONE element' but does not name that alternative directly.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

scout_clickA

Click an element by its ref from the latest scout_snapshot. Returns the outcome plus any oracle violations triggered. clicks=2 (or 3) probes IMPATIENT-USER behaviour: a rapid multi-click that fires the same state-changing request twice means the control is not guarded against double submission (button stays enabled, endpoint not idempotent) — use it on every important submit/create button once; the result says explicitly whether duplicates fired.

ParametersJSON Schema
NameRequiredDescriptionDefault
refYesElement ref, e.g. e12
taskNoWhat you are DOING right now, in a few words: the action, not the acceptance criteria. "Filtering the documents register by status", "Filling the deviation form with invalid dates", "Signing in as QA_Team" — NOT "§2.4 filtering narrows the set and the filter is reflected in the URL", which is what you are CHECKING, not what you are doing. Naming the item you are on is fine ("§2.4: filtering the documents register"); keep the rest to what a colleague would see over your shoulder. It stays set until you pass a different one, so a batch costs a few words, not one per call. Required on the tools that act unless a journey or an earlier call already set one.
clicksNo1 = normal; 2-3 = rapid repeated clicks (double-submit probe)
sessionNoTarget this session directly instead of the active one — pass it explicitly when dispatching to MULTIPLE sessions in one turn (e.g. two scout_click calls with different `session`), which then run CONCURRENTLY rather than queueing. Omit for single-session sequential use.
objectiveNoOld name for `task` (2.0). Prefer `task`.

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations exist, so the description must carry behavior disclosure. It explains that repeated clicks can fire the same state-changing request twice and that the result reports whether duplicates occurred, plus oracle violations. It does not warn that a click can navigate away or trigger irreversible side effects, but it does characterize the tool as state-changing and outcome-reporting.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The first sentence states purpose, the second summarizes output, and the remaining clause explains the non-obvious clicks behavior. It is dense but every element is informative; the long clicks explanation earns its place because it conveys a testing strategy rather than restating the schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description at least explains that the response includes an outcome and oracle violations, and it covers the key usage twist (double-click probing). It leaves the exact result shape and handling of task/session parameters to the schema, which is acceptable for a relatively simple interaction tool, though a bit more on failure modes would help.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema already covers all parameters, but the description adds real value: ref is tied to the latest scout_snapshot, and clicks is explained as a deliberate double-submit probe with interpretation guidance. This goes beyond the schema's generic '1 = normal; 2-3 = rapid repeated clicks'.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The first sentence names a concrete action and resource: click an element identified by a ref from the latest scout_snapshot. It also states the return content, distinguishing it from sibling actions like hover, type, or navigate without relying on the tool name alone.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit operational guidance: use clicks=2/3 on every important submit/create button as a double-submit probe. It does not name alternatives or exclusions, but the 'latest scout_snapshot' constraint and button-specific recommendation provide adequate context for when to invoke it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

scout_closeA

Close a session's browser (memory persists on disk). Default: the DEFAULT session. Pass session to close a specific one, or all=true to close every live session at the end of a multi-role run. A session that is a lane of a parallel run (named by scout_lane_brief or scout_lane_report) is not closed until its lane report has been accepted by scout_lane_report, since folding needs the session attached; the refusal names every such lane.

ParametersJSON Schema
NameRequiredDescriptionDefault
allNoClose every live session
forceNoClose even a lane whose report has not been folded yet. Its decisions are then lost: nothing is kept for calibration and nothing checks its defects were filed. Prefer folding it with scout_lane_report first.
sessionNoSession to close (default: the default session)

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden, and it delivers: it discloses persistence, default/all behavior, lane-close refusal, and the destructive consequence of force. It also mentions that refusal names every such lane, so the agent knows what to expect. No contradictions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, front-loaded with the core action and default behavior. The lane sentence is long but carries necessary exception logic; no filler or repeated schema content.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a state-changing tool with no annotations and no output schema, it covers the main behaviors, including edge cases (lanes, force). It doesn't describe return values or error handling beyond the refusal, but that's a minor gap for a close operation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema already documents all three parameters at 100% coverage, so the baseline is 3. The description adds operational context: when to use all, what force destroys, and the default session behavior. This is meaningful added value, though the schema itself already covers force well.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Close a session's browser.' It also states the key side effect (memory persists on disk) and distinguishes default, specific, and all sessions, which separates it from sibling tools like scout_session or scout_snapshot.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives explicit operational guidance: default session, session parameter for a specific one, and all=true for multi-role runs. It also explains when not to close—lane sessions awaiting acceptance—and directs the agent to scout_lane_report for folding. It doesn't enumerate all alternatives, but the key routing is present.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

scout_coverageA

Show exploration coverage: states visited across all runs, which elements remain unexercised, which options of a dropdown used this run no session has chosen yet, and which forms seen this run no session has submitted with every text field blank. Use to decide where to explore next and when the level's budget is satisfied.

ParametersJSON Schema
NameRequiredDescriptionDefault
sessionNoTarget this session directly instead of the active one — pass it explicitly when dispatching to MULTIPLE sessions in one turn (e.g. two scout_click calls with different `session`), which then run CONCURRENTLY rather than queueing. Omit for single-session sequential use.

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It clearly implies a read-only analysis operation ('show', 'use to decide') and details the specific data it surfaces. However, it does not explicitly state that it is non-destructive, nor does it mention permissions, rate limits, or any side effects—gaps that matter for a tool with zero annotation coverage.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences with no fluff. The first sentence front-loads the purpose with a comprehensive list of coverage dimensions, and the second sentence provides actionable guidance. Every phrase earns its place, making it easy to scan and parse.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's simplicity (one optional param, no output schema), the description adequately explains what the tool reports and when to use it. The only gap is the lack of information about the output format (e.g., text, table, JSON), but since the tool is clearly a reporting utility, an agent can infer the response type. This is a minor omission rather than a critical one.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema covers the single optional parameter `session` with a detailed explanation of when to pass it explicitly. Since schema description coverage is 100%, the baseline is 3. The tool description itself adds little beyond the schema's parameter documentation, only referencing 'across all runs' without elaborating on the session parameter's behavior.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('show') and a clear resource ('exploration coverage'), then enumerates exactly what is covered: visited states, unexercised elements, unchosen dropdown options, and blank-submitted forms. It also states its decision-making purpose, distinguishing it from sibling tools like scout_scan or scout_screenshot.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly says 'Use to decide where to explore next and when the level's budget is satisfied,' giving clear when-to-use context. However, it does not mention when NOT to use it or name alternative tools, so it stops short of the explicit exclusion and alternative guidance that would merit a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

scout_crawlA

Engine-side route sweep in ONE call: visits each path (default: all known routes not yet visited), records states into coverage memory, and returns a per-route health summary (HTTP status, element count, oracle violations, dead-ends, auth-redirects). Navigation-only — safe in read-only mode. Use this FIRST for broad coverage; explore interactively only where it flags problems or where journeys matter.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathsNoPaths to visit, e.g. ['/orders','/settings']. Omit to crawl all unvisited known routes.
sessionNoTarget this session directly instead of the active one — pass it explicitly when dispatching to MULTIPLE sessions in one turn (e.g. two scout_click calls with different `session`), which then run CONCURRENTLY rather than queueing. Omit for single-session sequential use.

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden and largely succeeds: it discloses the mutation profile ('Navigation-only — safe in read-only mode'), a side effect (records into coverage memory), and the return contents. It does not cover rate limits, auth requirements, or whether the crawl runs synchronously, which would be useful but are not critical for safe invocation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three dense sentences with zero filler: capability and outcome are front-loaded, safety follows, then usage strategy. Every clause earns its place, and the key differentiator ('Use this FIRST') is placed at the end as actionable guidance.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no output schema, the description compensates by enumerating the per-route health summary fields. Combined with full schema descriptions for both parameters and explicit usage guidance, almost everything an agent needs to invoke it correctly is present. Minor omissions: behavior when no unvisited routes exist and any timing/asynchronicity expectations.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the schema already documents both parameters. The description adds mild conceptual context (e.g., default behavior of visiting unvisited routes, purpose of recording into coverage memory), but does not substantially extend beyond the schema's own parameter descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource ('Engine-side route sweep... visits each path'), enumerates concrete outcomes (records states into coverage memory, returns per-route health summary with HTTP status, element count, oracle violations, dead-ends, auth-redirects). It differentiates from siblings like scout_scan and scout_navigate by being a broad one-call sweep versus interactive exploration.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly instructs 'Use this FIRST for broad coverage; explore interactively only where it flags problems or where journeys matter.' This gives both a when-to-use and a when-not-to-use, with the alternative approach (interactive exploration) implied, so an agent can sequence its actions correctly.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

scout_design_auditA

Computed-style design audit of the current page — a design connoisseur's read WITHOUT screenshots. Measurable defects (⚠): WCAG contrast, tiny targets, clipped text, aspect-distorted images, horizontal overflow, missing keyboard-focus indicators (sampled with real Tab presses). Craft suggestions (→): line measure and line-height rhythm, spacing-scale adherence, typography entropy, palette discipline (gray census, accent hue families, pure-#000 body text), elevation/control consistency, heading structure, indistinguishable links, and AI-slop tells (gradient text, glassmorphism, side-stripe borders, neon glows, violet gradients, identical card grids). Ends with a SYSTEM SUMMARY of design-system coherence. Run once per representative page; the → tier is improvement feedback — file genuine opportunities as ux-polish findings with the concrete numbers, not just defects.

ParametersJSON Schema
NameRequiredDescriptionDefault
sessionNoTarget this session directly instead of the active one — pass it explicitly when dispatching to MULTIPLE sessions in one turn (e.g. two scout_click calls with different `session`), which then run CONCURRENTLY rather than queueing. Omit for single-session sequential use.

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden and does a good job: it declares the audit is computed-style without screenshots, samples keyboard focus with 'real Tab presses', and lists both measurable defects and craft suggestions. It does not address side effects or permissions, but as a read-only audit the implied behavior is clear.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long but densely informative, with clear structure: purpose, defect types, suggestion types, summary, and usage note. It front-loads the core identity and every subsequent clause adds practical detail. A little trimming would be possible, but the length serves the complexity of the tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has no output schema and no annotations, the description compensates well by enumerating what the audit detects, the structure (defects vs. suggestions vs. system summary), and how to use the results. It does not specify the exact return format, but the content coverage is sufficient for an agent to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The only parameter, session, is already fully documented in the input schema with clear guidance on when to pass it explicitly. The description adds no additional parameter semantics, which is acceptable because schema coverage is 100%.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: a 'computed-style design audit of the current page', and explicitly distinguishes itself with 'WITHOUT screenshots'. It enumerates distinct audit categories (WCAG contrast, tiny targets, clipped text, etc.), making its purpose unmistakable compared to siblings like scout_screenshot or scout_scan.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives clear usage context: 'Run once per representative page' and instructs that improvement-tier findings should be filed as ux-polish findings with concrete numbers rather than just as defects. It does not explicitly name alternative tools or exclusion criteria, but the context is specific enough to guide correct invocation.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

scout_findingA

Record a structured finding (bug, UX issue, or improvement). Deduplicates across runs; automatically captures the recent action trace as the repro. Use for anything worth reporting: crashes, oracle violations you confirmed, dead ends, confusing UX, permission leaks, missing testids — and design-audit improvement opportunities (ux-polish) with their concrete measurements.

ParametersJSON Schema
NameRequiredDescriptionDefault
titleYesOne-line summary of the defect
detailYesWhat happened, what was expected, and the evidence
sessionNoTarget this session directly instead of the active one — pass it explicitly when dispatching to MULTIPLE sessions in one turn (e.g. two scout_click calls with different `session`), which then run CONCURRENTLY rather than queueing. Omit for single-session sequential use.
categoryYesPick the closest — use 'other' only when nothing fits
evidenceNoCanonical machine signature for dedup, e.g. 'GET /api/reports/dashboard 403' or 'widget dashboard-summary-widget shows 0'. Same bug re-found later should produce the same string.
severityYes
conventionNoOnly for a WORTH-A-LOOK finding: the observation is real, and it is a defect only under a convention of this project you cannot see. Name that convention, e.g. 'a 4px spacing scale' or 'test ids on every control'. The report lists it under "Worth a look", apart from the defects, and does not count it as one. Not for "I could not tell": leave that unfiled or look closer. Omit for a defect.

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does well: it discloses deduplication across runs, automatic capture of the recent action trace as the repro, and the distinction between a real defect and a convention-dependent 'worth a look' finding. It never covers permissions, whether findings can be edited/deleted, or how the record surfaces in reports, which keeps it short of a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with the core action, no filler. The second sentence is a long run-on enumeration of reportable categories, which is information-dense but slightly bloated rather than wasteful.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 7-parameter ingestion tool with no output schema, the description covers the key behavioral contracts (dedup, auto-trace capture, convention handling) plus usage triggers. It is complete enough to call confidently; only return behavior and lifecycle of the record are unaddressed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 86%, so the schema already documents nearly all parameters, and the baseline is 3. The description adds some conceptual framing (ux-polish needs 'concrete measurements', oracle violations must be 'confirmed') but never maps to specific fields like evidence's dedup-signature format or the session concurrency parameter, which the schema already handles.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Starts with a specific verb+resource ('Record a structured finding') and immediately scopes it to bugs, UX issues, or improvements, which distinguishes it from sibling tools like scout_note or scout_report. It also names concrete artifact types (crashes, permission leaks, missing testids), so the agent knows exactly what kind of object this tool produces.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly enumerates when to use it ('anything worth reporting: crashes, oracle violations you confirmed, dead ends, confusing UX...'), which gives strong context. It even routes a common edge case ('I could not tell') away from filing, but it never names a concrete alternative tool (e.g. scout_note) for lighter-weight capture, so it stops short of full when/when-not guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

scout_hoverA

Hover an element by ref like a user pausing the pointer on it, and report what it reveals: tooltips/popovers (diffed against pre-hover state), any other new page text that appeared (labelled as possibly unrelated on busy pages), the title attribute, and aria-describedby text — each item truncated to 300 chars. Hovering does not count as exercising the element. Use on badges, icons, truncated text, and error indicators BEFORE concluding an element 'does nothing' — hover-gated UI is invisible to snapshots and clicks.

ParametersJSON Schema
NameRequiredDescriptionDefault
refYesElement ref, e.g. e12
sessionNoTarget this session directly instead of the active one — pass it explicitly when dispatching to MULTIPLE sessions in one turn (e.g. two scout_click calls with different `session`), which then run CONCURRENTLY rather than queueing. Omit for single-session sequential use.

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries full behavioral disclosure responsibility. It explains that results are diffed against pre-hover state, unrelated text may be labelled, items are truncated to 300 characters, and hover does not count as exercising the element. This gives a clear and honest model of the tool's behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three dense sentences with no filler: the first defines behavior and outputs, the second states a key caveat, and the third provides usage guidance. The most important information is front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Despite having no output schema and no annotations, the description explains the result composition, truncation, and possible noise. Combined with complete parameter schema and explicit usage conditions, an agent has everything it needs to invoke and interpret the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%: both ref and session already have full descriptions, including the concurrency behavior of session. The tool description adds little per-parameter detail, but that is acceptable because the schema already handles parameter semantics.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb and resource: hover an element by ref and report what it reveals. It enumerates concrete outputs (tooltips, new text, title attribute, aria-describedby) that clearly distinguish it from sibling tools like scout_click and scout_snapshot.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly says to use it on badges, icons, truncated text, and error indicators before concluding an element 'does nothing', which is strong when-to-use guidance. It also notes that hovering does not count as exercising the element, implying it is not a substitute for interaction, though it does not name scout_click as the explicit alternative.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

scout_journeyA

Measure how EASY a real task is, not just whether it works — the question pass/fail e2e suites never answer. Wrap one user goal: scout_journey {action:'start', goal:'Create an order'}, perform it the way a first-time user would (navigate by CLICKING through the UI, not by jumping to a known deep URL — a shortcut invalidates the measurement), then scout_journey {action:'end', completed:true|false}. Returns interaction cost (clicks, navigations, distinct screens, elapsed), the actual path taken, and friction signals: BACKTRACKS (returning to a screen already left — the clearest sign the next step wasn't discoverable), screen count, and over-interaction. Run it on each module's primary journey; an abandoned journey is a high-severity finding.

ParametersJSON Schema
NameRequiredDescriptionDefault
goalNoFor start: the user-facing task, e.g. 'Create an order and assign it'
noteNoFor end: what made it hard or easy, in one line
actionYes'start' before attempting the task, 'end' when done or blocked
sessionNoTarget this session directly instead of the active one — pass it explicitly when dispatching to MULTIPLE sessions in one turn (e.g. two scout_click calls with different `session`), which then run CONCURRENTLY rather than queueing. Omit for single-session sequential use.
completedNoFor end: did the user actually achieve the goal? false is a strong finding.

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries full behavioral burden and does so well: it discloses that the tool actually performs the journey by clicking, that shortcuts invalidate the measurement, what metrics are returned, and how to interpret abandonment. The live side-effect potential is implied by the 'Create an order' example, but the 'perform it' wording makes the behavior clear.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four dense sentences cover purpose, lifecycle, execution rule, return metrics, and usage guidance. There is no filler; the opening contrast with pass/fail e2e orients the agent, and the lifecycle example is front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Even without an output schema, the description enumerates the returned interaction metrics and friction signals, defines the invalidating shortcut, and states the severity of an abandoned journey. Session handling is covered by the schema's detailed session parameter, so an agent has enough to invoke and interpret the tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, and the description adds a concrete usage template pairing action:'start' with goal and action:'end' with completed, plus a real example goal. It doesn't add separate semantics for session/note, but those are already well documented in the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific measurement goal — task ease rather than pass/fail — and defines a clear start/end journey lifecycle with concrete signals (backtracks, interaction cost, over-interaction). This clearly distinguishes scout_journey from execution-oriented siblings like scout_click or scout_navigate.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly says when to use it: wrap a user goal with start/end and run it on each module's primary journey. It also emphasizes the correct execution method (clicking through UI, no deep-link shortcuts) and treats abandoned journeys as high-severity. However, it doesn't name alternative tools or state when not to use it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

scout_lane_briefA

For a run split across parallel agents (lanes). Divides the app's known routes between N lanes and returns each lane's session name, the objective to attach it with, and the routes it owns — whole modules per lane, balanced by route count, so no two lanes audit the same area and none is left unopened. Call it after the first crawl, when route knowledge is complete. Touches no browser; pass the briefs to your lane agents, then use scout_lane_report for what they hand back.

ParametersJSON Schema
NameRequiredDescriptionDefault
goalNoWhat the whole run is for; each lane's objective is written against it
lanesYesHow many lanes to split across (1–8)
routesNoRoutes to split. Omit to split every route this project knows about.
sessionNoTarget this session directly instead of the active one — pass it explicitly when dispatching to MULTIPLE sessions in one turn (e.g. two scout_click calls with different `session`), which then run CONCURRENTLY rather than queueing. Omit for single-session sequential use.
runMinutesNoHow long the lanes will run, in minutes (default 60). When they attach by a saved role, the brief is refused if that login will not last this long plus the margin.
expiryMarginMinutesNoHow long past the run a saved role's login must still last, in minutes (default 10).

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does well: it discloses that the tool 'touches no browser' and that the caller must pass the briefs to lane agents itself. It also explains the balancing guarantee (no overlap, none left unopened) and the return shape, which is needed since there is no output schema. It doesn't cover failure modes or what happens if lanes exceed available routes.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four sentences, purpose front-loaded, zero filler. It is dense but every clause adds routing or behavioral information. Slightly long, and the return-value clause could be trimmed, but nothing is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 6-param tool with no annotations and no output schema, the description supplies the missing pieces: what it returns, that it has no browser side effects, and the ordering constraint relative to the crawl. Remaining gaps are minor — no explicit mention of what a refused brief looks like or how to handle lane counts exceeding route counts.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. The description adds genuine meaning beyond the schema by explaining the splitting algorithm ('whole modules per lane, balanced by route count'), which tells the agent how the `lanes` and `routes` inputs actually partition work rather than merely restating their types.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource — it divides an app's known routes across N parallel lanes and returns each lane's session name, objective, and owned routes. The scope ('whole modules per lane, balanced by route count, no two lanes audit the same area') is precise enough to separate it from scout_crawl or scout_lane_report without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit timing: 'Call it after the first crawl, when route knowledge is complete.' It also routes the agent forward, naming scout_lane_report as the counterpart for what lanes hand back. It stops short of stating when not to use it (e.g. a single-agent run with no parallel split).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

scout_lane_reportA

For a run split across parallel agents (lanes). Without reply: returns the paragraph to put in a lane's prompt, telling it to hand back ONE typed JSON object (verdicts, severities and categories from closed sets, a calibrated confidence per decision, routes covered, what blocked it). With reply: parses what the lane handed back and returns the one-line fold (defects, highs, unsure, mean confidence, routes) or the reason it was refused, to relay to the lane once. Touches no browser.

ParametersJSON Schema
NameRequiredDescriptionDefault
laneYesThe lane's name, as used in its session
replyNoThe text the lane handed back; omit to get the instruction instead

TDQS

A4.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It adds a helpful trait ('Touches no browser') and explains the two output behaviors, including the failure case ('reason it was refused'). However, it does not explicitly state whether it has any state-changing side effects or mention permissions/rate limits, leaving some ambiguity for a tool with no annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is information-dense and well-structured: it opens with the run context, then branches into the two modes with concrete outcomes, and closes with a safety note. Every sentence earns its place, though the density requires careful reading. It is slightly long but not wasteful.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, so the description must explain return values, and it does: it details the paragraph content for the no-reply mode and the one-line fold or refusal reason for the reply mode. It also explains the input contract for the lane. Missing are details like pagination or error semantics, but for a tool of this complexity it is reasonably complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description adds value beyond the schema by explicitly linking the two parameters: omitting `reply` yields the instruction, while providing it triggers parsing. This conditional relationship is not obvious from the schema alone, so the description enhances parameter understanding.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb+resource pair: it returns a lane prompt paragraph when `reply` is omitted and parses a lane's reply when `reply` is provided. The opening context 'For a run split across parallel agents (lanes)' clearly frames this as a lane-specific tool, distinguishing it from sibling tools like scout_report or scout_lane_brief.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit usage conditions: use it when a run is split across lanes, and it precisely states the two modes based on whether `reply` is omitted or provided. It does not name alternatives or exclusions, but the context is clear enough that an agent can decide when to invoke this tool over siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

scout_navigateA

Navigate to a URL or a path relative to the attached base URL (e.g. '/orders'). Also supports 'back' via scout_back.

ParametersJSON Schema
NameRequiredDescriptionDefault
taskNoWhat you are DOING right now, in a few words: the action, not the acceptance criteria. "Filtering the documents register by status", "Filling the deviation form with invalid dates", "Signing in as QA_Team" — NOT "§2.4 filtering narrows the set and the filter is reflected in the URL", which is what you are CHECKING, not what you are doing. Naming the item you are on is fine ("§2.4: filtering the documents register"); keep the rest to what a colleague would see over your shoulder. It stays set until you pass a different one, so a batch costs a few words, not one per call. Required on the tools that act unless a journey or an earlier call already set one.
targetYesAbsolute URL or path like /settings
sessionNoTarget this session directly instead of the active one — pass it explicitly when dispatching to MULTIPLE sessions in one turn (e.g. two scout_click calls with different `session`), which then run CONCURRENTLY rather than queueing. Omit for single-session sequential use.
objectiveNoOld name for `task` (2.0). Prefer `task`.

TDQS

A3.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

There are no annotations, so the description alone must disclose behavior. It only says 'Navigate' and mentions an 'attached base URL' without explaining side effects, page-load waiting, session/state changes, auth requirements, or return value. For a tool that changes the current page, this is a significant behavioral gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences with no fluff. The core action is front-loaded, and the back-navigation note is a useful addition. Every word earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with no annotations and no output schema, the description is too thin. It does not explain what 'attached base URL' means, what happens after navigation, or what the agent should expect in return. With 28 siblings, it also lacks guidance on when this tool is the right choice compared to other navigation or action tools.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. The description adds the concept of an attached base URL and a '/orders' example, which slightly enriches the `target` parameter meaning. However, it adds nothing about `task`, `session`, or `objective` beyond what the schema already provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Navigate') and a precise resource ('a URL or a path relative to the attached base URL'), with a concrete example. It also distinguishes itself from the sibling scout_back by explicitly routing 'back' navigation to that tool. An agent can clearly tell what this tool does and how it differs from at least one close sibling.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives a clear exclusion by stating that 'back' should go via scout_back. However, it does not mention when to prefer this over other navigation-related siblings like scout_crawl or scout_scroll. The intended use is implied (direct URL/path navigation) but not explicitly contrasted with alternatives beyond the back case.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

scout_noteA

Cumulative WRITTEN knowledge about the tested app — .scenescout/ASSUMPTIONS.md, in prose a human can read and correct. memory.json stores coverage; this stores UNDERSTANDING, so every run starts smarter than the last. READ it at the start of every session ({action:'read'}). ADD durable learnings as you go ({action:'add', section, note}): what the app is for (app-model), who each role is and what they're FOR — infer the persona from what the role can see and do, e.g. 'qa-role = reviewer: approves orders, cannot administer' (roles), UI patterns the app follows (conventions), rules discovered the hard way like 'an order can only ship once approved' (constraints), fragile areas worth re-testing every run (risks), domain terms (glossary), and how to get the app testable at all — the command that regenerates an expired login state, what has to be running (setup), which the engine reads back to you the next time a storage state has expired. Notes are dated, attributed to the acting role, and deduplicated. Do NOT record session-specific facts (ids, counts) — only durable knowledge.

ParametersJSON Schema
NameRequiredDescriptionDefault
noteNoFor add: the learning, one or two sentences, written for a future reader with no context
actionYes'read' the accumulated knowledge, or 'add' one durable learning
sectionNoFor add: which knowledge section this belongs to
sessionNoTarget this session directly instead of the active one — pass it explicitly when dispatching to MULTIPLE sessions in one turn (e.g. two scout_click calls with different `session`), which then run CONCURRENTLY rather than queueing. Omit for single-session sequential use.

TDQS

A4.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With zero annotations, the description carries the full burden and discloses the cumulative/persistent nature ('every run starts smarter than the last'), the file location, and the side effects of adds: 'Notes are dated, attributed to the acting role, and deduplicated.' It stops short of exhaustive because edge cases like first-run behavior when ASSUMPTIONS.md does not yet exist, and the exact shape of the read response, are left implicit.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Every sentence earns its place and the purpose is front-loaded, but the whole description is one dense paragraph with a ~70-word parenthetical section list and a dangling clause about engine read-back at the end. The length is proportionate to the tool's conceptual richness (7 sections, 2 actions, content policy), but the lack of structural breaks makes it harder for an agent to scan.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no annotations and no output schema, the description alone must cover purpose, storage location, usage timing, section semantics, content policy, and behavioral traits — and it covers all of these with concrete examples. Nothing an agent needs to invoke read/add correctly is missing; the only minor gaps (exact read return shape, first-run file creation) are self-evident for this kind of tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so baseline is 3; the description adds real value by giving a concrete example for every section enum — e.g., roles: 'qa-role = reviewer: approves orders, cannot administer' and constraints: 'an order can only ship once approved.' It also reinforces note-format guidance and session concurrency semantics, meaningfully elevating what the schema states.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The opening sentence states the resource and function precisely: 'Cumulative WRITTEN knowledge about the tested app — .scenescout/ASSUMPTIONS.md.' It further disambiguates from coverage tracking with 'memory.json stores coverage; this stores UNDERSTANDING,' and the twin actions read/add are made explicit. An agent can distinguish this from scout_coverage or scout_finding without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives direct operational directives: 'READ it at the start of every session' and 'ADD durable learnings as you go,' plus an explicit when-not rule: 'Do NOT record session-specific facts (ids, counts) — only durable knowledge.' It also names the storage-state trigger ('reads back to you the next time a storage state has expired') and explains the alternative store for coverage, so an agent knows exactly when this tool is the right one.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

scout_playbookA

Return the SceneScout testing method: setup order, write modes, how to explore, what counts as done, how to report. Call this ONCE before the first scout_attach in a conversation, then follow it. If this client offers a SceneScout skill, load that instead — it is the same text, so never read both. Takes no input and touches no browser.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden, and it states 'Takes no input and touches no browser,' making the read-only, side-effect-free nature clear. It also warns that calling it again is unnecessary and that it duplicates the skill text.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tightly packed sentences: content, usage timing/alternative, and input/side-effect constraints. No filler, and the most decision-relevant information appears first.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a no-input, no-output-schema tool, the description fully covers what the tool returns, when to invoke it, how it relates to the skill alternative, and that it has no browser side effects. Nothing an agent needs to call or follow it is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, so the schema already says nothing is needed. The description adds the explicit confirmation 'Takes no input,' which aligns with the empty schema, earning the zero-parameter baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb and resource: 'Return the SceneScout testing method' and enumerates its contents (setup order, write modes, exploration, done criteria, reporting). This clearly distinguishes it from the sibling execution tools like scout_attach and scout_report.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly says when to call it: 'ONCE before the first scout_attach in a conversation, then follow it.' It also names an alternative and provides a rule: if a SceneScout skill is offered, load that instead and 'never read both.'

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

scout_pressA

Press a keyboard key (e.g. Escape, Tab, Enter) — useful for closing modals and testing keyboard navigation.

ParametersJSON Schema
NameRequiredDescriptionDefault
keyYes
taskNoWhat you are DOING right now, in a few words: the action, not the acceptance criteria. "Filtering the documents register by status", "Filling the deviation form with invalid dates", "Signing in as QA_Team" — NOT "§2.4 filtering narrows the set and the filter is reflected in the URL", which is what you are CHECKING, not what you are doing. Naming the item you are on is fine ("§2.4: filtering the documents register"); keep the rest to what a colleague would see over your shoulder. It stays set until you pass a different one, so a batch costs a few words, not one per call. Required on the tools that act unless a journey or an earlier call already set one.
sessionNoTarget this session directly instead of the active one — pass it explicitly when dispatching to MULTIPLE sessions in one turn (e.g. two scout_click calls with different `session`), which then run CONCURRENTLY rather than queueing. Omit for single-session sequential use.
objectiveNoOld name for `task` (2.0). Prefer `task`.

TDQS

A4.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden. It states the core behavior and even hints at consequences (modal dismissal, keyboard navigation), but it does not disclose details such as whether the press targets the active page, whether it dispatches keydown/keyup, or how unsupported keys are handled. This is acceptable for a simple action but not richly transparent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single front-loaded sentence with no filler. It immediately states the action, gives examples, and then explains when it is useful — every word earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple key-press tool with a well-described parameter schema, the description is largely sufficient. It covers what the tool does and typical use cases. The only gap is an explicit relationship to sibling tools like scout_type, but this is minor given the overall clarity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 75%, and the tool description adds essential meaning to the undocumented `key` parameter by providing valid examples. The other parameters already have detailed schema descriptions, so the description's contribution is focused and useful for the one parameter that needs it.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action: 'Press a keyboard key' with concrete examples (Escape, Tab, Enter) and signals the intended resource (keyboard). It differentiates the tool from close/click/type siblings by focusing on key-based interaction rather than mouse or text entry, making intent unmistakable.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear use cases: 'closing modals and testing keyboard navigation.' It does not explicitly mention when to avoid this tool or name alternatives like scout_type, but the context is clear enough for an agent to choose it appropriately for key-press actions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

scout_reportA

Generate the final markdown report — findings, page quality scores (worst first), role capability matrix, oracle rollup, and the GAP LEDGER (an explicit list of what was NOT tested). Writes the full document to .scenescout/report.md and returns a bounded SUMMARY (full reports exceed client token limits). Gates by level: 'minimal' needs all routes visited + ≥1 design audit; 'medium' additionally needs several routes audited; 'extensive' REFUSES while the gap ledger is non-empty — that refusal is the completeness guarantee: an extensive report only generates when nothing known is left untested. force=true overrides (only when the user capped the budget).

ParametersJSON Schema
NameRequiredDescriptionDefault
forceNoGenerate even though gates are unmet (only when the user capped the budget)
levelNoWhich completion contract to enforce — match the level the run was asked formedium
historyNoHow much of the history to print. 'index' lists findings from earlier runs, and resolved ones, as a row each: id, severity, age, title. 'full' prints every one in full as before — on one project that was 1.75 MB against 113 KB, nearly half of it findings already fixed. Use 'full' when handing the document to someone who has no access to the memory.index
sessionNoTarget this session directly instead of the active one — pass it explicitly when dispatching to MULTIPLE sessions in one turn (e.g. two scout_click calls with different `session`), which then run CONCURRENTLY rather than queueing. Omit for single-session sequential use.

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It discloses that the tool writes a file to .scenescout/report.md, returns a bounded summary because full reports exceed token limits, can refuse based on gates, and that extensive mode's refusal is intentional as a completeness guarantee. This is unusually transparent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Every clause carries operational information and the critical output behavior is front-loaded. The length is justified by the number of gates and modes, though a later reference to 'oracle rollup'/role capability matrix relies on domain jargon.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no output schema, the description supplies the needed return behavior (bounded summary) and side effect (full doc at .scenescout/report.md). Gate logic, force semantics, history modes, and multi-session behavior are all covered, leaving no obvious gap for calling the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

All four parameters are already fully described in the schema (100% coverage), so the baseline is 3. The tool description adds meaningful context beyond the schema: it expands level into concrete gate requirements, explains why force should only be used on a budget cap, gives a concrete size example for history='full', and notes that separate session values run concurrently instead of queueing.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Clearly names the deliverable ('final markdown report'), enumerates its contents (findings, page quality scores worst-first, capability matrix, oracle rollup, gap ledger), and pinpoints the exact write target `.scenescout/report.md`. This is distinguishable from siblings like `scout_lane_report` because it is the aggregate final report, not a lane or coverage view.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides explicit level-dependent gate conditions ('minimal' needs all routes visited + ≥1 design audit; 'medium' adds several routes audited; 'extensive' refuses on a non-empty gap ledger) and a strict condition for force=true (only when the user capped the budget). It does not name alternatives such as scout_lane_report, so it stops short of a full when-not-to-use comparison.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

scout_requestA

Call the app's own API as this session, with the UI bypassed — the check that turns a hidden or disabled control into a proven refusal. A button that is not shown proves nothing; the same action refused by the server does. The fetch runs IN the page, so it carries the session's cookies and replays the Authorization header the app itself last sent, and it passes through the same interception the write policy is enforced on: in safe-write a mutation on a record this session did not create is refused here exactly as it would be for a click, and that refusal is the engine's safety net, not a finding. Returns the status line, the timing, the headers that decide whether two responses are truly identical (content-type, location, www-authenticate, retry-after, cache-control), and the body. Unlike a shell call, every request is recorded in the run's trail and its signature is what a finding should quote. Paths are fenced to the attached origin: use another session to reach another host.

ParametersJSON Schema
NameRequiredDescriptionDefault
bodyNoRequest body, sent as application/json unless a content-type header is given
pathYesPath on the attached origin, e.g. /api/things/12, or a full URL on that same origin
taskNoWhat you are DOING right now, in a few words: the action, not the acceptance criteria. "Filtering the documents register by status", "Filling the deviation form with invalid dates", "Signing in as QA_Team" — NOT "§2.4 filtering narrows the set and the filter is reflected in the URL", which is what you are CHECKING, not what you are doing. Naming the item you are on is fine ("§2.4: filtering the documents register"); keep the rest to what a colleague would see over your shoulder. It stays set until you pass a different one, so a batch costs a few words, not one per call. Required on the tools that act unless a journey or an earlier call already set one.
methodNoDefault GET
headersNoExtra headers. One given here wins over the app's own, which is how a session tests a different or absent credential.
sessionNoTarget this session directly instead of the active one — pass it explicitly when dispatching to MULTIPLE sessions in one turn (e.g. two scout_click calls with different `session`), which then run CONCURRENTLY rather than queueing. Omit for single-session sequential use.
objectiveNoOld name for `task` (2.0). Prefer `task`.

TDQS

A4.3/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden, and it delivers richly: it discloses in-page execution carrying session cookies and replaying the app's last Authorization header, passage through the write-policy interception layer, the interpretation of safe-write refusals as a safety net rather than a finding, the exact response composition (status line, timing, identity-determining headers, body), run-trail recording, and origin fencing. This goes well beyond what annotations would normally convey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long (~180 words) but densely packed: every sentence earns its place, from the front-loaded purpose through the safety-net clarification, the return-value specification, and the trail-signature guidance. It is structured with the core purpose first and operational caveats following. Slightly verbose, but the tool's complexity — arbitrary methods including DELETE, write-policy interaction, and findings-grade evidence — justifies nearly every clause.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a high-complexity tool with no annotations and no output schema, the description covers the critical bases: purpose, authentication context, write-policy interception, response shape, run-trail recording, and scope restriction. Minor gaps remain — no mention of rate limits, network-error behavior, or side effects of mutation methods beyond the safe-write note — but an agent has what it needs to invoke this tool correctly and interpret results.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the structured fields already document all seven parameters and the baseline is 3. The description adds some contextual enrichment — e.g., "Paths are fenced to the attached origin" clarifies the `path` parameter's bounds, and the Authorization-header replay explains what `headers` overrides — but most of its content concerns tool behavior, not parameter-level semantics. The schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource — "Call the app's own API as this session, with the UI bypassed" — and immediately establishes its distinct role among the scout_* siblings. It explicitly contrasts its proof value with UI actions: "A button that is not shown proves nothing; the same action refused by the server does," which differentiates it from scout_click and scout_navigate without ambiguity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives a clear when-to-use rationale (verify server-side refusal when a control is hidden or disabled), contrasts with shell calls ("Unlike a shell call, every request is recorded in the run's trail"), and provides when-not guidance: safe-write refusals are "the engine's safety net, not a finding," and paths are fenced to the attached origin ("use another session to reach another host"). It stops short of explicitly naming sibling alternatives such as scout_click as the preferred path when a control is visible.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

scout_resolveA

Mark a finding as resolved (by its id, shown when recorded and in the report). Resolved findings move to the report's green ✅ Resolved section, and reopen automatically as flagged REGRESSIONS if re-found later. Use when the user says a bug is fixed, or when re-testing shows the evidence no longer reproduces.

ParametersJSON Schema
NameRequiredDescriptionDefault
idNoAlias for `findingId`.
sessionNoTarget this session directly instead of the active one — pass it explicitly when dispatching to MULTIPLE sessions in one turn (e.g. two scout_click calls with different `session`), which then run CONCURRENTLY rather than queueing. Omit for single-session sequential use.
findingIdNoFinding id, e.g. a1b2c3d4e5

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and meets it by disclosing behavioral consequences: resolved findings move to the green Resolved section and reopen automatically as REGRESSIONS if re-found later. This adds meaningful context beyond the schema, though it does not address potential side effects like irreversibility or permissions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is three concise sentences with no filler. It front-loads the core action, then explains behavior and usage conditions. Every sentence contributes useful information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers purpose, usage triggers, and behavioral outcomes, which is sufficient for a simple action tool with no output schema and no required parameters. It does not explicitly explain the return value, but given the tool's simplicity and the strong schema coverage, the remaining gap is minor.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the parameters are already well-documented. The description adds minor value by hinting where to find the id ('shown when recorded and in the report'), but it does not elaborate on parameter formats beyond the schema, hence a baseline score.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb-resource pair ('Mark a finding as resolved') and clearly identifies the target (by id). It distinguishes itself from sibling tools by explaining the outcome (moving to the Resolved section and auto-reopening as REGRESSIONS), which is unique to this tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly states when to use it: when the user says a bug is fixed, or when re-testing shows the evidence no longer reproduces. It does not provide when-not-to-use guidance or name alternatives, but the usage context is clear and actionable.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

scout_run_planA

Execute up to 20 actions in ONE call — use for mechanical sequences (fill a form, walk a wizard) so each step doesn't cost a round-trip. Targets resolve at execution time by semantic locator: 'testid=…', 'text=…', 'label=…' or 'role=button[name=Save]' (never snapshot refs). The same steps, saved to .scenescout/flows/.json with expect-text / expect-url / expect-request steps added, are replayed by scenescout check on every pull request. An upload step attaches a file as scout_upload does (target required — the file input or the control that opens its chooser; value = a fixture kind or a project-relative path). The plan ABORTS at the first NEW oracle violation, policy refusal, or failed step, returning a transcript of how far it got; repeats of already-reported violations do not abort (they stay logged for the report).

ParametersJSON Schema
NameRequiredDescriptionDefault
taskNoWhat you are DOING right now, in a few words: the action, not the acceptance criteria. "Filtering the documents register by status", "Filling the deviation form with invalid dates", "Signing in as QA_Team" — NOT "§2.4 filtering narrows the set and the filter is reflected in the URL", which is what you are CHECKING, not what you are doing. Naming the item you are on is fine ("§2.4: filtering the documents register"); keep the rest to what a colleague would see over your shoulder. It stays set until you pass a different one, so a batch costs a few words, not one per call. Required on the tools that act unless a journey or an earlier call already set one.
stepsYes
sessionNoTarget this session directly instead of the active one — pass it explicitly when dispatching to MULTIPLE sessions in one turn (e.g. two scout_click calls with different `session`), which then run CONCURRENTLY rather than queueing. Omit for single-session sequential use.
objectiveNoOld name for `task` (2.0). Prefer `task`.

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden—and it delivers. It discloses execution-time target resolution, the prohibition on snapshot refs, the persistence and replay behavior, the upload semantics, the abort condition on new oracle violations/policy refusals/failed steps, the non-abort on repeated violations, and the transcript return. This is comprehensive behavioral disclosure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but well organized: purpose, locator semantics, replay behavior, upload specifics, and abort semantics each get a sentence. No filler is present, though the length is substantial; a short example or a tighter split would make it even easier to scan.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is complex, but the description covers the execution model, target resolution, replay purpose, upload behavior, and failure semantics, and it states that a transcript is returned. With no output schema, a bit more detail on the transcript's structure would help, but the core context an agent needs to invoke the tool correctly is present.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 75%, and the description adds meaning beyond the schema for key parameters: target locator syntax ('testid=…', 'text=…', 'label=…', 'role=button[name=Save]') and upload value semantics (fixture kind or project-relative path). It does not dwell on parameters already well described in the schema, like replace, pressEnter, and session, which is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Execute up to 20 actions in ONE call.' It clearly identifies the intended use case (mechanical sequences like filling a form or walking a wizard) and distinguishes itself from the individual scout_* sibling tools by emphasizing batching and the round-trip cost saving. The upload-step comparison to scout_upload further differentiates it.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly states when to use this tool: 'use for mechanical sequences (fill a form, walk a wizard) so each step doesn't cost a round-trip.' It also explains the replay workflow with scenescout check. It does not explicitly list exclusions or name alternative tools beyond scout_upload, so it falls short of a full 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

scout_scanA

Scan a project directory to discover the frontend workspace, framework, routes, dev command, Playwright auth storage states, and testid conventions. Run this first.

ParametersJSON Schema
NameRequiredDescriptionDefault
projectPathYesAbsolute path to the project root

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the burden of behavioral disclosure. It conveys a read-only reconnaissance behavior by saying 'scan a project directory to discover...' but does not mention whether state is persisted, what the return shape looks like, or if any environmental prerequisites exist. This leaves some behavioral ambiguity.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single efficient sentence plus the useful sequencing directive 'Run this first.' It frontloads the action and follows with a concrete list of discovery targets, containing no filler or redundant repetition of the tool name.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-parameter tool with no annotations and no output schema, the description provides the essential call context: what to pass, what will be discovered, and that it should be invoked first. It does not explicitly describe the return format, but the enumerated discovery targets partially compensate and make the tool callable.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The only parameter, projectPath, is already fully documented in the schema as 'Absolute path to the project root,' and the description adds no additional constraints, defaults, or format details. The discovered artifacts listed in the description concern tool output rather than parameter semantics, so the schema remains the primary source of parameter meaning.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('scan') and object ('project directory') and enumerates the concrete artifacts it uncovers: frontend workspace, framework, routes, dev command, Playwright auth storage states, and testid conventions. This makes its role as a project discovery tool unmistakable and distinguishes it from execution-oriented siblings. The directive 'Run this first' also makes its purpose immediately clear.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly instructs the agent to run this tool first, establishing clear sequencing relative to the many scout_* sibling tools. It does not name specific alternatives or conditions for choosing another tool, but the 'first' directive gives the agent enough contextual guidance for initial invocation.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

scout_screenshotA

Take a JPEG screenshot of the current viewport. LAST RESORT: geometry issues are in scout_snapshot and style/contrast/spacing issues are in scout_design_audit — images that failed to load are listed in scout_snapshot under BROKEN IMAGES — use a screenshot only for pixel-native content (a canvas, visual gestalt) that computed data cannot capture.

ParametersJSON Schema
NameRequiredDescriptionDefault
sessionNoTarget this session directly instead of the active one — pass it explicitly when dispatching to MULTIPLE sessions in one turn (e.g. two scout_click calls with different `session`), which then run CONCURRENTLY rather than queueing. Omit for single-session sequential use.

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the burden, and it does well by disclosing the JPEG format, viewport scope, and the 'LAST RESORT' nature of the tool. It could be slightly more explicit about side effects or what the caller receives, but the core behavior is transparent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is one dense, front-loaded sentence with no filler. It packs the main action first and the routing caveats after, though the long em-dash chain is slightly less scannable than a short two-sentence structure would be.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description is complete enough for selecting and invoking the tool: it says what is captured, in what format, and when not to use it. With no output schema, it slightly underspecifies the result format or delivery mechanism, but that is a minor gap for a simple screenshot tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the sole optional `session` parameter is fully documented in the schema. The description adds no parameter-specific detail, but none is required beyond what the schema already provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Take a JPEG screenshot of the current viewport.' It also clearly distinguishes itself from siblings by naming what belongs in scout_snapshot and scout_design_audit instead.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

This is exemplary routing guidance. It tells the agent to use this tool only for pixel-native content that computed data cannot capture, and explicitly directs geometry issues to scout_snapshot, style/contrast/spacing issues to scout_design_audit, and broken-image checks to the BROKEN IMAGES section of scout_snapshot.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

scout_scrollA

Scroll like a user — real apps hide their bugs below the fold. Reports the resulting position (px and %), and explicitly flags SCROLL LOCKED: scrollable content exists but the page will not move (the classic leaked modal scroll-lock that silently cuts users off from everything below the fold — snapshots also detect this passively as an OVERLAY line). Without target it scrolls the page, falling back to the largest scrollable pane on app-shell layouts. Pass target to scroll ONE region instead (a sidebar nav, a dialog body, a table pane): the page-level pick is the LARGEST scroll port, so a smaller region beside it never moves and its content looks truncated when it is only scrolled away — never call a nav item missing without scrolling its own container first. Use before judging a long page: the design audit measures at the current scroll position, so scroll + re-snapshot/re-audit deep sections; scroll also triggers lazy-loaded content whose failures then surface as oracle violations.

ParametersJSON Schema
NameRequiredDescriptionDefault
byNoScroll by px instead (positive = down). Default 600 when neither given.
toNoJump to an edge
targetNoScroll ONE region instead of the page: "testid=…", "text=…" or "label=…". Scrolls that element's nearest scrollable ancestor.
sessionNoTarget this session directly instead of the active one — pass it explicitly when dispatching to MULTIPLE sessions in one turn (e.g. two scout_click calls with different `session`), which then run CONCURRENTLY rather than queueing. Omit for single-session sequential use.

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full behavioral burden, and it is unusually rich: it discloses the fallback to the largest scrollable pane on app-shell layouts, that smaller regions beside it never move, that it reports SCROLL LOCKED, and that scrolling triggers lazy-loaded content. This goes well beyond what the schema alone reveals.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long and dense, but nearly every clause carries behavioral or workflow value, and it is front-loaded with the core action. Some rhetorical framing ('real apps hide their bugs below the fold') could be trimmed, and the parenthetical caveats make it a wall of text rather than a scannable definition, but it is efficient for the complexity it documents.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given 4 optional parameters, no output schema, and no annotations, the description covers the essential missing context: result format (px and %), SCROLL LOCKED detection, fallback behavior, target-region semantics, and why scrolling matters for audits and lazy-loaded content. The schema handles the remaining parameter details, so nothing critical is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. The description adds real value on top by explaining the target semantics in depth — single region, nearest scrollable ancestor, and the 'smaller region beside it' trap. It does not redundantly restate by/to/session, leaving the schema to cover those.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific behavior ('Scroll like a user'), the resource acted on (page or one target region), and the outcome (reports position in px and %, flags SCROLL LOCKED when scroll is blocked). It clearly differentiates itself from sibling tools like scout_snapshot and scout_design_audit by being the scroll-position-changing action that precedes them.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives explicit workflow guidance: use before judging a long page, then re-snapshot/re-audit deep sections, because the design audit measures at the current scroll position. It also warns against calling a nav item missing without scrolling its own container first, which is a concrete when-to-use versus when-to-target instruction.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

scout_selectC

Select an option in a by ref.

ParametersJSON Schema
NameRequiredDescriptionDefault
refYes
taskNoWhat you are DOING right now, in a few words: the action, not the acceptance criteria. "Filtering the documents register by status", "Filling the deviation form with invalid dates", "Signing in as QA_Team" — NOT "§2.4 filtering narrows the set and the filter is reflected in the URL", which is what you are CHECKING, not what you are doing. Naming the item you are on is fine ("§2.4: filtering the documents register"); keep the rest to what a colleague would see over your shoulder. It stays set until you pass a different one, so a batch costs a few words, not one per call. Required on the tools that act unless a journey or an earlier call already set one.
valueYesOption value or label
sessionNoTarget this session directly instead of the active one — pass it explicitly when dispatching to MULTIPLE sessions in one turn (e.g. two scout_click calls with different `session`), which then run CONCURRENTLY rather than queueing. Omit for single-session sequential use.
objectiveNoOld name for `task` (2.0). Prefer `task`.

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It does not disclose any behaviors beyond the literal act of selecting, such as whether it waits for options to load, how it handles missing values, or any side effects. It also does not mention that it may navigate or change application state.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Extremely concise, single sentence, front-loaded with the primary action. Only one sentence, but it is not verbose. However, it might be too sparse given the tool's complexity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description is minimal for a tool that manipulates a UI element. It lacks essential context such as how to identify the select, how values are matched, and any error conditions. Given there is no output schema and no annotations, the agent is left with too much uncertainty.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 80%, so most parameters are described in the schema. However, the description adds nothing about parameter meanings. 'ref' is ambiguous, and 'value' could be either an option value or label, which is already in the schema but not elaborated in the description. With partial coverage, the description should compensate but does not.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Select') and resource ('an option in a <select>') but is very terse. It does not distinguish itself clearly from siblings like scout_click or scout_press, and 'by ref' is vague without explaining what 'ref' refers to.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool versus alternatives like scout_click or scout_type. The description lacks any context about use cases, prerequisites, or exclusions. An agent would not know when to choose this over sibling tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

scout_sessionA

List live sessions, or set which one is the DEFAULT (used by any tool call that omits session). Prefer passing session directly on each tool call for multi-role work — that's what lets concurrent dispatch happen; scout_session is for sequential convenience (skip repeating session on every call) and for checking what's live. Both browsers stay live and authenticated regardless of which is default — re-snapshot a session after a break to see what changed while it was away.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameNoSession to make the default; omit to list sessions
paceMsNoChange how fast this session acts, mid-run: a floor between actions in milliseconds, for when a person is watching and needs to keep up. 0 restores full speed. With `name`, applies to that session; without, to every live session — which is what 'slow everything down so I can follow' means.
sessionNoAlias for `name`.

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations supplied, the description carries full behavioral burden and meets it: it discloses default-session persistence, that switching default does not deauthenticate other sessions ('Both browsers stay live and authenticated'), and the effect of paceMs on live sessions. It also hints at re-snapshotting after a break, adding operational behavior beyond the schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three dense sentences, front-loaded with the core action, then usage guidance, then behavioral caveats. Every sentence earns its place and no tautology or padding.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 3-optional-parameter tool with no output schema, this is nearly complete: purpose, default behavior, persistence, and speed control are covered. It does not describe the shape of the listing result or edge cases like unknown session names, but the stated behavior is enough for an agent to decide confidently.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% so the schema already documents each parameter. The description still adds value by explaining that the default is what is used when a tool call omits session, and by contrasting direct session passing with default-based convenience. Minor: it could restate paceMs semantics, but schema already covers them.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description opens with a precise dual function: 'List live sessions, or set which one is the DEFAULT' and explains what DEFAULT means. This distinguishes scout_session from sibling action tools like scout_snapshot or scout_navigate while making its resource ('sessions') explicit.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit decision guidance: prefer passing session on each call for multi-role/concurrent work, and use scout_session for sequential convenience or checking what's live. It names the context that should steer an agent away from this tool, which is exactly what a usage guideline should do.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

scout_snapshotA

Capture the current page state: URL, state fingerprint, interactable elements with refs (e1, e2, …), geometry issues, coverage, and oracle violations since the last action. Re-snapshotting the same route returns a DIFF (refs stay stable). Cheap — prefer this over screenshots.

ParametersJSON Schema
NameRequiredDescriptionDefault
fullNoForce a full element list instead of a diff
sessionNoTarget this session directly instead of the active one — pass it explicitly when dispatching to MULTIPLE sessions in one turn (e.g. two scout_click calls with different `session`), which then run CONCURRENTLY rather than queueing. Omit for single-session sequential use.

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the behavioral disclosure burden. It clearly discloses the diff behavior, stable refs, cheapness, and the 'since the last action' scope. It does not explicitly state whether the operation is read-only or discuss failure modes, but 'capture' implies non-mutation and the disclosed traits cover the most important behaviors.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and front-loaded, with the primary purpose stated in the first sentence and useful behavioral details following. Each sentence adds distinct value—purpose, diff behavior, and cost/alternative—with no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers what the snapshot contains, the diff behavior, the temporal scope ('since the last action'), and the important recommendation against screenshots. No output schema is present, so the listed content items serve as a sufficient description of what the agent will receive.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the input schema already documents both parameters in detail, including the full-diff toggle and session targeting semantics. The description itself adds no parameter-level meaning, which is acceptable given the schema's thoroughness; baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a strong verb and resource: 'Capture the current page state' and lists concrete outputs (URL, state fingerprint, interactable elements, refs, geometry issues, coverage, oracle violations). It also distinguishes itself from the screenshot sibling by stating it is cheaper and preferred, making the tool's identity clear.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly tells the agent to prefer this over screenshots and explains that re-snapshotting the same route returns a diff, which guides when to call it again. It does not enumerate exclusions relative to other scout_* siblings, but the guidance it gives is actionable and sufficient for the main decision.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

scout_typeA

Type into a text input/textarea/composer by ref, the way a real user does: if the field already holds content (e.g. an @-mention chip a menu click inserted), the text is APPENDED at the end — preserving that content — and the result reports what was already there (a separating space is added only at a word-to-word boundary). Pass replace=true to clear the field first (correcting a previous entry); an empty textValue always clears. Appending fires input events but not keydown, so keydown-driven triggers (slash/mention menus) will not react to appended text. Use for both valid values and boundary/fuzz values (empty, very long, unicode, script tags).

ParametersJSON Schema
NameRequiredDescriptionDefault
refYesElement ref, e.g. e12
taskNoWhat you are DOING right now, in a few words: the action, not the acceptance criteria. "Filtering the documents register by status", "Filling the deviation form with invalid dates", "Signing in as QA_Team" — NOT "§2.4 filtering narrows the set and the filter is reflected in the URL", which is what you are CHECKING, not what you are doing. Naming the item you are on is fine ("§2.4: filtering the documents register"); keep the rest to what a colleague would see over your shoulder. It stays set until you pass a different one, so a batch costs a few words, not one per call. Required on the tools that act unless a journey or an earlier call already set one.
valueNoAlias for `textValue`.
replaceNoClear the field before typing instead of appending to existing content
sessionNoTarget this session directly instead of the active one — pass it explicitly when dispatching to MULTIPLE sessions in one turn (e.g. two scout_click calls with different `session`), which then run CONCURRENTLY rather than queueing. Omit for single-session sequential use.
objectiveNoOld name for `task` (2.0). Prefer `task`.
textValueNoText to type
pressEnterNoPress Enter after typing

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full disclosure burden and exceeds it. It details append semantics, the separating-space rule, empty-text clearing, replace behavior, the exact event side effects (input fires, keydown does not), and the fact that the result reports pre-existing content. This is exemplary behavioral documentation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but every clause earns its place: purpose, append behavior, replace semantics, clearing rule, event caveat, and usage scope. It is front-loaded with the core action and then systematically covers edge cases. Length is justified by the behavioral complexity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with no annotations and no output schema, the description covers the full operational surface: how content is entered, when it is cleared, what side effects occur, and what the result reports. Nothing an agent needs to invoke it correctly or anticipate its behavior is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. The description adds meaningful context beyond the schema by explaining the interaction between textValue, replace, and empty-string clearing, and by clarifying which event behaviors apply. It enriches the parameters without restating their schema descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb ('Type') and resource ('text input/textarea/composer by ref'), and immediately distinguishes the tool's core behavior — appending vs. replacing — which separates it from sibling input tools like scout_click or scout_press. The purpose is unambiguous and action-oriented.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly states a broad use case ('valid values and boundary/fuzz values') and gives a critical constraint (keydown-driven triggers won't react to appended text), which effectively tells the agent when behavior may not match expectations. It does not name a sibling alternative (e.g., 'use scout_press for keydown-sensitive actions'), but the guidance is sufficiently clear for selection.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

scout_uploadA

Attach a file to an upload control the way a user does. ref is either a visible (snapshots list these with role file) or the button/label/dropzone that opens the file chooser — the chooser is intercepted and answered, which is how the hidden input behind a styled 'Choose file' control is reached. Omit ref to target the page's only file input, hidden or not (snapshots disclose hidden ones on a FILE INPUTS line). Nothing needs to exist on disk: a small VALID fixture (real PDF/PNG structure) is generated in memory, its kind inferred from the input's accept attribute or chosen with fixture; filePath uploads a real file but must live inside the attached project (fenced like navigation is fenced to the origin); name overrides the filename for boundary tests (wrong extension vs accept, very long, unicode). The result names the input, how the file reached it, flags a file that violates accept (a mismatch the app then accepts is a validation finding), warns if the app cleared the input after selection, and says whether a state-changing request fired on selection — if none did, click the form's submit, or check the next snapshot for a client-side rejection.

ParametersJSON Schema
NameRequiredDescriptionDefault
refNoElement ref of the file input OR of the control that opens the file chooser; omit when the page has exactly one file input
nameNoFilename override (default scenescout-fixture.<kind>, or the disk file's own name)
taskNoWhat you are DOING right now, in a few words: the action, not the acceptance criteria. "Filtering the documents register by status", "Filling the deviation form with invalid dates", "Signing in as QA_Team" — NOT "§2.4 filtering narrows the set and the filter is reflected in the URL", which is what you are CHECKING, not what you are doing. Naming the item you are on is fine ("§2.4: filtering the documents register"); keep the rest to what a colleague would see over your shoulder. It stays set until you pass a different one, so a batch costs a few words, not one per call. Required on the tools that act unless a journey or an earlier call already set one.
fixtureNoGenerated fixture kind; default: inferred from the input's accept attribute (pdf when there is none, or none we can generate)
sessionNoTarget this session directly instead of the active one — pass it explicitly when dispatching to MULTIPLE sessions in one turn (e.g. two scout_click calls with different `session`), which then run CONCURRENTLY rather than queueing. Omit for single-session sequential use.
filePathNoA real file to upload — absolute or relative to the project; must be inside the attached project. Exclusive with fixture.
objectiveNoOld name for `task` (2.0). Prefer `task`.

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries full behavioral burden and succeeds remarkably. It discloses the chooser interception, in-memory fixture generation, project-fencing of `filePath`, accept-violation flagging, warning if the input is cleared, and detection of state-changing requests. This goes far beyond a simple 'uploads a file' summary.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long but dense with operationally relevant detail; nearly every clause adds value. It is front-loaded with the core purpose and then proceeds from ref targeting to fixture/file selection to post-selection behavior. A few parentheticals could be trimmed, but the length is justified by the tool's complexity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity, zero annotations, and no output schema, the description is remarkably complete. It covers the input-targeting strategy, file source options, validation behavior, state-change detection, and necessary follow-up actions. An agent has enough context to invoke the tool correctly and interpret its result without additional structured metadata.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, but the description adds substantial meaning beyond the schema: `ref` resolution through styled controls, fixture kind inference from the accept attribute, project-fencing semantics for `filePath`, and `name` use for boundary tests. It also explains the result fields tied to these parameters, making parameter behavior concrete.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Attach a file to an upload control the way a user does.' It clearly distinguishes the tool's scope from general file operations by detailing `ref` targeting of file inputs, hidden inputs, and chooser-opening controls. The distinctive behavior of intercepting the chooser makes it unambiguous among sibling tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives extensive usage context: when to omit `ref`, how to distinguish fixture-based uploads from `filePath` uploads, and what to do after a selection that fires no state-changing request. It does not explicitly name alternative sibling tools or state when not to use this tool, but the parameter-level guidance is clear enough for correct invocation.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

scout_verifyA

Re-test findings earlier runs left open. With no arguments, returns the open findings in the order to re-test them — worst route first, grouped so a route is walked once — each with its evidence and repro steps. Pass ids to narrow it to specific findings. After re-testing one, call again with id and verdict to record what you saw: "gone" resolves it, "present" stamps it confirmed so the report stops calling it unverified, "changed" keeps it open and says the behaviour differs. Use after a fix wave, or at the start of a run against an app this project has tested before.

ParametersJSON Schema
NameRequiredDescriptionDefault
idNoThe finding being verified. Omit to get the worklist.
idsNoNarrow the worklist to these finding ids.
noteNoWhat you saw, in a sentence. Shown in the report beside the verdict.
sessionNoTarget this session directly instead of the active one — pass it explicitly when dispatching to MULTIPLE sessions in one turn (e.g. two scout_click calls with different `session`), which then run CONCURRENTLY rather than queueing. Omit for single-session sequential use.
verdictNoWhat the re-test found: "gone", "present" or "changed". Requires id.

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden, and it delivers: it discloses ordering (worst route first), grouping, returned content (evidence and repro steps), and the state effects of each verdict—'gone' resolves, 'present' confirms, 'changed' keeps open. This is unusually transparent about side effects.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Every sentence earns its place: the tool's action, the worklist behavior, narrowing, verdict recording, and usage timing. The information is front-loaded with the core purpose and flows logically without repetition or fluff.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with five parameters, no annotations, and no output schema, the description is remarkably complete. It covers what the worklist returns, how to narrow it, how to record verdicts, and what each verdict means, leaving no critical gap for an agent trying to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. The description adds meaning beyond the schema by explaining the interaction between id, ids, and verdict—including that a verdict is recorded by calling again with id and verdict—and clarifies what the verdict values do to report state.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: re-test findings left open by previous runs and record verdicts. It distinguishes itself from siblings like scout_scan or scout_resolve by describing a specific verification workflow with 'gone', 'present', and 'changed' outcomes.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives explicit when-to-use context: 'Use after a fix wave, or at the start of a run against an app this project has tested before.' It also explains the call pattern for listing, narrowing, and recording verdicts. It does not mention alternatives or when-not-to-use, so it stops short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 4 tool updatesv3.14.1
    • Changedscout_attach4 fields changed
      • addedInput schema / properties / actionTimeoutMs
        Added value: +{
        +  "description": "How long one click, keystroke, hover or pick may take, in ms. Default: the SCENESCOUT_ACTION_TIMEOUT_MS environment variable, else 5000. Raise it only when timeouts come from a loaded machine rather than the app.",
        +  "maximum": 120000,
        +  "minimum": 1000,
        +  "type": "integer"
        +}
      • addedInput schema / properties / navTimeoutMs
        Added value: +{
        +  "description": "How long a page may take to load, in ms. Default: the SCENESCOUT_NAV_TIMEOUT_MS environment variable, else 20000 (crawled pages 15000; a value set here applies to them too). Raise it only when timeouts come from a loaded machine rather than the app.",
        +  "maximum": 300000,
        +  "minimum": 1000,
        +  "type": "integer"
        +}
      • addedInput schema / properties / role
        Added value: +{
        +  "description": "Sign in as a role whose login was saved with `scenescout login <url> --role <name>` (kept in the project's .scenescout/auth/). Every session given the same role gets its own browser built from that one login. Not with storageStatePath. No saved login for the role: the attach is refused and names the command to run — ask the user to run it, since it opens a window for them to sign in.",
        +  "maxLength": 40,
        +  "type": "string"
        +}
      • changedInput schema / properties / storageStatePath / description
        Previous value: -"Optional Playwright storage-state JSON path for authenticated exploration"New value: +"Optional Playwright storage-state JSON path for authenticated exploration. Not with `role`."
    • Addedscout_capture
    • Changedscout_finding1 field changed
      • addedInput schema / properties / convention
        Added value: +{
        +  "description": "Only for a WORTH-A-LOOK finding: the observation is real, and it is a defect only under a convention of this project you cannot see. Name that convention, e.g. 'a 4px spacing scale' or 'test ids on every control'. The report lists it under \"Worth a look\", apart from the defects, and does not count it as one. Not for \"I could not tell\": leave that unfiled or look closer. Omit for a defect.",
        +  "maxLength": 160,
        +  "minLength": 1,
        +  "type": "string"
        +}
    • Changedscout_lane_brief2 fields changed
      • addedInput schema / properties / expiryMarginMinutes
        Added value: +{
        +  "description": "How long past the run a saved role's login must still last, in minutes (default 10).",
        +  "maximum": 240,
        +  "minimum": 0,
        +  "type": "integer"
        +}
      • addedInput schema / properties / runMinutes
        Added value: +{
        +  "description": "How long the lanes will run, in minutes (default 60). When they attach by a saved role, the brief is refused if that login will not last this long plus the margin.",
        +  "maximum": 1440,
        +  "minimum": 1,
        +  "type": "integer"
        +}
  2. 1 tool updatev3.11.1
    • Changedscout_run_plan1 field changed
      • changedInput schema / properties / steps / items / properties / target / description
        Previous value: -"testid=…, text=…, label=… (or a path for navigate; 'top'/'bottom'/±px for scroll; for upload: the file input or the control that opens its chooser)"New value: +"testid=…, text=…, label=…, role=<role>[name=\"…\"] (or a path for navigate; 'top'/'bottom'/±px for scroll; for upload: the file input or the control that opens its chooser)"
  3. 1 tool updatev3.10.0
    • Changedscout_close1 field changed
      • addedInput schema / properties / force
        Added value: +{
        +  "default": false,
        +  "description": "Close even a lane whose report has not been folded yet. Its decisions are then lost: nothing is kept for calibration and nothing checks its defects were filed. Prefer folding it with scout_lane_report first.",
        +  "type": "boolean"
        +}
  4. 1 tool updatev3.9.0
    • Changedscout_attach1 field changed
      • addedInput schema / properties / trustedEmbeds
        Added value: +{
        +  "description": "Origins of embedded frames (e.g. \"https://pay.example.com\") whose writes out of the app may go out — ONLY when the user named them, typically a provider in test mode, and only in safe-write mode. Never add one yourself. Hostile input, repeated-click probes and uploads stay refused in them.",
        +  "items": {
        +    "maxLength": 200,
        +    "type": "string"
        +  },
        +  "maxItems": 10,
        +  "type": "array"
        +}
  5. 16 tool updatesv3.4.0
    • Changedscout_attach4 fields changed
      • addedInput schema / properties / objective
        Added value: +{
        +  "description": "This session's objective: the whole remit you were given, in one sentence (\"Admin lane: §2 registers, §7 plan gating\", \"Approve and reject orders as a manager\"). It sits above the task, which is what the session is doing at any moment. Shown to whoever is watching the run; worth setting whenever more than one session is live.",
        +  "maxLength": 300,
        +  "type": "string"
        +}
      • addedInput schema / properties / paceMs
        Added value: +{
        +  "description": "A floor between actions, in milliseconds, for when a person is watching and needs to keep up — following a flow, taking notes, demonstrating. Default 0: as fast as the page allows, which is what a run wants otherwise. Changeable mid-run with scout_session {paceMs}.",
        +  "maximum": 60000,
        +  "minimum": 0,
        +  "type": "integer"
        +}
      • addedInput schema / properties / record
        Added value: +{
        +  "default": false,
        +  "description": "Keep a frame of the page after every action, under .scenescout/recordings/, and show it beside that step in report.html. Off by default: a recording is pictures of the app under test sitting in the project folder. Turn it on for QA work, where the run is evidence and not only a report.",
        +  "type": "boolean"
        +}
      • addedInput schema / properties / task
        Added value: +{
        +  "description": "What this session is doing right now, shown under its objective from the moment it appears, e.g. Signing in and taking stock. Passed alone it is read as the 2.0 spelling of `objective`. Defaults to a placeholder so a fresh card never reads as idle.",
        +  "maxLength": 300,
        +  "type": "string"
        +}
    • Changedscout_back2 fields changed
      • addedInput schema / properties / objective
        Added value: +{
        +  "description": "Old name for `task` (2.0). Prefer `task`.",
        +  "maxLength": 120,
        +  "type": "string"
        +}
      • addedInput schema / properties / task
        Added value: +{
        +  "description": "What you are DOING right now, in a few words: the action, not the acceptance criteria. \"Filtering the documents register by status\", \"Filling the deviation form with invalid dates\", \"Signing in as QA_Team\" — NOT \"§2.4 filtering narrows the set and the filter is reflected in the URL\", which is what you are CHECKING, not what you are doing. Naming the item you are on is fine (\"§2.4: filtering the documents register\"); keep the rest to what a colleague would see over your shoulder. It stays set until you pass a different one, so a batch costs a few words, not one per call. Required on the tools that act unless a journey or an earlier call already set one.",
        +  "maxLength": 120,
        +  "type": "string"
        +}
    • Changedscout_click2 fields changed
      • addedInput schema / properties / objective
        Added value: +{
        +  "description": "Old name for `task` (2.0). Prefer `task`.",
        +  "maxLength": 120,
        +  "type": "string"
        +}
      • addedInput schema / properties / task
        Added value: +{
        +  "description": "What you are DOING right now, in a few words: the action, not the acceptance criteria. \"Filtering the documents register by status\", \"Filling the deviation form with invalid dates\", \"Signing in as QA_Team\" — NOT \"§2.4 filtering narrows the set and the filter is reflected in the URL\", which is what you are CHECKING, not what you are doing. Naming the item you are on is fine (\"§2.4: filtering the documents register\"); keep the rest to what a colleague would see over your shoulder. It stays set until you pass a different one, so a batch costs a few words, not one per call. Required on the tools that act unless a journey or an earlier call already set one.",
        +  "maxLength": 120,
        +  "type": "string"
        +}
    • Addedscout_lane_brief
    • Addedscout_lane_report
    • Changedscout_navigate2 fields changed
      • addedInput schema / properties / objective
        Added value: +{
        +  "description": "Old name for `task` (2.0). Prefer `task`.",
        +  "maxLength": 120,
        +  "type": "string"
        +}
      • addedInput schema / properties / task
        Added value: +{
        +  "description": "What you are DOING right now, in a few words: the action, not the acceptance criteria. \"Filtering the documents register by status\", \"Filling the deviation form with invalid dates\", \"Signing in as QA_Team\" — NOT \"§2.4 filtering narrows the set and the filter is reflected in the URL\", which is what you are CHECKING, not what you are doing. Naming the item you are on is fine (\"§2.4: filtering the documents register\"); keep the rest to what a colleague would see over your shoulder. It stays set until you pass a different one, so a batch costs a few words, not one per call. Required on the tools that act unless a journey or an earlier call already set one.",
        +  "maxLength": 120,
        +  "type": "string"
        +}
    • Changedscout_note1 field changed
      • changedInput schema / properties / section / enum
        Previous value: -[
        -  "app-model",
        -  "roles",
        -  "conventions",
        -  "constraints",
        -  "risks",
        -  "glossary"
        -]New value: +[
        +  "app-model",
        +  "roles",
        +  "conventions",
        +  "constraints",
        +  "risks",
        +  "glossary",
        +  "setup"
        +]
    • Changedscout_press2 fields changed
      • addedInput schema / properties / objective
        Added value: +{
        +  "description": "Old name for `task` (2.0). Prefer `task`.",
        +  "maxLength": 120,
        +  "type": "string"
        +}
      • addedInput schema / properties / task
        Added value: +{
        +  "description": "What you are DOING right now, in a few words: the action, not the acceptance criteria. \"Filtering the documents register by status\", \"Filling the deviation form with invalid dates\", \"Signing in as QA_Team\" — NOT \"§2.4 filtering narrows the set and the filter is reflected in the URL\", which is what you are CHECKING, not what you are doing. Naming the item you are on is fine (\"§2.4: filtering the documents register\"); keep the rest to what a colleague would see over your shoulder. It stays set until you pass a different one, so a batch costs a few words, not one per call. Required on the tools that act unless a journey or an earlier call already set one.",
        +  "maxLength": 120,
        +  "type": "string"
        +}
    • Changedscout_report1 field changed
      • addedInput schema / properties / history
        Added value: +{
        +  "default": "index",
        +  "description": "How much of the history to print. 'index' lists findings from earlier runs, and resolved ones, as a row each: id, severity, age, title. 'full' prints every one in full as before — on one project that was 1.75 MB against 113 KB, nearly half of it findings already fixed. Use 'full' when handing the document to someone who has no access to the memory.",
        +  "enum": [
        +    "index",
        +    "full"
        +  ],
        +  "type": "string"
        +}
    • Addedscout_request
    • Changedscout_run_plan2 fields changed
      • addedInput schema / properties / objective
        Added value: +{
        +  "description": "Old name for `task` (2.0). Prefer `task`.",
        +  "maxLength": 120,
        +  "type": "string"
        +}
      • addedInput schema / properties / task
        Added value: +{
        +  "description": "What you are DOING right now, in a few words: the action, not the acceptance criteria. \"Filtering the documents register by status\", \"Filling the deviation form with invalid dates\", \"Signing in as QA_Team\" — NOT \"§2.4 filtering narrows the set and the filter is reflected in the URL\", which is what you are CHECKING, not what you are doing. Naming the item you are on is fine (\"§2.4: filtering the documents register\"); keep the rest to what a colleague would see over your shoulder. It stays set until you pass a different one, so a batch costs a few words, not one per call. Required on the tools that act unless a journey or an earlier call already set one.",
        +  "maxLength": 120,
        +  "type": "string"
        +}
    • Changedscout_select2 fields changed
      • addedInput schema / properties / objective
        Added value: +{
        +  "description": "Old name for `task` (2.0). Prefer `task`.",
        +  "maxLength": 120,
        +  "type": "string"
        +}
      • addedInput schema / properties / task
        Added value: +{
        +  "description": "What you are DOING right now, in a few words: the action, not the acceptance criteria. \"Filtering the documents register by status\", \"Filling the deviation form with invalid dates\", \"Signing in as QA_Team\" — NOT \"§2.4 filtering narrows the set and the filter is reflected in the URL\", which is what you are CHECKING, not what you are doing. Naming the item you are on is fine (\"§2.4: filtering the documents register\"); keep the rest to what a colleague would see over your shoulder. It stays set until you pass a different one, so a batch costs a few words, not one per call. Required on the tools that act unless a journey or an earlier call already set one.",
        +  "maxLength": 120,
        +  "type": "string"
        +}
    • Changedscout_session1 field changed
      • addedInput schema / properties / paceMs
        Added value: +{
        +  "description": "Change how fast this session acts, mid-run: a floor between actions in milliseconds, for when a person is watching and needs to keep up. 0 restores full speed. With `name`, applies to that session; without, to every live session — which is what 'slow everything down so I can follow' means.",
        +  "maximum": 60000,
        +  "minimum": 0,
        +  "type": "integer"
        +}
    • Changedscout_type2 fields changed
      • addedInput schema / properties / objective
        Added value: +{
        +  "description": "Old name for `task` (2.0). Prefer `task`.",
        +  "maxLength": 120,
        +  "type": "string"
        +}
      • addedInput schema / properties / task
        Added value: +{
        +  "description": "What you are DOING right now, in a few words: the action, not the acceptance criteria. \"Filtering the documents register by status\", \"Filling the deviation form with invalid dates\", \"Signing in as QA_Team\" — NOT \"§2.4 filtering narrows the set and the filter is reflected in the URL\", which is what you are CHECKING, not what you are doing. Naming the item you are on is fine (\"§2.4: filtering the documents register\"); keep the rest to what a colleague would see over your shoulder. It stays set until you pass a different one, so a batch costs a few words, not one per call. Required on the tools that act unless a journey or an earlier call already set one.",
        +  "maxLength": 120,
        +  "type": "string"
        +}
    • Changedscout_upload2 fields changed
      • addedInput schema / properties / objective
        Added value: +{
        +  "description": "Old name for `task` (2.0). Prefer `task`.",
        +  "maxLength": 120,
        +  "type": "string"
        +}
      • addedInput schema / properties / task
        Added value: +{
        +  "description": "What you are DOING right now, in a few words: the action, not the acceptance criteria. \"Filtering the documents register by status\", \"Filling the deviation form with invalid dates\", \"Signing in as QA_Team\" — NOT \"§2.4 filtering narrows the set and the filter is reflected in the URL\", which is what you are CHECKING, not what you are doing. Naming the item you are on is fine (\"§2.4: filtering the documents register\"); keep the rest to what a colleague would see over your shoulder. It stays set until you pass a different one, so a batch costs a few words, not one per call. Required on the tools that act unless a journey or an earlier call already set one.",
        +  "maxLength": 120,
        +  "type": "string"
        +}
    • Addedscout_verify
  6. 2 tool updatesv1.2.0
    • Changedscout_attach1 field changed
      • addedInput schema / properties / browser
        Added value: +{
        +  "description": "Browser to drive. Default: the SCENESCOUT_BROWSER environment variable, else chromium. firefox and webkit must be downloaded first (scenescout install --browser-only --browsers firefox). Use them for a cross-browser pass; stay on chromium otherwise.",
        +  "enum": [
        +    "chromium",
        +    "firefox",
        +    "webkit"
        +  ],
        +  "type": "string"
        +}
    • Addedscout_playbook
  7. 48 tool updatesv1.1.0
    • Removedft_attach
    • Removedft_back
    • Removedft_click
    • Removedft_close
    • Removedft_coverage
    • Removedft_crawl
    • Removedft_design_audit
    • Removedft_finding
    • Removedft_hover
    • Removedft_journey
    • Removedft_navigate
    • Removedft_note
    • Removedft_press
    • Removedft_report
    • Removedft_resolve
    • Removedft_run_plan
    • Removedft_scan
    • Removedft_screenshot
    • Removedft_scroll
    • Removedft_select
    • Removedft_session
    • Removedft_snapshot
    • Removedft_type
    • Removedft_upload
    • Addedscout_attach
    • Addedscout_back
    • Addedscout_click
    • Addedscout_close
    • Addedscout_coverage
    • Addedscout_crawl
    • Addedscout_design_audit
    • Addedscout_finding
    • Addedscout_hover
    • Addedscout_journey
    • Addedscout_navigate
    • Addedscout_note
    • Addedscout_press
    • Addedscout_report
    • Addedscout_resolve
    • Addedscout_run_plan
    • Addedscout_scan
    • Addedscout_screenshot
    • Addedscout_scroll
    • Addedscout_select
    • Addedscout_session
    • Addedscout_snapshot
    • Addedscout_type
    • Addedscout_upload
  8. 24 tool updatesv0.23.3
    • First observedft_attach
    • First observedft_back
    • First observedft_click
    • First observedft_close
    • First observedft_coverage
    • First observedft_crawl
    • First observedft_design_audit
    • First observedft_finding
    • First observedft_hover
    • First observedft_journey
    • First observedft_navigate
    • First observedft_note
    • First observedft_press
    • First observedft_report
    • First observedft_resolve
    • First observedft_run_plan
    • First observedft_scan
    • First observedft_screenshot
    • First observedft_scroll
    • First observedft_select
    • First observedft_session
    • First observedft_snapshot
    • First observedft_type
    • First observedft_upload

TDQS

A3.6/5.0

Scored across 30 tools

Disambiguation4/5

Most tools target clearly distinct actions (click, type, select, press, hover, scroll, upload, navigate) and the descriptions go out of their way to separate near-neighbors like scout_snapshot (state data), scout_capture (single-element PNG), scout_screenshot (last-resort viewport image) and scout_design_audit (computed styles). A couple of boundaries remain fuzzy — scout_back vs the 'back' mode of scout_navigate, and scout_snapshot vs scout_capture for visual inspection — so an agent could occasionally pick the wrong one, but nothing is genuinely redundant.

Naming Consistency5/5

Every tool uses the same scout_ prefix with a snake_case verb or verb_noun form (scout_attach, scout_capture, scout_run_plan, scout_design_audit), and the noun chosen matches the resource/action consistently across the set. There are no mixed conventions or stray camelCase names.

Tool Count3/5

At 30 tools this is well above the comfortable 3-15 band and sits in the 'heavy' zone, so an agent faces a large surface to learn. The breadth is partly justified by the genuinely broad domain (attach, scan, crawl, interact, audit, findings, coverage, lanes, reporting), and each tool maps to a real capability with little filler, but it is still on the heavy side of appropriate.

Completeness4/5

The surface covers the full test lifecycle: setup (scout_scan, scout_playbook), navigation/exploration (scout_navigate, scout_snapshot, scout_crawl, scout_run_plan), interaction (scout_click, scout_type, scout_select, scout_press, scout_hover, scout_scroll, scout_upload), analysis (scout_design_audit, scout_journey, scout_coverage), findings CRUD (scout_finding, scout_verify, scout_resolve), and reporting (scout_report, scout_lane_brief, scout_lane_report). Minor gaps exist — no explicit wait/assertion tool and no way to delete or list findings outside the report — but these are workable around.

Maintenance

ActivityMaintained
ResponsivenessResponsive

Related MCP Connectors

Related MCP Servers