SceneScout
Server Configuration
Describes the environment variables required to run the server.
| Name | Required | Description | Default |
|---|---|---|---|
No arguments | |||
Instructions
Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.
This server publishes no instructions, or was last inspected before Glama recorded them.
Capabilities
Features and capabilities supported by this server
Protocol revision2025-11-25
| Capability | Details |
|---|---|
| tools | {
"listChanged": true
} |
| prompts | {} |
| resources | {
"listChanged": true
} |
Tools
Functions exposed to the LLM to take actions
| Name | Description |
|---|---|
| scout_playbookA | Return the SceneScout testing method: setup order, write modes, how to explore, what counts as done, how to report. Call this ONCE before the first scout_attach in a conversation, then follow it. If this client offers a SceneScout skill, load that instead — it is the same text, so never read both. Takes no input and touches no browser. |
| scout_intakeA | Ask the start-of-run questions (the site's address, whether and how to sign in, what to check, whether the site holds real data) when the person gave no settings. Where this client can show a form, the person answers there and this returns the settings to use: the scout_login and scout_attach calls, in order. Otherwise, or when they decline or close the form, it returns the questions for you to ask in chat. Never asks for a password. Touches no browser. |
| scout_lane_briefA | For a run split across parallel agents (lanes). Divides the app's known routes between N lanes and returns each lane's session name, the |
| scout_lane_reportA | For a run split across parallel agents (lanes). Without |
| scout_scanA | Scan a project directory to discover the frontend workspace, framework, routes, dev command, Playwright auth storage states, and testid conventions. Run this first. |
| scout_attachA | Launch a browser and attach to a running web app. First attach in this conversation and you have read neither the SceneScout skill nor scout_playbook? Call scout_playbook before this. Write policy is enforced at the NETWORK layer: mode='observe' blocks EVERY request that is not a GET (login and token refresh excepted, and POSTs the user named in readPosts) — choose it for a target that holds real data, where even an ordinary form submission would create a record; mode='read-only' (default) blocks destructive-labeled elements AND all PUT/PATCH/DELETE + destructive POSTs, but lets ordinary form POSTs through; mode='safe-write' allows creating data and permits updates/deletes ONLY on resources this session created (use when the user wants create/edit flows tested); mode='destructive' allows everything — ONLY when the user explicitly confirmed a disposable/seeded environment. Pass |
| scout_loginA | Open a visible browser window for the USER to sign in to the app as a role, and save that sign-in for scout_attach { role }. Use it when an attach is refused because no sign-in is saved for the role, or the saved one has expired. Tell the user first, in plain words: "A browser window is opening. Sign in there as you normally would; it closes by itself once you are in." The window saves once the user is back on the app with a new session (a round trip through a single sign-on provider is followed, not taken for the end) and closes. Never type credentials into it yourself. Returns once signed in and saved, or after waitSeconds with the window still open: then call scout_login again with the same role to keep waiting. Closing the window saves nothing. Needs a desktop: on a machine with no display, ask the user to run |
| scout_sessionA | List live sessions, or set which one is the DEFAULT (used by any tool call that omits |
| scout_snapshotA | Capture the current page state: URL, state fingerprint, a one-line summary of the main area's heading and text, interactable elements with refs (e1, e2, …) and their state (pressed, selected, checked, expanded, current), what the page announces (alert and status regions, by their text), geometry issues, coverage, and oracle violations since the last action. Re-snapshotting a route returns a DIFF against its last snapshot, even after visiting elsewhere, and another tab of the same screen diffs against that screen's last tab; refs stay stable, including across a search or filter that rewrites only the query string. On a dense page, says what past the element cap was cut. Cheap — prefer this over screenshots. |
| scout_crawlA | Engine-side route sweep in ONE call: visits each path (default: all known routes not yet visited), records states into coverage memory, and returns a per-route health summary (HTTP status, element count, what the main area holds, oracle violations, dead-ends, auth-redirects, and ERROR-VIEW or STILL-LOADING for a main area showing only an alert or a loading placeholder). Navigation-only — safe in read-only mode. Use this FIRST for broad coverage; explore interactively only where it flags problems or where journeys matter. |
| scout_run_planA | Execute up to 20 actions in ONE call — use for mechanical sequences (fill a form, walk a wizard) so each step doesn't cost a round-trip. Targets resolve at execution time by semantic locator: 'testid=…', 'text=…', 'label=…' or 'role=button[name=Save]' (never snapshot refs). The same steps, saved to .scenescout/flows/.json with expect-text / expect-url / expect-request steps added, are replayed by |
| scout_clickA | Click an element by its ref from the latest scout_snapshot. Returns the outcome plus any oracle violations triggered. clicks=2 (or 3) probes IMPATIENT-USER behaviour: a rapid multi-click that fires the same state-changing request twice means the control is not guarded against double submission (button stays enabled, endpoint not idempotent) — use it on every important submit/create button once; the result says explicitly whether duplicates fired. |
| scout_typeA | Type into a text input/textarea/composer by ref, the way a real user does: if the field already holds content (e.g. an @-mention chip a menu click inserted), the text is APPENDED at the end — preserving that content — and the result reports what was already there (a separating space is added only at a word-to-word boundary). Pass replace=true to clear the field first (correcting a previous entry); an empty textValue always clears. Appending fires input events but not keydown, so keydown-driven triggers (slash/mention menus) will not react to appended text. Use for both valid values and boundary/fuzz values (empty, very long, unicode, script tags). |
| scout_uploadA | Attach a file to an upload control the way a user does. |
| scout_hoverA | Hover an element by ref like a user pausing the pointer on it, and report what it reveals: tooltips/popovers (diffed against pre-hover state), any other new page text that appeared (labelled as possibly unrelated on busy pages), the title attribute, and aria-describedby text — each item truncated to 300 chars. Hovering does not count as exercising the element. Use on badges, icons, truncated text, and error indicators BEFORE concluding an element 'does nothing' — hover-gated UI is invisible to snapshots and clicks. |
| scout_selectA | Select an option in a by ref. The value is matched against the options before anything is picked: an exact value, an exact label, either ignoring case, then a label it starts with. A value matching no option, or several, is refused at once with the options listed. |
| scout_navigateB | Navigate to a path on the attached origin (e.g. '/orders') or a full URL on it. A path resolves from the origin's root, whatever page the session attached on. Also supports 'back' via scout_back. |
| scout_requestA | Call the app's own API as this session, with the UI bypassed — the check that turns a hidden or disabled control into a proven refusal. A button that is not shown proves nothing; the same action refused by the server does. The fetch runs IN the page, so it carries the session's cookies and replays the Authorization header the app itself last sent, and it passes through the same interception the write policy is enforced on: in safe-write a mutation on a record this session did not create is refused here exactly as it would be for a click, and that refusal is the engine's safety net, not a finding. Returns the status line, the timing, the headers that decide whether two responses are truly identical (content-type, location, www-authenticate, retry-after, cache-control), and the body — its first 2000 characters, or the part named by select (one JSON value by path) or offset/limit (a window of characters). Unlike a shell call, every request is recorded in the run's trail and its signature is what a finding should quote. Paths are fenced to the attached origin: use another session to reach another host. |
| scout_networkA | List the data requests (fetch and XHR) the current page made since its document loaded: method, path, status and time, oldest first, each marked with the route it was sent from when a client-side route change moved the page since. Use it when the page and the server seem to disagree — an empty list where the API has data, a stale value after a save — to tell a request that failed, one still pending and one that never ran apart. Read-only: it lists what the browser already saw and sends nothing. Credentials in query strings are redacted. A full page load starts a new list; scout_request's own calls are marked. |
| scout_backB | Go back in browser history (tests back-button resilience). |
| scout_scrollA | Scroll like a user — real apps hide their bugs below the fold. Reports the resulting position (px and %), and explicitly flags SCROLL LOCKED: scrollable content exists but the page will not move (the classic leaked modal scroll-lock that silently cuts users off from everything below the fold — snapshots also detect this passively as an OVERLAY line). Without |
| scout_pressA | Press a keyboard key (e.g. Escape, Tab, Enter) — useful for closing modals and testing keyboard navigation. |
| scout_design_auditA | Computed-style design audit of the current page — a design connoisseur's read WITHOUT screenshots. Measurable defects (⚠): WCAG contrast, tiny targets, clipped text, aspect-distorted images, horizontal overflow, missing keyboard-focus indicators (sampled with real Tab presses). Craft suggestions (→): line measure and line-height rhythm, spacing-scale adherence, typography entropy, palette discipline (gray census, accent hue families, pure-#000 body text), elevation/control consistency, heading structure, indistinguishable links, and AI-slop tells (gradient text, glassmorphism, side-stripe borders, neon glows, violet gradients, identical card grids). Ends with a SYSTEM SUMMARY of design-system coherence. Run once per representative page; the → tier is improvement feedback — file genuine opportunities as ux-polish findings with the concrete numbers, not just defects. |
| scout_journeyA | Measure how EASY a real task is, not just whether it works — the question pass/fail e2e suites never answer. Wrap one user goal: scout_journey {action:'start', goal:'Create an order'}, perform it the way a first-time user would (navigate by CLICKING through the UI, not by jumping to a known deep URL — a shortcut invalidates the measurement), then scout_journey {action:'end', completed:true|false}. Returns interaction cost (clicks, navigations, distinct screens, elapsed), the actual path taken, and friction signals: BACKTRACKS (returning to a screen already left — the clearest sign the next step wasn't discoverable), screen count, and over-interaction. Run it on each module's primary journey; an abandoned journey is a high-severity finding. |
| scout_noteA | Cumulative WRITTEN knowledge about the tested app — .scenescout/ASSUMPTIONS.md, in prose a human can read and correct. memory.json stores coverage; this stores UNDERSTANDING, so every run starts smarter than the last. READ it at the start of every session ({action:'read'}). ADD durable learnings as you go ({action:'add', section, note}): what the app is for (app-model), who each role is and what they're FOR — infer the persona from what the role can see and do, e.g. 'qa-role = reviewer: approves orders, cannot administer' (roles), UI patterns the app follows (conventions), rules discovered the hard way like 'an order can only ship once approved' (constraints), fragile areas worth re-testing every run (risks), domain terms (glossary), and how to get the app testable at all — the command that regenerates an expired login state, what has to be running (setup), which the engine reads back to you the next time a storage state has expired. Notes are dated, attributed to the acting role, and deduplicated. Do NOT record session-specific facts (ids, counts) — only durable knowledge. |
| scout_screenshotA | Take a JPEG screenshot of the current viewport. LAST RESORT: geometry issues are in scout_snapshot and style/contrast/spacing issues are in scout_design_audit — images that failed to load are listed in scout_snapshot under BROKEN IMAGES — use a screenshot only for pixel-native content (a canvas, visual gestalt) that computed data cannot capture. |
| scout_captureA | Save a PNG of ONE element — its bounds plus a margin, from a real screenshot — under .scenescout/captures/, to show a person how it looks. Give the element's ref from the latest scout_snapshot. Not a way to judge design: scout_design_audit measures it. |
| scout_findingA | Record a structured finding (bug, UX issue, or improvement). Deduplicates across runs; automatically captures the recent action trace as the repro, and a picture of what it is about (the element |
| scout_statusA | Show the person running you how the run stands: each session's objective and current task, open findings by severity, coverage, and the live view's address. Call it when the user wants to watch or asks how the run is going. A host that renders MCP Apps shows a pane that keeps itself up to date; every other host gets the same as text, and you pass the |
| scout_status_pollA | Called by the scout_status pane every few seconds to refresh itself. Not for the agent: call scout_status instead. |
| scout_coverageA | Show exploration coverage: states visited, which elements remain unexercised, which options of a dropdown used this run no session has chosen yet, and which forms seen this run no session has submitted with every text field blank. Use to decide where to explore next and when the level's budget is satisfied. In a parallel run it shows this session's own work by default — the routes it reached this run and the forms it saw — so one lane is not handed another's gaps; scope:"project" shows every session's, each form and route tagged with the sessions that saw it. |
| scout_reportA | Generate the final markdown report — findings, page quality scores (worst first), role capability matrix, oracle rollup, and the GAP LEDGER (an explicit list of what was NOT tested). Writes the full document to .scenescout/report.md and returns a bounded SUMMARY (full reports exceed client token limits). Gates by level: 'minimal' needs all routes visited + ≥1 design audit; 'medium' additionally needs several routes audited; 'extensive' REFUSES while the gap ledger is non-empty — that refusal is the completeness guarantee: an extensive report only generates when nothing known is left untested. force=true overrides (only when the user capped the budget). |
| scout_resolveA | Mark a finding as resolved (by its id, shown when recorded and in the report). Resolved findings move to the report's green ✅ Resolved section, and reopen automatically as flagged REGRESSIONS if re-found later. Use when the user says a bug is fixed, or when re-testing shows the evidence no longer reproduces. |
| scout_verifyA | Re-test findings earlier runs left open. With no arguments, returns the open findings in the order to re-test them — worst route first, grouped so a route is walked once — each with its evidence and repro steps. Pass ids to narrow it to specific findings. After re-testing one, call again with id and verdict to record what you saw: "gone" resolves it, "present" stamps it confirmed so the report stops calling it unverified, "changed" keeps it open and says the behaviour differs. Use after a fix wave, or at the start of a run against an app this project has tested before. |
| scout_ticketsA | Read the tickets or acceptance criteria the person gave this run, pasted ( |
| scout_criterionA | Record whether one acceptance criterion of a ticket read with scout_tickets passed, failed or was not tested, with how sure you are. The link from a criterion to the findings that show it is YOUR judgement, stated with a confidence — never matched on words. A "fail" names the findings that show it (file them with scout_finding first); "not-tested" says why in untestedBecause. Recording the same criterion again from the same session replaces your earlier verdict. Touches no browser. |
| scout_closeA | Close a session's browser (memory persists on disk). Default: the DEFAULT session. Pass session to close a specific one, or all=true to close every live session at the end of a multi-role run. A session that is a lane of a parallel run (named by scout_lane_brief or scout_lane_report) is not closed until its lane report has been accepted by scout_lane_report, since folding needs the session attached; the refusal names every such lane. |
Prompts
Interactive templates invoked by user choice
| Name | Description |
|---|---|
| explore | Start an exploratory test session: loads the SceneScout method and states the target. |
| live | Return the loopback live-view URL for the current session. Takes no arguments and no password. |
| login | Call scout_login for a role and wait the way that tool waits. The person signs in in the window it opens. Takes the role, and the app URL when you have it. Never a password. |
Resources
Contextual data attached and managed by the client
| Name | Description |
|---|---|
| scenescout-status | The run's sessions, open findings by severity, coverage and the live view's address, refreshed every few seconds. |
TDQS
Scored across 37 tools
Most tools have sharply distinct purposes, and the descriptions explicitly defuse the riskiest overlaps (scout_screenshot vs scout_capture vs scout_design_audit; scout_status vs scout_status_poll; scout_request vs scout_network; scout_crawl vs scout_run_plan). A few pairs remain close enough that an agent must read carefully to pick correctly, e.g. scout_intake/scout_playbook/scout_note all being 'read this at the start' guidance, and the two lane tools being mirror halves of one workflow.
Every tool uses the same snake_case convention with a uniform `scout_` prefix and a noun/verb suffix describing the resource or action (scout_click, scout_type, scout_finding, scout_report). There are no camelCase or stylistic deviations, so the set is fully predictable.
37 tools is well past the 25-tool threshold the rubric treats as excessive, and several are low-value or redundant for the agent (scout_status_poll exists only to refresh a pane, scout_intake/scout_playbook/scout_note overlap as upfront reading). The domain is genuinely broad, but the surface would benefit from consolidation into fewer composite tools.
The set covers the full lifecycle of the stated purpose: setup (scan, attach, login, session), exploration (crawl, navigate, snapshot, network, interaction verbs), analysis (design audit, journey, coverage), knowledge (note, tickets, criterion), and reporting (finding, verify, resolve, report). Nothing obvious is missing, and report gating plus a gap ledger shows dead-ends are explicitly handled.