Skip to main content
Glama

Server Configuration

Describes the environment variables required to run the server.

NameRequiredDescriptionDefault

No arguments

Instructions

Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.

This server publishes no instructions, or was last inspected before Glama recorded them.

Capabilities

Features and capabilities supported by this server

Protocol revision2025-11-25

CapabilityDetails
tools
{
  "listChanged": true
}
prompts
{}
resources
{
  "listChanged": true
}

Tools

Functions exposed to the LLM to take actions

NameDescription
scout_playbookA

Return the SceneScout testing method: setup order, write modes, how to explore, what counts as done, how to report. Call this ONCE before the first scout_attach in a conversation, then follow it. If this client offers a SceneScout skill, load that instead — it is the same text, so never read both. Takes no input and touches no browser.

scout_intakeA

Ask the start-of-run questions (the site's address, whether and how to sign in, what to check, whether the site holds real data) when the person gave no settings. Where this client can show a form, the person answers there and this returns the settings to use: the scout_login and scout_attach calls, in order. Otherwise, or when they decline or close the form, it returns the questions for you to ask in chat. Never asks for a password. Touches no browser.

scout_lane_briefA

For a run split across parallel agents (lanes). Divides the app's known routes between N lanes and returns each lane's session name, the objective to attach it with, and the routes it owns — whole modules per lane, balanced by route count, so no two lanes audit the same area and none is left unopened. Call it after the first crawl, when route knowledge is complete. Touches no browser; pass the briefs to your lane agents, then use scout_lane_report for what they hand back.

scout_lane_reportA

For a run split across parallel agents (lanes). Without reply: returns the paragraph to put in a lane's prompt, telling it to hand back ONE typed JSON object (verdicts, severities and categories from closed sets, a calibrated confidence per decision, routes covered, what blocked it). With reply: parses what the lane handed back and returns the one-line fold (defects, highs, unsure, mean confidence, routes) or the reason it was refused, to relay to the lane once. Touches no browser.

scout_scanA

Scan a project directory to discover the frontend workspace, framework, routes, dev command, Playwright auth storage states, and testid conventions. Run this first.

scout_attachA

Launch a browser and attach to a running web app. First attach in this conversation and you have read neither the SceneScout skill nor scout_playbook? Call scout_playbook before this. Write policy is enforced at the NETWORK layer: mode='observe' blocks EVERY request that is not a GET (login and token refresh excepted, and POSTs the user named in readPosts) — choose it for a target that holds real data, where even an ordinary form submission would create a record; mode='read-only' (default) blocks destructive-labeled elements AND all PUT/PATCH/DELETE + destructive POSTs, but lets ordinary form POSTs through; mode='safe-write' allows creating data and permits updates/deletes ONLY on resources this session created (use when the user wants create/edit flows tested); mode='destructive' allows everything — ONLY when the user explicitly confirmed a disposable/seeded environment. Pass role to sign in with a login the user saved by scenescout login <url> --role <name>, or a Playwright storage-state JSON as storageStatePath. Pass session to keep MULTIPLE roles alive at once (one browser each, genuinely concurrent) for collaboration testing — target each directly with every tool's session param, or use scout_session to set which one is the default; coverage and findings merge into one project memory.

scout_loginA

Open a visible browser window for the USER to sign in to the app as a role, and save that sign-in for scout_attach { role }. Use it when an attach is refused because no sign-in is saved for the role, or the saved one has expired. Tell the user first, in plain words: "A browser window is opening. Sign in there as you normally would; it closes by itself once you are in." The window saves once the user is back on the app with a new session (a round trip through a single sign-on provider is followed, not taken for the end) and closes. Never type credentials into it yourself. Returns once signed in and saved, or after waitSeconds with the window still open: then call scout_login again with the same role to keep waiting. Closing the window saves nothing. Needs a desktop: on a machine with no display, ask the user to run scenescout login <url> --role <name> where they can see the window.

scout_sessionA

List live sessions, or set which one is the DEFAULT (used by any tool call that omits session). Prefer passing session directly on each tool call for multi-role work — that's what lets concurrent dispatch happen; scout_session is for sequential convenience (skip repeating session on every call) and for checking what's live. Both browsers stay live and authenticated regardless of which is default — re-snapshot a session after a break to see what changed while it was away.

scout_snapshotA

Capture the current page state: URL, state fingerprint, a one-line summary of the main area's heading and text, interactable elements with refs (e1, e2, …) and their state (pressed, selected, checked, expanded, current), what the page announces (alert and status regions, by their text), geometry issues, coverage, and oracle violations since the last action. Re-snapshotting a route returns a DIFF against its last snapshot, even after visiting elsewhere, and another tab of the same screen diffs against that screen's last tab; refs stay stable, including across a search or filter that rewrites only the query string. On a dense page, says what past the element cap was cut. Cheap — prefer this over screenshots.

scout_crawlA

Engine-side route sweep in ONE call: visits each path (default: all known routes not yet visited), records states into coverage memory, and returns a per-route health summary (HTTP status, element count, what the main area holds, oracle violations, dead-ends, auth-redirects, and ERROR-VIEW or STILL-LOADING for a main area showing only an alert or a loading placeholder). Navigation-only — safe in read-only mode. Use this FIRST for broad coverage; explore interactively only where it flags problems or where journeys matter.

scout_run_planA

Execute up to 20 actions in ONE call — use for mechanical sequences (fill a form, walk a wizard) so each step doesn't cost a round-trip. Targets resolve at execution time by semantic locator: 'testid=…', 'text=…', 'label=…' or 'role=button[name=Save]' (never snapshot refs). The same steps, saved to .scenescout/flows/.json with expect-text / expect-url / expect-request steps added, are replayed by scenescout check on every pull request. An upload step attaches a file as scout_upload does (target required — the file input or the control that opens its chooser; value = a fixture kind or a project-relative path). The plan ABORTS at the first NEW oracle violation, policy refusal, or failed step, returning a transcript of how far it got; repeats of already-reported violations do not abort (they stay logged for the report). For a sweep of independent steps (tabs, filters, pages) pass onViolation "continue": a new error status or its console echo is listed on its step's line and the plan goes on; a failed step, a policy refusal or any other violation still stops it. A select step whose value names no option fails at once, listing the options.

scout_clickA

Click an element by its ref from the latest scout_snapshot. Returns the outcome plus any oracle violations triggered. clicks=2 (or 3) probes IMPATIENT-USER behaviour: a rapid multi-click that fires the same state-changing request twice means the control is not guarded against double submission (button stays enabled, endpoint not idempotent) — use it on every important submit/create button once; the result says explicitly whether duplicates fired.

scout_typeA

Type into a text input/textarea/composer by ref, the way a real user does: if the field already holds content (e.g. an @-mention chip a menu click inserted), the text is APPENDED at the end — preserving that content — and the result reports what was already there (a separating space is added only at a word-to-word boundary). Pass replace=true to clear the field first (correcting a previous entry); an empty textValue always clears. Appending fires input events but not keydown, so keydown-driven triggers (slash/mention menus) will not react to appended text. Use for both valid values and boundary/fuzz values (empty, very long, unicode, script tags).

scout_uploadA

Attach a file to an upload control the way a user does. ref is either a visible (snapshots list these with role file) or the button/label/dropzone that opens the file chooser — the chooser is intercepted and answered, which is how the hidden input behind a styled 'Choose file' control is reached. Omit ref to target the page's only file input, hidden or not (snapshots disclose hidden ones on a FILE INPUTS line). Nothing needs to exist on disk: a small VALID fixture (real PDF/PNG structure) is generated in memory, its kind inferred from the input's accept attribute or chosen with fixture; filePath uploads a real file but must live inside the attached project (fenced like navigation is fenced to the origin); name overrides the filename for boundary tests (wrong extension vs accept, very long, unicode). The result names the input, how the file reached it, flags a file that violates accept (a mismatch the app then accepts is a validation finding), warns if the app cleared the input after selection, and says whether a state-changing request fired on selection — if none did, click the form's submit, or check the next snapshot for a client-side rejection.

scout_hoverA

Hover an element by ref like a user pausing the pointer on it, and report what it reveals: tooltips/popovers (diffed against pre-hover state), any other new page text that appeared (labelled as possibly unrelated on busy pages), the title attribute, and aria-describedby text — each item truncated to 300 chars. Hovering does not count as exercising the element. Use on badges, icons, truncated text, and error indicators BEFORE concluding an element 'does nothing' — hover-gated UI is invisible to snapshots and clicks.

scout_selectA

Select an option in a by ref. The value is matched against the options before anything is picked: an exact value, an exact label, either ignoring case, then a label it starts with. A value matching no option, or several, is refused at once with the options listed.

scout_navigateB

Navigate to a path on the attached origin (e.g. '/orders') or a full URL on it. A path resolves from the origin's root, whatever page the session attached on. Also supports 'back' via scout_back.

scout_requestA

Call the app's own API as this session, with the UI bypassed — the check that turns a hidden or disabled control into a proven refusal. A button that is not shown proves nothing; the same action refused by the server does. The fetch runs IN the page, so it carries the session's cookies and replays the Authorization header the app itself last sent, and it passes through the same interception the write policy is enforced on: in safe-write a mutation on a record this session did not create is refused here exactly as it would be for a click, and that refusal is the engine's safety net, not a finding. Returns the status line, the timing, the headers that decide whether two responses are truly identical (content-type, location, www-authenticate, retry-after, cache-control), and the body — its first 2000 characters, or the part named by select (one JSON value by path) or offset/limit (a window of characters). Unlike a shell call, every request is recorded in the run's trail and its signature is what a finding should quote. Paths are fenced to the attached origin: use another session to reach another host.

scout_networkA

List the data requests (fetch and XHR) the current page made since its document loaded: method, path, status and time, oldest first, each marked with the route it was sent from when a client-side route change moved the page since. Use it when the page and the server seem to disagree — an empty list where the API has data, a stale value after a save — to tell a request that failed, one still pending and one that never ran apart. Read-only: it lists what the browser already saw and sends nothing. Credentials in query strings are redacted. A full page load starts a new list; scout_request's own calls are marked.

scout_backB

Go back in browser history (tests back-button resilience).

scout_scrollA

Scroll like a user — real apps hide their bugs below the fold. Reports the resulting position (px and %), and explicitly flags SCROLL LOCKED: scrollable content exists but the page will not move (the classic leaked modal scroll-lock that silently cuts users off from everything below the fold — snapshots also detect this passively as an OVERLAY line). Without target it scrolls the page, falling back to the largest scrollable pane on app-shell layouts. Pass target to scroll ONE region instead (a sidebar nav, a dialog body, a table pane): the page-level pick is the LARGEST scroll port, so a smaller region beside it never moves and its content looks truncated when it is only scrolled away — never call a nav item missing without scrolling its own container first. Use before judging a long page: the design audit measures at the current scroll position, so scroll + re-snapshot/re-audit deep sections; scroll also triggers lazy-loaded content whose failures then surface as oracle violations.

scout_pressA

Press a keyboard key (e.g. Escape, Tab, Enter) — useful for closing modals and testing keyboard navigation.

scout_design_auditA

Computed-style design audit of the current page — a design connoisseur's read WITHOUT screenshots. Measurable defects (⚠): WCAG contrast, tiny targets, clipped text, aspect-distorted images, horizontal overflow, missing keyboard-focus indicators (sampled with real Tab presses). Craft suggestions (→): line measure and line-height rhythm, spacing-scale adherence, typography entropy, palette discipline (gray census, accent hue families, pure-#000 body text), elevation/control consistency, heading structure, indistinguishable links, and AI-slop tells (gradient text, glassmorphism, side-stripe borders, neon glows, violet gradients, identical card grids). Ends with a SYSTEM SUMMARY of design-system coherence. Run once per representative page; the → tier is improvement feedback — file genuine opportunities as ux-polish findings with the concrete numbers, not just defects.

scout_journeyA

Measure how EASY a real task is, not just whether it works — the question pass/fail e2e suites never answer. Wrap one user goal: scout_journey {action:'start', goal:'Create an order'}, perform it the way a first-time user would (navigate by CLICKING through the UI, not by jumping to a known deep URL — a shortcut invalidates the measurement), then scout_journey {action:'end', completed:true|false}. Returns interaction cost (clicks, navigations, distinct screens, elapsed), the actual path taken, and friction signals: BACKTRACKS (returning to a screen already left — the clearest sign the next step wasn't discoverable), screen count, and over-interaction. Run it on each module's primary journey; an abandoned journey is a high-severity finding.

scout_noteA

Cumulative WRITTEN knowledge about the tested app — .scenescout/ASSUMPTIONS.md, in prose a human can read and correct. memory.json stores coverage; this stores UNDERSTANDING, so every run starts smarter than the last. READ it at the start of every session ({action:'read'}). ADD durable learnings as you go ({action:'add', section, note}): what the app is for (app-model), who each role is and what they're FOR — infer the persona from what the role can see and do, e.g. 'qa-role = reviewer: approves orders, cannot administer' (roles), UI patterns the app follows (conventions), rules discovered the hard way like 'an order can only ship once approved' (constraints), fragile areas worth re-testing every run (risks), domain terms (glossary), and how to get the app testable at all — the command that regenerates an expired login state, what has to be running (setup), which the engine reads back to you the next time a storage state has expired. Notes are dated, attributed to the acting role, and deduplicated. Do NOT record session-specific facts (ids, counts) — only durable knowledge.

scout_screenshotA

Take a JPEG screenshot of the current viewport. LAST RESORT: geometry issues are in scout_snapshot and style/contrast/spacing issues are in scout_design_audit — images that failed to load are listed in scout_snapshot under BROKEN IMAGES — use a screenshot only for pixel-native content (a canvas, visual gestalt) that computed data cannot capture.

scout_captureA

Save a PNG of ONE element — its bounds plus a margin, from a real screenshot — under .scenescout/captures/, to show a person how it looks. Give the element's ref from the latest scout_snapshot. Not a way to judge design: scout_design_audit measures it.

scout_findingA

Record a structured finding (bug, UX issue, or improvement). Deduplicates across runs; automatically captures the recent action trace as the repro, and a picture of what it is about (the element ref names, else the viewport), kept for report.html and returned in this result as an image. Use for anything worth reporting: crashes, oracle violations you confirmed, dead ends, confusing UX, permission leaks, missing testids — and design-audit improvement opportunities (ux-polish) with their concrete measurements.

scout_statusA

Show the person running you how the run stands: each session's objective and current task, open findings by severity, coverage, and the live view's address. Call it when the user wants to watch or asks how the run is going. A host that renders MCP Apps shows a pane that keeps itself up to date; every other host gets the same as text, and you pass the Live view: address on. Takes no input and touches no browser.

scout_status_pollA

Called by the scout_status pane every few seconds to refresh itself. Not for the agent: call scout_status instead.

scout_coverageA

Show exploration coverage: states visited, which elements remain unexercised, which options of a dropdown used this run no session has chosen yet, and which forms seen this run no session has submitted with every text field blank. Use to decide where to explore next and when the level's budget is satisfied. In a parallel run it shows this session's own work by default — the routes it reached this run and the forms it saw — so one lane is not handed another's gaps; scope:"project" shows every session's, each form and route tagged with the sessions that saw it.

scout_reportA

Generate the final markdown report — findings, page quality scores (worst first), role capability matrix, oracle rollup, and the GAP LEDGER (an explicit list of what was NOT tested). Writes the full document to .scenescout/report.md and returns a bounded SUMMARY (full reports exceed client token limits). Gates by level: 'minimal' needs all routes visited + ≥1 design audit; 'medium' additionally needs several routes audited; 'extensive' REFUSES while the gap ledger is non-empty — that refusal is the completeness guarantee: an extensive report only generates when nothing known is left untested. force=true overrides (only when the user capped the budget).

scout_resolveA

Mark a finding as resolved (by its id, shown when recorded and in the report). Resolved findings move to the report's green ✅ Resolved section, and reopen automatically as flagged REGRESSIONS if re-found later. Use when the user says a bug is fixed, or when re-testing shows the evidence no longer reproduces.

scout_verifyA

Re-test findings earlier runs left open. With no arguments, returns the open findings in the order to re-test them — worst route first, grouped so a route is walked once — each with its evidence and repro steps. Pass ids to narrow it to specific findings. After re-testing one, call again with id and verdict to record what you saw: "gone" resolves it, "present" stamps it confirmed so the report stops calling it unverified, "changed" keeps it open and says the behaviour differs. Use after a fix wave, or at the start of a run against an app this project has tested before.

scout_ticketsA

Read the tickets or acceptance criteria the person gave this run, pasted (text) or from a file (path), and keep them so the report answers each criterion: passed, failed or not tested. Recognises Given/When/Then scenarios, checklists, numbered and "AC1:" criteria, and lists under an "Acceptance criteria" heading; several tickets may be given at once. A ticket with no recognisable criteria is reported as such, never guessed at. Returns each criterion's id (AC1, AC2, …) to judge it by with scout_criterion. Touches no browser.

scout_criterionA

Record whether one acceptance criterion of a ticket read with scout_tickets passed, failed or was not tested, with how sure you are. The link from a criterion to the findings that show it is YOUR judgement, stated with a confidence — never matched on words. A "fail" names the findings that show it (file them with scout_finding first); "not-tested" says why in untestedBecause. Recording the same criterion again from the same session replaces your earlier verdict. Touches no browser.

scout_closeA

Close a session's browser (memory persists on disk). Default: the DEFAULT session. Pass session to close a specific one, or all=true to close every live session at the end of a multi-role run. A session that is a lane of a parallel run (named by scout_lane_brief or scout_lane_report) is not closed until its lane report has been accepted by scout_lane_report, since folding needs the session attached; the refusal names every such lane.

Prompts

Interactive templates invoked by user choice

NameDescription
exploreStart an exploratory test session: loads the SceneScout method and states the target.
liveReturn the loopback live-view URL for the current session. Takes no arguments and no password.
loginCall scout_login for a role and wait the way that tool waits. The person signs in in the window it opens. Takes the role, and the app URL when you have it. Never a password.

Resources

Contextual data attached and managed by the client

NameDescription
scenescout-statusThe run's sessions, open findings by severity, coverage and the live view's address, refreshed every few seconds.

TDQS

A3.9/5.0

Scored across 37 tools

Disambiguation4/5

Most tools have sharply distinct purposes, and the descriptions explicitly defuse the riskiest overlaps (scout_screenshot vs scout_capture vs scout_design_audit; scout_status vs scout_status_poll; scout_request vs scout_network; scout_crawl vs scout_run_plan). A few pairs remain close enough that an agent must read carefully to pick correctly, e.g. scout_intake/scout_playbook/scout_note all being 'read this at the start' guidance, and the two lane tools being mirror halves of one workflow.

Naming Consistency5/5

Every tool uses the same snake_case convention with a uniform `scout_` prefix and a noun/verb suffix describing the resource or action (scout_click, scout_type, scout_finding, scout_report). There are no camelCase or stylistic deviations, so the set is fully predictable.

Tool Count2/5

37 tools is well past the 25-tool threshold the rubric treats as excessive, and several are low-value or redundant for the agent (scout_status_poll exists only to refresh a pane, scout_intake/scout_playbook/scout_note overlap as upfront reading). The domain is genuinely broad, but the surface would benefit from consolidation into fewer composite tools.

Completeness5/5

The set covers the full lifecycle of the stated purpose: setup (scan, attach, login, session), exploration (crawl, navigate, snapshot, network, interaction verbs), analysis (design audit, journey, coverage), knowledge (note, tickets, criterion), and reporting (finding, verify, resolve, report). Nothing obvious is missing, and report gating plus a gap ledger shows dead-ends are explicitly handled.

Maintenance

ActivityMaintained
ResponsivenessResponsive