twine-play-mcp
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@twine-play-mcpopen my Twine game and pick the first choice"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
twine-play-mcp
English · 简体中文
An MCP server that lets AI agents play, test and QA Twine / interactive-fiction HTML games.
The agent reads the current passage, sees numbered choices, clicks them, watches story variables, screenshots the game, and can save/restore state to explore branches — all through a small, token-friendly tool surface instead of a generic browser automation API.
Agent ──MCP(stdio)──> twine-play-mcp ──Playwright──> headless Chrome
│ │
│ static server (127.0.0.1) │ injected bridge
└────────> game HTML <──────────┘Why not a generic browser MCP?
Generic browser MCPs make the model guess DOM selectors, dump whole pages into context and have no notion of "story state". This server adds a semantic layer:
Passage view: text as Markdown, passage name, format/version, story metadata
Numbered choices with target passage names (and external-link blocking)
Story variables (SugarCube
State.variables) with safe depth/size capsNative state: SugarCube
Engine.backward/forward,Save.base64snapshotsFormat detection: SugarCube first, DOM fallback for Harlowe / Snowman / Chapbook / unknown
Tracker blocking and quiet console/network capture for clean playtesting
Related MCP server: Playwright MCP
Requirements
Node.js >= 20 (developed on 26)
Google Chrome installed (uses
channel: 'chrome'; no 200 MB browser download)Linux/macOS/Windows
Install
npm install -g twine-play-mcp # or: npx twine-play-mcpNo build step, no browser download — the package ships the compiled server and the page bridge, and drives the Chrome you already have.
Then point it at any published Twine HTML file (or a folder containing the game + assets):
twine-play-mcp # MCP server on stdioFrom source (for development)
git clone https://github.com/adorablelovelymia/twine-play-mcp.git
cd twine-play-mcp
npm install
npm run build # compiles to dist/ and copies the page bridge
# sanity checks (optional)
npm run spike # 17 end-to-end checks against a real SugarCube game
npm run smoke # spawns the MCP server over stdio and drives it with the MCP SDKClient configuration
The snippets below use npx, so no global install is required. If you installed globally,
replace "npx" + "twine-play-mcp" with "twine-play-mcp" on its own.
OpenCode (~/.config/opencode/opencode.json)
{
"$schema": "https://opencode.ai/config.json",
"mcp": {
"twine-play": {
"type": "local",
"command": ["npx", "-y", "twine-play-mcp"],
"enabled": true
}
}
}Claude Desktop / Cursor / any mcpServers client
{
"mcpServers": {
"twine-play": {
"command": "npx",
"args": ["-y", "twine-play-mcp"]
}
}
}Environment variables:
Variable | Purpose |
| Chrome executable if |
| Folder where browser downloads are captured (default |
| Live-view server port (default |
Tools
Tool | What it does |
| Open a local HTML file/folder or URL; optional PRNG seed; returns first observation |
| Passage text, numbered choices, inputs (paginated windows of 40 with |
| Click by 1-based number or label; works for passage choices and dialog buttons; |
| Wait for ms / for text / for the DOM to settle |
| Fill inputs/selects/checkboxes by ref, or press a key |
| Find buttons/links/labels/inputs by visible text or input name; returns refs for |
| Click dialogs, sidebar and menus (by ref / CSS selector / visible text), including |
| Upload a local file into an |
| Browser file control (any game): copy a captured download to a path — by |
| List files captured into the tool's persistent download folder (survives sessions/restarts), with name, size, time and absolute path |
| Inspect or discover UI panels outside the passage (mod GUIs, backstage); lists buttons/inputs and file inputs |
| Read story variables by dot path ( |
| Undo one passage (SugarCube |
| Restart from the beginning, optionally reseeding the PRNG |
| Named in-session snapshots for branch exploration |
| PNG of the viewport (canvas/visual games, visual QA); pass |
| Give the user eyes on the real page: local URL streaming JPEG frames (~1/s) of the actual Playwright tab + passage/step/journal; works headless; |
| JS exceptions, console errors and HTTP failures captured from the page |
| Action history: passages visited, choices taken, coverage counts |
| Session management |
Watching the page (live view)
live_view(game_id) starts a tiny local server (once per MCP process, port 4571+) and returns a URL
like http://127.0.0.1:4571/v/game_abc. Open it in any browser (or pass open: true) to watch the
actual tab the agent is driving — a JPEG frame about once per second, plus passage, step, engine
state, recent actions and the passage text. It works with headless games, frames are captured only
while somebody is watching, and closing the game stops it. For a raw browser window instead, open the
game with headless: false (open_game).
Handy combo: if your client can show a web page in a side pane (e.g. OpenCode's Review pane /
browser.tabs.open), point it at the live-view URL and you can follow along while the agent plays.
Pick one view (agents). To keep the user's screen clean, show a running game through exactly one channel — never stack them:
Default:
live_view— hand the URL to the user, or passopen: trueonce to launch it for them. Repeated calls reuse the same view and never open another tab.Only on explicit request:
open_game(headless: false)when the user asks for a real browser window. Don't add a live view on top; a headed window is already visible.screenshotis a one-shot visual check, not a stream — don't loop it to "show" the game.
If a view (live view tab or headed window) is already open, reuse it instead of starting a second
one. The MCP server ships this same policy in its instructions field, so MCP clients can pass it
to the model automatically; the tool descriptions repeat it where it matters (live_view,
open_game.headless, screenshot).
Agent ergonomics
Output: every play tool returns a formatted text observation (a string). Pass
format:"json"to receive a JSON string instead (JSON.parseit) withpassage,text,choices[{n,label,target}],inputs[{ref,kind,label,checked}],inputsTotal,dialog,status.Inputs are paginated, not truncated: a header like
Inputs (41-80 of 140)plusinputs_offset=80means everything is reachable — no silent hard cap.Label matching:
click_ui(text)andfind_ui(text)understand SugarCube<<radiobutton>>/<<checkbox>>labels, so options like "Jet black" or combat moves like "Punch" are clickable by text.Errors are compact: failures return
ERROR: code — message, aHint, the current passage and the available choices — never a full observation dump.Variables: observations do not embed variable blobs by default; use
get_variablesfor the keys you care about.include_variables:trueis still available when you want the (truncated) dump.Dialogs: dialog buttons appear as numbered choices tagged
[dialog], and aDialog buttons:line lists them; checkbox labels are shown on the input line.
Format support
Format | Detect | Text/choices | Variables | Passage name | Back | Snapshots |
SugarCube 2.21+ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ |
Harlowe 3 | ✅ | ✅ | — (engine internals are private) | — | ✅ sidebar undo | — |
Snowman 2 | ✅ | ✅ | ✅ | ✅ | — | ✅ state JSON |
Chapbook 1 | ✅ | ✅ | ✅ | ✅ | — | ✅ |
Unknown HTML | generic | ✅ DOM heuristics | — | — | — | — |
Everything degrades gracefully: an unknown or exotic format still plays with the generic DOM
path; format-specific tools report unsupported instead of failing.
Play-session example (what the agent sees)
[sugarcube 2.37.3 · step 3 · engine=idle · passage: 069]
You squeeze through the narrow gap...
Choices (2):
1. Go deeper -> 070
2. Check the mirror
Status:
Resistance: 500/500
Variables: {"resistance":500,"pleasure":0,"degradation":0,...}How it works
src/bridge/bridge.jsis injected into every page (addInitScript) and exposeswindow.__twineMCP: format detection, passage/choice extraction, click/fill helpers, settle-waiting, snapshots and seeding. All server calls go through this bridge only.src/session.tsowns the browser, oneBrowserContextper game (isolated saves) and a tiny static server so local games run onhttp://127.0.0.1(localStorage works).src/render.tsturns observations into compact Markdown for the model.Choices get temporary
data-twmcp-refattributes; the server prefers real Playwright clicks and falls back to DOM clicks for exotic macro-generated links.Spoiler policy: only what a player can see is returned. No passage lists or source dumps are exposed.
Complex games
Real games are not just passages and links. The MCP handles the awkward parts:
Modal dialogs (SugarCube
#ui-dialog, content gates, settings): their text appears as aDialog:block and their buttons/inputs are numbered like choices, so the agent can tick a consent checkbox (interact) and clickEnter(choose).Iframes: mod managers and dev panels often live in a child frame.
inspect_uidiscovers them (marked[iframe]),click_uiby text andupload_filesearch every frame.File workflows: uploads go through
upload_file(mod.zip, save import) — by clicking a trigger (trigger_text/trigger_selector, e.g.#saves-import) or pointing at an<input type=file>directly. Downloads go the other way through a persistent download folder: every browser download is captured there (TWMCP_DOWNLOAD_DIRoverrides the location),list_downloadsshows the contents, anddownload_filecopies one anywhere (path, default<cwd>/downloads/<name>) — either by clicking the game's export button or byname/indexafterwards. No manual temp-folder copying, and files surviveclose_gameand MCP restarts.DoL case study:
test/fixturesaside,scripts/dol-mcp-test.tsdrives Degrees of Lewdity end to end — consent gate → importingModI18N.mod.zipandGameOriginalImagePack.mod.zipthrough the in-game ModLoader GUI → page reload → importing a real.savethrough the SAVES dialog → several turns of normal play.
Tests
npm run spike # 17 checks against a real SugarCube 2.37 game (play, back, snapshot, screenshot)
npm run formats # 4 compiled fixtures: SugarCube 2.30, Harlowe 3.1, Snowman 2.0, Chapbook 1.0
npm run smoke # spawns the built MCP server and drives the tools over stdio
npm run clarity # agent-ergonomics regression on DoL character creation (pagination, labels, variables)
npm run fixtures # rebuild test/fixtures/compiled/*.html with Tweego (see test/fixtures/build.sh)
npx tsx scripts/dol-mcp-test.ts # Degrees of Lewdity: gate, mod import, save import, playscripts/inspect.ts <fixture> dumps the DOM/story-format internals of a game — handy when
adding a new adapter.
Status / roadmap
M1: SugarCube adapter, generic DOM fallback, observation/choice/input/wait/screenshot, snapshots, backtracking, console+network QA capture, stdio MCP, spike + smoke tests
M2: Harlowe / Chapbook / Snowman adapters verified against compiled fixtures
M2: play journal (
get_journal) for run summaries, resuming and QA coverageM3: dialogs/iframe-aware UI control (
click_ui,inspect_ui,upload_file) — verified on Degrees of Lewdity (mod import + save import + play)M4: agent ergonomics — input pagination + totals,
find_uilabel search,get_variables,format:"json", compact errors (driven by a naive-agent playtest that stalled on DoL character creation)M4: file workflows both ways —
upload_filefor mods/saves,download_file+list_downloadsfor a persistent browser download folder (no temp-folder copying)M5: live view — watch the real page in any browser (~1 fps frames + passage/step/journal); headed mode via
open_game(headless: false)M5: npm packaging — published as
twine-play-mcpM3:
click_atfor canvas games, spoiler-gated story-map analysis
Safety notes
Page scripts run in Chrome's sandbox; the bridge never exposes Node to the page.
No arbitrary
evaltool is exposed to the agent.Analytics/tracker hosts are blocked by default (
block_trackers: falseto disable).External links are blocked unless
allow_external: trueis passed.
中文文档
完整中文版见 README-zh.md(工具一览、客户端配置、格式支持、实时视图策略等均已翻译)。
一句话:这是一个让 AI agent 游玩 / 测试 Twine 文字游戏的 MCP 服务——无头 Chrome + 页面桥,
22 个工具,本地游戏经内置静态服务器以 http://127.0.0.1 打开(保证存档可用),
live_view 让你用任意浏览器实时看到 AI 正在操作的真实页面。
npm install -g twine-play-mcp # 或 npx twine-play-mcpAvailable Tools
22 toolsbackGo back one stepB
Undo the last passage navigation when the story format supports it (SugarCube: Engine.backward). Returns the new observation (text; format:"json" for JSON).
| Name | Required | Description | Default |
|---|---|---|---|
| format | No | Output format: "text" (default, human-readable) or "json" (JSON string for programmatic use). | |
| game_id | Yes | Session id returned by open_game. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries full burden. It discloses the format-support constraint and returns the new observation, which is useful behavioral context. However, it doesn't state what happens if unsupported (error message? no-op?), whether back at the start is a no-op, or whether the action is idempotent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single concise sentence with the key constraint (format support) front-loaded and return behavior appended. No waste, though the parenthetical engine reference slightly interrupts flow.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple two-param undo tool with no annotations and no output schema, the description covers the essentials: what it does, when it works, and what it returns. But it leaves out error behavior and edge cases (e.g., already at start), which an agent would need to call it correctly in all cases.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both parameters are documented in the schema. The description adds no meaning beyond what's there, mentioning only the format return type. Baseline 3 when schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific action ('undo the last passage navigation') with the resource (story navigation). It names the underlying engine call (SugarCube: Engine.backward), clarifying exactly what happens. It does not explicitly distinguish itself from siblings like restart or load_state, though 'undo last passage' is fairly distinct.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use it (to undo a navigation, and only when the story format supports it), but doesn't name alternatives or state when not to use it. The format-support caveat is a helpful condition, but no sibling routing is provided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
chooseClick a choiceA
Click a passage choice by its 1-based number from the last observation, or by (partial) label text. Numbered choices include dialog buttons (tagged [dialog] in observations) — so choose() answers dialogs too. Waits for the game to settle and returns the new observation. Pass expected to guard against clicking the wrong link.
| Name | Required | Description | Default |
|---|---|---|---|
| choice | Yes | 1-based choice number from the last observation, or label text (case-insensitive, may be partial if unambiguous). | |
| format | No | Output format: "text" (default, human-readable) or "json" (JSON string for programmatic use). | |
| game_id | Yes | Session id returned by open_game. | |
| expected | No | Substring the clicked choice label must contain; fails safely if it does not match. | |
| allow_external | No | Allow following links that leave the game (default false, they are blocked). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral burden and does disclose key traits: it blocks until the game settles and returns the new observation, and that expected fails safely rather than clicking the wrong link. It omits error behavior and timeout details, but the synchronous-settle and return semantics are genuinely valuable additions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four short sentences, front-loaded with the core action and matching rule before the dialog note and safety guard. Minor restatement of the choice/label mechanics that the schema also covers, but no serious padding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a five-parameter tool with no output schema and no annotations, the description covers the action, matching rule, wait/return behavior, and the primary safety guard. Adequate, though it could note more on what happens on timeout or an unmatched expected value.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so every parameter (choice, format, game_id, expected, allow_external) is already documented. The description paraphrases the choice and expected semantics but adds no syntax or format detail beyond the schema, so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (click) and resource (a passage choice), and specifies the mechanism: 1-based number from the last observation or partial label text. The note that numbered choices include [dialog] buttons successfully carves it apart from UI-level siblings like click_ui and interact.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explains the primary usage context (driving passage choices and answering dialogs) and offers the expected guard for safe clicking. It does not explicitly name alternatives or state when NOT to use it versus click_ui/interact, so it stops short of full routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
click_uiClick UI outside the passageA
Click dialogs, sidebar buttons and menus (SAVES, OPTIONS, ModLoader banner, modal buttons) by ref, CSS selector or visible text. Text matching covers -based controls too (SugarCube radio/checkbox options like "Jet black" or "Punch"); shortest match wins, exact=true for exact text. Returns the new observation (text; format:"json" for JSON). Use choose() for numbered passage choices and dialog buttons.
| Name | Required | Description | Default |
|---|---|---|---|
| ref | No | UI ref from an observation (e.g. "u1" or "x3"). | |
| text | No | Visible text of the target (case-insensitive, partial match; the shortest match wins). | |
| exact | No | Require an exact text match (default false). | |
| format | No | Output format: "text" (default, human-readable) or "json" (JSON string for programmatic use). | |
| game_id | Yes | Session id returned by open_game. | |
| selector | No | CSS selector, if you know the exact element. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does disclose meaningful behavior: shortest-match-wins resolution, exact=true semantics, label-based matching, and the return value (the new observation, text or JSON). It does not state side effects, whether the click mutates saved state, or failure behavior when no element matches, which are the remaining gaps.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Dense but front-loaded: the action and target types come first, then matching rules, then the return value, then the sibling routing. Every clause carries information, though the parenthetical control list and the dialog/button phrasing could be tightened slightly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, but the description states that it returns the new observation and how to request JSON, so return-value handling is covered. For a 6-parameter tool with one required arg and full schema coverage, this is complete enough to invoke correctly, missing only explicit side-effect/failure disclosure.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds resolution semantics not fully conveyed by the field docs: that text matching also covers <label>-based radio/checkbox options, that the shortest match wins, and that exact=true is opt-in. This is genuine extra meaning beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a concrete verb (click) plus the exact target classes (dialogs, sidebar buttons, menus, named controls) and the three addressing mechanisms (ref, CSS selector, visible text). It explicitly distinguishes itself from choose(), so an agent can separate it from the closest sibling without reading either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It names the alternative (choose()) and the condition that selects it (numbered passage choices and dialog buttons), and explains when text matching applies (label-based SugarCube controls). There is mild ambiguity between 'Click dialogs ... modal buttons' and 'use choose() for ... dialog buttons', but the routing rule itself is explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
close_gameClose a gameB
Close the browser tab and static server for a game session.
| Name | Required | Description | Default |
|---|---|---|---|
| game_id | Yes | Session id returned by open_game. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full disclosure burden. It does add real value by naming the two things destroyed (browser tab and static server), implying the session is torn down, but it is silent on whether in-flight state is lost, whether the session can be reopened, what errors occur for an unknown game_id, and whether this requires the game to be open. That is a meaningful gap for a teardown operation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler; the affected resources come immediately after the verb. Nothing here needs to be trimmed or reordered.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter tool with full schema coverage and no output schema, the description covers the basic action. But as an unannotated teardown tool it omits the information an agent most needs: effect on saved/unsaved state and how it relates to restart or open_game. Adequate, but with clear gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the schema already documents game_id as the id returned by open_game. The description adds nothing about the parameter (e.g. format, validity conditions for stale ids), so the baseline of 3 for fully-covered schemas is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('close') and the exact resources affected ('browser tab and static server for a game session'), which an agent can act on directly. It does not, however, distinguish itself from siblings like 'restart' or 'back', which could also terminate a session view, leaving some ambiguity about which teardown tool applies.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description says what gets closed but never says when to call it — no mention of end-of-session cleanup, no exclusion of alternatives such as 'restart' (which presumably keeps the session) or 'back'. The only usage signal is the indirect reference to 'open_game' in the schema, so the agent must infer the pairing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
download_fileSave a captured browser download to a fileA
Browser file control for any game: every download is captured into the tool's download folder (see list_downloads). Three modes: (a) pass trigger_text/trigger_ref/trigger_selector to click the game's download button and take that file; (b) pass name or index to take an already-captured file from the folder; (c) pass none of those to take the newest file in the folder. Without path the file stays in the download folder (path = its location); with path (a file or a directory) a copy is placed there too. Returns the path, size and the new observation.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | Name of an already-captured file from list_downloads (exact, suffix, or unique substring). | |
| path | No | Destination file or directory (default <cwd>/downloads/<name>). | |
| index | No | 1-based index from list_downloads (newest first). | |
| format | No | Output format: "text" (default, human-readable) or "json" (JSON string for programmatic use). | |
| game_id | Yes | Session id returned by open_game. | |
| timeout_ms | No | How long to wait for the download after clicking (default 30000). | |
| trigger_ref | No | Ref of the download button (alternative to trigger_text). | |
| trigger_text | No | Visible text of the button that starts the download (optional). | |
| trigger_selector | No | CSS selector of the download button (alternative to trigger_text). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and largely meets it: it discloses where downloads land, that path produces a copy (not a move), and that it returns path, size, and the new observation. It does not state whether the operation can fail/timeout behavior beyond timeout_ms, or whether re-invoking re-downloads versus reuses a captured file.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with purpose, then organized into labeled modes (a)/(b)/(c), so the reader can parse by branch. It is dense and somewhat run-on, but essentially every sentence earns its place; minor tightening on the path/return sentence would help.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 9-parameter, no-output-schema tool, the description supplies the missing return-value summary (path, size, observation) and the mode logic tying the optional parameters together, plus points to list_downloads for name/index sources. Nothing an agent needs to invoke it correctly is absent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the parameter baseline is 3, but the description adds genuine value beyond the schema by explaining how parameters combine into mutually exclusive modes (trigger_* vs name/index vs neither) and clarifying copy semantics of path, which the schema only describes as 'destination'.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('take/copy a captured browser download to a file') and enumerates three distinct operating modes. It names the sibling it depends on (list_downloads) so the agent can place it in the workflow without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit when-to-use branches: (a) trigger_* to click and capture, (b) name/index to take an already-captured file, (c) neither to take the newest file. It also states the condition for path (omit to leave in folder, supply to copy out). This is exactly the guidance an agent needs to pick parameters.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
find_uiFind a control by textA
Search visible controls (buttons, links, labels, inputs) by label text and/or input name; returns refs usable with click_ui(ref), interact(ref, value) and upload_file(ref). This is the fastest way to reach radio/checkbox options (e.g. "Jet black", "Punch") and any input beyond the 40-item observation window. Results are capped by limit; no game state changes.
| Name | Required | Description | Default |
|---|---|---|---|
| kind | No | Filter by kind: button, link, label, input, radio, checkbox, select, textarea. | |
| name | No | Filter by input name attribute (exact). | |
| text | No | Label text to search for (case-insensitive, partial by default). | |
| exact | No | Require an exact label match (default false). | |
| limit | No | Max matches to return (default 20). | |
| game_id | Yes | Session id returned by open_game. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries full burden. It does disclose that results are capped by limit and that there are 'no game state changes' (read-only), which is useful behavioral context. That said, it doesn't describe permission requirements, pagination, failure modes when no match is found, or return shape beyond 'refs'.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the core purpose and search criteria, then immediately covering return refs and usage context. Every clause adds value: use cases, cap behavior, and non-mutation guarantee.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only search tool with a rich 6-param schema and no output schema, the description is quite complete: it specifies inputs, return type (refs), downstream tools, result capping, and no side effects. It lacks explicit guidance on when to prefer over inspect_ui or how to handle empty results, but covers the essentials well.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the schema fully documents each parameter, establishing a baseline of 3. The description adds conceptual value by mentioning 'label text and/or input name' (mapping to text/name) and 'results are capped by limit', but does not add syntax, format, or interaction details beyond what the schema provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Search') and resource ('visible controls (buttons, links, labels, inputs)'), and names the exact criteria (label text and/or input name). It also names the ref-using siblings (click_ui, interact, upload_file), making it distinguishable from inspect_ui and observe without opening schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides clear positive context: 'fastest way to reach radio/checkbox options ... and any input beyond the 40-item observation window,' which implicitly tells when to prefer this over observe/inspect_ui. However, it never explicitly states when NOT to use it or contrasts it directly with inspect_ui, so it falls short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_console_errorsGet console errorsB
Return JavaScript errors/warnings captured from the page (useful for playtesting / QA).
| Name | Required | Description | Default |
|---|---|---|---|
| clear | No | Clear the buffer after reading (default false). | |
| game_id | Yes | Session id returned by open_game. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It hints that errors are 'captured' (an accumulating buffer) and the schema's 'clear' param reveals buffer semantics, but the description says nothing about whether errors persist across reads, whether output is bounded/truncated, or what the response shape looks like.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single tight sentence with the resource front-loaded and the usage hint appended. Nothing is wasted and nothing important is buried.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple two-parameter read tool this is close to adequate, but with no output schema the description should at least indicate the return shape (list of entries, fields like message/level/timestamp) and whether reads are cumulative. That gap leaves the agent guessing about results.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both parameters (clear, game_id) are already documented in the schema, and the description adds no syntax or format detail. Baseline 3 applies when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb (Return) and resource (JavaScript errors/warnings captured from the page), which is unambiguous and clearly distinct from siblings like get_journal or inspect_ui. It does not explicitly contrast with a sibling, but the resource is unique enough that the distinction is obvious.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The parenthetical '(useful for playtesting / QA)' implies the usage context but gives no explicit when-to-use vs when-not guidance and names no alternatives. Usage is only weakly implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_journalGet the play journalA
Return this session's action history: every passage visited and every choice taken (including back/load events). Useful to summarise a playthrough, resume a run, or report coverage for QA.
| Name | Required | Description | Default |
|---|---|---|---|
| clear | No | Clear the journal after reading (default false). | |
| limit | No | Show the last N entries (default 50). | |
| game_id | Yes | Session id returned by open_game. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It discloses what the journal contains (passages, choices, back/load events), but it completely omits the fact that the optional 'clear' parameter destructively wipes the journal after reading, and says nothing about permissions, reversibility, or side effects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with zero filler. The core return content is front-loaded, and the usage examples follow immediately, making it easy to skim.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read tool with no output schema, the description adequately explains what history is returned, but it misses the critical destructive side effect of the 'clear' parameter and does not describe the returned entry shape or pagination via 'limit'.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all three parameters (game_id, clear, limit). The description adds no parameter-level meaning beyond what the schema provides, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'Return this session's action history', and enumerates the content (passages, choices, back/load events). It clearly differentiates the tool's purpose from state-saving or UI-inspection siblings, but never explicitly names an alternative tool it should not be confused with.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives three clear usage contexts: summarising a playthrough, resuming a run, and reporting coverage for QA. No when-not-to-use or alternative-tool guidance is provided, but the positive guidance is specific and actionable.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_variablesRead story variablesA
Read specific story variables (SugarCube State.variables) by dot path, e.g. ["haircolour", "background", "player.background"]. Accepts "V.x", "variables.x" or plain "x". With no paths, returns a shallow summary of the top-level keys. Output is JSON.
| Name | Required | Description | Default |
|---|---|---|---|
| paths | No | Dot paths to read (default: top-level key summary). | |
| game_id | Yes | Session id returned by open_game. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description usefully discloses return format ("Output is JSON"), the no-paths default (shallow top-level summary), and the accepted prefix variants. It is silent on invalid-path handling and result-size limits, but for a simple read tool it adds meaningful context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three front-loaded sentences, each carrying distinct information (what, format variants, default/output). No filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter read tool with no output schema and no annotations, the description covers purpose, input format, default behavior, and output type. Only edge-case behavior (invalid paths, the 50-item cap) is left unstated.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3. The description goes further by documenting accepted input forms ("V.x", "variables.x", plain "x") that the schema does not mention, giving genuine added meaning.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (read) and resource (story variables / SugarCube State.variables) with concrete dot-path examples. It is clearly distinguable from sibling tools like get_journal or inspect_ui, which read different resources.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by the examples and the default-behavior note, but there is no explicit when-to-use, when-not-to-use, or routing to an alternative sibling. It leaves the agent to infer context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
inspect_uiInspect a UI panelA
Inspect DOM outside the passage. Pass a CSS selector to get its text, buttons (with refs usable in click_ui) and inputs (file inputs usable in upload_file). Without a selector, lists overlay panels (mod GUIs, dev panels) that contain buttons or file inputs.
| Name | Required | Description | Default |
|---|---|---|---|
| game_id | Yes | Session id returned by open_game. | |
| selector | No | CSS selector of the panel to inspect; omit to discover overlay panels. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does real work: it discloses the two behavioral modes, the shape of the return (text, buttons, inputs), and cross-tool contracts (refs usable in click_ui, file inputs usable in upload_file). It omits failure behavior for invalid selectors and whether inspection has side effects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences, front-loaded with the core action, then the two modes, then the return contents. No filler and every clause carries information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description must describe returns and it does reasonably well: text, buttons with refs, inputs, or a list of overlay panels. It stops short of noting limits such as whether results are truncated or how nested elements are represented.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so both parameters are already documented in the schema, including the 'omit to discover overlay panels' semantics that the description restates. The description adds the downstream purpose of button refs and file inputs but no selector syntax or format detail beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a concrete verb and resource plus a scope boundary ('Inspect DOM outside the passage') and enumerates what comes back (text, buttons with refs, inputs). It does not explicitly differentiate itself from the sibling find_ui, which is the only thing keeping it from a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit branching guidance for its two modes: pass a selector to inspect a known panel, omit it to discover overlay panels such as mod GUIs and dev panels. It does not state when to prefer this over find_ui or interact, so the alternatives half of the guidance is missing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
interactFill an input or press a keyA
Interact with input fields or the keyboard: set a value on an input/textarea/select by ref, or press a key (e.g. "Enter", "ArrowUp"). For radios/checkboxes pass value "true" or "false". Inputs are refs from the last observation (or find_ui). Returns the new observation; pass format:"json" for a JSON string.
| Name | Required | Description | Default |
|---|---|---|---|
| key | No | Keyboard key to press, e.g. "Enter", "a", "ArrowDown". | |
| ref | No | Input ref from the last observation (e.g. "i1"). Optional when only pressing a key. | |
| value | No | Value to set on the input/select. Radios/checkboxes: "true" or "false". | |
| format | No | Output format: "text" (default, human-readable) or "json" (JSON string for programmatic use). | |
| game_id | Yes | Session id returned by open_game. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It does disclose the return value (new observation) and the format:"json" option, which is genuinely useful. However, it never states that this mutates page/form state, whether it submits the form, or what happens on an invalid ref/key – notable gaps for a write-side tool with zero annotation coverage.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences, front-loaded with the two primary modes before edge cases and return behavior. No padding, though the radio/checkbox note somewhat duplicates the schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with 5 params and no output schema, the description covers usage mechanics, ref sourcing, an edge case, and the return format. It stops short of error/auth behavior, but an agent has enough to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all five parameters are already documented in the schema. The description largely restates schema content (key examples, radio/checkbox "true"/"false", ref sourcing from observation/find_ui) rather than adding new semantics. Baseline 3 is appropriate when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource with two clearly enumerated modes: set a value on an input/textarea/select by ref, or press a key. This distinguishes it from siblings like click_ui (clicking), choose (selection panels), and find_ui (locating refs) without the agent needing to open schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear context: inputs are refs from the last observation or find_ui, and it explains the radio/checkbox special case (pass "true"/"false"). It does not state when NOT to use this tool (e.g. vs click_ui for buttons, or choose for option lists), so there is no explicit alternative routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_downloadsList captured browser downloadsA
List files captured from the browser into the tool's download folder (any game; the folder persists across sessions and MCP restarts). Use download_file(name|index, path) to copy one anywhere. Shows name, size, capture time and the absolute folder path.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Max files to list (default 20, newest first). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does disclose two non-obvious traits: the folder is shared across 'any game' and persists across sessions and MCP restarts. It also enumerates what is returned. It stops short of stating ordering beyond what the schema implies, error behavior, or auth needs, keeping it out of the top band.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, zero filler, and the core purpose plus cross-game scope are front-loaded before the sibling pointer and return-value summary. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter read tool with no output schema, the description covers scope, persistence, the follow-up action, and the return fields (name, size, capture time, absolute folder path). Only minor gaps remain, such as ordering nuances beyond the schema and failure behavior on an empty folder.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the single 'limit' parameter is already documented with its default and 'newest first' ordering in the schema. The description adds nothing about limit or pagination, so the baseline 3 applies when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('List files captured from the browser into the tool's download folder') and immediately scopes it ('any game'), which distinguishes it from the per-game navigation tools in the sibling set. An agent can tell this apart from download_file, which is named as the companion action.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly routes the agent onward: 'Use download_file(name|index, path) to copy one anywhere,' which clarifies this is the discovery step before retrieval. It does not state when *not* to use it or any prerequisites, but the context is unambiguous for a read-only listing tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_gamesList open gamesA
List the currently open game sessions with their id, source, story title and step count.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden, and it does disclose the return fields (id, source, story title, step count), which helps an agent plan. However, it says nothing about read-only guarantees beyond the implicit 'List', or about ordering, pagination, or the empty-list case.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler; the verb and resource come first and the returned fields follow as useful detail. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter, read-only listing tool with no output schema, naming the returned fields is the key missing piece and the description supplies it. Only the absence of any note on ordering or empty results keeps it from being fully complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so per the rubric the baseline is 4. There are no arguments whose meaning the description would need to clarify, and it correctly adds no spurious parameter discussion.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description pairs a specific verb ('List') with a precise resource ('currently open game sessions') and even enumerates the returned fields, so the agent knows exactly what it retrieves. It does not explicitly contrast itself with siblings like get_journal or observe, but 'open game sessions' is distinct enough to avoid confusion.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to call this versus alternatives such as get_journal or observe, nor any prerequisites (e.g., call it to discover game ids before load_state). Usage is only weakly implied by the word 'open', which does not rise to explicit when-to-use guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
live_viewOpen a live view of the game pageA
Give the user eyes on the actual page the agent is controlling: starts a tiny local web server (once per MCP process) that streams JPEG screenshots of the real Playwright tab (~1/s) together with passage, step, engine state, recent actions and the passage text. Returns a URL like http://127.0.0.1:4571/v/game_abc — open it in any browser (works with headless games too). Frames are captured only while someone is watching. Set open:true to also launch the URL in the default browser. DISPLAY POLICY: this is the default — and usually the only — way to show a game. One view per game: do not also switch to a headed window or loop screenshot; repeated calls return the same URL and never launch a second browser tab.
| Name | Required | Description | Default |
|---|---|---|---|
| open | No | Also open the URL in the default browser on this machine (default false). Only pass true when the user wants you to open it — at most once per game. | |
| game_id | Yes | Session id returned by open_game. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does so well: it discloses that the server starts once per MCP process, frames are captured only while someone is watching, repeated calls return the same URL and never open a second tab, and it works with headless games. These are exactly the non-obvious traits an agent needs to avoid redundant or conflicting display actions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action and the returned URL before the policy details, and every sentence carries substantive information. It is slightly long and repeats the open:true behavior that the schema already states, so not perfectly economical.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter tool with no annotations and no output schema, the description covers what it does, what it returns, when to use it, exclusivity rules, and idempotency. Nothing an agent needs to invoke it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so both parameters are already documented. The description still adds operational meaning for 'open' ('also launch the URL in the default browser') and the constraint 'at most once per game', going slightly beyond the schema, though the core param semantics remain schema-driven.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: it starts a local web server that streams JPEG screenshots of the real Playwright tab along with passage/step/engine state. It distinguishes itself from the sibling 'screenshot' by being a continuous streaming view rather than a one-off capture, and names the return value (a URL).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The DISPLAY POLICY block explicitly states when to use it ('this is the default — and usually the only — way to show a game') and when not to ('do not also switch to a headed window or loop screenshot'). It also names the idempotent alternative behavior of repeated calls, so an agent has no ambiguity about routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
load_stateLoad an in-session snapshotB
Restore a snapshot created by save_state and return the resulting observation (text; format:"json" for JSON).
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | Snapshot name (default "default"). | |
| format | No | Output format: "text" (default, human-readable) or "json" (JSON string for programmatic use). | |
| game_id | Yes | Session id returned by open_game. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It discloses the return value shape and the format flag, but a restore operation inherently overwrites the current in-session state — the description never says the existing state is discarded or whether loading is reversible, which is the single most important behavioral fact for this tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One compact sentence with the primary action front-loaded and no filler. The trailing parenthetical about format is slightly redundant with the schema but keeps the sentence readable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With 100% schema coverage and only three parameters, the description is adequate for identifying the tool and its return type in the absence of an output schema. It is however missing the destructive/prerequisite context (state replacement, dependency on a prior save_state and a live game_id session) an agent needs to invoke it safely.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, and each parameter (name default 'default', format enum, game_id from open_game) is fully documented in the schema. The description's mention of 'format:"json" for JSON' merely restates the enum, adding no syntax or defaulting detail beyond the schema, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb+resource ('Restore a snapshot') and explicitly ties it to the sibling that creates those snapshots ('created by save_state'), so the agent can distinguish it from save_state without opening a schema. It stops short of a 5 only because it doesn't say what 'restore' affects (session/game state vs. some sub-resource).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Referencing save_state implies the ordering (save first, then load), which is useful context. However, there is no explicit when-to-use guidance, no precondition statement, and no mention of alternatives such as restart or open_game for reconstituting state.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
observeObserve current game stateA
Read the current passage: text, numbered choices, input fields, dialog state and status text. Returns a formatted text observation (string); pass format:"json" for a JSON string. Inputs are paginated in windows of 40: when truncated, the header says e.g. "Inputs (41-80 of 132)" — call again with inputs_offset=80. Use find_ui(text) to jump to a specific control, get_variables for specific story variables, and since_last=true when polling to save tokens.
| Name | Required | Description | Default |
|---|---|---|---|
| format | No | Output format: "text" (default, human-readable) or "json" (JSON string for programmatic use). | |
| game_id | Yes | Session id returned by open_game. | |
| since_last | No | If true and nothing changed, return only a short "no change" note (default true; ignored in json format). | |
| inputs_limit | No | How many inputs to list (default 40). | |
| inputs_offset | No | Skip this many inputs before listing (default 0). | |
| include_status | No | Include status/caption text (default true). | |
| max_text_chars | No | Cap passage text length (default 12000). | |
| include_variables | No | Embed the (truncated) story variables (default false; prefer get_variables for specific keys). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and largely succeeds: it discloses the return type (formatted text, or a JSON string via format), the 40-item pagination window, the exact truncation header format ('Inputs (41-80 of 132)'), and the token-saving behavior of since_last. It stops short of stating rate limits or the full shape of the observation, but the read-only nature is clear.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the read purpose, then return format, then pagination mechanics, then alternatives. Dense and each sentence is informative, though the pagination example adds some length; overall efficient with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an 8-parameter read tool with no output schema, the description explains the return values (text vs. JSON string) and the pagination behavior an agent must act on. It is complete enough to call correctly, though it could say slightly more about the JSON variant's structure.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3 and the schema already documents every parameter. The description adds cross-parameter workflow meaning: how inputs_offset pairs with the truncation header, that since_last is ignored in json format, and the token trade-off of polling versus fetching variables.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Read the current passage') and enumerates exactly what is read: text, numbered choices, input fields, dialog state, status text. It distinguishes itself from find_ui ('jump to a specific control') but does not differentiate from closer siblings like inspect_ui or live_view, leaving some ambiguity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly routes to alternatives with conditions: find_ui(text) for a specific control, get_variables for specific story variables, and since_last=true when polling to save tokens. Both when-to-use and sibling selection are covered without inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
open_gameOpen a Twine gameA
Open a Twine / interactive-fiction HTML game and return the first observation (passage text, numbered choices, inputs, dialog). Accepts a local .html file, a game folder (index.html or a single html is picked), or an http(s) URL. Local games are served over 127.0.0.1 so saves work. Returns a formatted text observation (string); pass format:"json" for a JSON string. Inputs are listed in windows of 40 (see inputs_offset/inputs_limit on observe) and find_ui(text) locates any control by label.
| Name | Required | Description | Default |
|---|---|---|---|
| seed | No | Seed SugarCube PRNG (State.prng) for reproducible runs; game must use SugarCube randomness to be deterministic. | |
| format | No | Output format: "text" (default, human-readable) or "json" (JSON string for programmatic use). | |
| source | Yes | Path to an .html file or folder, or an http(s) URL. | |
| headless | No | Run browser headless (default true). Set false ONLY when the user explicitly asks to watch a real browser window — and then do not also open a live_view for the same game. | |
| block_trackers | No | Block analytics/tracker requests for a quiet session (default true). | |
| wait_timeout_ms | No | How long to wait for the game to settle after load (default 20000). | |
| include_variables | No | Embed the (truncated) story variables in the observation (default false; prefer get_variables for specific keys). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does reasonably well: it discloses the return shape ('formatted text observation (string); pass format:"json" for a JSON string'), the fact that local games are served on 127.0.0.1 to enable saves, and that inputs are windowed in groups of 40. It omits session/lifecycle behavior (e.g., whether a prior game must be closed first) and error behavior on bad sources, so it is strong but not exhaustive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose and return value are front-loaded in the first clause, which is the most important information. The trailing sentence about 40-item input windows and find_ui is useful routing but sits awkwardly at the end and slightly dilutes the core message about opening a game.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 7-parameter entry-point tool with no output schema, the description covers the essentials: accepted sources, the returned observation and its two formats, and where to look for UI controls. It leaves unstated what happens on a failed load or whether close_game must precede a re-open, which keeps it from being fully complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all seven parameters (seed, format, source, headless, block_trackers, wait_timeout_ms, include_variables). The description restates the format parameter and references inputs_offset/inputs_limit from observe, adding only marginal meaning beyond the structured fields. Baseline 3 is correct.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb + resource ('Open a Twine / interactive-fiction HTML game') and states the immediate consequence ('return the first observation'), which separates it cleanly from list_games, close_game, restart and observe. It also names the accepted source kinds, so the agent knows exactly what this tool consumes versus its siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives clear context for when the tool applies (local .html file, folder with index.html, or http(s) URL) and notes that local games are served over 127.0.0.1 so saves work. It does not explicitly exclude anything or say when to prefer a sibling (e.g., list_games first to find a source), so it falls short of the 5-level 'when-not/alternatives' bar.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
restartRestart the gameB
Restart the story from the beginning (optionally with a new PRNG seed). Returns the first observation (text; format:"json" for JSON).
| Name | Required | Description | Default |
|---|---|---|---|
| seed | No | New SugarCube PRNG seed. | |
| format | No | Output format: "text" (default, human-readable) or "json" (JSON string for programmatic use). | |
| game_id | Yes | Session id returned by open_game. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It helpfully discloses the return value ('Returns the first observation') and the output format choice, but it never states that restarting discards current progress/state or what happens to existing save points - a key behavioral trait for a mutation tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One front-loaded sentence with a compact parenthetical; nothing is wasted. The return-value clause is slightly crammed onto the end but still earns its place by covering a missing output schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 3-parameter tool with no output schema, the description adequately covers the return value and the seed/format options. However, with no annotations it leaves the destructive/state-resetting nature of the operation and its relationship to save_state/load_state unexplained.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents game_id, seed, and format in detail. The description's 'optionally with a new PRNG seed' only confirms seed's optionality and the format note repeats the schema, so baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Restart the story from the beginning'), which is clear and distinct from siblings like load_state or back by the 'from the beginning' scoping phrase. It stops short of explicitly naming the sibling it contrasts with, so it lands just under a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied: an agent can infer you restart when you want to reset the story, and the 'from the beginning' wording hints at the distinction from load_state. There is no explicit when-to-use, when-not-to-use, or named alternative, so it does not reach a 4.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
save_stateSave an in-session snapshotA
Save the full game state under a name so you can branch: save -> try a path -> load_state -> try another path. Supported natively by SugarCube; other formats report unsupported.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | Snapshot name (default "default"). | |
| game_id | Yes | Session id returned by open_game. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description must carry behavioral burden. It usefully discloses format support: SugarCube is native and other formats 'report unsupported', which tells the agent about failure behavior. However it omits whether saving overwrites an existing name, where snapshots persist, and whether the operation is reversible beyond the implied load_state pairing.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two compact sentences, front-loaded with the core action and immediately followed by the branching rationale. No filler, and the format-support caveat is placed where it is most useful.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter tool with no output schema, the description covers purpose, workflow, and a compatibility caveat, which is nearly everything needed to invoke it. It leaves minor gaps around overwrite behavior and persistence, but nothing that would cause a mis-call.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, with both name (default \"default\") and game_id documented inline, so the schema already does the heavy lifting. The description mentions saving 'under a name' but adds no syntax, naming-convention, or collision semantics beyond what the schema provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Save the full game state under a name') and distinguishes itself from the load_state sibling by pairing with it in a branching workflow. An agent can immediately tell what the tool produces and how it differs from related state tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The 'save -> try a path -> load_state -> try another path' sequence gives explicit usage context and names the companion tool. It also flags a boundary condition (only SugarCube is natively supported), though it does not exhaustively state when not to save or how this compares to restart/back.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
screenshotScreenshot the gameA
Take a PNG screenshot of the game viewport. Useful for canvas/image-driven games and visual QA. Pass path to save it to a file (returns the path instead of the image). One-shot: this is not a live view — to let the user watch the game, call live_view once instead of taking screenshots repeatedly.
| Name | Required | Description | Default |
|---|---|---|---|
| path | No | Optional output file path; when set, the PNG is written there and only the path is returned. | |
| game_id | Yes | Session id returned by open_game. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations, so description carries the full burden. It discloses the one-shot nature (not a live view), the return behavior (image vs path when path is set), and output format (PNG). It doesn't cover auth, rate limits, or error conditions, but for a screenshot tool these are less critical and the key behavioral traits are well covered.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, front-loaded with the core action, then usage context, then the crucial distinction from live_view. No wasted words and the most important routing information is placed first.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 2-param read-only tool with full schema coverage and no output schema, the description provides all needed context: what it returns, how path changes behavior, and when to use it vs live_view. Missing only edge-case error behavior, which is minor.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents both parameters. The description adds the behavioral consequence of setting path (returns path instead of image), which is a meaningful addition beyond the schema's basic description. Baseline 3 is appropriate when schema is comprehensive, but the extra consequence is slightly helpful.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States the specific action (take a PNG screenshot) and target (game viewport), and distinguishes itself from the sibling live_view by naming it. An agent can tell exactly what this does without reading the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly states when to use it (canvas/image-driven games, visual QA) and when not to (not a live view; call live_view once to let the user watch). This directly answers the 'when vs alternatives' question with a named alternative.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
upload_fileUpload a file into the gameA
Upload a local file (mod .zip, exported .save, image) into an in the game. Provide trigger_text/trigger_ref for buttons that open a picker ("Load from File…", "Import"), or let it target the file input directly. If a hardcoded selector like #saves-import does not exist in the build, use find_ui("Load from File") and pass its ref as trigger_ref. Returns the new observation (text; format:"json" for JSON).
| Name | Required | Description | Default |
|---|---|---|---|
| ref | No | Ref of a file input from inspect_ui (alternative to selector). | |
| path | Yes | Absolute path of the file on the machine running this MCP server. | |
| format | No | Output format: "text" (default, human-readable) or "json" (JSON string for programmatic use). | |
| game_id | Yes | Session id returned by open_game. | |
| selector | No | CSS selector of an <input type=file> (used when no trigger is given). | |
| trigger_ref | No | Ref of the trigger button (alternative to trigger_text). | |
| trigger_text | No | Visible text of the button that opens the file picker. | |
| trigger_selector | No | CSS selector of the trigger button (e.g. "#saves-import"). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It discloses the return value ('Returns the new observation; format:"json" for JSON'), which is genuinely useful, but says nothing about permissions, what happens to existing game state on upload, failure behavior when no ref/selector/trigger resolves, or whether the upload triggers navigation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, no filler, and the core action is front-loaded in the first clause. The middle sentence is dense with three alternate targeting modes, but each clause carries information an agent needs.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an 8-parameter tool with no annotations and no output schema, the description covers the primary flow, the parameter interplay, the fallback path, and the return format. Remaining gaps are error/failure handling and side effects on game state, which are secondary but not negligible.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% so the baseline is 3, but the description goes beyond the schema by explaining the relationship among the four targeting parameters (trigger_* for picker-opening buttons vs selector/ref for the input itself) and the ordered fallback strategy. That is real added meaning over the flat per-parameter descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Upload a local file ... into an <input type=file> in the game') and enumerates the file kinds involved (mod .zip, exported .save, image). This clearly separates it from the sibling download_file, the inverse operation, without needing to open either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives concrete routing between the trigger parameters and the direct-target path ('Provide trigger_text/trigger_ref for buttons that open a picker ... or let it target the file input directly'), plus an explicit fallback when a hardcoded selector is missing ('use find_ui("Load from File") and pass its ref as trigger_ref'). It does not, however, state when this tool should be preferred over other siblings like interact or click_ui for the same UI.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
waitWait for the gameB
Wait for time (ms), for a text to appear, and/or for the page to settle, then return the new observation. Useful for timed passages and animations. Returns text; pass format:"json" for a JSON string.
| Name | Required | Description | Default |
|---|---|---|---|
| ms | No | Milliseconds to wait. | |
| format | No | Output format: "text" (default, human-readable) or "json" (JSON string for programmatic use). | |
| game_id | Yes | Session id returned by open_game. | |
| for_text | No | Wait until this text appears anywhere on the page (up to timeout_ms). | |
| timeout_ms | No | Timeout for for_text (default 15000). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations provided, so description carries full burden. It discloses return format (text/json) and waiting conditions, but omits timeout behavior for ms wait, error handling if text never appears, and whether the observation is returned as a snapshot. Adequate but incomplete for a wait tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Single sentence stating what it waits for and returns, followed by a brief note on use case and format parameter. Slightly dense but front-loaded and efficient.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with 5 parameters and no output schema, description covers the core waiting and return format but lacks details on timeout interaction, error cases, and how the observation is structured. Sufficient for basic invocation but missing edge-case guidance.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and parameters have descriptions, so baseline 3 is appropriate. Description mentions ms, for_text, and format, but adds no syntax or default details beyond schema; timeout_ms and game_id are not mentioned in description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific action (wait) with three conditions (time, text appearance, page settle) and the return (new observation). Distinguishes from siblings like observe (immediate) and screenshot, though doesn't explicitly name an alternative.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Mentions 'useful for timed passages and animations' which implies usage context, but doesn't specify when to choose wait vs observe or other tools, nor when to avoid waiting.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
22 tool updates
v0.1.3- First observed
back - First observed
choose - First observed
click_ui - First observed
close_game - First observed
download_file - First observed
find_ui - First observed
get_console_errors - First observed
get_journal - First observed
get_variables - First observed
inspect_ui - First observed
interact - First observed
list_downloads - First observed
list_games - First observed
live_view - First observed
load_state - First observed
observe - First observed
open_game - First observed
restart - First observed
save_state - First observed
screenshot - First observed
upload_file - First observed
wait
TDQS
Scored across 22 tools
Most tools target distinct actions (observe, choose, save/load, screenshots, files), but there are ambiguous pairs: choose() explicitly answers dialog buttons while click_ui() claims to 'click dialogs', and find_ui vs inspect_ui both locate controls. These overlaps require careful description reading to pick correctly.
All names are snake_case and follow a verb_noun pattern for resource-targeting tools (save_state, open_game, get_variables, list_downloads). The spread of bare verbs (observe, choose, wait, back, restart, find_ui, inspect_ui, click_ui) deviates slightly, though it reads as an intentional convention for acting on the current game state.
22 tools is on the heavy side for a game-automation server. Several functions arguably overlap (choose/click_ui, find_ui/inspect_ui, screenshot/live_view), suggesting some consolidation is possible, though most tools do earn their place.
The surface spans the full play lifecycle: open/close, observe, choose/interact, wait, back/restart, save/load state, screenshots/live view, file upload/download, variables, console errors, journal, and UI inspection. Minor gaps exist (no direct story-variable setter, no passage search/jump), but these are workaroundable via interact and find_ui.
Maintenance
Related MCP Connectors
Run, debug and inspect Playwright E2E tests from any AI agent: diagnostics, live DOM, selectors.
Hosted browser for AI agents: screenshots, post-JS DOM, console, WCAG. No install, no API key.
Build, version, review, and export websites, web apps, and games from a conversation.
Web scraping for agents. Point it at a URL and it returns the page as clean markdown, JavaScript-rendered pages included. Point it at a site and it maps the URLs or crawls the section you need in the background, a few pages at a time so results fit in the conversation. Search the web and read full pages, extract fields with a JSON schema you define (validated, never invented), read a store's catalogue or a blog's posts from the platform's own feed, and check whether a page has changed. Failed requests cost nothing. The free plan includes 1,500 credits a month.
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceEnables AI assistants to control web browsers through Playwright automation, providing 50+ tools for navigation, interaction, testing, accessibility audits, and visual testing across Chromium, Firefox, and WebKit.17 npmMIT
- AlicenseNot gradedqualityCmaintenanceProvides browser automation capabilities for LLMs using Playwright, leveraging structured accessibility snapshots to interact with web pages without needing vision models. It enables tasks like web navigation, data extraction, and automated testing through a lightweight and deterministic toolset.7 npmApache 2.0
- AlicenseNot gradedqualityCmaintenanceEnables AI assistants to play interactive fiction games (Glulx and Z-machine) through the Model Context Protocol, with automatic save/restore and optional journaling mode.MIT
- AlicenseNot gradedqualityCmaintenanceEnables AI coding agents to autonomously interact with and test web applications in a real browser, providing DOM/Accessibility tree extraction, runtime telemetry, screenshot capture, and Markdown test reports.98 npm1MIT