Skip to main content
Glama

twine-play-mcp

English · 简体中文

An MCP server that lets AI agents play, test and QA Twine / interactive-fiction HTML games.

The agent reads the current passage, sees numbered choices, clicks them, watches story variables, screenshots the game, and can save/restore state to explore branches — all through a small, token-friendly tool surface instead of a generic browser automation API.

Agent  ──MCP(stdio)──>  twine-play-mcp  ──Playwright──>  headless Chrome
                              │                               │
                              │  static server (127.0.0.1)    │  injected bridge
                              └────────> game HTML <──────────┘

Why not a generic browser MCP?

Generic browser MCPs make the model guess DOM selectors, dump whole pages into context and have no notion of "story state". This server adds a semantic layer:

  • Passage view: text as Markdown, passage name, format/version, story metadata

  • Numbered choices with target passage names (and external-link blocking)

  • Story variables (SugarCube State.variables) with safe depth/size caps

  • Native state: SugarCube Engine.backward/forward, Save.base64 snapshots

  • Format detection: SugarCube first, DOM fallback for Harlowe / Snowman / Chapbook / unknown

  • Tracker blocking and quiet console/network capture for clean playtesting

Related MCP server: Playwright MCP

Requirements

  • Node.js >= 20 (developed on 26)

  • Google Chrome installed (uses channel: 'chrome'; no 200 MB browser download)

  • Linux/macOS/Windows

Install

npm install -g twine-play-mcp   # or: npx twine-play-mcp

No build step, no browser download — the package ships the compiled server and the page bridge, and drives the Chrome you already have.

Then point it at any published Twine HTML file (or a folder containing the game + assets):

twine-play-mcp         # MCP server on stdio

From source (for development)

git clone https://github.com/adorablelovelymia/twine-play-mcp.git
cd twine-play-mcp
npm install
npm run build          # compiles to dist/ and copies the page bridge

# sanity checks (optional)
npm run spike          # 17 end-to-end checks against a real SugarCube game
npm run smoke          # spawns the MCP server over stdio and drives it with the MCP SDK

Client configuration

The snippets below use npx, so no global install is required. If you installed globally, replace "npx" + "twine-play-mcp" with "twine-play-mcp" on its own.

OpenCode (~/.config/opencode/opencode.json)

{
  "$schema": "https://opencode.ai/config.json",
  "mcp": {
    "twine-play": {
      "type": "local",
      "command": ["npx", "-y", "twine-play-mcp"],
      "enabled": true
    }
  }
}

Claude Desktop / Cursor / any mcpServers client

{
  "mcpServers": {
    "twine-play": {
      "command": "npx",
      "args": ["-y", "twine-play-mcp"]
    }
  }
}

Environment variables:

Variable

Purpose

TWMCP_CHROME_PATH

Chrome executable if channel: 'chrome' cannot find it

TWMCP_DOWNLOAD_DIR

Folder where browser downloads are captured (default ~/.cache/twine-play-mcp/downloads)

TWMCP_VIEW_PORT / TWMCP_VIEW_HOST

Live-view server port (default 4571, auto-increments if busy) and bind host (default 127.0.0.1)

Tools

Tool

What it does

open_game

Open a local HTML file/folder or URL; optional PRNG seed; returns first observation

observe

Passage text, numbered choices, inputs (paginated windows of 40 with inputs_offset), dialog, status bar; since_last saves tokens; format:"json" for structured output

choose

Click by 1-based number or label; works for passage choices and dialog buttons; expected guard; external links blocked by default

wait

Wait for ms / for text / for the DOM to settle

interact

Fill inputs/selects/checkboxes by ref, or press a key

find_ui

Find buttons/links/labels/inputs by visible text or input name; returns refs for click_ui/interact. The fastest way to reach radio/checkbox options (SugarCube macro labels)

click_ui

Click dialogs, sidebar and menus (by ref / CSS selector / visible text), including <label>-based controls and iframes

upload_file

Upload a local file into an <input type=file> (mod .zip, .save import) via trigger button or direct input, frame-aware

download_file

Browser file control (any game): copy a captured download to a path — by trigger_text/trigger_ref (click the game's download button), by name/index from the download folder, or the newest file by default. The folder copy stays

list_downloads

List files captured into the tool's persistent download folder (survives sessions/restarts), with name, size, time and absolute path

inspect_ui

Inspect or discover UI panels outside the passage (mod GUIs, backstage); lists buttons/inputs and file inputs

get_variables

Read story variables by dot path (V.hairlength), or a shallow top-level key summary; avoids dumping the whole variable state

back

Undo one passage (SugarCube Engine.backward)

restart

Restart from the beginning, optionally reseeding the PRNG

save_state / load_state

Named in-session snapshots for branch exploration

screenshot

PNG of the viewport (canvas/visual games, visual QA); pass path to save to disk

live_view

Give the user eyes on the real page: local URL streaming JPEG frames (~1/s) of the actual Playwright tab + passage/step/journal; works headless; open:true launches the default browser

get_console_errors

JS exceptions, console errors and HTTP failures captured from the page

get_journal

Action history: passages visited, choices taken, coverage counts

list_games / close_game

Session management

Watching the page (live view)

live_view(game_id) starts a tiny local server (once per MCP process, port 4571+) and returns a URL like http://127.0.0.1:4571/v/game_abc. Open it in any browser (or pass open: true) to watch the actual tab the agent is driving — a JPEG frame about once per second, plus passage, step, engine state, recent actions and the passage text. It works with headless games, frames are captured only while somebody is watching, and closing the game stops it. For a raw browser window instead, open the game with headless: false (open_game).

Handy combo: if your client can show a web page in a side pane (e.g. OpenCode's Review pane / browser.tabs.open), point it at the live-view URL and you can follow along while the agent plays.

Pick one view (agents). To keep the user's screen clean, show a running game through exactly one channel — never stack them:

  1. Default: live_view — hand the URL to the user, or pass open: true once to launch it for them. Repeated calls reuse the same view and never open another tab.

  2. Only on explicit request: open_game(headless: false) when the user asks for a real browser window. Don't add a live view on top; a headed window is already visible.

  3. screenshot is a one-shot visual check, not a stream — don't loop it to "show" the game.

If a view (live view tab or headed window) is already open, reuse it instead of starting a second one. The MCP server ships this same policy in its instructions field, so MCP clients can pass it to the model automatically; the tool descriptions repeat it where it matters (live_view, open_game.headless, screenshot).

Agent ergonomics

  • Output: every play tool returns a formatted text observation (a string). Pass format:"json" to receive a JSON string instead (JSON.parse it) with passage, text, choices[{n,label,target}], inputs[{ref,kind,label,checked}], inputsTotal, dialog, status.

  • Inputs are paginated, not truncated: a header like Inputs (41-80 of 140) plus inputs_offset=80 means everything is reachable — no silent hard cap.

  • Label matching: click_ui(text) and find_ui(text) understand SugarCube <<radiobutton>> / <<checkbox>> labels, so options like "Jet black" or combat moves like "Punch" are clickable by text.

  • Errors are compact: failures return ERROR: code — message, a Hint, the current passage and the available choices — never a full observation dump.

  • Variables: observations do not embed variable blobs by default; use get_variables for the keys you care about. include_variables:true is still available when you want the (truncated) dump.

  • Dialogs: dialog buttons appear as numbered choices tagged [dialog], and a Dialog buttons: line lists them; checkbox labels are shown on the input line.

Format support

Format

Detect

Text/choices

Variables

Passage name

Back

Snapshots

SugarCube 2.21+

✅

✅

✅ State.variables

✅

✅ Engine.backward

✅ Save.base64 (2.37+) / Save.deserialize (older)

Harlowe 3

✅

✅

— (engine internals are private)

—

✅ sidebar undo

—

Snowman 2

✅

✅

✅ story.state

✅

—

✅ state JSON

Chapbook 1

✅

✅

✅ engine.state.saveToObject()

✅ trail

—

✅ restoreFromObject

Unknown HTML

generic

✅ DOM heuristics

—

—

—

—

Everything degrades gracefully: an unknown or exotic format still plays with the generic DOM path; format-specific tools report unsupported instead of failing.

Play-session example (what the agent sees)

[sugarcube 2.37.3 · step 3 · engine=idle · passage: 069]
You squeeze through the narrow gap...

Choices (2):
  1. Go deeper -> 070
  2. Check the mirror

Status:
Resistance: 500/500
Variables: {"resistance":500,"pleasure":0,"degradation":0,...}

How it works

  • src/bridge/bridge.js is injected into every page (addInitScript) and exposes window.__twineMCP: format detection, passage/choice extraction, click/fill helpers, settle-waiting, snapshots and seeding. All server calls go through this bridge only.

  • src/session.ts owns the browser, one BrowserContext per game (isolated saves) and a tiny static server so local games run on http://127.0.0.1 (localStorage works).

  • src/render.ts turns observations into compact Markdown for the model.

  • Choices get temporary data-twmcp-ref attributes; the server prefers real Playwright clicks and falls back to DOM clicks for exotic macro-generated links.

  • Spoiler policy: only what a player can see is returned. No passage lists or source dumps are exposed.

Complex games

Real games are not just passages and links. The MCP handles the awkward parts:

  • Modal dialogs (SugarCube #ui-dialog, content gates, settings): their text appears as a Dialog: block and their buttons/inputs are numbered like choices, so the agent can tick a consent checkbox (interact) and click Enter (choose).

  • Iframes: mod managers and dev panels often live in a child frame. inspect_ui discovers them (marked [iframe]), click_ui by text and upload_file search every frame.

  • File workflows: uploads go through upload_file (mod .zip, save import) — by clicking a trigger (trigger_text / trigger_selector, e.g. #saves-import) or pointing at an <input type=file> directly. Downloads go the other way through a persistent download folder: every browser download is captured there (TWMCP_DOWNLOAD_DIR overrides the location), list_downloads shows the contents, and download_file copies one anywhere (path, default <cwd>/downloads/<name>) — either by clicking the game's export button or by name/index afterwards. No manual temp-folder copying, and files survive close_game and MCP restarts.

  • DoL case study: test/fixtures aside, scripts/dol-mcp-test.ts drives Degrees of Lewdity end to end — consent gate → importing ModI18N.mod.zip and GameOriginalImagePack.mod.zip through the in-game ModLoader GUI → page reload → importing a real .save through the SAVES dialog → several turns of normal play.

Tests

npm run spike      # 17 checks against a real SugarCube 2.37 game (play, back, snapshot, screenshot)
npm run formats    # 4 compiled fixtures: SugarCube 2.30, Harlowe 3.1, Snowman 2.0, Chapbook 1.0
npm run smoke      # spawns the built MCP server and drives the tools over stdio
npm run clarity    # agent-ergonomics regression on DoL character creation (pagination, labels, variables)
npm run fixtures   # rebuild test/fixtures/compiled/*.html with Tweego (see test/fixtures/build.sh)
npx tsx scripts/dol-mcp-test.ts   # Degrees of Lewdity: gate, mod import, save import, play

scripts/inspect.ts <fixture> dumps the DOM/story-format internals of a game — handy when adding a new adapter.

Status / roadmap

  • M1: SugarCube adapter, generic DOM fallback, observation/choice/input/wait/screenshot, snapshots, backtracking, console+network QA capture, stdio MCP, spike + smoke tests

  • M2: Harlowe / Chapbook / Snowman adapters verified against compiled fixtures

  • M2: play journal (get_journal) for run summaries, resuming and QA coverage

  • M3: dialogs/iframe-aware UI control (click_ui, inspect_ui, upload_file) — verified on Degrees of Lewdity (mod import + save import + play)

  • M4: agent ergonomics — input pagination + totals, find_ui label search, get_variables, format:"json", compact errors (driven by a naive-agent playtest that stalled on DoL character creation)

  • M4: file workflows both ways — upload_file for mods/saves, download_file + list_downloads for a persistent browser download folder (no temp-folder copying)

  • M5: live view — watch the real page in any browser (~1 fps frames + passage/step/journal); headed mode via open_game(headless: false)

  • M5: npm packaging — published as twine-play-mcp

  • M3: click_at for canvas games, spoiler-gated story-map analysis

Safety notes

  • Page scripts run in Chrome's sandbox; the bridge never exposes Node to the page.

  • No arbitrary eval tool is exposed to the agent.

  • Analytics/tracker hosts are blocked by default (block_trackers: false to disable).

  • External links are blocked unless allow_external: true is passed.

中文文档

完整中文版见 README-zh.md(工具一览、客户端配置、格式支持、实时视图策略等均已翻译)。

一句话:这是一个让 AI agent 游玩 / 测试 Twine 文字游戏的 MCP 服务——无头 Chrome + 页面桥, 22 个工具,本地游戏经内置静态服务器以 http://127.0.0.1 打开(保证存档可用), live_view 让你用任意浏览器实时看到 AI 正在操作的真实页面。

npm install -g twine-play-mcp   # 或 npx twine-play-mcp

Available Tools

22 tools
backGo back one stepB

Undo the last passage navigation when the story format supports it (SugarCube: Engine.backward). Returns the new observation (text; format:"json" for JSON).

ParametersJSON Schema
NameRequiredDescriptionDefault
formatNoOutput format: "text" (default, human-readable) or "json" (JSON string for programmatic use).
game_idYesSession id returned by open_game.

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries full burden. It discloses the format-support constraint and returns the new observation, which is useful behavioral context. However, it doesn't state what happens if unsupported (error message? no-op?), whether back at the start is a no-op, or whether the action is idempotent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single concise sentence with the key constraint (format support) front-loaded and return behavior appended. No waste, though the parenthetical engine reference slightly interrupts flow.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple two-param undo tool with no annotations and no output schema, the description covers the essentials: what it does, when it works, and what it returns. But it leaves out error behavior and edge cases (e.g., already at start), which an agent would need to call it correctly in all cases.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both parameters are documented in the schema. The description adds no meaning beyond what's there, mentioning only the format return type. Baseline 3 when schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific action ('undo the last passage navigation') with the resource (story navigation). It names the underlying engine call (SugarCube: Engine.backward), clarifying exactly what happens. It does not explicitly distinguish itself from siblings like restart or load_state, though 'undo last passage' is fairly distinct.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies when to use it (to undo a navigation, and only when the story format supports it), but doesn't name alternatives or state when not to use it. The format-support caveat is a helpful condition, but no sibling routing is provided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

chooseClick a choiceA

Click a passage choice by its 1-based number from the last observation, or by (partial) label text. Numbered choices include dialog buttons (tagged [dialog] in observations) — so choose() answers dialogs too. Waits for the game to settle and returns the new observation. Pass expected to guard against clicking the wrong link.

ParametersJSON Schema
NameRequiredDescriptionDefault
choiceYes1-based choice number from the last observation, or label text (case-insensitive, may be partial if unambiguous).
formatNoOutput format: "text" (default, human-readable) or "json" (JSON string for programmatic use).
game_idYesSession id returned by open_game.
expectedNoSubstring the clicked choice label must contain; fails safely if it does not match.
allow_externalNoAllow following links that leave the game (default false, they are blocked).

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the behavioral burden and does disclose key traits: it blocks until the game settles and returns the new observation, and that expected fails safely rather than clicking the wrong link. It omits error behavior and timeout details, but the synchronous-settle and return semantics are genuinely valuable additions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four short sentences, front-loaded with the core action and matching rule before the dialog note and safety guard. Minor restatement of the choice/label mechanics that the schema also covers, but no serious padding.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a five-parameter tool with no output schema and no annotations, the description covers the action, matching rule, wait/return behavior, and the primary safety guard. Adequate, though it could note more on what happens on timeout or an unmatched expected value.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so every parameter (choice, format, game_id, expected, allow_external) is already documented. The description paraphrases the choice and expected semantics but adds no syntax or format detail beyond the schema, so baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (click) and resource (a passage choice), and specifies the mechanism: 1-based number from the last observation or partial label text. The note that numbered choices include [dialog] buttons successfully carves it apart from UI-level siblings like click_ui and interact.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explains the primary usage context (driving passage choices and answering dialogs) and offers the expected guard for safe clicking. It does not explicitly name alternatives or state when NOT to use it versus click_ui/interact, so it stops short of full routing guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

click_uiClick UI outside the passageA

Click dialogs, sidebar buttons and menus (SAVES, OPTIONS, ModLoader banner, modal buttons) by ref, CSS selector or visible text. Text matching covers -based controls too (SugarCube radio/checkbox options like "Jet black" or "Punch"); shortest match wins, exact=true for exact text. Returns the new observation (text; format:"json" for JSON). Use choose() for numbered passage choices and dialog buttons.

ParametersJSON Schema
NameRequiredDescriptionDefault
refNoUI ref from an observation (e.g. "u1" or "x3").
textNoVisible text of the target (case-insensitive, partial match; the shortest match wins).
exactNoRequire an exact text match (default false).
formatNoOutput format: "text" (default, human-readable) or "json" (JSON string for programmatic use).
game_idYesSession id returned by open_game.
selectorNoCSS selector, if you know the exact element.

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does disclose meaningful behavior: shortest-match-wins resolution, exact=true semantics, label-based matching, and the return value (the new observation, text or JSON). It does not state side effects, whether the click mutates saved state, or failure behavior when no element matches, which are the remaining gaps.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Dense but front-loaded: the action and target types come first, then matching rules, then the return value, then the sibling routing. Every clause carries information, though the parenthetical control list and the dialog/button phrasing could be tightened slightly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, but the description states that it returns the new observation and how to request JSON, so return-value handling is covered. For a 6-parameter tool with one required arg and full schema coverage, this is complete enough to invoke correctly, missing only explicit side-effect/failure disclosure.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds resolution semantics not fully conveyed by the field docs: that text matching also covers <label>-based radio/checkbox options, that the shortest match wins, and that exact=true is opt-in. This is genuine extra meaning beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a concrete verb (click) plus the exact target classes (dialogs, sidebar buttons, menus, named controls) and the three addressing mechanisms (ref, CSS selector, visible text). It explicitly distinguishes itself from choose(), so an agent can separate it from the closest sibling without reading either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It names the alternative (choose()) and the condition that selects it (numbered passage choices and dialog buttons), and explains when text matching applies (label-based SugarCube controls). There is mild ambiguity between 'Click dialogs ... modal buttons' and 'use choose() for ... dialog buttons', but the routing rule itself is explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

close_gameClose a gameB

Close the browser tab and static server for a game session.

ParametersJSON Schema
NameRequiredDescriptionDefault
game_idYesSession id returned by open_game.

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full disclosure burden. It does add real value by naming the two things destroyed (browser tab and static server), implying the session is torn down, but it is silent on whether in-flight state is lost, whether the session can be reopened, what errors occur for an unknown game_id, and whether this requires the game to be open. That is a meaningful gap for a teardown operation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler; the affected resources come immediately after the verb. Nothing here needs to be trimmed or reordered.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-parameter tool with full schema coverage and no output schema, the description covers the basic action. But as an unannotated teardown tool it omits the information an agent most needs: effect on saved/unsaved state and how it relates to restart or open_game. Adequate, but with clear gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and the schema already documents game_id as the id returned by open_game. The description adds nothing about the parameter (e.g. format, validity conditions for stale ids), so the baseline of 3 for fully-covered schemas is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('close') and the exact resources affected ('browser tab and static server for a game session'), which an agent can act on directly. It does not, however, distinguish itself from siblings like 'restart' or 'back', which could also terminate a session view, leaving some ambiguity about which teardown tool applies.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description says what gets closed but never says when to call it — no mention of end-of-session cleanup, no exclusion of alternatives such as 'restart' (which presumably keeps the session) or 'back'. The only usage signal is the indirect reference to 'open_game' in the schema, so the agent must infer the pairing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

download_fileSave a captured browser download to a fileA

Browser file control for any game: every download is captured into the tool's download folder (see list_downloads). Three modes: (a) pass trigger_text/trigger_ref/trigger_selector to click the game's download button and take that file; (b) pass name or index to take an already-captured file from the folder; (c) pass none of those to take the newest file in the folder. Without path the file stays in the download folder (path = its location); with path (a file or a directory) a copy is placed there too. Returns the path, size and the new observation.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameNoName of an already-captured file from list_downloads (exact, suffix, or unique substring).
pathNoDestination file or directory (default <cwd>/downloads/<name>).
indexNo1-based index from list_downloads (newest first).
formatNoOutput format: "text" (default, human-readable) or "json" (JSON string for programmatic use).
game_idYesSession id returned by open_game.
timeout_msNoHow long to wait for the download after clicking (default 30000).
trigger_refNoRef of the download button (alternative to trigger_text).
trigger_textNoVisible text of the button that starts the download (optional).
trigger_selectorNoCSS selector of the download button (alternative to trigger_text).

TDQS

A4.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and largely meets it: it discloses where downloads land, that path produces a copy (not a move), and that it returns path, size, and the new observation. It does not state whether the operation can fail/timeout behavior beyond timeout_ms, or whether re-invoking re-downloads versus reuses a captured file.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with purpose, then organized into labeled modes (a)/(b)/(c), so the reader can parse by branch. It is dense and somewhat run-on, but essentially every sentence earns its place; minor tightening on the path/return sentence would help.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 9-parameter, no-output-schema tool, the description supplies the missing return-value summary (path, size, observation) and the mode logic tying the optional parameters together, plus points to list_downloads for name/index sources. Nothing an agent needs to invoke it correctly is absent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the parameter baseline is 3, but the description adds genuine value beyond the schema by explaining how parameters combine into mutually exclusive modes (trigger_* vs name/index vs neither) and clarifying copy semantics of path, which the schema only describes as 'destination'.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('take/copy a captured browser download to a file') and enumerates three distinct operating modes. It names the sibling it depends on (list_downloads) so the agent can place it in the workflow without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit when-to-use branches: (a) trigger_* to click and capture, (b) name/index to take an already-captured file, (c) neither to take the newest file. It also states the condition for path (omit to leave in folder, supply to copy out). This is exactly the guidance an agent needs to pick parameters.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

find_uiFind a control by textA

Search visible controls (buttons, links, labels, inputs) by label text and/or input name; returns refs usable with click_ui(ref), interact(ref, value) and upload_file(ref). This is the fastest way to reach radio/checkbox options (e.g. "Jet black", "Punch") and any input beyond the 40-item observation window. Results are capped by limit; no game state changes.

ParametersJSON Schema
NameRequiredDescriptionDefault
kindNoFilter by kind: button, link, label, input, radio, checkbox, select, textarea.
nameNoFilter by input name attribute (exact).
textNoLabel text to search for (case-insensitive, partial by default).
exactNoRequire an exact label match (default false).
limitNoMax matches to return (default 20).
game_idYesSession id returned by open_game.

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries full burden. It does disclose that results are capped by limit and that there are 'no game state changes' (read-only), which is useful behavioral context. That said, it doesn't describe permission requirements, pagination, failure modes when no match is found, or return shape beyond 'refs'.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with the core purpose and search criteria, then immediately covering return refs and usage context. Every clause adds value: use cases, cap behavior, and non-mutation guarantee.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only search tool with a rich 6-param schema and no output schema, the description is quite complete: it specifies inputs, return type (refs), downstream tools, result capping, and no side effects. It lacks explicit guidance on when to prefer over inspect_ui or how to handle empty results, but covers the essentials well.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the schema fully documents each parameter, establishing a baseline of 3. The description adds conceptual value by mentioning 'label text and/or input name' (mapping to text/name) and 'results are capped by limit', but does not add syntax, format, or interaction details beyond what the schema provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Search') and resource ('visible controls (buttons, links, labels, inputs)'), and names the exact criteria (label text and/or input name). It also names the ref-using siblings (click_ui, interact, upload_file), making it distinguishable from inspect_ui and observe without opening schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides clear positive context: 'fastest way to reach radio/checkbox options ... and any input beyond the 40-item observation window,' which implicitly tells when to prefer this over observe/inspect_ui. However, it never explicitly states when NOT to use it or contrasts it directly with inspect_ui, so it falls short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_console_errorsGet console errorsB

Return JavaScript errors/warnings captured from the page (useful for playtesting / QA).

ParametersJSON Schema
NameRequiredDescriptionDefault
clearNoClear the buffer after reading (default false).
game_idYesSession id returned by open_game.

TDQS

B3.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden. It hints that errors are 'captured' (an accumulating buffer) and the schema's 'clear' param reveals buffer semantics, but the description says nothing about whether errors persist across reads, whether output is bounded/truncated, or what the response shape looks like.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single tight sentence with the resource front-loaded and the usage hint appended. Nothing is wasted and nothing important is buried.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple two-parameter read tool this is close to adequate, but with no output schema the description should at least indicate the return shape (list of entries, fields like message/level/timestamp) and whether reads are cumulative. That gap leaves the agent guessing about results.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both parameters (clear, game_id) are already documented in the schema, and the description adds no syntax or format detail. Baseline 3 applies when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb (Return) and resource (JavaScript errors/warnings captured from the page), which is unambiguous and clearly distinct from siblings like get_journal or inspect_ui. It does not explicitly contrast with a sibling, but the resource is unique enough that the distinction is obvious.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The parenthetical '(useful for playtesting / QA)' implies the usage context but gives no explicit when-to-use vs when-not guidance and names no alternatives. Usage is only weakly implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_journalGet the play journalA

Return this session's action history: every passage visited and every choice taken (including back/load events). Useful to summarise a playthrough, resume a run, or report coverage for QA.

ParametersJSON Schema
NameRequiredDescriptionDefault
clearNoClear the journal after reading (default false).
limitNoShow the last N entries (default 50).
game_idYesSession id returned by open_game.

TDQS

A3.5/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden. It discloses what the journal contains (passages, choices, back/load events), but it completely omits the fact that the optional 'clear' parameter destructively wipes the journal after reading, and says nothing about permissions, reversibility, or side effects.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences with zero filler. The core return content is front-loaded, and the usage examples follow immediately, making it easy to skim.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read tool with no output schema, the description adequately explains what history is returned, but it misses the critical destructive side effect of the 'clear' parameter and does not describe the returned entry shape or pagination via 'limit'.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all three parameters (game_id, clear, limit). The description adds no parameter-level meaning beyond what the schema provides, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'Return this session's action history', and enumerates the content (passages, choices, back/load events). It clearly differentiates the tool's purpose from state-saving or UI-inspection siblings, but never explicitly names an alternative tool it should not be confused with.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives three clear usage contexts: summarising a playthrough, resuming a run, and reporting coverage for QA. No when-not-to-use or alternative-tool guidance is provided, but the positive guidance is specific and actionable.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_variablesRead story variablesA

Read specific story variables (SugarCube State.variables) by dot path, e.g. ["haircolour", "background", "player.background"]. Accepts "V.x", "variables.x" or plain "x". With no paths, returns a shallow summary of the top-level keys. Output is JSON.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathsNoDot paths to read (default: top-level key summary).
game_idYesSession id returned by open_game.

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description usefully discloses return format ("Output is JSON"), the no-paths default (shallow top-level summary), and the accepted prefix variants. It is silent on invalid-path handling and result-size limits, but for a simple read tool it adds meaningful context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three front-loaded sentences, each carrying distinct information (what, format variants, default/output). No filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a two-parameter read tool with no output schema and no annotations, the description covers purpose, input format, default behavior, and output type. Only edge-case behavior (invalid paths, the 50-item cap) is left unstated.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. The description goes further by documenting accepted input forms ("V.x", "variables.x", plain "x") that the schema does not mention, giving genuine added meaning.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (read) and resource (story variables / SugarCube State.variables) with concrete dot-path examples. It is clearly distinguable from sibling tools like get_journal or inspect_ui, which read different resources.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied by the examples and the default-behavior note, but there is no explicit when-to-use, when-not-to-use, or routing to an alternative sibling. It leaves the agent to infer context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

inspect_uiInspect a UI panelA

Inspect DOM outside the passage. Pass a CSS selector to get its text, buttons (with refs usable in click_ui) and inputs (file inputs usable in upload_file). Without a selector, lists overlay panels (mod GUIs, dev panels) that contain buttons or file inputs.

ParametersJSON Schema
NameRequiredDescriptionDefault
game_idYesSession id returned by open_game.
selectorNoCSS selector of the panel to inspect; omit to discover overlay panels.

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does real work: it discloses the two behavioral modes, the shape of the return (text, buttons, inputs), and cross-tool contracts (refs usable in click_ui, file inputs usable in upload_file). It omits failure behavior for invalid selectors and whether inspection has side effects.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences, front-loaded with the core action, then the two modes, then the return contents. No filler and every clause carries information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description must describe returns and it does reasonably well: text, buttons with refs, inputs, or a list of overlay panels. It stops short of noting limits such as whether results are truncated or how nested elements are represented.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so both parameters are already documented in the schema, including the 'omit to discover overlay panels' semantics that the description restates. The description adds the downstream purpose of button refs and file inputs but no selector syntax or format detail beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a concrete verb and resource plus a scope boundary ('Inspect DOM outside the passage') and enumerates what comes back (text, buttons with refs, inputs). It does not explicitly differentiate itself from the sibling find_ui, which is the only thing keeping it from a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit branching guidance for its two modes: pass a selector to inspect a known panel, omit it to discover overlay panels such as mod GUIs and dev panels. It does not state when to prefer this over find_ui or interact, so the alternatives half of the guidance is missing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

interactFill an input or press a keyA

Interact with input fields or the keyboard: set a value on an input/textarea/select by ref, or press a key (e.g. "Enter", "ArrowUp"). For radios/checkboxes pass value "true" or "false". Inputs are refs from the last observation (or find_ui). Returns the new observation; pass format:"json" for a JSON string.

ParametersJSON Schema
NameRequiredDescriptionDefault
keyNoKeyboard key to press, e.g. "Enter", "a", "ArrowDown".
refNoInput ref from the last observation (e.g. "i1"). Optional when only pressing a key.
valueNoValue to set on the input/select. Radios/checkboxes: "true" or "false".
formatNoOutput format: "text" (default, human-readable) or "json" (JSON string for programmatic use).
game_idYesSession id returned by open_game.

TDQS

A3.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden. It does disclose the return value (new observation) and the format:"json" option, which is genuinely useful. However, it never states that this mutates page/form state, whether it submits the form, or what happens on an invalid ref/key – notable gaps for a write-side tool with zero annotation coverage.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences, front-loaded with the two primary modes before edge cases and return behavior. No padding, though the radio/checkbox note somewhat duplicates the schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation tool with 5 params and no output schema, the description covers usage mechanics, ref sourcing, an edge case, and the return format. It stops short of error/auth behavior, but an agent has enough to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all five parameters are already documented in the schema. The description largely restates schema content (key examples, radio/checkbox "true"/"false", ref sourcing from observation/find_ui) rather than adding new semantics. Baseline 3 is appropriate when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource with two clearly enumerated modes: set a value on an input/textarea/select by ref, or press a key. This distinguishes it from siblings like click_ui (clicking), choose (selection panels), and find_ui (locating refs) without the agent needing to open schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear context: inputs are refs from the last observation or find_ui, and it explains the radio/checkbox special case (pass "true"/"false"). It does not state when NOT to use this tool (e.g. vs click_ui for buttons, or choose for option lists), so there is no explicit alternative routing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_downloadsList captured browser downloadsA

List files captured from the browser into the tool's download folder (any game; the folder persists across sessions and MCP restarts). Use download_file(name|index, path) to copy one anywhere. Shows name, size, capture time and the absolute folder path.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMax files to list (default 20, newest first).

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does disclose two non-obvious traits: the folder is shared across 'any game' and persists across sessions and MCP restarts. It also enumerates what is returned. It stops short of stating ordering beyond what the schema implies, error behavior, or auth needs, keeping it out of the top band.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, zero filler, and the core purpose plus cross-game scope are front-loaded before the sibling pointer and return-value summary. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-parameter read tool with no output schema, the description covers scope, persistence, the follow-up action, and the return fields (name, size, capture time, absolute folder path). Only minor gaps remain, such as ordering nuances beyond the schema and failure behavior on an empty folder.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and the single 'limit' parameter is already documented with its default and 'newest first' ordering in the schema. The description adds nothing about limit or pagination, so the baseline 3 applies when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('List files captured from the browser into the tool's download folder') and immediately scopes it ('any game'), which distinguishes it from the per-game navigation tools in the sibling set. An agent can tell this apart from download_file, which is named as the companion action.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly routes the agent onward: 'Use download_file(name|index, path) to copy one anywhere,' which clarifies this is the discovery step before retrieval. It does not state when *not* to use it or any prerequisites, but the context is unambiguous for a read-only listing tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_gamesList open gamesA

List the currently open game sessions with their id, source, story title and step count.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden, and it does disclose the return fields (id, source, story title, step count), which helps an agent plan. However, it says nothing about read-only guarantees beyond the implicit 'List', or about ordering, pagination, or the empty-list case.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler; the verb and resource come first and the returned fields follow as useful detail. Every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter, read-only listing tool with no output schema, naming the returned fields is the key missing piece and the description supplies it. Only the absence of any note on ordering or empty results keeps it from being fully complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so per the rubric the baseline is 4. There are no arguments whose meaning the description would need to clarify, and it correctly adds no spurious parameter discussion.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description pairs a specific verb ('List') with a precise resource ('currently open game sessions') and even enumerates the returned fields, so the agent knows exactly what it retrieves. It does not explicitly contrast itself with siblings like get_journal or observe, but 'open game sessions' is distinct enough to avoid confusion.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to call this versus alternatives such as get_journal or observe, nor any prerequisites (e.g., call it to discover game ids before load_state). Usage is only weakly implied by the word 'open', which does not rise to explicit when-to-use guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

live_viewOpen a live view of the game pageA

Give the user eyes on the actual page the agent is controlling: starts a tiny local web server (once per MCP process) that streams JPEG screenshots of the real Playwright tab (~1/s) together with passage, step, engine state, recent actions and the passage text. Returns a URL like http://127.0.0.1:4571/v/game_abc — open it in any browser (works with headless games too). Frames are captured only while someone is watching. Set open:true to also launch the URL in the default browser. DISPLAY POLICY: this is the default — and usually the only — way to show a game. One view per game: do not also switch to a headed window or loop screenshot; repeated calls return the same URL and never launch a second browser tab.

ParametersJSON Schema
NameRequiredDescriptionDefault
openNoAlso open the URL in the default browser on this machine (default false). Only pass true when the user wants you to open it — at most once per game.
game_idYesSession id returned by open_game.

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does so well: it discloses that the server starts once per MCP process, frames are captured only while someone is watching, repeated calls return the same URL and never open a second tab, and it works with headless games. These are exactly the non-obvious traits an agent needs to avoid redundant or conflicting display actions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core action and the returned URL before the policy details, and every sentence carries substantive information. It is slightly long and repeats the open:true behavior that the schema already states, so not perfectly economical.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a two-parameter tool with no annotations and no output schema, the description covers what it does, what it returns, when to use it, exclusivity rules, and idempotency. Nothing an agent needs to invoke it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so both parameters are already documented. The description still adds operational meaning for 'open' ('also launch the URL in the default browser') and the constraint 'at most once per game', going slightly beyond the schema, though the core param semantics remain schema-driven.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: it starts a local web server that streams JPEG screenshots of the real Playwright tab along with passage/step/engine state. It distinguishes itself from the sibling 'screenshot' by being a continuous streaming view rather than a one-off capture, and names the return value (a URL).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The DISPLAY POLICY block explicitly states when to use it ('this is the default — and usually the only — way to show a game') and when not to ('do not also switch to a headed window or loop screenshot'). It also names the idempotent alternative behavior of repeated calls, so an agent has no ambiguity about routing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

load_stateLoad an in-session snapshotB

Restore a snapshot created by save_state and return the resulting observation (text; format:"json" for JSON).

ParametersJSON Schema
NameRequiredDescriptionDefault
nameNoSnapshot name (default "default").
formatNoOutput format: "text" (default, human-readable) or "json" (JSON string for programmatic use).
game_idYesSession id returned by open_game.

TDQS

B3.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It discloses the return value shape and the format flag, but a restore operation inherently overwrites the current in-session state — the description never says the existing state is discarded or whether loading is reversible, which is the single most important behavioral fact for this tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One compact sentence with the primary action front-loaded and no filler. The trailing parenthetical about format is slightly redundant with the schema but keeps the sentence readable.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With 100% schema coverage and only three parameters, the description is adequate for identifying the tool and its return type in the absence of an output schema. It is however missing the destructive/prerequisite context (state replacement, dependency on a prior save_state and a live game_id session) an agent needs to invoke it safely.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and each parameter (name default 'default', format enum, game_id from open_game) is fully documented in the schema. The description's mention of 'format:"json" for JSON' merely restates the enum, adding no syntax or defaulting detail beyond the schema, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a specific verb+resource ('Restore a snapshot') and explicitly ties it to the sibling that creates those snapshots ('created by save_state'), so the agent can distinguish it from save_state without opening a schema. It stops short of a 5 only because it doesn't say what 'restore' affects (session/game state vs. some sub-resource).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Referencing save_state implies the ordering (save first, then load), which is useful context. However, there is no explicit when-to-use guidance, no precondition statement, and no mention of alternatives such as restart or open_game for reconstituting state.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

observeObserve current game stateA

Read the current passage: text, numbered choices, input fields, dialog state and status text. Returns a formatted text observation (string); pass format:"json" for a JSON string. Inputs are paginated in windows of 40: when truncated, the header says e.g. "Inputs (41-80 of 132)" — call again with inputs_offset=80. Use find_ui(text) to jump to a specific control, get_variables for specific story variables, and since_last=true when polling to save tokens.

ParametersJSON Schema
NameRequiredDescriptionDefault
formatNoOutput format: "text" (default, human-readable) or "json" (JSON string for programmatic use).
game_idYesSession id returned by open_game.
since_lastNoIf true and nothing changed, return only a short "no change" note (default true; ignored in json format).
inputs_limitNoHow many inputs to list (default 40).
inputs_offsetNoSkip this many inputs before listing (default 0).
include_statusNoInclude status/caption text (default true).
max_text_charsNoCap passage text length (default 12000).
include_variablesNoEmbed the (truncated) story variables (default false; prefer get_variables for specific keys).

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and largely succeeds: it discloses the return type (formatted text, or a JSON string via format), the 40-item pagination window, the exact truncation header format ('Inputs (41-80 of 132)'), and the token-saving behavior of since_last. It stops short of stating rate limits or the full shape of the observation, but the read-only nature is clear.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the read purpose, then return format, then pagination mechanics, then alternatives. Dense and each sentence is informative, though the pagination example adds some length; overall efficient with no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For an 8-parameter read tool with no output schema, the description explains the return values (text vs. JSON string) and the pagination behavior an agent must act on. It is complete enough to call correctly, though it could say slightly more about the JSON variant's structure.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3 and the schema already documents every parameter. The description adds cross-parameter workflow meaning: how inputs_offset pairs with the truncation header, that since_last is ignored in json format, and the token trade-off of polling versus fetching variables.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Read the current passage') and enumerates exactly what is read: text, numbered choices, input fields, dialog state, status text. It distinguishes itself from find_ui ('jump to a specific control') but does not differentiate from closer siblings like inspect_ui or live_view, leaving some ambiguity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly routes to alternatives with conditions: find_ui(text) for a specific control, get_variables for specific story variables, and since_last=true when polling to save tokens. Both when-to-use and sibling selection are covered without inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

open_gameOpen a Twine gameA

Open a Twine / interactive-fiction HTML game and return the first observation (passage text, numbered choices, inputs, dialog). Accepts a local .html file, a game folder (index.html or a single html is picked), or an http(s) URL. Local games are served over 127.0.0.1 so saves work. Returns a formatted text observation (string); pass format:"json" for a JSON string. Inputs are listed in windows of 40 (see inputs_offset/inputs_limit on observe) and find_ui(text) locates any control by label.

ParametersJSON Schema
NameRequiredDescriptionDefault
seedNoSeed SugarCube PRNG (State.prng) for reproducible runs; game must use SugarCube randomness to be deterministic.
formatNoOutput format: "text" (default, human-readable) or "json" (JSON string for programmatic use).
sourceYesPath to an .html file or folder, or an http(s) URL.
headlessNoRun browser headless (default true). Set false ONLY when the user explicitly asks to watch a real browser window — and then do not also open a live_view for the same game.
block_trackersNoBlock analytics/tracker requests for a quiet session (default true).
wait_timeout_msNoHow long to wait for the game to settle after load (default 20000).
include_variablesNoEmbed the (truncated) story variables in the observation (default false; prefer get_variables for specific keys).

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does reasonably well: it discloses the return shape ('formatted text observation (string); pass format:"json" for a JSON string'), the fact that local games are served on 127.0.0.1 to enable saves, and that inputs are windowed in groups of 40. It omits session/lifecycle behavior (e.g., whether a prior game must be closed first) and error behavior on bad sources, so it is strong but not exhaustive.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose and return value are front-loaded in the first clause, which is the most important information. The trailing sentence about 40-item input windows and find_ui is useful routing but sits awkwardly at the end and slightly dilutes the core message about opening a game.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 7-parameter entry-point tool with no output schema, the description covers the essentials: accepted sources, the returned observation and its two formats, and where to look for UI controls. It leaves unstated what happens on a failed load or whether close_game must precede a re-open, which keeps it from being fully complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all seven parameters (seed, format, source, headless, block_trackers, wait_timeout_ms, include_variables). The description restates the format parameter and references inputs_offset/inputs_limit from observe, adding only marginal meaning beyond the structured fields. Baseline 3 is correct.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb + resource ('Open a Twine / interactive-fiction HTML game') and states the immediate consequence ('return the first observation'), which separates it cleanly from list_games, close_game, restart and observe. It also names the accepted source kinds, so the agent knows exactly what this tool consumes versus its siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives clear context for when the tool applies (local .html file, folder with index.html, or http(s) URL) and notes that local games are served over 127.0.0.1 so saves work. It does not explicitly exclude anything or say when to prefer a sibling (e.g., list_games first to find a source), so it falls short of the 5-level 'when-not/alternatives' bar.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

restartRestart the gameB

Restart the story from the beginning (optionally with a new PRNG seed). Returns the first observation (text; format:"json" for JSON).

ParametersJSON Schema
NameRequiredDescriptionDefault
seedNoNew SugarCube PRNG seed.
formatNoOutput format: "text" (default, human-readable) or "json" (JSON string for programmatic use).
game_idYesSession id returned by open_game.

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It helpfully discloses the return value ('Returns the first observation') and the output format choice, but it never states that restarting discards current progress/state or what happens to existing save points - a key behavioral trait for a mutation tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One front-loaded sentence with a compact parenthetical; nothing is wasted. The return-value clause is slightly crammed onto the end but still earns its place by covering a missing output schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 3-parameter tool with no output schema, the description adequately covers the return value and the seed/format options. However, with no annotations it leaves the destructive/state-resetting nature of the operation and its relationship to save_state/load_state unexplained.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents game_id, seed, and format in detail. The description's 'optionally with a new PRNG seed' only confirms seed's optionality and the format note repeats the schema, so baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Restart the story from the beginning'), which is clear and distinct from siblings like load_state or back by the 'from the beginning' scoping phrase. It stops short of explicitly naming the sibling it contrasts with, so it lands just under a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is only implied: an agent can infer you restart when you want to reset the story, and the 'from the beginning' wording hints at the distinction from load_state. There is no explicit when-to-use, when-not-to-use, or named alternative, so it does not reach a 4.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

save_stateSave an in-session snapshotA

Save the full game state under a name so you can branch: save -> try a path -> load_state -> try another path. Supported natively by SugarCube; other formats report unsupported.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameNoSnapshot name (default "default").
game_idYesSession id returned by open_game.

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description must carry behavioral burden. It usefully discloses format support: SugarCube is native and other formats 'report unsupported', which tells the agent about failure behavior. However it omits whether saving overwrites an existing name, where snapshots persist, and whether the operation is reversible beyond the implied load_state pairing.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two compact sentences, front-loaded with the core action and immediately followed by the branching rationale. No filler, and the format-support caveat is placed where it is most useful.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a two-parameter tool with no output schema, the description covers purpose, workflow, and a compatibility caveat, which is nearly everything needed to invoke it. It leaves minor gaps around overwrite behavior and persistence, but nothing that would cause a mis-call.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, with both name (default \"default\") and game_id documented inline, so the schema already does the heavy lifting. The description mentions saving 'under a name' but adds no syntax, naming-convention, or collision semantics beyond what the schema provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Save the full game state under a name') and distinguishes itself from the load_state sibling by pairing with it in a branching workflow. An agent can immediately tell what the tool produces and how it differs from related state tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The 'save -> try a path -> load_state -> try another path' sequence gives explicit usage context and names the companion tool. It also flags a boundary condition (only SugarCube is natively supported), though it does not exhaustively state when not to save or how this compares to restart/back.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

screenshotScreenshot the gameA

Take a PNG screenshot of the game viewport. Useful for canvas/image-driven games and visual QA. Pass path to save it to a file (returns the path instead of the image). One-shot: this is not a live view — to let the user watch the game, call live_view once instead of taking screenshots repeatedly.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathNoOptional output file path; when set, the PNG is written there and only the path is returned.
game_idYesSession id returned by open_game.

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations, so description carries the full burden. It discloses the one-shot nature (not a live view), the return behavior (image vs path when path is set), and output format (PNG). It doesn't cover auth, rate limits, or error conditions, but for a screenshot tool these are less critical and the key behavioral traits are well covered.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, front-loaded with the core action, then usage context, then the crucial distinction from live_view. No wasted words and the most important routing information is placed first.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 2-param read-only tool with full schema coverage and no output schema, the description provides all needed context: what it returns, how path changes behavior, and when to use it vs live_view. Missing only edge-case error behavior, which is minor.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents both parameters. The description adds the behavioral consequence of setting path (returns path instead of image), which is a meaningful addition beyond the schema's basic description. Baseline 3 is appropriate when schema is comprehensive, but the extra consequence is slightly helpful.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States the specific action (take a PNG screenshot) and target (game viewport), and distinguishes itself from the sibling live_view by naming it. An agent can tell exactly what this does without reading the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly states when to use it (canvas/image-driven games, visual QA) and when not to (not a live view; call live_view once to let the user watch). This directly answers the 'when vs alternatives' question with a named alternative.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

upload_fileUpload a file into the gameA

Upload a local file (mod .zip, exported .save, image) into an in the game. Provide trigger_text/trigger_ref for buttons that open a picker ("Load from File…", "Import"), or let it target the file input directly. If a hardcoded selector like #saves-import does not exist in the build, use find_ui("Load from File") and pass its ref as trigger_ref. Returns the new observation (text; format:"json" for JSON).

ParametersJSON Schema
NameRequiredDescriptionDefault
refNoRef of a file input from inspect_ui (alternative to selector).
pathYesAbsolute path of the file on the machine running this MCP server.
formatNoOutput format: "text" (default, human-readable) or "json" (JSON string for programmatic use).
game_idYesSession id returned by open_game.
selectorNoCSS selector of an <input type=file> (used when no trigger is given).
trigger_refNoRef of the trigger button (alternative to trigger_text).
trigger_textNoVisible text of the button that opens the file picker.
trigger_selectorNoCSS selector of the trigger button (e.g. "#saves-import").

TDQS

A4.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It discloses the return value ('Returns the new observation; format:"json" for JSON'), which is genuinely useful, but says nothing about permissions, what happens to existing game state on upload, failure behavior when no ref/selector/trigger resolves, or whether the upload triggers navigation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, no filler, and the core action is front-loaded in the first clause. The middle sentence is dense with three alternate targeting modes, but each clause carries information an agent needs.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For an 8-parameter tool with no annotations and no output schema, the description covers the primary flow, the parameter interplay, the fallback path, and the return format. Remaining gaps are error/failure handling and side effects on game state, which are secondary but not negligible.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% so the baseline is 3, but the description goes beyond the schema by explaining the relationship among the four targeting parameters (trigger_* for picker-opening buttons vs selector/ref for the input itself) and the ordered fallback strategy. That is real added meaning over the flat per-parameter descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Upload a local file ... into an <input type=file> in the game') and enumerates the file kinds involved (mod .zip, exported .save, image). This clearly separates it from the sibling download_file, the inverse operation, without needing to open either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives concrete routing between the trigger parameters and the direct-target path ('Provide trigger_text/trigger_ref for buttons that open a picker ... or let it target the file input directly'), plus an explicit fallback when a hardcoded selector is missing ('use find_ui("Load from File") and pass its ref as trigger_ref'). It does not, however, state when this tool should be preferred over other siblings like interact or click_ui for the same UI.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

waitWait for the gameB

Wait for time (ms), for a text to appear, and/or for the page to settle, then return the new observation. Useful for timed passages and animations. Returns text; pass format:"json" for a JSON string.

ParametersJSON Schema
NameRequiredDescriptionDefault
msNoMilliseconds to wait.
formatNoOutput format: "text" (default, human-readable) or "json" (JSON string for programmatic use).
game_idYesSession id returned by open_game.
for_textNoWait until this text appears anywhere on the page (up to timeout_ms).
timeout_msNoTimeout for for_text (default 15000).

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations provided, so description carries full burden. It discloses return format (text/json) and waiting conditions, but omits timeout behavior for ms wait, error handling if text never appears, and whether the observation is returned as a snapshot. Adequate but incomplete for a wait tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Single sentence stating what it waits for and returns, followed by a brief note on use case and format parameter. Slightly dense but front-loaded and efficient.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with 5 parameters and no output schema, description covers the core waiting and return format but lacks details on timeout interaction, error cases, and how the observation is structured. Sufficient for basic invocation but missing edge-case guidance.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and parameters have descriptions, so baseline 3 is appropriate. Description mentions ms, for_text, and format, but adds no syntax or default details beyond schema; timeout_ms and game_id are not mentioned in description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific action (wait) with three conditions (time, text appearance, page settle) and the return (new observation). Distinguishes from siblings like observe (immediate) and screenshot, though doesn't explicitly name an alternative.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Mentions 'useful for timed passages and animations' which implies usage context, but doesn't specify when to choose wait vs observe or other tools, nor when to avoid waiting.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 22 tool updatesv0.1.3
    • First observedback
    • First observedchoose
    • First observedclick_ui
    • First observedclose_game
    • First observeddownload_file
    • First observedfind_ui
    • First observedget_console_errors
    • First observedget_journal
    • First observedget_variables
    • First observedinspect_ui
    • First observedinteract
    • First observedlist_downloads
    • First observedlist_games
    • First observedlive_view
    • First observedload_state
    • First observedobserve
    • First observedopen_game
    • First observedrestart
    • First observedsave_state
    • First observedscreenshot
    • First observedupload_file
    • First observedwait

TDQS

A3.7/5.0

Scored across 22 tools

Disambiguation4/5

Most tools target distinct actions (observe, choose, save/load, screenshots, files), but there are ambiguous pairs: choose() explicitly answers dialog buttons while click_ui() claims to 'click dialogs', and find_ui vs inspect_ui both locate controls. These overlaps require careful description reading to pick correctly.

Naming Consistency4/5

All names are snake_case and follow a verb_noun pattern for resource-targeting tools (save_state, open_game, get_variables, list_downloads). The spread of bare verbs (observe, choose, wait, back, restart, find_ui, inspect_ui, click_ui) deviates slightly, though it reads as an intentional convention for acting on the current game state.

Tool Count3/5

22 tools is on the heavy side for a game-automation server. Several functions arguably overlap (choose/click_ui, find_ui/inspect_ui, screenshot/live_view), suggesting some consolidation is possible, though most tools do earn their place.

Completeness4/5

The surface spans the full play lifecycle: open/close, observe, choose/interact, wait, back/restart, save/load state, screenshots/live view, file upload/download, variables, console errors, journal, and UI inspection. Minor gaps exist (no direct story-variable setter, no passage search/jump), but these are workaroundable via interact and find_ui.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    D
    maintenance
    Enables AI assistants to control web browsers through Playwright automation, providing 50+ tools for navigation, interaction, testing, accessibility audits, and visual testing across Chromium, Firefox, and WebKit.
    17 npm
    MIT
  • A
    license
    Not graded
    quality
    C
    maintenance
    Provides browser automation capabilities for LLMs using Playwright, leveraging structured accessibility snapshots to interact with web pages without needing vision models. It enables tasks like web navigation, data extraction, and automated testing through a lightweight and deterministic toolset.
    7 npm
    Apache 2.0
  • A
    license
    Not graded
    quality
    C
    maintenance
    Enables AI assistants to play interactive fiction games (Glulx and Z-machine) through the Model Context Protocol, with automatic save/restore and optional journaling mode.
    MIT