Skip to main content
Glama
Redseb
by Redseb

run_playtest

Play the game headless from a script of steps to verify runtime behavior validators cannot check: doors, NPC dialogue, choice branches, invisible walls, or battle outcomes.

Instructions

Play the game headless from a script of steps and report what happened — the runtime check validators cannot do (does the door transfer, does the NPC say the right thing, does the choice branch, is there an invisible wall). One browser session runs the whole script; each step reports its outcome and the run stops at the first failing step. Steps: load {mapId,x,y,direction?,party?,level?,gold?,switches?,variables?,selfSwitches?,items?,equip?,encounters?} starts a fresh game there (random encounters off unless encounters:true); startEvent {eventId} starts a map event as if triggered and lets it run until it shows text or goes idle (waiting out any transfer — reported as transferredTo); advanceText {maxMs?} presses OK until the event is idle, STOPPING at an open choice list/battle — returns the map message lines shown and any open choices (text shown during a battle comes back separately as battleLines); choose {index} picks a 0-based choice; walk {direction, steps?} walks tile by tile, reporting where it ended (to); if a tile refused entry, stoppedAt (the player's tile) and blockedTile (the refused one); when a step fires a touch event (a door) it stops, waits out the transfer and reports transferredTo {mapId,x,y}, plus eventRunning/messageOpen if the event is still going; press {button, times?}; wait {ms}; autoBattle {troopId?, canEscape?, canLose?, maxMs? (default 60000)} fights (a started or new battle) on auto AI until it ends, returning the battle's message lines (troop events, victory text). Battles are fast-forwarded (20 engine frames per drawn frame: same logic, same odds, a 10-turn boss fight in seconds) unless the run sets realtime: true; screenshot {name?} saves a PNG; eval {script} evaluates a JS expression in the game page and returns its value (e.g. "$gameSwitches.value(3)"). Reported text reads as the message window shows it: \V[n]/\N[n]/\P[n]/\G expanded, control codes (\C[n], \I[n], ., | …) removed. Every step result carries ok; the response ends with finalState (scene, map, position, gold, party) and problems (console/page errors, HTTP 404s). Max 200 steps. Read-only: never writes the project. A run can take MINUTES (booting ~5-10 s, long cutscenes, realtime battles): clients should raise their request timeout, or send a progressToken with resetTimeoutOnProgress — the tool then sends a progress notification per step and every 5 s during long ones. Split long scripts into several runs if your client can do neither. Needs the optional dependency playwright-core and a Chromium from the Playwright cache (npx playwright install chromium-headless-shell) or the RPGMAKER_MCP_CHROMIUM env var.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
outNoDirectory for screenshot PNGs. Default: <os tmpdir>/rpgmaker-mz-mcp/renders.
stepsYesThe script, run in order.
inlineNoAlso return the PNG(s) as image content in the response (default false: paths only — read the file to view it).
realtimeNoPlay battles at real-time speed (default false: battles are fast-forwarded). Real time takes ~10-20x longer — raise autoBattle maxMs to match.

Schema Changelog

Changes observed during successful MCP inspections.

  1. Addedv1.4.0

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and delivers: read-only ('never writes the project'), run duration in minutes, progress-token behavior, the 200-step cap, stop-at-first-failure semantics, battle fast-forwarding, and the external playwright-core/Chromium dependency. This is unusually thorough behavioral disclosure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the purpose and the differentiator, then flows into step semantics and operational notes. It is long, but the density is justified by the 9-action step union and the minute-scale async behavior; almost every sentence carries actionable detail.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Covers what an agent needs to call it correctly: script semantics, timeouts/progress handling, dependency setup, output shape (ok per step, finalState, problems), and failure behavior. Nothing material is missing for a complex no-output-schema tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds meaningful semantics not obvious from the schema: what each step action does (startEvent runs until text/idle, advanceText stops at choice/battle and returns battleLines separately), how reported text is decoded (\V[n] expansion, control codes stripped), and the fast-forward vs realtime distinction.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('play the game headless from a script of steps') and immediately positions it against siblings: it is 'the runtime check validators cannot do'. An agent can distinguish it from validate_project and render_map without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly frames itself as the runtime complement to validators and gives conditions for use (checking door transfers, NPC text, choice branches, invisible walls). It also advises splitting long scripts and raising timeouts. It does not name a specific sibling alternative for edge cases, so not quite a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Deploy Server

Other Tools