rpgmaker-mcp
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@rpgmaker-mcprender Map001 as a PNG preview and list its events"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
rpgmaker-mcp
Give your AI agent eyes inside RPG Maker MZ. An MCP server that renders your maps to real images, edits them cell by cell while you watch, writes event logic in named commands instead of magic numbers, and then boots the actual game to prove it works.

永宁镇 · 青瓦水乡 (map #6, 44×32): a two-cell-thick city wall with a south
gate, three roof/wall material pairs, a stone plaza with statues and pillars, a
canal crossed by three bridges, a lotus pond and a market stall. Built
cell-by-cell through this server's tools; the image is render_map output at
scale 1, composited from the project's own tileset PNGs — autotiles, wall
shadows, z-order and event sprites included.

The same build as a live session: the agent narrates each tool batch on the
left; the observation console on the right shows map #6 at step 9 of 9 — nine
paint steps with their cell counts (1,408 + 429 + 335 + …), pause /
single-step / replay controls and a per-cell inspector.
Why this exists
RPG Maker MZ stores an entire game as JSON under data/. That is great for
tooling and terrible for agents: a map is a flat array of thousands of integers,
autotiles are shape numbers, and event scripts are numeric opcodes. Editing that
by text is a guessing game, and the agent never finds out it built a wall with no
door until a human opens the editor.
This server closes the loop:
data/Map006.json ──► paint / events / step editor ──► composited PNG preview
img/tilesets/* (your project's own tiles) (browser, 127.0.0.1)
data/System.json ──► event_* builders (named MZ codes) ──► real playtestEvery write is validated, every render is faithful to the engine's own tile rules, and every claim below is reproducible with a command in this repo.
Related MCP server: RPG Maker MZ MCP Server
What you get
1. The agent sees the map, not the numbers
render_mapcomposites your project's own tileset PNGs into a map image — autotile adjacency, wall shadows, z-order and events included. Crop a region, scale it, overlay a coordinate grid, event markers or the region layer.tileset_catalogreports which Tileset mode a map uses and which image file actually sits in each A1–A5 / B–E slot, so the agent never assumes a palette.tile_paletteshows what a sheet looks like with its tile ids;tile_infoprobes one id for its sheet, autotile role, the paired A4 wall-top/wall-side bases, and warns when a B–E tile is one corner of a 2×2 / 3×3 composite.inspect_cellreturns all six layers of one cell (4 visual + shadow + region).
2. Editing you can watch, one cell at a time
open_editorstarts a step session;putground,put_event,move_event,set_event_imageeach save and push an exact diff to the observer, which draws it and acknowledges. Pause, single-step, speed selection and replay of the last session are built into the console.Batch tools (
paint_tiles,place_building,stamp_region) do rectangles, building footprints with roof/wall pairing, and cross-map region copies.Mistakes are cheap: SHA-256 revision checks on every write, a cross-process writer lock, a backup before each change, and
edit_history/undo_map_edit.
3. Wall shadows that match the editor
Autotile walls cast shadows in MZ through layer 4 bit masks, and getting them
wrong produces the classic "striped wall" bug. paint_tiles runs an
editor-parity shadow reconcile by default (autoShadow): wall bodies stay
0, the first ground cell right of a wall gets 5, stale masks are reclaimed
when their wall is removed, hand-painted masks are left alone, and the result
reports shadowCells. Writes that would fight the reconcile are refused instead
of silently dropped.
4. Event logic as named commands — 46 of them
event_show_text, event_show_choices, event_battle, event_give_items,
event_if, event_switches, event_move_route, event_transfer_player,
event_screen_fade, event_set_weather, event_shop, event_play_se … each
one appends real MZ command codes to an event page in a single transaction, with
parameters validated against the engine's Game_Interpreter layout (including
the MV↔MZ differences: Play SE is 250, Play ME is 249, choices take five
parameters). Anything not covered falls through to event_raw_commands.
5. Playtest for real, without touching your project
playtest_startruns the full game in a local browser throughplugin/MZVisualBridge.js, injected via the test HTTP response — your project's plugin list is not modified. Thenruntime_control(start / move / interact / input / teleport / reload),runtime_capturefor screenshots,runtime_statusfor switches, variables and self-switches.native_playtest_startdoes the same in a private copy of the licensed NW.js runtime on Windows (~320 MB, copied once, verified, isolated config dir).The browser channel blocks external network and WebSocket traffic.
6. Nothing proprietary ships here
The repo contains only original code. No RPG Maker core scripts, tiles,
characters, music, fonts, NW.js binaries, demo projects or runtime tokens. The
renderer reads your installation at runtime; the server binds its observer and
playtest endpoints to 127.0.0.1 with a private token.
The 78 tools
Group | Tools |
Project |
|
Observation |
|
Map painting |
|
Events |
|
Event logic (46) |
|
Step editor |
|
Verify / undo |
|
Playtest |
|
Full parameter tables: docs/TOOLS.md.
Quick start
Requirements: Node.js 20+, an RPG Maker MZ project folder (the one with
data/, img/, js/), and a locally installed Edge / Chrome / Chromium. No
build step, no browser download.
git clone https://github.com/twrsm666/rpgmaker-mcp
cd rpgmaker-mcp
npm ci
node src/server.js --project "/absolute/path/to/your-mz-project" --engine "/absolute/path/to/RPG Maker MZ"--project is the game project (contains data/System.json), not the engine
install. If the project already has js/rmmz_core.js, design previews work
without --engine. Browser auto-detection covers Windows, common Linux paths
and macOS Chrome; override with --browser or RPG_MCP_BROWSER.
Flag | Effect |
| MZ project directory (required) |
| local MZ installation directory |
| Chromium-family executable to render with |
| observer port (default: random) |
| refuse map file writes |
| observer only, no stdio MCP |
| enable playtest channel and runtime tools |
MCP client configuration
Copy and adapt examples/mcp-config.example.json:
{
"mcpServers": {
"rpg-maker-mz": {
"command": "node",
"args": [
"/absolute/path/to/rpgmaker-mcp/src/server.js",
"--project", "/absolute/path/to/your-project",
"--engine", "/absolute/path/to/RPG Maker MZ",
"--live-bridge"
]
}
}
}project_info / preview_focus return the observer URL with its private token;
the same URL is printed to stderr (stdout carries only MCP JSON-RPC). Keep those
out of public repos.
Editing, in practice
Visual writes carry an expectedSheet assertion so a tile id can never land on
the wrong sheet (tileId: 0 clears a cell and takes no sheet):
{"mapId": 6, "expectedRevision": "<latest>", "rectangles": [
{"x": 5, "y": 20, "w": 30, "h": 2, "layer": 0, "tileId": 2336, "expectedSheet": "A1"}
]}A step session, one call per visible step:
const editor = await visualEditor(mcpClient, 6, { holdMs: 300 });
await editor.putground(10, 8, 0, 2816, "Meadow ground", "A2");
await editor.put_event({ x: 12, y: 8, name: "Guide", text: "Welcome.",
image: { characterName: "People1", characterIndex: 0 } });
await editor.move_event(1, 13, 8);
await editor.close();Coordinates are zero-based. Layers 0–3 are MZ's visual stack (not
"ground/interior/dungeon" categories); layer 4 is the shadow mask and layer 5 the
region id. presentation.status=rendered means the browser confirmed drawing
that revision; no_observer / pending_or_paused mean it did not, and saved
data is never rolled back.
Event logic stacks the same way — build the event, then append behaviour:
{"mapId": 6, "expectedRevision": "<latest>", "eventId": 7,
"condition": {"type": "gold", "amount": 500, "test": ">="},
"thenCommands": [{"code": 101, "parameters": ["", 0, 0, 2, ""]},
{"code": 401, "parameters": ["You are rich!"]}]}Gallery
The two images at the top of this page are from one real 0.5.0 session. For reference, this is what the previous TypeScript renderer produced — kept so the two lines can be compared:

Testing
No engine or copyrighted assets needed:
npm ci
npm run check # source-level invariants
npm test # node --test, 57 pass without an engine
npm run audit:releaseGitHub Actions runs check + test on Node 20/22, Windows and Linux. With a
licensed MZ install, the full suites unlock:
$env:RPG_MCP_ENGINE = "D:\Tools\RPG Maker MZ"
npm run demo -- --engine "$env:RPG_MCP_ENGINE"
npm run verify # stdio tools end to end
npm run verify:ui # observer UI
npm run verify:steps # step editor / pause / replay
npm run verify:runtime # browser playtest
npm run verify:nw # native NW.js playtestEverything runs against throwaway copies under .work/; screenshots and reports
land in verification/. Both are gitignored.
Known boundaries
The native MZ editor does not share memory. Save and close the editor before MCP writes; reopen afterwards.
The design canvas does not run plugins or event logic — lighting, custom drawing, real passability and battles are verified in the playtest.
There is no "upload a script, auto-clear the game, get a report" tool; the runtime input/move tools are primitives you orchestrate.
Step replay lives in the current server process only.
Encrypted assets are unsupported; validated baseline is MZ 1.8.x on Windows with local browser rendering.
Static passability analysis is approximate — playtest for truth.
One server per project; separate servers share only the on-disk write lock.
Documentation
docs/TOOLS.md— every tool, its inputs and its guarantees.docs/VALIDATION.md— what was measured on which build.docs/TROUBLESHOOTING.md— the 37-item field log of MZ automation traps (headless focus gating, autotile shapes, ★ passability, MV↔MZ opcode differences, NW.js exit codes…). Hard-won, all reproduced.docs/tool-defects.md— the honest ledger of known open defects.docs/legacy-0.4.2/— the previous TypeScript line: README, changelog, acceptance runbook, review response, release checklist.
License
MIT for the code in this repository only — see LICENSE and
THIRD_PARTY_NOTICES.md. Third-party dependencies and
all RPG Maker files keep their own licenses. This project is not affiliated with
or endorsed by RPG Maker / Gotcha Gotcha Games / Kadokawa.
Available Tools
65 toolsadd_commandsAppend event commandsA
Append raw {code, indent, parameters} commands to a page, keeping the terminating code-0 entry last. Use command_catalog first to learn the parameter layout of a code. The page's block structure is re-read after the append, and anything whose indent does not match the branch it sits under comes back in warnings.
| Name | Required | Description | Default |
|---|---|---|---|
| at | No | Insert before this list index instead of appending before the terminator | |
| mapId | Yes | ||
| eventId | Yes | ||
| commands | Yes | ||
| pageIndex | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the load and does reasonably well: it discloses the terminator-handling invariant (code-0 stays last), the post-append re-read of block structure, and the warnings-on-indent-mismatch return behavior. It still omits failure behavior for invalid codes and any permission/validation notes, so it is not fully transparent for a mutation tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, all substantive, with the core action front-loaded and the prerequisite and return behavior following in logical order. No filler or restatement of the title.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema and no annotations, so the description must cover everything; it explains the warnings return but omits what happens on malformed codes, whether writes are atomic, and how mapId/eventId/pageIndex scope the operation. Adequate for the happy path, incomplete for an agent handling edge cases.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 20%, so the description must compensate: it mirrors the required command fields and adds real meaning to `indent` via the branch-matching rule. But mapId, eventId, and pageIndex are unexplained (only `at` is documented in the schema), leaving several parameters undocumented in both places.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (append) and resource (raw {code, indent, parameters} commands to a page), which is precise and unambiguous. It does not name the obvious sibling set_commands or explain how append differs from replacing a page's command list, so sibling differentiation is left implicit.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives one actionable prerequisite (call command_catalog first to learn a code's parameter layout), which is genuinely useful. However, it never says when to choose this over set_commands or decode_commands, nor any conditions/exclusions for use.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
assert_in_gameRun a list of assertions in the running gameA
Hand over a list of {label, expression} and get a pass/fail report back, the way a test runner does it, instead of writing a throwaway script for every playtest. Each expression is evaluated in game scope and has to be truthy; give one a timeoutMs to poll it until it holds, which is what a dialog opening, a transfer landing or a battle ending needs. A failing assertion does not stop the rest unless stopOnFailure is set, and whatever the game logged while the list ran comes back with the result, so a failure arrives with its own reason attached. Needs Allow Eval on the plugin.
| Name | Required | Description | Default |
|---|---|---|---|
| pollMs | No | ||
| assertions | Yes | ||
| stopOnFailure | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only supply idempotentHint=false, so the description carries most of the burden and delivers: failure does not abort remaining assertions unless stopOnFailure is set, the game's log is returned alongside results, and it declares the 'Allow Eval' prerequisite. It omits polling cadence (pollMs default) and any throughput/rate limits.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core mechanism and then layered with behavior and prerequisites. It is a dense paragraph but every sentence adds information; it could be trimmed slightly but wastes little.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 3-parameter, no-output-schema tool it covers the essentials: what goes in, what comes back (pass/fail report plus logs), failure handling, and the Allow Eval prerequisite. Only pollMs behavior and the 60-item cap remain unexplained.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Top-level schema coverage is 0%, so the description must compensate, and it explains timeoutMs (poll until truthy) and stopOnFailure meaningfully. But pollMs is never mentioned in the description and the nested label/expression semantics are largely left to the schema's inline descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'Hand over a list of {label, expression} and get a pass/fail report back, the way a test runner does it.' This clearly distinguishes it from siblings like live_eval (single evaluation), live_wait, and validate_game. An agent can identify it as the batched in-game assertion tool without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives strong context: 'instead of writing a throwaway script for every playtest' and concrete scenarios ('a dialog opening, a transfer landing or a battle ending'). However, it does not explicitly name sibling alternatives such as live_wait or live_eval and when to prefer them, so it stops short of the top score.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
batchApply several tool calls or none of themADestructiveIdempotent
Run a list of {tool, args} calls in order as one transaction. If a step fails — or is rejected before it runs, because the arguments do not match what that tool declares — every file the batch wrote goes back to the bytes it had before, so a bad argument on step four never leaves a half-built map behind. Each step's arguments are validated against that tool's own schema first, so the whole list is checked for the obvious mistakes before anything is written. Reading tools may be included and their results come back per step. undo_writes, rollback_data and a nested batch are refused as steps: they move the write journal this transaction rolls back against. Steps that talked to the running game are named in the reply, because a file rollback cannot un-press a key — follow it with live_reload or a new game.
| Name | Required | Description | Default |
|---|---|---|---|
| steps | Yes | ||
| stopOnFailure | No | Stop at the first failing step (default), or run the rest and still roll everything back |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only supply idempotentHint/destructiveHint; the description carries the real behavioral load by disclosing the rollback guarantee ('every file the batch wrote goes back to the bytes it had before'), pre-flight schema validation of each step, the refusal list, and the critical limitation that game-side effects cannot be undone ('a file rollback cannot un-press a key') plus the live_reload remedy. That is exactly the context annotations cannot express.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core transaction behavior, and each subsequent sentence (rollback scope, validation, refusals, game-side caveat) adds distinct operational value. It is dense with em-dash clauses and could be tightened, but there is little genuine filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an orchestrating tool with no output schema, the description covers failure semantics, validation timing, excluded steps, and the shape of the reply ('results come back per step', 'steps that talked to the running game are named in the reply'). Only the stopOnFailure behavior and exact reply structure are left to the schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 50%, but the description compensates by explaining steps semantics well: ordering ('in order'), transaction scope, per-step validation against each tool's schema, and per-step results. It does not, however, explain what stopOnFailure actually changes (stop at first failure vs continue-and-roll-back), which the schema description carries alone.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb, resource and mechanism: 'Run a list of {tool, args} calls in order as one transaction.' Combined with the title 'Apply several tool calls or none of them', an agent immediately knows this is the atomic multi-call executor and can distinguish it from every single-purpose sibling (create_map, set_tiles, etc.).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives clear context on composition: reading tools may be included, results come back per step, and it explicitly excludes undo_writes, rollback_data and nested batch as steps with a stated reason ('they move the write journal this transaction rolls back against'). What is missing is an explicit statement of when to prefer batch over issuing the same calls individually.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
block_structureExplain event block structureB
Report the engine-verified rules for nested command lists: which codes take a deeper indent, which continuation codes repeat, and where a block ends. Given a page it also returns the same warnings the writers return — the places where this list's indents will not do what they look like they mean.
| Name | Required | Description | Default |
|---|---|---|---|
| mapId | No | ||
| eventId | No | ||
| pageIndex | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full disclosure burden. It usefully reveals that the output includes `warnings` matching what writers return and that the rules are 'engine-verified', which frames the tool as an authoritative read. However it says nothing about side effects, permissions, or whether the call is purely read-only (only implied by the verb 'report').
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the core capability and then the warning behavior. The em-dash clause in the second sentence is slightly convoluted but the text is tightly scoped with little waste.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and no annotations, the description must do more work, and it does explain the semantic content of the return (rules plus warnings). But it omits the shape of the output and leaves all three parameters unexplained, so an agent still lacks what it needs to call the tool confidently.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% and the three parameters (mapId, eventId, pageIndex) are entirely undocumented in the schema. The description's only nod to them is the vague 'Given a page', which loosely maps to pageIndex but leaves mapId and eventId meaning unexplained, so it does not compensate for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (report) and resource (engine-verified rules for nested command lists), and even samples the content (which codes indent deeper, which continuation codes repeat, where a block ends). It is clear on its own, though it does not explicitly contrast itself with sibling tools like command_catalog or decode_commands that also deal with command structure.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit when-to-use or when-not-to-use guidance, and no named alternative among the many command-structure siblings. The phrase 'Given a page it also returns warnings' hints that a page context is optional rather than a requirement, but the agent is left to infer the conditions for calling this.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
check_assetsFind asset references that do not resolveARead-only
Scan every image and audio name the project data mentions — database fields (character, face, battler, title1/title2, battlebacks, tileset names, animation effect), map parallax/bgm/bgs fields, and Play BGM/BGS/ME/SE, Show Picture and Play Movie commands inside events — and report the ones with no file on disk plus the ones whose spelling differs from the file only in case. This is the cheap way to find both before a playtest does: a missing image stops the engine's game loop and paints an error into a DOM panel that a canvas screenshot cannot see, and a missing audio file says nothing at all. Icons are not covered: MZ draws them by index from img/system/IconSet.png rather than by name.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnlyHint annotation, it discloses exactly what is scanned, that it reports two categories of finding (missing file, case-only mismatch), why the findings matter (missing image halts the game loop and paints into a DOM panel invisible to screenshots; missing audio is silent), and an explicit scope boundary for icons.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The purpose and scan scope are front-loaded in the first sentence, followed by the rationale and the icon exclusion. Every sentence carries information; there is no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only tool with no output schema, the description does state what is reported (unresolved references plus case-mismatch findings), which covers the substantive return content. It omits the result shape and the limit parameter, leaving minor gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the sole parameter (limit, 1-500) is never mentioned in the description, so it fails to compensate for the coverage gap. The agent must guess that limit caps the number of reported results.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb (scan/report) and resource (image and audio asset references), and enumerates the exact sources checked: database fields, map parallax/bgm/bgs, and event commands. An agent can immediately distinguish it from write-side siblings like import_asset.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives clear when-to-use context ('the cheap way to find both before a playtest does') and an explicit when-not exclusion ('Icons are not covered'). It stops short of naming a sibling tool as the alternative, so it is slightly below full routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
clear_eventsTake the events off a mapADestructiveIdempotent
Remove every event on a map in one transaction, or only the ones whose name or character sheet you name. This is the call that makes a rebuild script possible: without it, re-painting a map's event layer means one remove_event per event and a half-finished map if the script dies halfway, and a build that runs twice ends up with two of every NPC. dryRun answers with what would go, which is also the quickest way to see a map's cast. The map is rendered in the reply, so you can look at the empty room rather than trust the count.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | Only events with this exact name | |
| mapId | Yes | ||
| dryRun | No | ||
| eventId | No | Only this one event | |
| characterName | No | Only events drawn from this character sheet, e.g. "!Crystal" |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true and idempotentHint=true, so the safety profile is covered. The description adds non-annotated behavior: the whole removal happens 'in one transaction' (atomicity), dryRun returns a preview rather than mutating, and the map is rendered in the reply. It does not discuss undo/rollback paths despite undo_writes and rollback_data existing as siblings.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The core action is front-loaded in the first sentence, and each following sentence adds distinct information (rationale, dryRun, return rendering). The rebuild-script narrative is slightly padded, but nothing is repeated and no sentence is pure filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description usefully states that the map is rendered in the reply and that dryRun returns a preview, and it covers atomicity and filtering. Remaining gaps are edge cases: invalid mapId behavior, filter combination rules, and whether the clear is recoverable via the rollback/undo siblings.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 60%; mapId and dryRun have no schema descriptions at all. The description compensates by explaining dryRun ('answers with what would go') and by describing the filter semantics (by name, by character sheet, or all), which is more than the schema conveys. It still doesn't clarify how multiple filters combine (AND vs OR) or what eventId does when combined with name/characterName.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource with scope: 'Remove every event on a map in one transaction, or only the ones whose name or character sheet you name.' It is immediately distinguishable from the sibling remove_event (single-event) and find_events (query-only), without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly frames the use case against the alternative: without this call a rebuild means 'one remove_event per event and a half-finished map if the script dies halfway.' It also routes the inspection case to dryRun ('the quickest way to see a map's cast'), so both the do-this and do-that-instead paths are named.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
command_catalogLook up event command codesA
Search the command dictionary derived from the engine's Game_Interpreter. Returns each code's parameter names in order, so event commands can be built without guessing. Ask for exact codes to also get the engine body, which shows how each parameter is consumed (a name like operateValueArg1 alone does not say whether the amount is params[1] or params[2]).
| Name | Required | Description | Default |
|---|---|---|---|
| codes | No | Exact codes to fetch, e.g. [125] — note this is plural and takes an array | |
| query | No | Substring match against labels, method names and parameter names |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does notably well: it discloses that results are parameter names in order, that exact `codes` additionally returns the engine body, and why that body matters (disambiguating which params index is consumed). It does not state read-only nature explicitly, but read behavior is implied by 'Search/Returns'.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, front-loaded with the core purpose and then the reason the richer mode matters. The parenthetical example earns its place by clarifying the value of the engine body. No filler, though the final clause is slightly dense.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, so the description must convey return shape, and it does: parameter names in order, plus engine body when `codes` is supplied. For a 2-param, all-optional lookup tool this is sufficient to call correctly, with only minor gaps around error/empty-result behavior.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3, but the description adds genuine meaning beyond the schema: it explains that passing exact `codes` changes the return shape to include the engine body, and warns that a name like `operateValueArg1` alone does not disambiguate params[1] vs params[2]. That is semantics the schema does not convey.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'Search the command dictionary derived from the engine's Game_Interpreter.' An agent understands this is a lookup of event command codes. However, it never names its closest sibling (decode_commands) or otherwise differentiates itself, so the agent must infer the boundary.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description does distinguish the two modes ('Ask for exact `codes` to also get the engine body'), which is useful usage signal. But it gives no explicit when-to-use guidance relative to alternatives like decode_commands, leaving the selection to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
copy_eventCopy an event, wholeA
Duplicate an event onto another cell or another map: every page with its conditions, graphic, trigger and full command list, deep-cloned, plus the note. The copy gets a fresh id and its own self switches (those are keyed by event id at runtime), so a chest copy starts closed even if the original is open. Both copies are live afterwards, which matters for an autorun or parallel-process event. The source is not modified.
| Name | Required | Description | Default |
|---|---|---|---|
| x | Yes | ||
| y | Yes | ||
| name | No | Defaults to the source's name | |
| mapId | Yes | ||
| eventId | Yes | ||
| toMapId | No | Defaults to the map the source is on |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does substantial work: deep-cloning semantics, fresh id, self switches keyed by event id so a copied chest starts closed, both copies remaining live, and the source being unmodified. It stops short of covering permissions, failure modes, or what the call returns, so it is strong but not exhaustive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action, then layered with behavioral detail in efficiently packed sentences. Dense but each sentence adds real information; no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a six-parameter mutation tool with no annotations and no output schema, the description covers the behavioral contract well (clone depth, id/switch reset, liveness, source immutability). The gap is parameter-level detail for the four undocumented required fields.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 33% (name and toMapId documented; mapId, eventId, x, y bare). The description implies destination semantics via 'another cell or another map' but never maps its language onto x, y, or toMapId explicitly, so the four required parameters remain semantically thin.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource ('Duplicate an event onto another cell or another map') and immediately enumerates what is copied (pages, conditions, graphic, trigger, command list, note). This clearly distinguishes it from siblings like place_event or add_commands, which create or extend rather than deep-clone.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage context is implied rather than stated: the note that 'both copies are live afterwards, which matters for an autorun or parallel-process event' signals a relevant scenario, but there is no explicit when-to-use versus place_event, and no prerequisites or exclusions are named.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_database_entryAdd a database rowADestructive
Claim a row in a database table the way the editor does. MZ keeps these as 1-based arrays in which the row's id is its array index, and it never removes one: a fresh project ships unused slots as complete rows with an empty name, and those names are what keeps an id from being reused under a saved game. So this takes the first blank slot, and only grows the table when none is left — pass no id and it behaves like the editor's add button. The field shape comes from the blank slot itself, else from a blank sibling, else from copyFrom (or the last row, which the reply reports as basedOn because you have just copied a real entry's stats, icon and all). fields is then merged over that, shallowly, exactly as patch_database_entry does. System.json has no rows and MapInfos is the map tree, so those two refuse; use patch_database_entry and create_map. There is no delete: blanking a name is what removal means here, and it is a patch.
| Name | Required | Description | Default |
|---|---|---|---|
| id | No | Row to claim (default: the first blank slot, else a new one at the end) | |
| table | Yes | Actors, Classes, Skills, Items, Weapons, Armors, Enemies, Troops, States, Animations, Tilesets or CommonEvents | |
| fields | No | Fields to set on the new row, e.g. {name, iconIndex, description} | |
| copyFrom | No | Copy this existing row's fields as the starting point |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With destructiveHint=true and idempotentHint=false already declared, the description goes well beyond them: it explains the 1-based array model, that rows are never removed, that blank-slot names prevent id reuse, that it only grows the table when no blank slot exists, and that `fields` merges shallowly. This is rich behavioral context an agent cannot get from the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core concept and no filler sentences, but the middle is a long run-on covering fallback logic and merge semantics that could be tightened. Informative and structured, yet slightly dense for a single-tool description.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex, mutation-heavy tool with no output schema, it covers the schema model, the default/growth path, the copy/merge behavior, what the reply reports (basedOn), and the refusal cases with alternatives. Nothing an agent needs to invoke it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds meaning: it clarifies the default-id path (first blank slot, else append), the `copyFrom`-vs-last-row fallback, and that the merge is shallow 'exactly as patch_database_entry does'. It adds value beyond the schema, though most parameter facts are already documented there.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Claim a row in a database table') and immediately frames it against the editor's add button. It explicitly distinguishes itself from siblings patch_database_entry and create_map, including the exact tables each handles. An agent can tell it apart without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit when/when-not: omit `id` to mimic the editor's add button, System.json and MapInfos tables refuse and must use patch_database_entry / create_map, and removal is a patch, not a delete. Alternatives are named with the conditions that select them.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_mapCreate a mapC
Create an empty map file plus its MapInfos entry.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | ||
| width | Yes | ||
| height | Yes | ||
| parentId | No | ||
| tilesetId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It does disclose one meaningful side effect: the tool also creates a MapInfos entry, which is useful. However, it omits whether an existing file is overwritten, what defaults are applied, permission requirements, or what happens on name collision.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler, and the key artifact (empty map file) is stated immediately. It is efficient, though arguably terse to the point of under-specification given the parameter gaps.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a file-creating tool with no annotations, no output schema, and five undocumented parameters, the description is not complete enough. It states the outcome but leaves callers without the usage and parameter detail they need to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Five parameters with 0% schema description coverage and the description adds nothing about any of them. It never explains width/height bounds, the role of tilesetId (required), parentId semantics, or whether name is optional/auto-generated, so an agent gets no help from the description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb (Create) and resource (an empty map file plus its MapInfos entry), which is more informative than the title alone. It clearly distinguishes from deletion/list siblings, but it never explains how it differs from the closely related 'make_map' sibling, leaving that ambiguity unresolved.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use create_map versus alternatives such as make_map, or whether it should precede set_map_properties/set_tiles. No prerequisites, ordering, or exclusions are given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
decode_commandsDecode an event's commandsC
Render a page's command list as readable lines using the engine-derived command dictionary.
| Name | Required | Description | Default |
|---|---|---|---|
| mapId | Yes | ||
| eventId | Yes | ||
| pageIndex | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. 'Render ... as readable lines' implies a non-mutating read, but nothing states that explicitly, nor does it mention permissions, whether the operation fails for missing maps/events, or what the returned lines actually contain.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler, which is the right size for a simple tool. It is efficient rather than padded, though the terseness is part of what leaves the gaps above.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, no annotations, and zero parameter documentation, the description should be doing much more work. It never explains the shape of the decoded output, how mapId/eventId are sourced, or how this differs from the adjacent command tools.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for three parameters, so the description must compensate and largely does not: mapId and eventId are never mentioned, and pageIndex is only obliquely gestured at by 'a page's command list'. The agent gets no indication of what these identifiers reference or what happens when pageIndex is omitted.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a concrete verb ('Render ... as readable lines') and a specific resource ('a page's command list'), so the agent knows this is a read/transform operation rather than a mutation. It does not, however, differentiate itself from look-alike siblings such as command_catalog or show_text, so the boundary between them must be inferred.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use statement, no prerequisite, and no named alternative. The agent must guess that this is the tool for inspecting an already-written command list rather than for authoring (add_commands/set_commands) or enumerating available commands (command_catalog).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
delete_mapDelete a map and its entry in the treeADestructive
Remove data/MapNNN.json and clear its MapInfos slot — the other half of create_map, and the call that lets a build script tidy the prototypes it made on the way to the version it wanted. It refuses while some event on another map still transfers the player there, because that is a door into a map that is gone, and it lists those events so the caller can decide; force deletes anyway and still lists them. Both files are backed up first, so undo_writes puts the map file and its tree entry back together.
| Name | Required | Description | Default |
|---|---|---|---|
| force | No | Delete even though events lead here; the reply names them | |
| mapId | Yes | ||
| dryRun | No | Report what would go, and what points here, without writing |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Goes well beyond the destructiveHint annotation: it specifies that both files are backed up first, that undo_writes restores them together, that the operation refuses while inbound transfer events exist, and that those events are listed so the caller can decide. This is exactly the kind of pre-conditions and reversibility detail annotations cannot convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three dense sentences front-loaded with the destructive action and its scope, then refusal semantics, then backup/undo. Information-rich with minimal waste, though the final sentence packs backup and undo into one clause that could be split.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive, non-idempotent mutation with no output schema and no annotations covering pre-conditions, the description supplies the missing pieces: reversibility via backups/undo_writes, the inbound-transfer refusal, the force override, and the dry-run probe. An agent has everything needed to call it safely.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 67%: `force` and `dryRun` are documented in the schema, but `mapId` is not. The description's discussion of force/refusal semantics adds behavioral meaning beyond the terse schema text, partially compensating for the undocumented mapId, though it doesn't state mapId's integer/tree-slot meaning explicitly.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (remove/delete) and two concrete resources (`data/MapNNN.json` and its `MapInfos` tree slot), and explicitly positions itself as the inverse of `create_map`, a named sibling. An agent can distinguish it from `clear_events`, `make_map`, or `remove_event` without inspecting schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Describes a real usage context (build scripts tidying prototypes) and the refusal condition with transfer events, plus the `force` escape hatch and the `dryRun` report-without-writing path. It doesn't enumerate when to prefer a sibling like rollback_data, but the operational context is clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
describe_tilesSay what a tile id isARead-only
The tile dictionary, read out of the project the way the engine reads it. For each id: which slot it belongs to (A1–A4 are autotile shape groups of 48, A5/B–E are fixed), which image that slot is bound to, which layer MZ draws it on, and what its entry in the tileset's 8192-entry flags array means — open or blocked from which of the four directions, whether 0x10 is set so those direction bits do nothing at all (much of the stock tileset ships that way), ladder, bush, counter, damage floor, the vehicle bits, the terrain tag — plus how many cells of a map actually use it. Ask four ways: mapId for every tile a map really paints, which is the fastest way to learn an unfamiliar tileset; tiles for the ids in your hand; slot to walk one slot one base pattern at a time; or neither, for the slot table alone. contactSheet also draws the answer: a labelled sheet, one 3x3 block per id rendered by the same code path as a real map, with the id and the engine's own passability verdict printed under each block, because "which id is the water" is a question a picture settles faster than a table. The sheet is drawn on a scratch map that this call creates and deletes again; no tileset row and no real map is touched.
| Name | Required | Description | Default |
|---|---|---|---|
| slot | No | Walk one slot, one base pattern at a time | |
| limit | No | Cap the list (default 120) | |
| mapId | No | Describe every tile this map paints, with cell counts | |
| tiles | No | Describe exactly these ids | |
| tilesetId | No | Which tileset's flags to read (default: the map's own) | |
| contactSheet | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With readOnlyHint already declaring safety, the description adds substantial context beyond annotations: it explains the flags array semantics (directional passability, the 0x10 quirk in stock tilesets, ladder/bush/counter/vehicle/terrain bits) and discloses that contactSheet creates and deletes a scratch map, touching no tileset row or real map. That side-effect disclosure is exactly what annotations cannot convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with what the tool reads, then the query modes, then the contact sheet caveat. Every sentence carries information, but the dense em-dash prose runs long and could be split for scannability.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description carries the full burden of describing return values — and it does, enumerating every field of an entry and what the contact sheet draws. For a 6-parameter tool with a nested object, nothing an agent needs to call it correctly appears to be missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is already 83%, so the baseline is 3, but the description adds real meaning: it explains the slot grouping (A1–A4 autotile shape groups of 48, A5/B–E fixed), what mapId vs tiles vs slot each retrieve, and that tilesetId defaults to the map's own tileset. Minor gaps remain, such as the interaction between limit and the various modes.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource — reading the tile dictionary as the engine reads it — and enumerates exactly what each entry carries (slot, bound image, layer, flags semantics, usage counts). It is clearly distinguishable from siblings like tileset_slots, inspect_cell and render_map.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly lays out four query modes and when each applies: mapId 'for every tile a map really paints, which is the fastest way to learn an unfamiliar tileset', tiles for ids in hand, slot to walk a slot pattern-by-pattern, or neither for the slot table alone. contactSheet is also positioned as the answer for visual questions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
enable_pluginSwitch a plugin file onADestructiveIdempotent
Put a plugin that is already sitting in js/plugins/ into js/plugins.js — the same edit the editor's plugin manager makes, and the step patch_plugin cannot do because it only reaches entries that are already listed. The parameters come from the file's own @param/@default header, so nothing runs on a value the plugin never declares, and parameters overrides single keys on top of that. Re-running it on a plugin that is already listed switches it on and merges your parameters instead of adding a second entry. A game already running read js/plugins.js at boot, so it needs a restart — and an editor with the plugin manager open overwrites this file on save.
| Name | Required | Description | Default |
|---|---|---|---|
| file | No | Set when the file is not <name>.js. A trailing .js is accepted and stripped, so either spelling works | |
| name | Yes | The plugin's name, which is the file name in js/plugins without .js | |
| status | No | Pass false to switch it off without removing the entry | |
| parameters | No | Values to put on top of the header's @default ones |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare idempotentHint=true and destructiveHint=true, so the bar is lower; the description adds real context — parameter defaults come from the file's own header, re-running merges rather than duplicating, and a running game needs a restart while an open editor manager will overwrite the file on save. It stops short of stating exactly what this tool overwrites in plugins.js, which keeps it from a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Dense but front-loaded with the core action, and each clause (patch_plugin contrast, header sourcing, idempotent merge, restart, editor race) carries distinct information. Slightly packed, but no sentence is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a file-mutating tool with no output schema and partial annotations, the description covers the important operational caveats (restart, editor overwrite, merge semantics). It doesn't describe the return value, which matters somewhat given no output schema exists, but the essential call-correctly information is present.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds meaning the schema doesn't: parameters are seeded from the file's `@param`/`@default` header and only override single keys on top of that, and it explains why nothing runs on undeclared values. That goes beyond the schema's field-level text.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a precise verb and target (put a plugin file into `js/plugins.js`) and immediately differentiates itself from the sibling `patch_plugin`, which 'only reaches entries that are already listed.' An agent can tell exactly what this does and how it differs from neighbors without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It names the alternative (`patch_plugin`) and the condition that selects this tool over it, plus clarifies re-run behavior on an already-listed plugin. That covers when-to-use and the alternative in one pass.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
find_eventsSearch events across mapsB
Find events by name, by command code used, or by graphic. Answers questions like 'which events call common event 12' without reading every map by hand.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | Substring of the event name | |
| limit | No | ||
| mapId | No | Restrict the search to one map | |
| commandCode | No | ||
| characterName | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden, and it discloses little. It never states that this is a read-only sweep, whether it searches all maps by default, how the limit is applied or what the results look like. 'Without reading every map by hand' is the only hint at breadth, and it is implicit.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, zero filler, and the search axes are front-loaded before the illustrative use case. The example sentence earns its place by clarifying commandCode semantics rather than padding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With 5 optional parameters, no annotations and no output schema, an agent still lacks the default limit, the behavior when no filter is supplied, and the result shape. The description covers the core search concept adequately but leaves these operational gaps for a tool whose structured fields do not fill them.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 40% (name and mapId documented; commandCode, characterName and limit bare), so the description must compensate. It does partially: the common-event example gives real meaning to the undocumented commandCode parameter, and 'by name' aligns with the name filter, but characterName and limit remain unexplained in both places.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (find) and resource (events) plus the search axes (name, command code, graphic), which cleanly separates it from the mutating siblings place_event, remove_event, copy_event and clear_events. The claimed 'by graphic' axis has no corresponding parameter in the schema, which slightly muddies the precision of the stated scope.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The example question ('which events call common event 12') implies when the tool is useful, but there is no explicit when-to-use, when-not-to-use, or named alternative (e.g. get_map / inspect_cell for a single-map read). Usage must be inferred from the illustrative case.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
fix_projectMake a project that did not come from the editor bootableADestructiveIdempotent
Write the System.json keys the engine reads without a fallback and this project does not have. An installation ships data/newdata as a third way to start a project besides the editor, and that template has no advanced.windowOpacity: the game stops on the title screen's first window with a stack that names clamp and nothing about the missing key, and a project copied from the template is otherwise indistinguishable. The values come out of your own installation's template — they belong to the engine, so this package does not carry a copy of them; advanced.windowOpacity is the one key the template itself lacks, and 192 is what the editor writes. It also puts switches and variables back into the array shape the engine indexes, and writes only what is genuinely absent, so running it on a project the editor made changes nothing. It does not invent maps, events or a start position: validate_game reports those and set_startup writes them. dryRun answers with the plan alone.
| Name | Required | Description | Default |
|---|---|---|---|
| dryRun | No | Report what would be written and change nothing | |
| template | No | A System.json to take the missing values from. Default: <install>/data/newdata/data/System.json, found through RMMZ_CORESCRIPT_ROOT |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare idempotentHint=true and destructiveHint=true, so the safety profile is framed. The description adds real value on top: only genuinely absent keys are written, values come from the installation template rather than the package, and dryRun returns a plan. It explains provenance and the write-avoidance behavior the annotations do not.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is an undivided run-on narrative of several hundred words for a two-parameter tool. Key facts (dryRun behavior, sibling alternatives) are buried mid-paragraph rather than front-loaded, forcing the agent to extract signal from dense prose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a repair tool with no output schema, the description covers what gets written, what does not (maps, events, start position), the no-op case, and what dryRun returns. It is functionally complete, though the density makes it harder to consume than its content warrants.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so both parameters are documented, setting the baseline at 3. The description still adds meaning: dryRun 'answers with the plan alone' and template values come from your installation, clarifying where the default and the written values originate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb+resource: it writes the System.json engine keys a project is missing without a fallback. It also implicitly separates itself from siblings by noting it does not invent maps/events/start position. The core purpose is recoverable but buried in a long narrative, so it takes effort to parse.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It names alternatives and the selecting condition: validate_game reports missing maps/events/start position and set_startup writes them, while this tool handles only engine keys. It also clarifies that running on an editor-made project changes nothing, which helps an agent decide when it is a no-op.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_mapInspect a mapA
Read a map's structure: dimensions, tileset, every event with position, graphic, trigger and page conditions. Optionally dump a rectangular slice of one tile layer as ids.
| Name | Required | Description | Default |
|---|---|---|---|
| mapId | Yes | Map id, e.g. 1 for Map001.json | |
| layerDump | No | Return a grid slice of one layer (0-3 tiles, 4 shadow bits, 5 region id) | |
| includeCommands | No | Include decoded event command lists (default false) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden. It discloses what is returned and notes that layer dumps and command decoding are optional, which is useful, but it does not mention side effects, permission needs, or output size concerns for a read-only inspection tool. Adequate but not rich.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two compact sentences, front-loading the core read behavior and then the optional slice dump. No wasted words, though separating the optional behavior more clearly would slightly improve structure.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, so the description appropriately enumerates returned contents. It covers the main read path and the optional dump mode, but omits any guidance on when the optional modes are preferable or how large outputs behave, leaving a small completeness gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents mapId, layerDump, and includeCommands fully. The description adds only a brief restatement of the optional layer-dump behavior and does not extend parameter meaning beyond the schema, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (read) and resource (a map's structure) and enumerates the exact contents returned: dimensions, tileset, events with position, graphic, trigger, and page conditions. This clearly distinguishes it from siblings like inspect_cell or render_map, which focus on narrower or visual targets.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies a read/inspection context but does not state when to use this tool versus alternatives such as inspect_cell, find_events, or render_map. There are no stated prerequisites, exclusions, or named alternatives, so usage is only weakly guided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
import_assetPut a file where the engine will load itADestructive
MZ has no importer and no per-asset metadata: the folder and the base name are the whole registration. This copies a file into one of the folders the engine's own loaders read (img/characters, img/faces, img/sv_actors, img/sv_enemies, img/enemies, img/battlebacks1/2, img/titles1/2, img/tilesets, img/parallaxes, img/pictures, img/animations, img/system, audio/bgm, audio/bgs, audio/me, audio/se, movies), checks the extension belongs there, refuses a name that differs from an existing file only by case (the filesystem folds those together and the second one would never load), and writes through the same backup journal as every other write, so undo_writes takes it back — including deleting a file that did not exist before. The reply gives the value to put in the data field (the name without extension) and which fields already name it, which is the half of check_assets that was missing a file. Overwriting needs overwrite: true; the bytes being replaced stay recoverable in the backup list.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | Asset name without extension (default: the source file's own base name) | |
| folder | Yes | Target folder, with or without its img/ or audio/ prefix | |
| source | Yes | File to bring in: an absolute path, or project-relative when moving between folders | |
| overwrite | No | Replace the file that is already there |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare destructiveHint=true and idempotentHint=false; the description adds substantial behavior beyond that — extension validation, case-folding name refusal, writing through the backup journal, undo_writes recovering it including deleting a previously-nonexistent file, and recoverability of overwritten bytes. This is exactly the context annotations can't carry.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
It's a dense single paragraph, but front-loaded with the core fact (no importer, folder+name is registration) before the mechanics. The long folder enumeration is bulky yet necessary for correctness; otherwise there's little waste.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description still explains the reply (the value for the data field and which fields already name the asset), covers validation and rollback behavior, and needs nothing more for an agent to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the schema already documents all four parameters (baseline 3). The description adds meaning on top: `name` defaults to the source base name and overwriting requires `overwrite: true`, clarifying the mutation semantics of two params.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource (copies a file into folders the engine's loaders read) and enumerates the exact target folders, so the agent knows precisely what happens. It also distinguishes itself from siblings by referencing check_assets and undo_writes and framing itself as the missing half of asset checking.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives clear conditions: the folder and base name are the whole registration, overwriting needs `overwrite: true`, and the write is reversible via undo_writes. It doesn't explicitly state when NOT to use it or name a competing import path, but the context is strong enough to route the agent correctly.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
inspect_cellInspect one map cellA
Read every tile layer, shadow bits, region id, terrain tag and the event standing on a single cell.
| Name | Required | Description | Default |
|---|---|---|---|
| x | Yes | ||
| y | Yes | ||
| mapId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden, and it does establish this as a pure read and discloses the complete returned payload, which is the key behavioral fact. It omits edge-case behavior (out-of-bounds coordinates, nonexistent mapId) and any error or permission notes.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One front-loaded sentence with the verb first and zero filler. Every clause earns its place by naming a returned data category.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, and the description usefully enumerates the return contents, which is appropriate. However, for a 3-parameter tool with zero schema descriptions it leaves parameter meaning and usage context unaddressed, so it is only minimally complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% and none of the three required parameters have descriptions. 'Single cell' weakly implies x/y identify the cell, but the coordinate space, bounds, and the meaning of mapId are never explained, so the description does not compensate for the schema gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb ('Read') plus a precise resource ('a single cell') and an enumeration of exactly what is retrieved: tile layers, shadow bits, region id, terrain tag, and the event on the cell. This distinguishes it clearly from bulk siblings like get_map, render_map, and describe_tiles.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No statement of when to use this versus alternatives such as get_map (whole map) or describe_tiles (tileset metadata), and no preconditions. The agent must infer from the name alone that this is the single-cell inspection tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
link_mapsJoin two maps with doorsADestructiveIdempotent
Write both ends of a connection in one call: a door event on each map that transfers the player to the other, with a graphic (!Door1, a tile, or nothing for an invisible exit), the door sound, the fade, and the direction they arrive facing. a and b are the doorway cells, and a player coming through one does not stop on it: they arrive at the first open cell beside the far door (above it, then right, left, below) facing away, and the reply's landed says which cell that was. Pass land on an end to choose it yourself — including the doorway cell itself, which is what the transfer onto the threshold used to do. A door walled in on all four sides has nowhere beside it, so it keeps the doorway and the answer carries a warning. The reason it exists is the check it does first: a player-touch door only fires when the player can stand on its cell, and a destination cell that is blocked means the player arrives stuck inside a wall with no way back — both of which a transferred playtest surfaces minutes later as 'nothing happens'. requires writes a locked pair of pages: page 0 refuses, page 1 (conditioned on the switch or item) does the transfer. twoWay: false leaves the far end alone — no door is written there, so its coordinates are the destination itself and nothing shifts. Both maps come back rendered.
| Name | Required | Description | Default |
|---|---|---|---|
| a | Yes | One end of the link | |
| b | Yes | The other end | |
| fade | No | 0 none, 1 black (default), 2 white | |
| sound | No | SE played on the way through; default Door1, pass null for silence | |
| twoWay | No | Default true: an event at each end, each leading to the other | |
| graphic | No | A sheet name (!Door1), {tileId: n}, or "none" for an invisible exit. Default !Door1 | |
| replace | No | ||
| requires | No | Gate both doors behind this condition. A condition is one of {switch: id, value?}, {selfSwitch: "A".."D", value?}, {variable: id, op?: "=="|">="|"<="|">"|"<"|"!=", value: n or {variable: id}} (op defaults to ">="), {item: id}, {weapon: id}, {armor: id}, {gold: n} (at least n), {actor: id, state?|skill?}, {button: "ok", pressed?}, or {code: 111, parameters: [...]} to write the engine's own shape. | |
| direction | No | Which way the player faces on arrival; by default they face away from the door they came through, and 0 keeps the direction they were walking (the engine ignores a 0 and leaves the facing alone) | |
| lockedMessage | No | What the locked door says instead of transferring |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only say idempotent/destructive; the description adds the real behavioral payload: the landing algorithm (above, right, left, below, facing away), the `landed` reply field, the warning when a door is walled in on all four sides, the `requires` two-page locked structure, and that both maps are re-rendered. This is well beyond what the annotations convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core write, then descends into landing rules and rationale. Dense but every sentence carries information; the rationale paragraph is slightly long but earns its place by explaining the blocked-destination failure mode.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With 10 parameters, nested objects, and no output schema, the description carries the return-value burden and does so by explaining the reply's `landed` field and warning behavior. Remaining gap is the undocumented `replace` parameter and the absence of any overwrite/destructive warning despite destructiveHint=true.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 90%, so baseline is 3, but the description meaningfully extends it: `a`/`b` as doorway cells, `land` override including the doorway cell itself, `direction` arrival-facing, `sound`, `graphic`, and `twoWay` semantics. Only `replace` goes unexplained, which keeps it from a 5.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Opens with a precise verb+resource+scope: 'Write both ends of a connection in one call: a door event on each map that transfers the player to the other.' An agent can distinguish this from place_event, set_event_page, or map_connectivity without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explains the rationale for existing (the blocked-destination check that a live playtest surfaces late) and when to reach for `requires` (locked pairs) or `twoWay: false` (leave the far end alone). It gives clear context but never explicitly names an alternative sibling tool for the non-linked cases.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_backupsList write backupsA
Every write copies the previous file into .rpgmaker-mcp/backups. List them for a data file, or pass a project-relative path like js/plugins.js for one of the few files outside data/ that the tools also write.
| Name | Required | Description | Default |
|---|---|---|---|
| table | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, and it does disclose a genuinely useful behavioral fact: every write produces a backup copy in .rpgmaker-mcp/backups, including a few files outside data/. It does not say what a listed entry looks like, ordering, or whether backups can be restored, which is a notable gap for a no-annotation tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, no filler, and the mechanism/location is front-loaded before the scoping instructions. Every clause carries information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple optional-parameter list tool with no output schema and no annotations, the description covers origin, location, and scoping. It still leaves the return shape and the relationship to sibling restore/undo tools unstated, which is a meaningful gap given the absence of structured metadata.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate. It does: the lone parameter is named 'table' but the description reveals it also accepts a project-relative path such as js/plugins.js, and that it is optional (omitting it lists all). This adds meaning the bare schema string type does not convey.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (list) and resource (write backups), and explains the mechanism that creates them (every write copies the previous file into .rpgmaker-mcp/backups). This is enough to separate it from siblings like write_history, rollback_data, and undo_writes, though it never names them explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explains how to scope the listing (a data file, or a project-relative path like js/plugins.js), which implies when to use it. However, it gives no explicit guidance on when to prefer this over rollback_data, undo_writes, or write_history, so the agent must infer the boundary.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_mapsList mapsA
List every map with its size, tileset, event count and encounter settings.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. 'List' strongly implies a read-only, whole-collection operation and it discloses the returned fields, but it says nothing about pagination, ordering, or size/rate limits for a potentially large collection.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler. The verb and scope lead, and the returned fields follow compactly. Every word earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, no annotations, and no parameters, the description is the sole carrier of contract information, and listing the exact returned fields covers the key need. It could still clarify ordering or scale for large projects, but nothing essential is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so there is no parameter semantics to convey; the baseline for a parameterless tool is 4. The description appropriately focuses on the output fields instead.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('List') and resource ('every map'), and enumerates the returned fields (size, tileset, event count, encounter settings). This clearly distinguishes it as an enumeration tool versus the singular get_map sibling, though it does not name that sibling explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'List every map' implies the usage context (enumerating all maps), but there is no explicit when-to-use guidance, no mention of alternatives like get_map for a single map, and no prerequisites. Usage is only implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_pluginsList configured pluginsARead-only
Read js/plugins.js and report every entry in load order: enabled or not, its description, and the parameter values the editor stored. Each entry is cross-checked against the @param blocks in the plugin's own file, so undeclared names a key the plugin never mentions (a typo, or a leftover from an older version) and unset lists declared parameters that are running on their @default. fileExists: false means the plugin is switched on but its script is not in js/plugins/, which stops the game at boot.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
readOnlyHint=true already covers safety, and the description adds substantial context beyond it: it defines the cross-check behavior against @param blocks, the meaning of 'undeclared', 'unset' (running on @default), and 'fileExists: false' (plugin enabled but script missing, which halts the game at boot). That is real behavioral disclosure an agent could not infer from annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two dense sentences, front-loaded with the primary action and then the diagnostic semantics. Every clause earns its place, though the second sentence is long enough that it could be split for scannability.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description must carry the return contract, and it does: it enumerates the reported fields and explains the edge-case values (undeclared, unset, fileExists:false) including the boot-failure consequence. Nothing an agent needs to interpret the result is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so the baseline is 4. The description correctly implies a no-argument, whole-file scan and doesn't invent parameters that don't exist.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource: 'Read js/plugins.js and report every entry in load order' with the exact fields surfaced (enabled, description, stored params). This distinguishes it cleanly from read_plugin_source (single file) and enable_plugin/patch_plugin (mutations) without the reader opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by the scope (report every entry in the plugins file), but the description never names an alternative or a when-not condition, e.g. 'to inspect one plugin's source use read_plugin_source' or 'to change a setting use patch_plugin'. Adequate but leaves routing to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
live_diagnosticsRead the running game's errors and consoleARead-only
Return what the live game has logged: uncaught exceptions, entries from the engine's own error screen, failed image/audio/data loads, and console.warn/error output. The plugin keeps a rolling 200-entry buffer and merges repeats, so a per-frame throw arrives once with a repeat count. Pass the cursor from a previous reply as since to fetch only what is new, or clear before an action so whatever appears afterwards was caused by it, or full to get the recorded call stack with each entry. This keeps answering after the game has crashed (SceneManager.stop() freezes the render loop; stopped reports it), which is exactly when the file layer cannot tell you what went wrong.
| Name | Required | Description | Default |
|---|---|---|---|
| full | No | Include each error's call stack, up to 24 frames. Off by default to keep the reply small. | |
| clear | No | Empty the buffer after reading it | |
| limit | No | ||
| since | No | Only return entries newer than this cursor (default 0) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare readOnlyHint; the description adds the rolling 200-entry buffer with repeat-merge semantics, call-stack retrieval behavior, and the critical fact that it keeps answering after SceneManager.stop() freezes the render loop, with `stopped` reporting that state. This is rich behavior beyond the annotation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads what gets returned, then parameter recipes, then the crash-resilience fact. Dense but every clause earns its place. The parenthetical about SceneManager.stop() is slightly compressed, but overall well-structured.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read tool with no output schema and rich annotations-light context, the description covers return content, dedup behavior, cursor semantics, and post-crash resilience. An agent has everything it needs to invoke and interpret results correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 75% (full, clear, since documented; limit is not). The description explains the intent of `since` and `clear` with concrete workflows and reinforces `full`, adding operational meaning beyond the schema. `limit` remains unexplained in both places, keeping this from a 5.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (return logged output) and resource (live game logs), and enumerates exactly what's captured: uncaught exceptions, error screen entries, failed loads, console.warn/error. This distinguishes it from live_status, validate_game, and check_assets.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit operational recipes: pass cursor as `since` for incremental reads, use `clear` before an action to attribute new errors, use `full` for stacks. Closes with a 'when to reach for this' statement — after a crash, when the file layer can't help — which is a strong usage cue.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
live_dialogRead and answer what the running game is showingA
The dialog layer of a live game, driven by intent instead of counted key presses. read says what the game is waiting for and what it is showing: the message text, the choice options with each one's enabled state and where the cursor is, the number pad's digits, and whether the window is actually able to take a key right now. answer picks an option by index or types a number, dismiss presses on until the game stops waiting, and cancel takes the cancel branch. Every one of them ends by reading the game again, so the reply tells you what came next.
Why this is a tool and not three live_key calls: $gameMessage.isChoice() turns true the moment a choice is queued, while Window_ChoiceList is still fading in, and the engine only moves the cursor when Window_Selectable.isCursorMovable() holds — which needs the window open and active. A cursor key sent in those frames is dropped, and the answer comes back as the first option whatever you meant. Measured on a real playtest: it cost a run. So answer waits for the window to be able to take keys, presses one edge at a time, and re-reads the cursor after every press instead of assuming it moved; the reply says how many presses it took and refuses an option the game has switched off, because that press does nothing and an empty wait looks like a bug.
| Name | Required | Description | Default |
|---|---|---|---|
| index | No | answer: which option, 0-based — the engine's own order, same as `make_choice_scene` options | |
| limit | No | dismiss/cancel: how many presses to try before giving up (default 12) | |
| action | No | Default read | |
| number | No | answer: the value to type into a number-input window | |
| waitMs | No | How long to wait for the window to be takeable (default 8000) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With only idempotentHint=false and destructiveHint=false in annotations, the description carries the burden and delivers: it discloses that answer waits for the window to be takeable, presses one edge at a time, re-reads the cursor after each press, reports the press count, and refuses disabled options. It omits what a failure/timeout reply looks like and the return format, so it falls just short of fully exhaustive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the purpose and the action list, then a rationale paragraph justifying the design. Everything is relevant, but the second paragraph is dense and longer than strictly needed to route an agent, so it is efficient rather than maximal.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, and the description compensates by detailing what read returns (message text, options with enabled state and cursor position, number-pad digits, window takeability) and what every action's reply contains. Nothing needed to call it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds real meaning by tying parameters to action branches (index/number for answer, presses for dismiss/cancel) and confirming 0-based engine ordering. It does not mention limit or waitMs, leaving those entirely to the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb set (read/answer/dismiss/cancel) over a specific resource (the live game's dialog layer) and enumerates exactly what each action does. It also explicitly distinguishes itself from the sibling live_key, so an agent can tell them apart without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Names the alternative ('three live_key calls') and explains precisely when to prefer this tool: the choice-queued/cursor-unmovable race where raw cursor keys are dropped. It also routes between its own four actions (read to observe, answer/dismiss/cancel to act).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
live_evalEvaluate in the running gameA
Run a JavaScript expression inside the live game and return its value. Requires the plugin's Allow Eval parameter, a configured RMMZ_LIVE_TOKEN, and a running playtest. Anything the engine exposes works: $gameVariables.value(3), $gamePlayer._x, $gameSwitches.value(12). An expression that returns a promise is waited for and the settled value comes back, so one call can watch a frame counter, a transfer or a scene change instead of polling for it. The game gives a promise 20000ms to settle and the call waits for the same 20000ms by default, so a promise that needs 15s is answered, not cut off; raise timeoutMs to 60000 for a longer wait, but the plugin still abandons the promise itself at 20s. Two engine behaviors bite here, so they are worth knowing before a result looks wrong. Calling an API that moves the character directly ($gamePlayer.moveStraight(6)) reaches some of what a key does and not the rest: the party step count and the touch triggers happen inside the step itself, but the encounter counter only ticks in Game_Player.updateNonmoving, which is a frame or two later, so a fast scripted walk can outrun it — live_move holds the arrow instead and gets all of it. And $gamePlayer.performTransfer() run by hand rebuilds the map from the previous map's $dataMap, because loading the new file is Scene_Map's job: the player then stands "on" the new map with the old map's events running under them. reserveTransfer(...) and wait is the way.
| Name | Required | Description | Default |
|---|---|---|---|
| timeoutMs | No | How long to wait (default 20000, the same ceiling the plugin gives a promise) | |
| expression | Yes | A single expression; wrap statements in (() => { ... })() |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden and does so: it documents the promise-settling model, the 20000ms plugin ceiling versus the timeoutMs wait, and two concrete engine pitfalls (encounter counter lag on scripted walk; performTransfer using the stale $dataMap). This is far beyond what the schema provides.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the one-line purpose, then prerequisites, then the timeout semantics, then the pitfalls. It is long, but for an eval tool the two engine-behavior paragraphs earn their place by preventing misread results; only minor phrasing ('the way') is loose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema and no annotations, yet the description explains what comes back (the settled value of a promise), the timeout behavior, prerequisites, and the main failure modes. Nothing an agent needs to invoke this correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds real meaning: it explains the 20000ms default and the 60000 ceiling for timeoutMs, and notes that the plugin still abandons the promise at 20s regardless. It also implies expression is a single expression, reinforcing the schema's IIFE guidance.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'Run a JavaScript expression inside the live game and return its value.' This unambiguously distinguishes it from siblings like live_move, live_key, and assert_in_game, which act on the game rather than evaluating arbitrary expressions.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly states prerequisites (Allow Eval parameter, RMMZ_LIVE_TOKEN, running playtest) and routes the agent to alternatives with conditions: use live_move instead of a scripted moveStraight, and reserveTransfer+wait instead of performTransfer. This is when-to-use and when-not guidance grounded in named alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
live_keySend a key to the running gameA
Press a game key in the live game (Ok=13/32, Cancel=27, Shift=16, arrow keys 37-40) to advance dialogs or drive input. Each press is held for a counted number of engine frames and re-asserted on every one of them, and the call answers once the game has run them all, so a slow or headless frame rate cannot swallow it and neither can the browser taking focus. The reply carries the measured frame rate it was sized with and whether the player's cell actually changed. To walk somewhere, prefer live_move: one press is one step only if the player happens to be standing still when it lands. Needs the RMMZLiveBridge plugin, not Allow Eval.
| Name | Required | Description | Default |
|---|---|---|---|
| pulses | No | How many short presses to send (default 3) | |
| keyCode | Yes | ||
| holdFrames | No | Engine frames each press is held for (default 2) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does most of it well: it explains the frame-hold and re-assert mechanism, that the call blocks until the game has consumed all frames, that slow/headless frame rates or browser focus loss cannot swallow the input, and that the reply carries the measured frame rate and whether the player's cell changed. It does not warn that a key press can trigger irreversible in-game actions (menu confirmations, selections), which is the main behavioral risk for a live-input tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action and key table, and every sentence carries information. The third sentence is dense and slightly overloaded (frame rate, focus loss, reply contents in one breath), but nothing is padding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with no annotations and no output schema, the description covers purpose, alternative routing, dependency, blocking semantics and reply contents, which is nearly everything an agent needs. It is silent on failure modes such as the plugin being absent or no live session running.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 67%; keyCode has no schema description, and the description compensates with a concrete mapping (Ok=13/32, Cancel=27, Shift=16, arrows 37-40). It also gives meaning to the frame-count behavior behind holdFrames and pulses, though it never names those parameters directly or reconciles them against the documented defaults.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Press a game key in the live game') and immediately differentiates from the sibling that could be confused with it by naming live_move and the condition that selects it. The parenthetical key-code table makes the exact resource concrete rather than generic.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says when to use it (advance dialogs or drive input) and when to use the alternative instead ('To walk somewhere, prefer live_move: one press is one step only if the player happens to be standing still when it lands'). It also states the prerequisite ('Needs the RMMZLiveBridge plugin, not Allow Eval'), which is exactly the kind of routing guidance an agent needs.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
live_moveWalk the player the way a key doesA
Hold an arrow key until the player has arrived cells cells away, and answer with where they really got to and what stopped them. Use this instead of live_key whenever the intent is go-there rather than press-this: a direction press only moves the player on a frame when they are standing still, so a key press is not one step, and this waits for the steps themselves. Because it drives the engine's own input path, the things that only happen when a person walks happen here too - the party's step count, the encounter roll, an event set to player-touch. A wall, a dialog, a running event or a scene change ends the walk early and says which. Needs the RMMZLiveBridge plugin, not Allow Eval.
| Name | Required | Description | Default |
|---|---|---|---|
| cells | No | How many cells to walk (default 1) | |
| direction | Yes | Numpad direction: 2 down, 4 left, 6 right, 8 up | |
| timeoutMs | No | Give up after this long (default 8000) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations exist, so the description carries the full burden and does so: it discloses that it drives the engine's own input path, the side effects that follow (step count, encounter roll, player-touch events), the four conditions that end the walk early, and the plugin prerequisite (RMMZLiveBridge, not Allow Eval). This is exactly the behavioral context annotations would otherwise provide.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core behavior and return shape before the rationale, and most sentences earn their place. The live_key comparison and side-effect sentences are valuable but dense; a slightly tighter phrasing would lose nothing.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description still tells the agent what comes back ('where they really got to and what stopped them'), what interrupts the walk, and what environment is required. Nothing needed to invoke it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and every parameter is documented in-schema (cells, numpad direction, timeoutMs), so the baseline is 3. The description reinforces cells semantics but adds no format detail beyond the schema, which is appropriate when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Opens with a concrete verb+resource behavior ('Hold an arrow key until the player has arrived `cells` cells away') and states what it answers with. It explicitly contrasts itself with live_key, so an agent can distinguish it from its closest sibling without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit routing rule: 'Use this instead of live_key whenever the intent is go-there rather than press-this,' and justifies it with the frame-timing constraint on key presses. This is a when-to-use plus a named alternative and the condition that selects it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
live_pauseFreeze or resume the running gameAIdempotent
Stop the engine's update loop without leaving the scene, the way Unity's pause works: SceneManager.updateMain is skipped, so nothing moves, no event interpreter advances and no input is sampled, while the last frame stays on screen and live_screenshot / live_eval / live_diagnostics keep answering. Keys sent with live_key while paused are dropped, because Input is only sampled inside the update loop. Resume with paused: false.
| Name | Required | Description | Default |
|---|---|---|---|
| paused | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare idempotentHint, so the description carries the burden and does so richly: it discloses exactly what halts (SceneManager.updateMain, event interpreter, input sampling), what persists (last frame, screenshot/eval/diagnostics), and the non-obvious side effect that live_key input is dropped while paused.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The behavioral consequence is front-loaded in the first sentence and every subsequent clause adds distinct, load-bearing information. No filler or redundancy despite its length.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with no output schema and minimal annotations, the description fully covers the state change, its scope, its side effects, and the resume path. An agent has everything needed to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% with a single boolean, so the description must compensate, and it does: 'Resume with paused: false' pins down the polarity of the parameter. It could add a touch more (e.g. default state), but the core semantics are covered.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource: stop/resume the engine's update loop. It explicitly contrasts with the surrounding live_* family by naming what still works (live_screenshot, live_eval, live_diagnostics) and what is dropped (live_key), so an agent can place it precisely among siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear context for when the effect applies (update loop skipped, input unsampled) and how to reverse it (paused: false). It stops short of explicitly routing to alternatives like live_step for frame advancement, so it is clear context without full when-not/alternative guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
live_reloadHot-reload the map the game is standing onA
Make the running playtest re-read the current map file and rebuild itself around it: tiles, autotiles, events and map size, with the player left where they are. This is the author-then-look loop — after set_tiles / place_event / set_event_page, call it instead of booting a new game and transferring. It returns once the reloaded map is actually live, so the next live_eval or live_screenshot sees the change. Event interpreters restart from the top and database tables (System, Actors, Items, Tilesets) are not re-read, so a new game is still the answer for those. Does not need Allow Eval.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries full burden. It discloses critical behavioral traits: the player stays in place, event interpreters restart from the top, database tables are not re-read, the call blocks until the reloaded map is live, and it does not require Allow Eval. This is unusually rich behavioral context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, front-loaded with the core action and effect. Every sentence adds distinct value: the action and scope, the workflow context and alternative, and the caveats (what changes propagate, what doesn't, and the Eval requirement). No waste.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter mutation tool with no annotations or output schema, the description covers what changes, what doesn't, the blocking behavior, and the prerequisites. An agent has everything needed to invoke it correctly in the right sequence.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Zero parameters, so baseline is 4. No parameter semantics needed. The description doesn't waste space on parameters because there are none.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('re-read the current map file and rebuild itself') and enumerates exactly what is affected: tiles, autotiles, events and map size. Distinguishes itself from siblings by naming set_tiles, place_event, and set_event_page as the tools that precede it.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly defines the workflow context ('author-then-look loop') and when to call it versus alternatives ('call it instead of booting a new game and transferring'). Also specifies when NOT to use it: database table changes (System, Actors, Items, Tilesets) require a new game.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
live_screenshotSee the frame the game is drawingARead-only
Grab the running game's own composited frame as a PNG: message windows, face graphics, fonts, weather, character sprites, tile blending. render_map shows what the data says; this shows what the player sees. It re-runs the engine's render pass (Graphics._app.render()) and reads the canvas in the same task, so it works without MZ enabling preserveDrawingBuffer. Video playback is a separate DOM element and is not part of the frame. Needs a running playtest.
| Name | Required | Description | Default |
|---|---|---|---|
| saveTo | No | Also write the PNG to this absolute path |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With readOnlyHint=true already covering the safety profile, the description adds substantial implementation context: it re-runs Graphics._app.render() and reads the canvas in the same task, which is why it works without preserveDrawingBuffer. It also discloses a known blind spot (DOM-based video is excluded) and the playtest dependency — all non-obvious traits not derivable from annotations or schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Five tight sentences, front-loaded with the core action, then differentiation, then mechanism, then caveats. Every sentence adds information; nothing is restated from the title or schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, but the description states the return artifact (a PNG of the composited frame), the mechanism, and the operational precondition, which is everything needed to call and interpret this tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Only one parameter (saveTo) at 100% schema description coverage, so the schema already defines its meaning and the description does not elaborate on it. Baseline 3 is appropriate when the schema carries the parameter explanation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource — 'Grab the running game's own composited frame as a PNG' — and enumerates exactly what is composited into it. It explicitly contrasts itself with the sibling render_map ('shows what the data says; this shows what the player sees'), so an agent can separate the two without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Names the alternative (render_map) and the condition that selects it, plus a hard precondition: 'Needs a running playtest.' It also carves out a scope exclusion (video playback is not part of the frame). It stops short of a full when/when-not matrix, but the routing signal is clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
live_sessionStart or stop a headless playtestADestructive
Own the whole loop from the tool surface: start serves the project on a loopback port, launches a windowless Chromium-family browser on it and waits for the RMMZLiveBridge plugin to report, so live_status / live_screenshot / assert_in_game have a game to answer from without anyone opening the editor. By default it then walks the title screen into a new game with the engine's own commandNewGame; a client that allows only sixty seconds per call should pass newGame: false and call boot afterwards, because a cold boot plus that walk can exceed sixty seconds. boot alone drives whatever the page is showing into a map, status reports what is running, reload reloads the page (what a freshly written plugin file needs), stop closes the browser and releases the port. The browser is the only process this tool starts, it is recorded in .rpgmaker-mcp/session.json, and stop kills that pid's tree only after confirming the recorded profile directory is still on its command line, so a recycled pid is never touched. It refuses to start a second browser while a game is already reporting, because two games on one bridge cannot be told apart. Reaching a map needs the plugin params Allow Eval and keepAwake: without the first the session comes up on the title screen and says so, without the second a headless page never has focus and MZ skips scene updates.
| Name | Required | Description | Default |
|---|---|---|---|
| force | No | start even though a game is already reporting to the bridge | |
| action | Yes | ||
| bootMs | No | How long the walk to the starting map may take (default 45000) | |
| browser | No | Absolute path to a browser executable; otherwise RMMZ_BROWSER, then Edge/Chrome/Chromium | |
| newGame | No | Walk to the starting map as part of this call (default true) | |
| gamePort | No | Loopback port for the project files (default 8080; moves on if busy) | |
| waitForBridgeMs | No | How long to wait for the plugin to report (default 45000) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only say openWorldHint=false and destructiveHint=true; the description adds far more: the browser is the sole process started, it is recorded in .rpgmaker-mcp/session.json, and `stop` kills only that pid's tree after confirming the recorded profile directory is still on its command line. It also discloses the Allow Eval / keepAwake prerequisites and their failure modes.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the primary `start` behavior and action list before the edge cases. It is long and dense for a single paragraph, but nearly every clause carries operational value, so the size is largely earned.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists and this is a stateful, destructive-capable session tool, yet the description covers start/stop lifecycle, prerequisites, session-file tracking, refusal behavior, and action semantics. An agent has enough to call it correctly without opening other docs.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 86%, so the schema already documents most parameters; the description still adds real meaning by explaining the newGame default walk, the 60-second call-budget interaction with boot, and why force exists. This goes beyond the schema text rather than merely restating it.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb set and resource: it 'owns the whole loop' and enumerates each action (start/boot/status/reload/stop) with what each does. It explicitly distinguishes itself from siblings by saying live_status / live_screenshot / assert_in_game need a game that this tool provides.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit when-to-use and alternatives: pass newGame:false and call `boot` afterwards under a 60s-per-call client, use `reload` for a freshly written plugin file, `boot` alone to drive to a map. It also states the refusal condition (won't start a second browser while a game is reporting).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
live_statusLive game statusA
Report what the running game last sent through the RMMZLiveBridge plugin: scene, map, player position, switches, variables and party. listening is about this process's own socket, and the call waits for the bind to report, so a port another copy of this server already holds reads as a listenError here rather than as health. Also explains how to enable the bridge.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does disclose meaningful behavior: the call waits for the bind to report, and a conflicting port surfaces as a listenError instead of a health signal. It also points at how to enable the bridge. It does not state what is returned when no game is running or the bridge is disabled, which is the main remaining gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose is front-loaded in the first sentence, with the socket caveat following. The second sentence is dense and slightly convoluted ("a port another copy of this server already holds reads as a listenError here rather than as health"), and the closing "Also explains how to enable the bridge" is meta-commentary rather than tool information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description must describe the return payload, and it does enumerate the reported fields (scene, map, position, switches, variables, party). Combined with the disclosed wait/listenError behavior and bridge-enabling pointer, an agent has enough to call it correctly, though the no-game/disabled-bridge case is unaddressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so the baseline is 4. The description's mention of `listening` refers to an output/status field, not an input, and it correctly does not invent parameter semantics that the empty schema does not support.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb ("Report") and resource ("what the running game last sent through the RMMZLiveBridge plugin") and enumerates the payload: scene, map, player position, switches, variables and party. This clearly separates it from action siblings like live_move or live_reload. It stops short of naming the closest alternatives (live_diagnostics, live_wait), so it does not fully differentiate by name.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It clarifies what `listening` means and how a port already held by another server copy manifests (a listenError rather than health), which is real usage-relevant context. However, it never states when to reach for this tool versus live_diagnostics or live_wait, so routing between siblings is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
live_stepAdvance the running game by exact framesA
Run N engine frames and stop again, for watching an animation cycle, a message window page by page, or a battle sequence frame by frame. Pauses first if the game is running, then calls SceneManager.updateMain exactly N times (1..600), so the reported advanced is the frame delta to compare against. A frame that throws is handed to the engine's own catchException, so the error shows up in live_diagnostics rather than killing the session silently.
| Name | Required | Description | Default |
|---|---|---|---|
| frames | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The only annotation is idempotentHint=false, so the description carries nearly the full burden and does so well: it discloses the pause-first behavior, the exact engine call invoked, the 1..600 bound, and that exceptions are routed to the engine's catchException so they surface in live_diagnostics rather than crashing the session. That is rich operational detail beyond anything in the structured fields.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three dense sentences with zero filler: purpose and use cases first, then mechanism, then error behavior. Each sentence adds information an agent needs, and the most decision-relevant content is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter live-control tool with no output schema, the description is nearly self-sufficient: it explains the mutation, the frame semantics, the `advanced` return delta, and failure handling. Remaining gaps are the default when frames is omitted and whether the session is left paused afterwards.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% and the single parameter has no description, so the description must compensate – and it does, explaining that frames is the N count passed to updateMain and bounding it at 1..600. It does not state what happens if frames is omitted, even though the parameter is optional (0 required), which is a small residual gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a precise verb and mechanism: run N engine frames, pause, then call SceneManager.updateMain exactly N times, and stop again. The use cases (animation cycle, message window paging, battle sequence) make the intent concrete. It does not explicitly distinguish itself from the sibling live_wait, which is the nearest alternative, so it stops short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear usage contexts – frame-by-frame watching of an animation, a message window page, or a battle sequence – which is more than implied guidance. However, it never states when NOT to use it or names an alternative such as live_wait or live_pause, leaving the agent to infer the frame-exact vs time/condition-wait split.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
live_waitWait for a game conditionB
Poll an expression in the running game until it is truthy, so a sequence (menu open, battle over, message finished) can be followed instead of guessed at.
| Name | Required | Description | Default |
|---|---|---|---|
| pollMs | No | ||
| timeoutMs | No | ||
| expression | Yes | Polled about every 250ms until truthy or timeout |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It discloses the core behavior (repeated polling until truthy) and that the tool blocks a sequence, but says nothing about what happens on timeout, whether errors are thrown or falsy returned, or default polling/timeout values.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single dense sentence with the core mechanism front-loaded and the rationale trailing. No filler, though the parenthetical examples could be tightened.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema, no annotations, and two of three parameters undocumented. For a blocking poll tool an agent needs to know the return value on success vs. timeout and the default/max wait, none of which the description supplies.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 33%: only 'expression' is documented (and its 250ms polling note is in the schema, not the description). pollMs and timeoutMs are completely undocumented in both the schema and the description, and the description does not compensate with defaults, units, or bounds.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: polling an expression in the running game until truthy. It is clearly distinct from live_eval (evaluate once) and assert_in_game by the polling-until-true semantics, though it never names those siblings explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'so a sequence (menu open, battle over, message finished) can be followed instead of guessed at' gives concrete usage scenarios, which is better than nothing. However, it names no alternatives (e.g., live_eval or assert_in_game) and states no when-not conditions, so routing is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
make_battleAuthor a fight in one callADestructiveIdempotent
Both rows a battle needs — the Enemies row and the Troops row that holds it — from one ask, plus the region that rolls it if you want it met by walking. create_database_entry will write either row, but it has to be handed the whole thing, and the shapes here are the ones that fail quietly: a drop is {kind, dataId, denominator} where kind 0 is an empty slot and denominator: 3 means one chance in three (30 would be read as one in 30, not 30%); a trait is {code, dataId, value} — 11 element rate, 12 debuff rate, 14 state resist, 21 param, 22 hit and evasion, 31 attack element, 32 attack state, 63 collapse type — where the key is value, unlike an item effect, which uses value1/value2; and an action naming a skill that has no Skills row is a battle that throws on the turn the foe decides to act. So names are resolved against the project's own tables, what is not there is refused, and what was written comes back in words. Idempotent by name: a build script that runs twice leaves one foe, not one foe per run.
| Name | Required | Description | Default |
|---|---|---|---|
| foe | No | ||
| zone | No | Roll this troop in a region — `make_encounter_zone` for this one group | |
| troop | No | ||
| dryRun | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare idempotentHint and destructiveHint; the description reinforces idempotency by name ('leaves one foe, not one foe per run') and adds real behavioral detail beyond the annotations — names resolved against project tables, missing rows refused, and the specific failure mode where an action naming a skill with no Skills row 'throws on the turn the foe decides to act.' It does not explicitly warn about the destructive overwrite of existing rows, so not a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The purpose is front-loaded, but the prose is dense and parenthetical, with long run-on sentences carrying embedded code enumerations. Each clause is informative, but the trait-code list and the item-effect aside make it heavier than needed for a description whose job is selection and invocation.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex nested tool with no output schema, the description covers the two-row output scope, the tricky encodings, resolution/refusal behavior, and idempotency, and briefly notes output comes back 'in words.' Combined with a fairly rich schema, this is nearly complete, with the main gap being the destructive-overwrite implications of rewriting an existing row.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With schema coverage at only 25%, the description compensates by documenting the failure-prone encodings the schema doesn't: the drop shape `{kind, dataId, denominator}` and its denominator semantics, the trait `{code, dataId, value}` codes (11/12/14/21/22/31/32/63), and the value-vs-value1/value2 distinction. This is meaningful semantics beyond the schema, though it references a 'trait' shape not obviously present as a top-level property.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb+resource ('Author a fight') and precisely what gets created: 'Both rows a battle needs — the Enemies row and the Troops row that holds it — from one ask.' It explicitly distinguishes itself from the sibling create_database_entry, which 'has to be handed the whole thing.' An agent can tell what this tool produces versus alternatives without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It names the alternative (create_database_entry) and the condition selecting this tool over it, and hints at the zone/region use case ('the region that rolls it if you want it met by walking'). It also frames the idempotency benefit for 'a build script that runs twice.' It stops short of explicit when-not-to-use exclusions, so a 4 rather than 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
make_chestAdd a chest that pays out onceADestructiveIdempotent
Place a chest: the closed graphic, the line it says, what it gives, and the opened page that replaces it afterwards. MZ has no chest command, so this writes the two pages the engine does read — page 0 hands out the contents and sets self switch A, page 1 is conditioned on self switch A and shows the open chest. Because a self switch is keyed to the event id, copy_event of this chest gives a second chest that is closed again. requires gates the payout with a Conditional Branch rather than a page, so a locked chest stays openable once the key condition arrives instead of silently becoming an unlocked one. The contents are checked before anything is written: an id that is a blank slot, or gold that would make this a tax, fails the call.
| Name | Required | Description | Default |
|---|---|---|---|
| x | Yes | ||
| y | Yes | ||
| name | No | ||
| mapId | Yes | ||
| sound | No | SE played on opening; default Chest1 | |
| graphic | No | Defaults to the !Chest sheet every MZ project ships, closed at index 0 and open at index 1 | |
| message | No | Said as it is opened, e.g. "Inside is a lantern and 40 gold." | |
| replace | No | ||
| contents | Yes | What the chest gives when opened | |
| requires | No | What the player must have or have done first, e.g. {switch: 4} or {item: 3}. A condition is one of {switch: id, value?}, {selfSwitch: "A".."D", value?}, {variable: id, op?: "=="|">="|"<="|">"|"<"|"!=", value: n or {variable: id}} (op defaults to ">="), {item: id}, {weapon: id}, {armor: id}, {gold: n} (at least n), {actor: id, state?|skill?}, {button: "ok", pressed?}, or {code: 111, parameters: [...]} to write the engine's own shape. | |
| lockedMessage | No | Said when `requires` is not met; default "It is locked." |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only carry idempotentHint and destructiveHint; the description goes well beyond them, disclosing the two-page self-switch architecture, that `requires` is implemented as a Conditional Branch so a locked chest stays openable, and that validation runs before any write (blank slot id or gold-as-tax fails the call). This is rich, non-obvious behavioral context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action, then details; four dense sentences where each carries real information. Slightly heavy on mechanism (self switch, copy_event) that could be trimmed, but nothing is redundant.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an 11-param tool with nested objects and no output schema, it covers what gets written, the payout-once behavior, validation failures, and the requires-gating semantics. It omits any statement of what the call returns and does not address the `replace` flag, leaving minor gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 55% and several properties (x, y, mapId, name, replace) lack descriptions, but the prose compensates by explaining the semantics of `requires`/`lockedMessage` interaction and the `contents` validation, adding meaning beyond the schema for the most complex nested params.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Opens with a concrete verb+resource: 'Place a chest' and enumerates the exact parts it writes (closed graphic, message line, contents, replaced open page). An agent can distinguish it immediately from siblings like place_event, make_npc, or make_shop, which handle different event types.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explains that MZ has no chest command so this tool is the intended path, and clarifies the copy_event interaction (a copied chest is closed again because self switches key off event id). It never states exclusions or names a sibling alternative for edge cases, so it's strong context without explicit when-not-to-use routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
make_choice_sceneAuthor an interactive beat in one callADestructiveIdempotent
Write an event's whole script from a step list: dialog, choice branches (each option can carry its own when gate), if over switches, variables, items, gold and buttons, loop with break, battle with win/escape/lose branches, shop, transfers, audio, screen effects, gameOver. This is the tool for a scene rather than a person: pass a cell to place a new event, or an existing eventId and pageIndex to rewrite that page. The compiler produces the indentation and the branch markers (402/403, 411/412, 413, 601-603) the engine reads, so a body cannot land outside the branch it belongs to — the mistake that otherwise shows up as a choice that always takes the first option. Checked before writing: unknown step names, conditions or commands that name a switch, variable or database row the game does not have, audio and face files the project lacks. One transaction, map rendered back.
| Name | Required | Description | Default |
|---|---|---|---|
| x | No | ||
| y | No | ||
| name | No | ||
| when | No | Condition for the page this script goes to | |
| image | No | The event's graphic. A characterName that is not in img/characters fails the call | |
| mapId | Yes | ||
| script | Yes | The page's whole script. Each step is an object with exactly one key: `say` (a string, a list of lines, or {lines, speaker?, faceName?, faceIndex?, background?, positionType?}), `scroll` ({lines, fast?, wait?}), `comment`, `choice` ({prompt?, options: [{label, then, when?, lockedMessage?}], cancel?: steps, defaultType?, positionType?, background?}), `if` ({when, then, else?}), `loop` (steps, closed by `break` or `goto`), `break`, `exit`, `label`, `goto`, `switch` ({id, value?}), `selfSwitch` ({letter, value?}), `variable` ({id, set?|add?|multiply?}), `gold` (a number; negative takes), `item`/`weapon`/`armor` ({id, count?, take?}), `heal`, `transfer` ({mapId, x, y, direction?, fade?}), `wait` (frames), `fade` ("out"|"in"), `flash` ({color?, duration?, wait?}), `weather` ({type, power, duration?, wait?}), `se`/`me`/`bgm` (a name or {name, volume?, pitch?, pan?}), `stopBgm`, `stopSe`, `animate` ({target?, animationId, wait?}), `balloon` ({target?, balloonId, wait?}), `moveRoute` ({target?, route: [route steps], repeat?, skippable?, wait?}), `battle` ({troopId, canEscape?, canLose?, win?, escape?, lose?}), `shop` ({goods: [{id, kind?, price?, purchaseOnly?}]}), `menu`, `saveScreen`, `gameOver`, `title`, `commonEvent`, `erase`, `script` (engine code, stored as 355/655) and `raw` (verbatim {code, parameters, indent?}). `script` and `raw` are escape hatches: they come back named in the reply, and a build that leans on them is reporting a shape this layer does not cover yet. | |
| eventId | No | Rewrite this event instead of placing a new one | |
| replace | No | With `name`: rebuild the event of that name instead of adding another (default true) | |
| trigger | No | 0 action button, 1 player touch, 2 event touch, 3 autorun (fires on entering the map), 4 parallel process. 0 and 1 are checked for a cell the player can stand on; 3 and 4 are not, so an opening cutscene can sit on a wall in the corner the way the editor's own do | |
| pageIndex | No | With eventId: which page to replace (default 0) | |
| priorityType | No | 0 below tiles (default: a zone or a cutscene should not block), 1 same as tiles, 2 above |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Adds substantial behavior beyond the annotations: pre-write validation (unknown steps, missing switches/variables/rows, missing audio and face files), automatic indentation and branch-marker emission, and the 'one transaction, map rendered back' atomicity guarantee. The 'rewrite that page' wording also discloses the overwrite that destructiveHint=true implies.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with purpose and the place-vs-rewrite decision, then the step catalog, then validation/transaction behavior. The step enumeration is long but each token names a real construct; slightly dense as a single block, but no filler sentences.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 12-param, no-output-schema mutation tool this covers purpose, modes, validation and atomicity, and notes that `script`/`raw` escape hatches 'come back named in the reply'. It stops short of describing the overall response payload or pagination/limits, but is close to complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema leaves `script` as an opaque array of single-key objects (additionalProperties {}), and the description compensates by enumerating the whole step vocabulary and the shape of key steps (choice options with `when`/`then`, battle win/escape/lose branches, loop closed by break/goto). This is meaning the schema does not carry.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Write an event's whole script from a step list') and enumerates the exact constructs it authors (dialog, choice branches, if, loop, battle, shop, etc.). It also positions itself against siblings by declaring 'the tool for a scene rather than a person', separating it from make_npc/make_chest.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear context for both modes ('pass a cell to place a new event, or an existing eventId and pageIndex to rewrite that page') and differentiates scene-level authoring from person-level authoring. It does not explicitly name nearby alternatives like add_commands/set_commands/place_event, so the routing is implied rather than spelled out.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
make_encounter_zoneTurn a region into a battle zoneADestructiveIdempotent
Paint a region, make the troops, and attach the encounter rows in one call. A zone is three writes across two files (layer 5 of the map, then the map's encounter list) plus a Troops row per group, and the row shapes are exactly the ones that break a game when they are guessed: MZ reads regionSet unconditionally, so a row without that array throws inside Scene_Map on its first roll and freezes the map mid-walk, and a row without weight makes the weight sum NaN so that no row on the map ever rolls. A troop entry either names an existing troopId or lists enemies (ids, or {id, level?, x?, y?}), and the row is cloned from a real troop so its member shape matches what the editor writes. Region 0 means the whole map. The reply calls out any row whose region no cell carries, because that is the mistake that makes a zone feel empty rather than broken.
| Name | Required | Description | Default |
|---|---|---|---|
| mapId | Yes | ||
| region | No | Paint the region on layer 5; omit to keep whatever is already painted | |
| troops | Yes | The battle groups that can appear here | |
| clearRegion | No | Erase this region id from the whole map first, so a re-run does not leave the old paint behind | |
| keepExisting | No | Add to the map's current encounter rows instead of replacing them | |
| encounterStep | No | Steps the player walks between encounter rolls on this map (default 20; MZ's new maps use 30) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
It goes well beyond the annotations by disclosing concrete failure modes: MZ reads regionSet unconditionally so a row without it throws inside Scene_Map and freezes the map, and a missing weight produces a NaN sum that blocks all rolls. It also explains replacement vs additive behavior and that the reply flags orphaned regions, all runtime consequences an agent cannot infer.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action and structured as what-it-does followed by why-the-shapes-matter warnings. The prose is dense and vivid ('freezes the map mid-walk'), which mostly earns its place as load-bearing caution, though a few phrasings could be tightened.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 6-parameter, nested, multi-file mutation with no output schema, the description covers inputs, failure modes, clearing, and even the reply's diagnostic callout. It omits explicit return structure, but the mention of what the reply reports partially fills that gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 83%, so the baseline is 3, but the description adds real semantics: region 0 means the whole map, a troop entry either names an existing troopId or lists enemies, and rows are cloned from a real troop so member shape matches the editor. These clarify nested-parameter intent beyond the schema descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening sentence names a specific composite action and its parts: 'Paint a region, make the troops, and attach the encounter rows in one call.' It clearly distinguishes itself from siblings like make_battle and make_map by describing the three coordinated writes across two files.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use it ('in one call' composite of region + troops + encounter rows) and hints at re-run behavior via clearRegion, but never names alternatives or states when NOT to use it versus make_battle or set_map_properties. Usage is inferable rather than explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
make_itemWrite a row the player can hold or castBDestructiveIdempotent
An Items, Weapons, Armors or Skills row in one call, with the effects and equipment rules spelled in words. The effect shape is the whole reason this exists: MZ stores an item effect as {code, dataId, value1, value2} — 11 recovers HP as mhp × value1 + value2, 12 MP, 13 TP, 21 and 22 add and remove a state, 31 and 32 add a buff and a debuff (dataId is the parameter, value1 the turns), 33 and 34 remove them, 41 is the special (0 = escape), 44 runs a common event — while a trait on the same row is {code, dataId, value} with one key. Send the trait's value to an effect and the engine computes mhp × undefined, floors it to NaN, and adds NaN to the actor's HP: a number the bar cannot draw and the save file now has to carry. This compiles the words into the two shapes and refuses a state, parameter or common event the project does not have. The type ids are the other quiet failure: an item whose itypeId is not 1 or 2 is never listed — MZ has no item-type name list, the menu reads those two numbers and $dataSystem.itemCategories decides whether the tab is open at all — and a weapon whose wtypeId matches no weapon type cannot be equipped by anyone, so both are checked against System.json and said out loud. Idempotent by name, like the rest of the layer: re-running a build script rewrites the same row instead of adding one.
| Name | Required | Description | Default |
|---|---|---|---|
| id | No | Rewrite this row instead of making one | |
| icon | No | Index into img/system/IconSet.png | |
| name | No | ||
| note | No | ||
| price | No | ||
| scope | No | 0 none, 1 one enemy, 2 all enemies, 3 random, 7 self, 8 one ally, 9 all allies, 10 party, 11 dead ally | |
| speed | No | Action order bonus, the engine's own field | |
| table | No | Default Items | |
| damage | No | ||
| dryRun | No | ||
| mpCost | No | Skills only | |
| params | No | Weapons and armors: the flat parameter change when equipped | |
| tpCost | No | Skills only | |
| tpGain | No | ||
| traits | No | Weapons and armors: the engine's own {code, dataId, value}, single value | |
| atypeId | No | Armors: System.armorTypes | |
| effects | No | ||
| etypeId | No | Weapons and armors: the equipment slot, System.equipTypes | |
| itypeId | No | Items: 1 for the Item tab, 2 for the Key Item tab. MZ keeps no item-type name list, so any other number leaves the row out of the menu | |
| repeats | No | ||
| stypeId | No | Skills only: the type from System.skillTypes, by name or id | |
| wtypeId | No | Weapons: System.weaponTypes — a row nothing names cannot be equipped | |
| copyFrom | No | ||
| message1 | No | Skills only: the battle log line, `X uses …` | |
| message2 | No | Skills only: the second battle log line | |
| occasion | No | 0 any time, 1 menu only, 2 battle only, 3 never — never is not the same as unbuyable | |
| consumable | No | ||
| animationId | No | ||
| description | No | ||
| successRate | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare idempotentHint=true and destructiveHint=true; the description reinforces idempotency ('re-running a build script rewrites the same row instead of adding one') and adds substantial extra context: validation/refusal of unknown states, parameters and common events, and the NaN failure mode from mismatching trait/effect shapes. It stops short of explaining reversibility or the destructive profile explicitly, so 4 rather than 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose is front-loaded in the opening clause, which is good, but the body is a long single block dense with engine lore (NaN arithmetic, item-type menu behavior) that is educational yet costly for a tool-selection summary. It is not padded with filler, but it is longer than the selection decision requires.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a high-complexity 30-param, nested-object mutation tool with no output schema, the description covers the effect/trait compilation contract, validation refusals, and idempotency. It omits return/result behavior and the dryRun path, but the core failure modes an agent must avoid are addressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 30 parameters, 57% schema description coverage and nested objects, the schema does much of the work. The description adds genuine conceptual value by explaining the {code,dataId,value1,value2} effect shape versus the {code,dataId,value} trait shape and the itypeId/wtypeId validity rules, but leaves many uncovered params (dryRun, copyFrom, note, repeats) to the schema, so baseline 3 is right.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (write/make) and resource (a row across Items/Weapons/Armors/Skills), and the framing 'Write a row the player can hold or cast' is concrete. It explains what makes it distinct (compiling effects/traits into engine shapes), but never names the generic siblings like create_database_entry or patch_database_entry it presumably supersedes, so an agent must infer the routing.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description is rich on what the tool does and why, but gives no explicit when-to-use versus when-not guidance and no alternatives. There is no mention of dryRun or when to prefer this over the generic database tools, so the agent gets no routing signal beyond implied specialism.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
make_mapBuild a whole map in one callADestructiveIdempotent
One intent — "a 26x18 field of grass, with a stone floor and a wall ring and a gap for the door" — as one call. It makes or reuses the map, sets its tileset, resolves the entire paint plan, writes it in one pass, applies the passage flags the plan names, and answers with the rendered map. Before this the high-level layer stopped at content and every build script had to drop to create_map + set_tiles per map, which is where the depth mistakes came from.
The plan is resolved to one tile per cell, the last stroke winning, because that is how the engine decides passage: Game_Map.checkPassage reads tile layers 3 down to 0 and returns on the first tile whose flags say something, so a wall left on layer 3 underneath a floor tile still blocks the doorway — the map renders as a room and the player cannot walk out of it. Paint ground first, then what sits on it, and nothing is ever stacked. On a repaint the plan owns the cell's whole stack: every layer it does not name is zeroed, so last run's roof cannot survive under this run's floor.
Each tile still lands on the layer its own slot belongs to (A1/A2 → 1, A3 → 2, A4 → 3, A5/B–E → 0); that is set_tiles with layer omitted, so a forest is never buried under the grass again.
flags is handed to set_tileset_flags for this map's tileset in the same transaction, so "these ids are walls" is stated where the walls are painted. regions paints layer 5 for make_encounter_zone to roll on. The reply counts the walkable cells the finished terrain leaves, because a map with no walkable cell is not something the engine will report — it is a game that boots and never moves.
| Name | Required | Description | Default |
|---|---|---|---|
| fill | No | Tile id for every cell, under all the paint strokes | |
| find | No | Reuse the map whose tree name equals this, making it if there is none — the idempotent form for a build script | |
| name | No | The map's name in the map tree | |
| flags | No | `set_tileset_flags`' own `tiles` entries ({tileId|range, passable, blockFrom, bush, counter, terrainTag, …}), applied to this map's tileset | |
| mapId | No | Re-paint this map instead of making one | |
| paint | No | Rectangles and cell lists, applied in order over `fill`; later strokes win | |
| width | No | Required for a new map; ignored on a repaint (resizing means rebuilding the tile array) | |
| dryRun | No | Report the plan and write nothing | |
| height | No | ||
| regions | No | Painted on layer 5, on top of whatever tile is there | |
| parentId | No | Map-tree parent for a new map (default 1) | |
| tilesetId | No | Required for a new map; changing it on a repaint re-reads every tile id against the new flags | |
| properties | No | `set_map_properties`' own fields: displayName, disableDashing, battlebacks, parallax, bgm/bgs, encounterStep, encounters… | |
| clearRegions | No | Erase these region ids from the whole map first |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare idempotentHint and destructiveHint; the description goes well beyond by explaining the one-tile-per-cell, last-stroke-wins resolution, the layer-zeroing on repaint, the passage-flag behavior, the layer assignment per slot, and the transaction enclosing set_tileset_flags. This is rich behavioral disclosure. It docs the dryRun and the idempotent find semantics, though it does not state the destructive nature of repaint explicitly except through the zeroing behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long and somewhat dense, moving between a scenario story, engine internals, and per-parameter behavior. It is front-loaded with the one-call intent, which is good, but several sentences could be trimmed without losing value, and the engine-internals digression is longer than necessary for tool selection.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given 14 parameters, nested objects, no output schema, and destructive+idempotent annotations, the description covers the important behavioral and transactional details. It could be more complete on the output (it mentions the reply counts walkable cells), and on permissions, but for a high-level composite tool it is fairly complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 93%, so the schema already documents most parameters. The description adds meaning beyond the schema by explaining the resolution semantics (last stroke wins), paint ordering (ground first, then what sits on it), layer-zeroing on repaint, and how flags/regions/properties map to sibling tools. This is well above the baseline 3.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a concrete verb+resource ('makes or reuses the map, sets its tileset, resolves the paint plan, writes it in one pass, applies passage flags') and distinguishes itself from the low-level alternative ('create_map + set_tiles per map'). It is clear it is a high-level composite tool. It does not explicitly name siblings like create_map as alternatives by name, but it does reference the low-level pattern, which is close enough for a 4.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage context is clear: it is the high-level one-call form that replaces dropping to create_map + set_tiles per map, and it names the idempotent form via 'find'. It explains when to prefer this over the lower-level layer. It does not enumerate when-not-to-use, but the alternative framing is strong.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
make_npcAdd a character who talksADestructiveIdempotent
Place a non-player character in one call: graphic, movement, what it says, and an optional second page that takes over once a switch or self switch is set. The dialog compiles into the shapes the engine reads (Show Text 101 plus one 401 per line; a follow page conditioned on the self switch the first page then sets), and the whole thing runs as one transaction — if any part fails, nothing is written. Answers with the map rendered around the new NPC. The two defaults worth knowing: priorityType 1 (same as tiles), because a person should block their cell, and replace true, so re-running a build script updates this NPC instead of leaving a second copy of it. patrol writes the page's own movement route, which the editor's moveType 1 runs. For more than talk — choices, gates, battles — use script here or make_choice_scene.
| Name | Required | Description | Default |
|---|---|---|---|
| x | Yes | ||
| y | Yes | ||
| say | No | What the first page says | |
| name | No | The event's name; also how replace finds this NPC again | |
| note | No | ||
| image | No | The event's graphic. A characterName that is not in img/characters fails the call | |
| mapId | Yes | ||
| pages | No | Three or more pages, in engine order: the first whose condition holds is the one that runs. Given `pages`, it *is* the page list and every page carries its own say or script — a top-level say/script alongside `pages` is refused, because merging it into every page conditions page 0 on the last page's condition and leaves a mute NPC | |
| follow | No | The page that replaces the first once its condition holds | |
| patrol | No | A custom movement route, e.g. [{right: true}, {right: true}, {left: true}, {left: true}]. Route steps are {key: value}: down, left, right, up, lowerLeft, lowerRight, upperRight, upperLeft, random, toward, away, forward, backward; jump {dx, dy}; wait n; turn ("down"|"left"|"right"|"up"|"90right"|"90left"|"180"|"random90"|"random"|"toward"|"away"); switch {id, value}; speed 0..6; frequency 0..4; walkAnime, stepAnime, directionFix, through, transparent (true/false); image {characterName, characterIndex}; opacity 0..255; blend 0..4; se {name}; script "...". | |
| script | No | The first page's steps instead of plain lines. Each step is an object with exactly one key: `say` (a string, a list of lines, or {lines, speaker?, faceName?, faceIndex?, background?, positionType?}), `scroll` ({lines, fast?, wait?}), `comment`, `choice` ({prompt?, options: [{label, then, when?, lockedMessage?}], cancel?: steps, defaultType?, positionType?, background?}), `if` ({when, then, else?}), `loop` (steps, closed by `break` or `goto`), `break`, `exit`, `label`, `goto`, `switch` ({id, value?}), `selfSwitch` ({letter, value?}), `variable` ({id, set?|add?|multiply?}), `gold` (a number; negative takes), `item`/`weapon`/`armor` ({id, count?, take?}), `heal`, `transfer` ({mapId, x, y, direction?, fade?}), `wait` (frames), `fade` ("out"|"in"), `flash` ({color?, duration?, wait?}), `weather` ({type, power, duration?, wait?}), `se`/`me`/`bgm` (a name or {name, volume?, pitch?, pan?}), `stopBgm`, `stopSe`, `animate` ({target?, animationId, wait?}), `balloon` ({target?, balloonId, wait?}), `moveRoute` ({target?, route: [route steps], repeat?, skippable?, wait?}), `battle` ({troopId, canEscape?, canLose?, win?, escape?, lose?}), `shop` ({goods: [{id, kind?, price?, purchaseOnly?}]}), `menu`, `saveScreen`, `gameOver`, `title`, `commonEvent`, `erase`, `script` (engine code, stored as 355/655) and `raw` (verbatim {code, parameters, indent?}). `script` and `raw` are escape hatches: they come back named in the reply, and a build that leans on them is reporting a shape this layer does not cover yet. | |
| replace | No | Rebuild the NPC of this name on this map instead of adding a second one (default true) | |
| speaker | No | Name shown above the dialog box, for every page here | |
| through | No | ||
| trigger | No | 0 action button (an NPC's default), 1 player touch, 2 event touch, 3 autorun, 4 parallel process | |
| faceName | No | Face image from img/faces, for every page here | |
| moveType | No | 0 static, 1 custom route (pass patrol), 2 random, 3 toward the player, 4 away | |
| moveSpeed | No | ||
| walkAnime | No | ||
| directionFix | No | Stop the sprite turning as it moves — what a facing-specific shopkeeper needs | |
| priorityType | No | 0 below tiles (the player walks through), 1 same as tiles (blocks, the default), 2 above tiles | |
| moveFrequency | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only carry idempotent/destructive hints; the description adds the transaction/atomicity guarantee ('if any part fails, nothing is written'), the reply shape (map rendered around the new NPC), the two consequential defaults (priorityType 1, replace true), and the note that script/raw escape hatches are echoed back in the reply. This is substantial disclosure beyond the annotations and is consistent with idempotentHint=true and destructiveHint=true.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core behavior, then defaults, then the escape-hatch caveat. Sentences are dense but each carries distinct information (atomicity, defaults, alternatives). It is on the long side for a tool description, but little is filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex, deeply nested tool with no output schema, it covers the write semantics, the reply's rendered map, defaults, and the escape-hatch limitation. It omits error/validation detail beyond the characterName failure (which the schema already states) but is otherwise complete enough to invoke confidently.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 22 params at 64% coverage, the description adds real meaning the schema does not: the rationale and default for priorityType 1 and replace true, and that `patrol` writes the page's movement route run by moveType 1. It leaves other params (e.g. moveSpeed, moveFrequency, walkAnime) to the schema, so it does not fully compensate for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Opens with a specific verb+resource ('Place a non-player character in one call') and enumerates exactly what it creates: graphic, movement, dialogue, and an optional follow page. It is clearly separable from sibling place_event (raw event) and make_choice_scene (branching scenes).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a clear routing rule for adjacent needs — 'For more than talk — choices, gates, battles — use `script` here or make_choice_scene' — and explains the replace default as the build-script idiom. It does not exhaustively state when NOT to use it (e.g. vs place_event), so it stops short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
make_shopAdd a shopkeeper with goodsADestructiveIdempotent
One call for a working shop: the keeper's graphic, a greeting, the goods, and a line for after the window closes. MZ stores a shop as command 302 carrying the first good in its own parameters plus one 605 line per further good, each [kind, id, priceType, price, purchaseOnly] with kind 0 item / 1 weapon / 2 armor and priceType 0 meaning the row's own price. MZ has no purchase branch, so the commands after the goods run when the window closes whatever the player did — this says so in the reply instead of inventing a hook that does not exist. Every good is checked first: the row must exist and must not be a blank slot, and a good with no price whose item price is 0 is reported, because that is a free item. when adds the page that replaces the locked one, which is how a shop that opens after a quest is built.
| Name | Required | Description | Default |
|---|---|---|---|
| x | Yes | ||
| y | Yes | ||
| name | No | ||
| when | No | The shop opens only once this holds; the keeper still greets you before that | |
| goods | Yes | What is on the counter | |
| image | No | Default Actor2 index 1 facing up, which is the merchant sheet MZ ships | |
| mapId | Yes | ||
| denied | No | What the keeper says while the shop is still closed | |
| replace | No | ||
| speaker | No | ||
| faceName | No | ||
| farewell | No | Said once the shop window closes | |
| greeting | No | A line, a list of lines, or {lines, speaker?, faceName?, faceIndex?, background?, positionType?} |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare idempotentHint=true and destructiveHint=true, and the description adds substantial behavioral context beyond them. It explains the underlying MZ command encoding (302/605), that commands after the goods run when the shop window closes regardless of purchase, the validation of each good, free-item reporting, and that `when` replaces the locked page. No annotation contradiction is present.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The purpose is front-loaded, but the description is a dense, single paragraph mixing purpose, MZ implementation internals, and validation rules. Some details about command 302/605 are tangential to invocation, and the text could be more structured for a 13-parameter tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given 13 parameters, nested objects, and no output schema, the description covers the creation behavior, side effects, and key parameter semantics fairly well. However, it omits explanations for several optional parameters and does not describe what the tool returns, leaving some gaps for an agent to infer.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 46%, so the description must compensate. It explains the goods structure, validation rules, and `when` semantics, but leaves many parameters (name, replace, speaker, faceName, denied, image) largely unexplained. It also describes internal numeric kind/priceType values that do not match the schema's string enum, which may confuse rather than clarify.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific composite creation action: it builds a working shop with keeper graphic, greeting, goods, and farewell line. It distinguishes itself from siblings like make_chest and make_npc by focusing on the shop event and its goods array. An agent can identify the tool's purpose without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'One call for a working shop' implies the use case, and the explanation of `when` shows how to build conditional shops. However, there is no explicit when-to-use guidance versus alternatives such as make_npc or make_chest, and no exclusions or prerequisites are stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
map_connectivityMap reachability, across portalsARead-only
Flood-fill walkable cells from a start position and report which events the player can actually reach. Answers 'can the player get there' instead of trusting tile placement. An event blocks a cell only when one of its pages is same as tiles (the engine's isNormalPriority), which is why a chest stops you and a floor decal does not; conditions are not evaluated, so a page that only applies later still counts as blocking. An event is standable when its own cell is reachable and touchable when a neighbour is, which is all an action-button event needs; an event with an autorun or parallel page is automatic and needs no path at all. Transfer Player commands found on standable events are followed into their destination maps, so a game whose rooms are joined only by portals reads as one connected space instead of a pile of isolated cells — and a portal that lands on an impassable tile is called out, because that is a trap the player cannot escape. reports carries one entry per map walked, in the order they were reached.
| Name | Required | Description | Default |
|---|---|---|---|
| x | Yes | ||
| y | Yes | ||
| mapId | Yes | ||
| maxMaps | No | ||
| followTransfers | No | Walk the maps that transfers on this one lead to (default true) | |
| includeEventBlocking | No | Treat events that block the tile as walls (default true) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With only readOnlyHint available, the description carries the behavioral load and does so richly: blocking requires a same-as-tiles page (isNormalPriority), conditions are not evaluated, standable/touchable/automatic are defined, transfers are followed, and a portal landing on an impassable tile is flagged as a trap. This is exactly the extra context annotations cannot supply.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the core action and definition, then layers semantics efficiently; nearly every sentence earns its place. It is a dense single paragraph, and the run-on style around the portal/trap clause slightly hurts scannability.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex traversal tool with no output schema, the description covers reachability categories, transfer following, and the reports ordering well. It omits the entry shape and the maxMaps bound, which are the remaining gaps an agent would want before calling it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 33%, so the description must compensate. It effectively explains the blocking behavior behind includeEventBlocking and the transfer-following behind followTransfers, but maxMaps (the walk bound) and the coordinate/mapId semantics are never addressed, leaving one non-obvious parameter undocumented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'flood-fill walkable cells from a start position and report which events the player can actually reach.' It also frames the intent ('can the player get there') versus the naive alternative of trusting tile placement, which cleanly separates it from inspect_cell/get_map siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear context for when to reach for it — verifying actual reachability rather than tile placement, and following portals into destination maps. However, it never names an explicit alternative or exclusion (e.g. when to use inspect_cell instead), so the routing guidance is implied rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
patch_database_entryPatch a database entryA
Apply a shallow field patch to one entry of a database table (or to System.json's top-level keys). Nested objects and arrays must be passed complete, because MZ has no partial-update semantics on them — a nested object that arrives without a key the entry already had loses that key, and the reply names the keys it lost so the caller can read the entry first. System's switches and variables are the exception: pass either an array of names or an object keyed by id, and the array the engine needs is what gets written.
| Name | Required | Description | Default |
|---|---|---|---|
| id | No | Entry id; ignored for System | |
| patch | Yes | ||
| table | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does so well: it discloses that nested objects/arrays have no partial-update semantics, that missing keys are lost, and that the reply names lost keys. It also documents the switches/variables exception. It omits auth/permission requirements and idempotency, hence not a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action and scope; the remaining sentences each carry distinct behavioral information (loss semantics, reply behavior, exception). Slightly dense and long-winded in the nested-object sentence, but no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with no annotations and no output schema, the description covers the key risks (data loss on nested patches) and hints at the response (names lost keys). It could say more about failure modes or permissions, but the essentials for correct invocation are present.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 33% (only 'id' documented), so the description must compensate, and it does: it explains patch semantics for nested structures, the complete-pass requirement, the switches/variables alternative input shapes, and that id is ignored for System. It adds substantial meaning beyond the schema's bare types.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('apply a shallow field patch') and resource ('one entry of a database table') with scope qualifier ('or System.json's top-level keys'). This clearly distinguishes it from siblings like create_database_entry and read_database.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It implies a workflow ('read the entry first' when keys are lost) but never states when to choose this tool over read_database or create_database_entry, nor any exclusions. Usage context is only inferable from the semantics of 'patch'.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
patch_pluginChange a plugin's switch or parametersADestructive
Set one plugin's enabled status, description, or parameter values in js/plugins.js — the same edit as the editor's plugin manager, without reformatting the file: only the touched entry's object literal is rewritten. Parameter values are stored the way MZ stores them, so scalars become strings and struct or array values stay JSON. PluginManager.parameters is read once at boot, so a running playtest needs a restart to see this, and an editor with the plugin manager open will overwrite the file on save.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | ||
| status | No | ||
| parameters | No | Merged over the existing keys; pass null to drop one | |
| description | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare destructiveHint=true and idempotentHint=false, and the description substantially enriches that: only the touched entry is rewritten (no full-file reformat), scalars are coerced to strings while struct/array values stay JSON, PluginManager.parameters is read once at boot so a restart is required, and an open editor plugin manager will clobber the file on save. These are exactly the behavioral facts an agent needs before mutating the file.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The core action is front-loaded in the first clause, followed by the file-behavior nuance and then the operational caveats. Every sentence earns its place, with no redundant restatement of the title or schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with no output schema, the description covers the write scope, the coercion behavior, and the two critical side effects (boot-time parameter caching and editor-overwrite hazard). An agent has enough to invoke it correctly and anticipate the consequences.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 25% (only `parameters` is documented), so the description carries extra weight: it maps the three settable fields (enabled status, description, parameter values) to the parameters and explains the MZ storage coercion (scalars to strings, struct/array stays JSON). It does not clarify `name` semantics or the `status`/`description` params individually, so it compensates well but not fully.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: set one plugin's enabled status, description, or parameter values in `js/plugins.js`. It further distinguishes itself from write_plugin_source by noting it only rewrites the touched entry's object literal rather than reformatting the file. An agent can tell it apart from sibling plugin tools without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Frames the tool as 'the same edit as the editor's plugin manager' and gives clear situational context (running playtest needs a restart; editor open will overwrite on save). It does not, however, explicitly name when to prefer this over siblings like enable_plugin or write_plugin_source, leaving the alternative routing to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
place_eventPlace or update an eventB
Create an event at a cell, or move and rename an existing one by id.
| Name | Required | Description | Default |
|---|---|---|---|
| x | Yes | ||
| y | Yes | ||
| id | No | Existing event id to move; omit to create | |
| name | No | ||
| note | No | ||
| mapId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It says nothing about whether moving an event onto an occupied cell overwrites, what happens to unspecified fields on update, permission requirements, or reversibility. For a mutating tool with zero annotation coverage this is a significant gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence that covers both modes with zero filler. Nothing to trim.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
A 6-parameter mutation tool with no annotations, no output schema, and near-zero schema description coverage gets only one sentence. Return value, error behavior, note semantics, and the overwrite question are all unaddressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 17% (just id), so the description must compensate. It does map x/y to 'cell', id to 'move', and name to 'rename', but leaves note and mapId unexplained and gives no format or constraint detail for any of them.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource with dual-mode scope: 'Create an event at a cell, or move and rename an existing one by id.' That distinguishes it from remove_event, copy_event, and find_events. It stops short of naming those siblings explicitly, so it's clear but not fully differentiated.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Implies when each mode applies via the 'omit id to create' semantics and the schema's id description, but never names alternatives (e.g., copy_event, set_event_page) or states prerequisites like whether the map must be loaded. Usage is inferred rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
project_infoProject overviewA
Summarize the configured RPG Maker MZ project: maps, tilesets, database tables, switch and variable names, and which engine version the renderer uses.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. 'Summarize' implies a non-mutating read and the enumerated scope tells the agent what it covers, but it says nothing about whether this is expensive, whether it requires the project to be configured/loaded, or how it behaves on an empty or invalid project.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with the verb first and the enumerated contents following. Every clause carries information; there is no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter, low-complexity read tool with no output schema, the description's enumeration of returned categories (maps, tilesets, database tables, switch/variable names, engine version) effectively serves as the return-value contract. Only the absence of any usage context keeps it from being fully complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so there is nothing for the description to disambiguate. The baseline for a parameterless tool applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Summarize') and resource ('the configured RPG Maker MZ project') and enumerates the exact contents of the summary: maps, tilesets, database tables, switch/variable names, and engine version. This distinguishes it from narrower siblings like list_maps or read_database, though it never names those siblings explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit when-to-use guidance, no statement of prerequisites, and no routing to alternatives such as list_maps or read_database for targeted queries. The orientation use case is only implied by the word 'Summarize'.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
read_databaseRead a database tableA
Read data/.json. Returns entries compactly, or one entry in full. Covers Actors, Classes, Skills, Items, Weapons, Armors, Enemies, Troops, States, Animations, Tilesets, CommonEvents, System, MapInfos. Every table answers count and entries, so a caller can read one shape; System is a single object rather than a row table, so it answers value as well and entries holds that object as its one row.
| Name | Required | Description | Default |
|---|---|---|---|
| id | No | Return one entry instead of the whole table | |
| table | Yes | ||
| fields | No | Restrict the summary to these fields |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden, and it does disclose meaningful behavior: the compact-vs-full return mode, and the System special case where 'entries' holds a single object. It does not mention failure behavior for an invalid table name or whether reads are cached, which keeps it short of a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action and the return-shape summary before the table enumeration. The final sentence about System is somewhat convoluted and slightly overlong, but no sentence is pure filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, describing the return shapes (`count`, `entries`, `value`) is essential and is done. Valid table values, the single-entry mode, and the field-restriction mode are all covered; only error/edge-case behavior for an unknown table is absent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 67%, and the description compensates for the biggest gap: the `table` parameter has no description or enum in the schema, yet the description enumerates all valid table names. The `id` semantics ('one entry in full') are also restated and reinforced, though `fields` is not elaborated beyond the schema text.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Read data/<table>.json') and immediately names the exact file convention. It is clearly distinguishable from the write-oriented siblings patch_database_entry and create_database_entry, and the enumeration of covered tables pins down the scope precisely.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by the read-only framing and the table list, so an agent can infer this is the lookup tool. However, it never states when to prefer this over siblings such as find_events or project_info, nor does it name any alternative or exclusion.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
read_plugin_sourceRead a plugin's sourceARead-only
Return a window of lines from js/plugins/<name>.js, the only way to see what a plugin actually hooks before editing around it. Defaults to the first 200 lines; pass fromLine to page through a long file. declared carries the parsed @param blocks so a parameter name can be matched to the code that reads it.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | ||
| fromLine | No | ||
| lineCount | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, so the safety profile is covered. The description adds real value beyond that: the default 200-line window, the paging mechanism, and the fact that the result carries parsed @param blocks in a `declared` field.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences with zero banter; the resource and its defining property are front-loaded, and each remaining sentence adds a distinct fact (default size, paging, declared output).
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and three sparsely documented params, the description covers the essentials an agent needs: where the source lives, how to page it, and what the parsed output field contains. The only real gap is the size/limit of the `lineCount` parameter.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must carry the load and mostly does: it explains `name` through the path template, and `fromLine` as the paging control plus the 200-line default that implies `lineCount`. Only the `lineCount` parameter itself is left unnamed and unexplained.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Return a window of lines from js/plugins/<name>.js') and explicitly positions it against siblings by calling it 'the only way to see what a plugin actually hooks before editing around it', which distinguishes it from patch_plugin and write_plugin_source.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear context for use (reading plugin hooks before editing, paging long files via fromLine) and implicitly routes against the plugin-writing siblings, but never states an explicit 'when not to use' or names an alternative tool by name.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
remove_eventDelete an eventB
Delete an event from a map, leaving a null hole in the events array as the editor does.
| Name | Required | Description | Default |
|---|---|---|---|
| mapId | Yes | ||
| eventId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, but it does disclose one genuinely important trait: deletion leaves a null hole rather than compacting the array, so other event indices remain stable. It says nothing about error behavior for a missing eventId, whether the map must be loaded or a live session active, or permissions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no padding: verb, resource, and resulting state in order of importance. The trailing "as the editor does" is mildly vague but still conveys that behavior matches the editor's own delete.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-required-parameter mutation tool with no annotations and no output schema, the description covers the post-condition (null hole, no reindexing) but omits prerequisites, failure modes, and return behavior. It is adequate for the core action but not complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for both parameters, so the description must compensate. The "null hole in the events array" phrasing implicitly suggests eventId is an array index, which is useful, but mapId is left entirely unexplained and the index semantics are only implied, not stated.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb and resource ("Delete an event from a map"), which is unambiguous. It does not, however, distinguish itself from near siblings such as clear_events, place_event, or copy_event, so an agent must infer the routing itself.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no statement of when to use this tool versus alternatives like clear_events (bulk removal) or place_event/copy_event (creation). The only contextual hint is "as the editor does," which describes the resulting state, not the conditions for calling the tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
render_mapRender a map to PNGB
Composite a map exactly the way the engine does (autotiles, shadow bits, higher-tile z-order) and return it as an image, so layout can be verified visually. Overlays can show passability, region ids or terrain tags.
| Name | Required | Description | Default |
|---|---|---|---|
| mapId | Yes | ||
| scale | No | Pixels per tile / 48. Defaults to fitting a 2048px side | |
| saveTo | No | Also write the PNG to this absolute path | |
| overlay | No | Default none | |
| showGrid | No | ||
| onlyLayers | No | Restrict to these tile layers | |
| showEvents | No | Draw event markers and event graphics (default true) | |
| animationFrame | No | A1 water/waterfall animation frame to freeze |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, and it does disclose meaningful behavior: rendering is engine-accurate and the return is an image. It omits other behavioral facts, notably that saveTo writes a file to disk and whether the tool has any side effects, so it is only partially transparent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, no filler, with the core action and its engine-accurate compositing detail front-loaded before the overlay aside. Slightly terse given eight parameters, but every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an 8-parameter tool with no annotations and no output schema, the description covers the essence (engine-accurate render returned as an image, overlay semantics) but leaves gaps: image return format, the saveTo write side effect, and behavior of scale/grid/layer options are unaddressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 75%, so the schema already documents most parameters (scale, saveTo, overlay, onlyLayers, showEvents, animationFrame). The description's overlay sentence adds conceptual meaning by mapping overlays to passability/region/terrain, but says nothing about scale, saveTo, showGrid or layer restriction — a baseline 3.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('composite a map ... return it as an image') and adds the distinguishing detail that compositing mirrors engine behavior with autotiles, shadow bits and z-order. It does not, however, explicitly differentiate itself from near-siblings like live_screenshot or get_map, so an agent must infer the boundary.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The clause 'so layout can be verified visually' implies the intended use case (visual verification of layout), which is useful inferred context. There is no explicit when-to-use/when-not guidance and no mention of alternatives such as live_screenshot for a running game versus this offscreen render.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
rollback_dataRoll back a data fileADestructive
Restore a backed-up copy of one file. table takes the same identifier list_backups prints: a database name like System, a map as Map037 or 37, or a project-relative path like js/plugins.js. Without to this reverts the newest change (or the one before it, when the newest already matches what is on disk); with to it restores a specific backup number from list_backups, where 0 is the oldest. For undoing several of your own recent edits across files, prefer undo_writes.
| Name | Required | Description | Default |
|---|---|---|---|
| to | No | Backup number from list_backups (0 = oldest); omit for the newest change | |
| table | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true and idempotentHint=false, so the safety profile is covered structurally. The description goes beyond that by disclosing the backup-selection logic (newest vs one-before-when-disk-matches) and the 0=oldest ordering, but it never states that the current file content is overwritten or what happens on failure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, front-loaded with the core action, then the identifier format, then the `to` behavior, then the routing hint. Every clause carries information and none restates structured fields redundantly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple two-parameter destructive tool with annotations carrying the safety profile and no output schema, the description covers purpose, identifier formats, parameter branching, ordering, and the sibling alternative. Nothing an agent needs to call it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 50%, with `table` undocumented in the schema, yet the description fully compensates by enumerating the accepted identifier forms (database name, map as Map037/37, project-relative path) with concrete examples. It also restates the `to` semantics (backup number, 0=oldest, omit for newest) meaningfully.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb+resource ('Restore a backed-up copy of one file') and explicitly contrasts itself with siblings by referencing list_backups for identifiers and undo_writes for multi-file undo. An agent can disambiguate rollback_data from undo_writes without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It spells out the when-to-use branching: without `to` it reverts the newest change, with `to` it restores a specific backup, and it names the alternative ('prefer `undo_writes`') with the condition that selects it. This is explicit when/when-not/alternatives guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
set_commandsReplace a page's command listB
Overwrite an event page's whole command list, terminator included. Block levels are read the way the engine reads them: a body written at the same indent as its Conditional Branch or Loop is reported in warnings instead of being quietly kept.
| Name | Required | Description | Default |
|---|---|---|---|
| list | Yes | ||
| mapId | Yes | ||
| eventId | Yes | ||
| pageIndex | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations, so description carries full burden. It discloses two non-obvious behavioral traits: the terminator is included in the overwrite, and indent-level mismatches produce `warnings` rather than being silently accepted. It does not state whether the operation is destructive/irreversible or what happens to existing commands, but the warning behavior is valuable added context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Efficient two-sentence structure front-loading the core action, then the warning behavior. No filler, though it is a bit dense.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Mutation tool with no annotations, no output schema, and 0% schema description coverage. The description covers the warning behavior but leaves parameter semantics and destructive/permission aspects unexplained, which is inadequate for a complex 4-param write operation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% and description mentions no parameter names at all. The `list` array structure (code, indent, parameters) and mapId/eventId/pageIndex are entirely undocumented in prose. For a 4-parameter tool with 0% schema description coverage, the description fails to compensate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Clear verb+resource: 'Overwrite an event page's whole command list'. The terminator detail and sibling set_commands/add_commands distinction is not explicitly made, but the overwrite scope is clear from 'whole' and 'Overwrite'.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No explicit when-to-use guidance or alternatives named. The overwrite semantics imply this is a full replacement vs add_commands, but that routing is left to inference. No prerequisites or exclusions stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
set_event_pageConfigure an event pageC
Set a page's graphic, trigger, priority, movement and conditions. Adds the page when it does not exist yet. trigger: 0 action button, 1 player touch, 2 event touch, 3 autorun, 4 parallel. priorityType: 0 below tiles, 1 same as tiles, 2 above tiles.
| Name | Required | Description | Default |
|---|---|---|---|
| image | No | ||
| mapId | Yes | ||
| eventId | Yes | ||
| through | No | ||
| trigger | No | ||
| moveType | No | ||
| moveSpeed | No | ||
| pageIndex | No | Default 0 | |
| stepAnime | No | ||
| walkAnime | No | ||
| conditions | No | Setting a condition also enables its *Valid flag; pass an empty object to clear all conditions | |
| directionFix | No | ||
| priorityType | No | ||
| moveFrequency | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, and it does disclose one meaningful behavior: it creates the page if absent. It also decodes trigger and priorityType values. But it says nothing about permissions, return value, or whether unset properties are preserved or reset, which matters for a 14-parameter mutator.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences, front-loaded with the main action, followed by the upsert caveat and the enum legends. No filler. Slightly dense with value lists but each earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with 14 parameters, nested objects, no annotations, and no output schema, the description is too thin. Movement-related parameters (moveType/moveSpeed/moveFrequency) and the conditions object are unaddressed, leaving an agent to guess at a substantial portion of the call.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 14% across 14 parameters. The description documents two of them (trigger, priorityType) with their value meanings, but leaves moveType, moveSpeed, moveFrequency, through, directionFix, walkAnime, stepAnime, the image sub-object, and the conditions semantics entirely to an undocumented schema. Partial compensation for a large gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Set) and resource (an event page) and enumerates the configurable aspects: graphic, trigger, priority, movement, conditions. This clearly distinguishes it from place_event/remove_event/copy_event, though it doesn't explicitly say how it relates to those siblings. Strong but not sibling-routing.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The only contextual guidance is 'Adds the page when it does not exist yet,' which is an upsert behavior note rather than when-to-use guidance. There is no mention of when to prefer this over add_commands/set_commands, prerequisites (does the event need to exist first?), or exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
set_map_propertiesChange a map's own settingsADestructiveIdempotent
Set the map-level fields that are not tiles and not events: display name, tileset, the map-name banner, dashing, battlebacks, parallax (name, loop, auto-scroll, start offset), the map's own BGM/BGS and whether they autoplay, and the encounter table. Encounters are the reason this exists — list_maps reports them, and without a writer a playtest can be built that never meets a monster. Pass encounters as [{regionId, troopId, appearances}]: regionId 0 means the whole map, otherwise it matches the region ids painted on layer 5 with set_tiles, and appearances is the weight against the other rows. Omit clearEncounters to replace the list, set it true to empty it. Size is not here on purpose: resizing means rebuilding the tile array, and a wrong array length corrupts the map, so create the map at the size you want.
| Name | Required | Description | Default |
|---|---|---|---|
| bgm | No | ||
| bgs | No | ||
| name | No | Name in the map tree (MapInfos), not the in-game title | |
| note | No | ||
| mapId | Yes | ||
| tilesetId | No | ||
| encounters | No | ||
| parallaxSx | No | ||
| parallaxSy | No | ||
| scrollType | No | 0 screen, 1 continuous | |
| autoplayBgm | No | ||
| autoplayBgs | No | ||
| displayName | No | Map name shown on screen; empty falls back to the tree name | |
| parallaxName | No | ||
| parallaxShow | No | ||
| encounterStep | No | ||
| parallaxLoopX | No | ||
| parallaxLoopY | No | ||
| disableDashing | No | ||
| battleback1Name | No | ||
| battleback2Name | No | ||
| clearEncounters | No | ||
| specifyBattleback | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations give idempotentHint=true and destructiveHint=true, yet the description adds real context beyond them: encounters are replaced wholesale unless clearEncounters is set, and the size omission is justified by corruption risk ('a wrong array length corrupts the map'). This discloses overwrite semantics and a danger the annotations only hint at.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the scoping statement and field list, then the detail that matters (encounters). Every sentence carries content, including the closing size rationale, though the description is dense and could trim the field enumeration slightly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema required and a single required param, the description covers the complex nested encounters logic, the destructive replacement behavior, and the deliberate omission of resizing. A twenty-three-parameter tool could still say more about the non-encounter fields, so not full marks.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 13%, so the description must compensate. It does so for the hardest parameter, giving the encounters shape ([{regionId, troopId, appearances}]), the regionId=0 whole-map convention and its link to layer-5 regions painted by set_tiles, and the weight meaning. It also maps the other field groups by name, though per-field details (scrollType, encounterStep units) are left to the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Set) and a precisely scoped resource, then defines it by exclusion: 'map-level fields that are not tiles and not events.' The enumerated field list (display name, tileset, banner, dashing, battlebacks, parallax, BGM/BGS, encounter table) lets an agent separate this from set_tiles, set_event_page, and patch_database_entry without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explains why the tool exists ('Encounters are the reason this exists — list_maps reports them, and without a writer a playtest can be built that never meets a monster'), which tells the agent when it is needed. Names alternatives implicitly: set_tiles for painting regions, create_map for sizing, list_maps for reading. No explicit when-not beyond sizing, so not a full 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
set_startupSet how a new game opensADestructiveIdempotent
Write the opening state of the game in one call: the title on the splash and in the window, where New Game puts the player (map, cell, facing), who starts in the party, and the switch and variable name tables. patch_database_entry can do all of it and has to be handed advanced whole to touch one key of it, which is the shape of mistake this exists to prevent — a partial nested patch drops the keys you did not send. Idempotent: re-running a build script sets the same opening rather than appending. Reads back the start cell through the engine's own passability rules and refuses to place the player inside a wall, because that is a game that boots and never moves.
| Name | Required | Description | Default |
|---|---|---|---|
| title | No | The game title the engine shows and the save files carry | |
| startX | No | ||
| startY | No | ||
| troopId | No | ||
| switches | No | Names by index, or keyed by id — either way the array the engine indexes is what gets written | |
| variables | No | Names by index, or keyed by id | |
| startMapId | No | ||
| testBattle | No | Also set the troop the editor's test battle uses | |
| partyMembers | No | Actor ids, in the order the engine lines them up | |
| startDirection | No | 2 down, 4 left, 6 right, 8 up |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare idempotentHint=true and destructiveHint=true, and the description reinforces idempotency ('re-running a build script sets the same opening rather than appending') while adding behavior the annotations do not cover: it validates placement through the engine's passability rules and refuses to place the player inside a wall. It does not discuss failure modes for other params (invalid map/troop ids), keeping it below a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose is front-loaded in the first clause and every sentence carries information. The middle sentence about patch_database_entry is a touch winding, but the whole paragraph is efficient for a tool with this much behavioral nuance.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description still conveys the write's scope, its idempotent/validating behavior, and the sibling it replaces. The remaining gap is the 0-required-param case: it never says whether omitting a field leaves it untouched or clears it, which matters for a destructive write.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 60%, so the schema carries much of the load. The description maps conceptual groups to params (title, start map/cell/facing, party members, switch/variable tables), but adds no syntax or format detail for the undocumented params (startX, startY, startMapId, troopId), so it lands at the baseline for partial coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb+resource ('Write the opening state of the game in one call') and then enumerates exactly which facets it covers: title on the splash/window, New Game placement (map, cell, facing), starting party, and the switch/variable name tables. It is clearly distinguishable from siblings like set_map_properties or patch_database_entry.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly names the alternative (patch_database_entry) and explains why this tool exists instead: patch_database_entry must be handed `advanced` whole and a partial nested patch drops unset keys. This is a concrete when-to-use/when-not-to-use rule rather than implied usage.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
set_tilesPaint tilesA
Write tile ids on one map. Pass explicit cells, or a rectangle to fill. Layer 0-3 are tile layers, 4 holds shadow bits 0-15, 5 holds region ids. MZ puts A1/A2 ground on layer 1, A3 on 2, A4 upper walls on 3 and A5/B/C/D/E on 0 — leave layer out and each tile goes to the layer its own slot belongs to, which is the difference between a forest and a field of trees buried under the grass. Pass it to override, and painting a tile on a layer it does not belong to still renders, only at the wrong depth, so the reply says so.
| Name | Required | Description | Default |
|---|---|---|---|
| rect | No | Fill a rectangle with one tile id; autotile ids keep their shape ids, so a filled area becomes one contiguous autotile region | |
| cells | No | ||
| mapId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden and does substantial work: it decodes layer semantics (0-3 tile layers, 4 shadow bits, 5 region ids), the MZ slot-to-layer mapping, and the key trait that painting on a wrong layer 'still renders, only at the wrong depth, so the reply says so'. It omits permissions/auth, idempotency, and undo interaction, so not a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The purpose is front-loaded and the sentences are information-dense, each adding layer rules or behavior. The 'forest vs a field of trees buried under the grass' metaphor is vivid but somewhat verbose, making the paragraph longer than strictly necessary.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with no annotations and no output schema, the description supplies the essential context an agent needs: layer model, auto-layer inference, and the wrong-layer side effect. It is nearly complete, missing only permission/prerequisite notes and a fuller picture of the reply contents beyond the wrong-layer warning.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 33%, so the description must compensate — and it does for the critical `layer` parameter, giving meaning (0-3 tile layers, 4 shadow bits, 5 region ids) well beyond the bare integer 0-5 in the schema, plus the omit-to-infer behavior. It leaves tileId/rect/cells coordinate semantics largely to the schema, so it only partly closes the gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb+resource: 'Write tile ids on one map', immediately establishing a mutation of map tile data. It is clearly distinguishable in substance from read-oriented siblings like describe_tiles and flag-oriented set_tileset_flags, though it never names an alternative explicitly, keeping it short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives clear usage modes — 'Pass explicit cells, or a rectangle to fill' — and adds a conditional rule ('leave layer out and each tile goes to the layer its own slot belongs to'). What is missing is explicit routing against sibling tools (e.g. when to prefer describe_tiles or batch), so no exclusions are stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
set_tileset_flagsMake a tile a wall, a ladder, a bushADestructiveIdempotent
Edit what tiles do, without the editor. MZ keeps this as an 8192-entry flags array per tileset: 0x01/0x02/0x04/0x08 are impassable from Down/Left/Right/Up, 0x10 makes those four override the layers underneath, 0x20 ladder, 0x40 bush, 0x80 counter, 0x100 damage floor, 0x200/0x400/0x800 boat/ship/airship, and bits 12 up are the terrain tag. Passage is per direction, so passable: false is a fence and blockFrom: ["up"] is a ledge you can step onto but not climb back off. One bit has to be handled for you: MZ's 0x10 means "no effect on passage" — Game_Map.checkPassage skips the tile entirely — and much of the stock tileset (the pillars, trees and statues among others) ships with it set, so writing passage bits into one of those tiles without clearing it reports success and blocks nothing. Any passage change therefore clears 0x10 unless you pass overwrite: true on purpose, and the reply says when it did. For an A1-A4 autotile the id a map stores is one of 48 shapes of a base pattern and the engine reads the flags of that stored id, so a change is applied across the whole shape group unless shapes: false says otherwise; the reply reports which ids moved and what each flag became. This reaches every map that uses the tileset at once, which is the point and the danger: dryRun shows the before and after without writing, and validate_game re-checks the maps that now block.
| Name | Required | Description | Default |
|---|---|---|---|
| tiles | Yes | ||
| dryRun | No | Report what would change and write nothing | |
| tilesetId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Goes well beyond the idempotent/destructive annotations by disclosing non-obvious mechanics: the automatic clearing of 0x10 (with overwrite:true escape), the A1-A4 48-shape group propagation (with shapes:false), the global reach across maps, and what the reply reports. This is exactly the kind of hidden side effect an agent needs.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads 'Edit what tiles do, without the editor' effectively, but then delivers a single dense paragraph of bit-level detail. Every sentence is informative, yet the wall-of-text format and heavy parenthetical asides hurt scanability for an agent.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex mutation tool with no output schema and low schema coverage, the description supplies the missing pieces: flag semantics, the 0x10 gotcha, shape-group scope, global blast radius, dryRun preview, and follow-up validation. An agent has enough to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 33%, and the description compensates strongly, giving semantic meaning the schema lacks (passable:false as a 'fence' vs blockFrom:['up'] as a 'ledge'), explaining the overwrite/0x10 interaction and the shapes group default. It adds real interpretation beyond the raw bit descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource: editing tile flag/passability bits in an MZ tileset's 8192-entry flags array, and names concrete effects (wall, ladder, bush). The purpose is unmistakable, though it never names a sibling (e.g., set_tiles or describe_tiles) to draw the boundary explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear operational context: use dryRun to preview, validate_game to re-check affected maps, and warns the change reaches every map using the tileset. It does not explicitly state when to choose this over siblings like set_tiles or describe_tiles, so it stops short of full when/when-not guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
show_textAppend a Show Text dialogA
Add a Show Text dialog to an event page. MZ stores a dialog as one command 101 carrying [faceName, faceIndex, background, positionType, speakerName] followed by one command 401 per text line, each holding its line in parameters[0]; this tool builds that shape so lines can never end up as null in $gameMessage._texts.
| Name | Required | Description | Default |
|---|---|---|---|
| lines | Yes | One entry per dialog line; escape codes like \N[1] work | |
| mapId | Yes | ||
| indent | No | Use the surrounding block's indent when inserting inside a branch | |
| eventId | Yes | ||
| faceName | No | Face image name from img/faces, default empty | |
| faceIndex | No | ||
| pageIndex | No | ||
| background | No | 0 window, 1 dim, 2 transparent | |
| speakerName | No | ||
| positionType | No | Engine default is 2 |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It usefully discloses the exact emitted structure (command 101 plus one 401 per line) and a failure mode it avoids, but says nothing about append-vs-overwrite semantics, permissions, or whether existing page content is preserved.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded purpose in the first sentence, with the technical data-model detail confined to a second sentence. It is dense but every clause is relevant to correct invocation; no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 10-parameter, annotation-free tool with no output schema, the description covers the tricky data-shape concern well but leaves the relationship between eventId and pageIndex, and the append/overwrite behavior, unexplained. Adequate but with clear gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 50% and the description compensates by naming the command 101 payload order [faceName, faceIndex, background, positionType, speakerName] and clarifying that each line lands in parameters[0]. This adds real meaning beyond the partially-documented schema, though it does not cover mapId/eventId/pageIndex.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'Add a Show Text dialog to an event page,' and explains the underlying MZ command shape. An agent can tell what it produces, though it does not explicitly distinguish itself from siblings like add_commands, set_commands, or live_dialog.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied: the closing note that this 'builds that shape so lines can never end up as null' hints at why to prefer it over raw add_commands, but there is no explicit when-to-use, when-not, or named alternative. The agent must infer the routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
tileset_slotsList tileset slotsB
Show which image is bound to each A1-A5/B-E slot of a tileset and the first tile id of each slot, so tile ids can be chosen without guessing.
| Name | Required | Description | Default |
|---|---|---|---|
| tilesetId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the burden. It usefully discloses what the response contains (image bound to each slot, first tile id per slot), which matters since there is no output schema, but it says nothing about read-only safety, behavior on an unknown tilesetId, or whether missing slots are reported.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One tight sentence with the outcome front-loaded and the rationale trailing; nothing is wasted. It is perhaps slightly under-specified rather than over-long, but structurally sound.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and no annotations, the description does the minimum by sketching the return content, but it leaves the sole parameter undocumented and omits error/edge behavior. Adequate for a simple read, but not complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There is one required parameter, tilesetId, with 0% schema description coverage, and the description never explains it — no mention of valid id ranges, how to discover an id, or what an invalid id does. The single sentence about slots does not compensate for the undocumented input.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Names a specific verb (Show) and resource (tileset slot bindings A1-A5/B-E) plus the returned datum (first tile id per slot). It is clear what the tool does, though it never distinguishes itself from the closest sibling describe_tiles, which likely also reports tile information.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The clause 'so tile ids can be chosen without guessing' implies the use case (resolving tile ids before set_tiles or place_event), but there is no explicit when-to-use guidance or statement of how this differs from describe_tiles or list_maps. Usage is inferable rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
undo_writesUndo recent writesADestructive
Step the last writes back, newest first, each file restored to the bytes it had before that write — which is what Ctrl+Z stands in for here, since the editor keeps its undo stack in a process this server cannot reach. A write that created a file removes it again. Pass steps to undo several, or since with the index a previous write_history reported to land exactly back at that point whatever was written in between; check write_history first, because one tool call can touch more than one file and a half-undone operation is worse than the mistake you meant to take back.
| Name | Required | Description | Default |
|---|---|---|---|
| since | No | Undo back to this journal length, the `index` write_history reported, instead of counting steps | |
| steps | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true and idempotentHint=false, so the bar is lower, but the description adds real substance: it discloses that undoing a file-creating write deletes the file, and warns that a half-undone operation is worse than the original mistake. It does not cover error behavior or what happens if the journal is shorter than requested.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Core purpose is front-loaded in the first clause, and the Ctrl+Z aside earns its place by explaining why a server-side undo exists at all. The remaining text is one dense, comma-chained block that takes effort to parse, but there is little outright filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, so the description carries the return-behavior burden, and it does tell the agent what state results (files restored, created files removed). Combined with the `write_history` prerequisite and the multi-file hazard warning, an agent has enough to call it correctly, though failure modes and the neither-parameter default remain unstated.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 50% (`steps` is undocumented in the schema), and the description compensates by explaining both parameters: `steps` as a count of writes to reverse, `since` as a journal index from `write_history` that lands exactly at a prior point regardless of intervening writes. It doesn't state the default when neither is supplied, leaving the single-step default implicit in 'the last writes'.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('step the last writes back, newest first, each file restored to the bytes it had before that write') and pins down the scope of what 'undo' means here. The Ctrl+Z analogy plus the note that the server cannot reach the editor's undo stack cleanly separates it from rollback_data and list_backups in the sibling set.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly names both control paths and the condition that selects each: `steps` to undo several, `since` with the `index` a previous `write_history` reported to land exactly at a point 'whatever was written in between'. It also states the prerequisite ('check `write_history` first') and the reason (one call can touch more than one file).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
validate_gameAudit the game before playing itARead-only
Read-only check of the things that only show up in play. System: the start position on a map that exists, inside it, on a cell the player can stand, with a party of at least one named actor; the switch and variable tables in the array shape the engine indexes, and every id an event names that is past their end; and the keys the engine dereferences without a fallback (advanced.windowOpacity and friends), which is how a project copied from the shipped data/newdata template turns out to die on its own title screen. Events: transfer destinations (checked against the destination map's own size and passability), every database id a command reads, encounter rows (a real troop, a regionSet that is painted, a non-zero weight), pages whose indent does not match the branch they sit under, and the graphics and audio the project does not have — including the picture a Show Picture command names, which nothing else reads, and which only ever loads as <name>.png. Also tile ids painted from a slot the tileset has no image for, and reachability from the start through the portals that exist. Each problem comes back as {severity, where, what, fix} and fix names the call that clears it. This is the gate a build script should pass before a playtest starts, and it never writes.
| Name | Required | Description | Default |
|---|---|---|---|
| from | No | Where reachability starts; default System's start position | |
| audio | No | Also check referenced audio files (default true) | |
| limit | No | Cap on reported problems (default 120) | |
| mapIds | No | Restrict the walk to these maps (default: every map in the project) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With readOnlyHint already declaring safety, the description still adds real behavioral context: it 'never writes,' it returns each problem as {severity, where, what, fix}, and `fix` names the call that clears it. It also explains that a project copied from data/newdata can die on its title screen, giving the agent a model of what failure looks like.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the one-line purpose before the detail and organizes the enumeration with 'System:' and 'Events:' labels, so it is navigable despite its length. The clauses are dense and run long, but for a validator covering this many checks most of them carry information an agent needs.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, so the description carries the return-shape burden and does so ({severity, where, what, fix}), plus read-only semantics and the playtest-gate use case. For a complex 4-param audit tool, nothing essential is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents `from`, `audio`, `limit`, and `mapIds`. The description alludes to start-position reachability and audio checking but adds no syntax, defaults, or format detail beyond the schema, so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Opens with a specific verb+resource ('Read-only check of the things that only show up in play') and then enumerates the exact classes of defects audited (System start/switch-variable tables/missing engine keys, transfer destinations, encounter rows, page indent, missing graphics/audio, tileset slots, reachability). This is clearly distinguishable from siblings like check_assets, map_connectivity, or fix_project.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly positions the tool as 'the gate a build script should pass before a playtest starts,' which gives clear when-to-use context, and implies it precedes fix_project by noting each fix names the clearing call. It does not, however, explicitly contrast itself with sibling validators such as check_assets or map_connectivity, so the routing guidance is strong but not exhaustive.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
write_historyList this session's writesARead-only
Every file this server process has written, newest first, with the backup each one can be reverted to. This is the ledger undo_writes steps through; it lives in memory, so it starts empty when the server restarts even though the backup files are still on disk (list_backups sees those).
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare readOnlyHint=true; the description adds the crucial behavioral facts that the ledger is in-memory, starts empty after a server restart, and that backups persist on disk anyway. It also discloses ordering (newest first) and the backup-per-entry payload, which no structured field conveys.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the core purpose and scope before the memory/persistence caveat. Every clause carries load: ordering, revertable backups, relationship to undo_writes, and the restart caveat.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description usefully describes the return shape (files plus their backups, newest first) and the persistence caveat. The only gap is that the limit parameter governing result size is left unexplained.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The single 'limit' parameter has 0% schema description coverage and is not mentioned anywhere in the description, so there is no guidance on its purpose, default, or 1-200 bound. Only the implicit 'newest first' hints at ordering; the parameter itself is undocumented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: every file this server process has written, ordered newest first, each with its revertable backup. It explicitly differentiates itself from siblings, calling out that it is the ledger undo_writes steps through and that list_backups sees the on-disk files instead.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Clearly positions the tool against undo_writes (which steps through this ledger) and list_backups (which reads backups on disk), so the agent can infer when each is appropriate. It stops short of an explicit 'use this when / do not use when' statement, but the routing context is strong.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
write_plugin_sourceWrite a plugin's sourceADestructive
Create or replace js/plugins/<name>.js. This is Unity's manage_script equivalent, with one difference that matters for the loop: MZ plugins are plain JavaScript loaded at boot, so there is no compile step to wait for — the file is live the moment a game starts. The write is journaled, so undo_writes takes it back, and js/plugins.js is not touched unless enable is set. With enable the plugin is added to the editor's list (or switched on if it is already there) and its parameters filled from the @default values in the header it was just given, which means the header has to be written first for the parameters to be anything. A game already running will not pick any of this up: restart it.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | File name without extension; no path separators | |
| text | Yes | ||
| enable | No | ||
| parameters | No | Parameter values to store when enabling, overriding the header defaults |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare destructive=true and idempotent=false; the description adds substantial beyond that: the write is journaled and reversible via undo_writes, js/plugins.js is untouched unless enable is set, enable adds/switches the plugin and fills parameters from @default header values, the header must be written first, and a running game requires a restart. This is rich, non-obvious behavioral disclosure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the action and target path, then organized around the gotchas (no compile step, journaling, enable behavior, restart). Dense but nearly every clause carries operational signal; slightly long, though little is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive multi-parameter write with no output schema and a nested parameters object, the description covers reversibility, side effects on js/plugins.js, the header-order dependency, and the restart requirement. It could be marginally clearer that text is the complete file contents, but the essentials for correct invocation are present.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 50% (text and enable are undescribed), so the description carries real weight: it defines the enable semantics (list insertion vs. toggling on) and describes how parameters are populated from header defaults and overridden. name and text are inferable from the path reference and the 'source file' framing, giving an overall above-baseline lift.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Opens with a specific verb+resource and exact target path: 'Create or replace `js/plugins/<name>.js`'. An agent immediately knows this is a full-file write of a plugin source, distinct from read_plugin_source or patch_plugin. It also anchors the concept by comparing to Unity's manage_script.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives strong operating context ('MZ plugins are plain JavaScript loaded at boot, so there is no compile step') and explains the enable-to-list path, but never routes the agent explicitly between this and siblings like patch_plugin, enable_plugin, or read_plugin_source. Usage is implied rather than stated as when/when-not.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
65 tool updates
v0.4.2- First observed
add_commands - First observed
assert_in_game - First observed
batch - First observed
block_structure - First observed
check_assets - First observed
clear_events - First observed
command_catalog - First observed
copy_event - First observed
create_database_entry - First observed
create_map - First observed
decode_commands - First observed
delete_map - First observed
describe_tiles - First observed
enable_plugin - First observed
find_events - First observed
fix_project - First observed
get_map - First observed
import_asset - First observed
inspect_cell - First observed
link_maps - First observed
list_backups - First observed
list_maps - First observed
list_plugins - First observed
live_diagnostics - First observed
live_dialog - First observed
live_eval - First observed
live_key - First observed
live_move - First observed
live_pause - First observed
live_reload - First observed
live_screenshot - First observed
live_session - First observed
live_status - First observed
live_step - First observed
live_wait - First observed
make_battle - First observed
make_chest - First observed
make_choice_scene - First observed
make_encounter_zone - First observed
make_item - First observed
make_map - First observed
make_npc - First observed
make_shop - First observed
map_connectivity - First observed
patch_database_entry - First observed
patch_plugin - First observed
place_event - First observed
project_info - First observed
read_database - First observed
read_plugin_source - First observed
remove_event - First observed
render_map - First observed
rollback_data - First observed
set_commands - First observed
set_event_page - First observed
set_map_properties - First observed
set_startup - First observed
set_tiles - First observed
set_tileset_flags - First observed
show_text - First observed
tileset_slots - First observed
undo_writes - First observed
validate_game - First observed
write_history - First observed
write_plugin_source
TDQS
Scored across 65 tools
While each tool targets a distinct operation, the server has layered overlapping abstractions that are hard to tell apart: create_map vs make_map, set_tiles vs set_tileset_flags, place_event vs make_npc vs make_choice_scene, make_chest, link_maps, make_shop, make_encounter_zone, make_battle, and make_item all ultimately write events or database rows; add_commands vs write_plugin_source vs set_commands vs make_choice_scene write scripts; and many live_* tools overlap (live_eval, live_key, live_move, live_step). An agent choosing between the high-level builders and the low-level primitives has no clear rulebook.
Mostly snake_case is consistent, but verb styles are mixed: imperative verbs (get, set, create, delete, list, add, remove, copy, place, clear), 'make_' builders (make_map, make_npc, make_battle, make_item, make_choice_scene, make_chest, make_shop, make_encounter_zone), and nouns (project_info, command_catalog, block_structure, tile_slots) coexist, with several live_* tools. The pattern is readable but not predictable.
65 tools is far beyond typical MCP guidance and forces an agent to navigate layers of abstractions for the same underlying write operations. The high-level builders could collapse many of the primitive tools, reducing cognitive load.
The surface covers the RPG Maker MZ lifecycle end to end: map creation, tile painting, event scripting, database CRUD, plugin management, asset validation, backup/undo, validation, and live runtime control. No obvious domain gap remains; even asset import and rollback are present.
Maintenance
Related MCP Connectors
AI game assets for agents: consistent sprites, 2D animations, tiles, maps, music and engine exports.
Nifty's MCP server — exposes tasks, projects, messages, and files as tools for AI agents.
Your org's AI agents, tasks, runs, search, and brain files as MCP tools and resources.
Drive a live Cinevva game session: edit game files, import CC0 assets, preview changes.
Related MCP Servers
- AlicenseBqualityDmaintenanceEnables AI agents to directly manipulate RPG Maker MZ projects through natural language commands, allowing creation and modification of game assets like items, weapons, enemies, maps, and plugins without manually editing game files.36140 npm2MIT
- AlicenseCqualityDmaintenanceEnables AI models to develop and automate RPG Maker MZ projects by creating maps, events, and plugins through natural language commands. It provides comprehensive tools for database management, asset integrity checks, and direct map tile manipulation.2810 npm1ISC
- AlicenseAqualityDmaintenanceEnables creating RPG Maker MZ games using natural language through AI assistance, with tools for project management, map creation, event systems, and batch operations.10140 npm1MIT
- AlicenseBqualityDmaintenanceEnables AI assistants to act as co-developers for RPG Maker MV projects, providing full database CRUD, map and event editing, plugin management, playtest control, and automatic backups.4151 npm2MIT