Skip to main content
Glama

Run a list of assertions in the running game

assert_in_game

Runs truthy expression assertions in a live RPG Maker MZ playtest and returns a pass/fail report. Polls until dialogs, transfers, or battles complete instead of requiring a throwaway script.

Instructions

Hand over a list of {label, expression} and get a pass/fail report back, the way a test runner does it, instead of writing a throwaway script for every playtest. Each expression is evaluated in game scope and has to be truthy; give one a timeoutMs to poll it until it holds, which is what a dialog opening, a transfer landing or a battle ending needs. A failing assertion does not stop the rest unless stopOnFailure is set, and whatever the game logged while the list ran comes back with the result, so a failure arrives with its own reason attached. Needs Allow Eval on the plugin.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
pollMsNo
assertionsYes
stopOnFailureNo

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.4.2

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only supply idempotentHint=false, so the description carries most of the burden and delivers: failure does not abort remaining assertions unless stopOnFailure is set, the game's log is returned alongside results, and it declares the 'Allow Eval' prerequisite. It omits polling cadence (pollMs default) and any throughput/rate limits.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core mechanism and then layered with behavior and prerequisites. It is a dense paragraph but every sentence adds information; it could be trimmed slightly but wastes little.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 3-parameter, no-output-schema tool it covers the essentials: what goes in, what comes back (pass/fail report plus logs), failure handling, and the Allow Eval prerequisite. Only pollMs behavior and the 60-item cap remain unexplained.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Top-level schema coverage is 0%, so the description must compensate, and it explains timeoutMs (poll until truthy) and stopOnFailure meaningfully. But pollMs is never mentioned in the description and the nested label/expression semantics are largely left to the schema's inline descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: 'Hand over a list of {label, expression} and get a pass/fail report back, the way a test runner does it.' This clearly distinguishes it from siblings like live_eval (single evaluation), live_wait, and validate_game. An agent can identify it as the batched in-game assertion tool without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives strong context: 'instead of writing a throwaway script for every playtest' and concrete scenarios ('a dialog opening, a transfer landing or a battle ending'). However, it does not explicitly name sibling alternatives such as live_wait or live_eval and when to prefer them, so it stops short of the top score.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.