Skip to main content
Glama

Test & verify

test

Run scripted gameplay scenarios, lint projects, execute GUT/gdUnit4 suites, and compare screenshot baselines to verify Godot games and catch regressions.

Instructions

Verify the game works. Scenarios are scripted playtests (play → input → wait → assert → screenshot) that you can save under res://tests/scenarios and re-run as regression tests. Also project lint (broken references, missing shapes/textures, script errors), GUT/gdUnit4 unit tests, and screenshot baselines.

Actions:

  • scenario: {scenario: {name?, scene?, steps: [...], keep_running?, stop_on_failure?=true, mute?=true}, save_as?: 'name'} run a playtest (audio muted unless mute=false). Steps: {wait_for: expr} {wait_node: path} {wait_signal: {path, signal}} {wait: s} {input: [input steps]} {assert: expr, message?} {expect_node: path} {expect_no_node: path} {eval: expr} {call: {path, method, args}} {set: {path, props}} {sample: {...}} {screenshot: label} {compare_baseline: 'res://tests/baselines/x.png'} {assert_no_errors: true} {time_scale: n}. Expressions run with the scene root as self.

  • run_all: {} run every saved scenario in res://tests/scenarios.

  • list: {} saved scenarios and detected unit test frameworks.

  • lint: {path?} check all scenes (missing dependencies, broken refs, missing collision shapes/textures) and all scripts (compile errors). Run before playing.

  • unit: {framework?: 'gut'|'gdunit4', dir?='res://test', filter?} run unit tests headless.

  • compare_image: {image_path | image_b64, baseline, threshold?=0.01, update?} compare against a baseline PNG (records it if missing).

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
dirNoTest directory.
nameNoSaved scenario name (res://tests/scenarios/<name>.json).
actionYesWhat to do. See the tool description for each action's parameters.
save_asNoAlso save the scenario under this name.
baselineNoBaseline PNG path.
scenarioNoScenario object {name?, scene?, steps: [...]}.
frameworkNo'gut' or 'gdunit4'.

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.1.0

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare non-read-only, non-destructive, closed-world, so the safety profile is partly covered. The description still adds real behavioral context the annotations do not: audio is muted by default ('audio muted unless mute=false'), stop_on_failure defaults to true, scenarios can be persisted via save_as, and compare_image records a baseline if missing. These are meaningful side-effect disclosures.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The opening summary is front-loaded, and the per-action bullet list with nested step types is dense but well organized. The exhaustive step enumeration is lengthy yet justified given the schema treats steps as free-form; only minor trimming is possible.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a broad test/verify tool with no output schema, the description never explains what a run returns — pass/fail results, assertion output, error reporting, or how a failed scenario surfaces. The input side is thoroughly covered, but the result side is a notable gap for a verification tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds substantial meaning beyond the schema — notably the entire step sub-language (wait_for, wait_node, wait_signal, assert, expect_node, screenshot, etc.) and the note that expressions run with the scene root as self. This elaborates the opaque `scenario` object that the schema only marks as additionalProperties.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a concrete verb+resource ('Verify the game works') and then enumerates each of the six actions with what it does (scripted playtests, project lint, unit tests, baseline comparison). It is clearly distinguishable from siblings like `run`, `game`, and `input`. It falls short of 5 only because it is a kitchen-sink tool bundling several distinct capabilities rather than one focused function.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is one sequencing hint ('Run before playing' for lint) and scenario explanation, but no explicit when-to-use/when-not guidance and no routing to alternatives among the many siblings. Usage is largely implied by the action names rather than stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.