emptysock-mcp
The server exposes MCP tools for working with EmptySock game projects: managing save files, inspecting GMS2 projects, reading story graphs, validating VisualScript graphs, estimating battle damage, and querying a connected live game's physics, scene, actor, and navmesh systems over a local WebSocket bridge.
Save files: read, write, delete, and list sandboxed JSON save slots under
SAVE_BASE_DIR.GMS2: inspect a
.yypproject to get project name, asset counts, object names, and script names;emptysock_layer_inforeturns LayerSystem reference docs.Story Graph: export and parse a scene's
.storyGraph.jsoninto nodes, edges, and start node.VisualScript: structurally validate node graphs — duplicate IDs, dangling connections, unreachable nodes, and missing branch false-paths.
Battle: estimate physical damage using the default
BattleSystemformula.Live game bridge (requires a connected game): navmesh pathfinding/nearest node, 2D physics raycast/overlap/body state, scene entity listing/info/component retrieval/entity creation, and actor messaging/broadcast/inbox/list.
Particle emitter config: a stub — reads/writes a default-merged config but does not touch a live emitter.
Security/ops: Zod input validation, path-traversal protection, rate limiting, and JSON audit logging to stderr.
Caveat: although the schema lists
physics_raycast_3d, the README explicitly says it is not actually registered as a tool.
Provides read-only integration with GameMaker Studio 2, allowing inspection of .yyp project files and returning a JSON summary with project name, asset counts, and lists of object and script names.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@emptysock-mcpFind a path from (10,10) to (50,50) on map level1."
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
emptysock-mcp
A Model Context Protocol server for the EmptySock game engine. It hands Claude Desktop, AI agents, and the Claude API a set of tools for poking at an EmptySock project: reading save files, validating Story Graphs and VisualScript graphs, importing GameMaker Studio 2 projects, estimating battle damage, and (eventually, see below) querying a live running game's physics and scene state.
Some of these tools do real work against real files on disk today. A few are honest placeholders waiting on a live connection to an actual game process. The table below tells you which is which — no tool here pretends to be more finished than it is.
Requirements
Node.js 20+
npm 9+
Related MCP server: Hayba
Installation
git clone https://github.com/eleferrets/emptysock-mcp.git
cd emptysock-mcp
npm install
npm run buildConfiguration
Copy the example env file and fill in whatever you need:
cp .env.example .envVariable | Required | What it does |
| No | The one directory the save tools are allowed to touch. Everything |
| No | The directory project asset files live under. |
| No | How many calls a single tool can take in one rate-limit window before it starts saying no. Default |
| No | The length of that window, in milliseconds. Default |
| No | Port the live-bridge WebSocket server binds to on |
Never commit
.env. It's gitignored for a reason — keep secrets in your CI/CD secret manager instead.
Running the server
stdio (recommended for local use and Claude Desktop)
npm run dev # development — tsx, no build step
# or after building:
node dist/server.jsThe server talks over stdin/stdout. There's no network port and no auth layer to configure, because there's nothing listening for anyone to break into.
Claude Desktop
Add the server to your Claude Desktop config (~/Library/Application Support/Claude/claude_desktop_config.json on macOS):
{
"mcpServers": {
"emptysock": {
"command": "node",
"args": ["/absolute/path/to/emptysock-mcp/dist/server.js"],
"env": {
"SAVE_BASE_DIR": "/absolute/path/to/your/saves"
}
}
}
}Restart Claude Desktop and the EmptySock tools show up in the tool picker.
What this server actually talks to (read this before assuming a tool does more than it does)
This server hosts a plain WebSocket bridge on 127.0.0.1:EMPTYSOCK_BRIDGE_PORT (default 7777, see lib/bridge.ts) that a live game or the IDE preview dials into as the client. Once connected, tool calls relay real EngineQuery/EngineQueryResult envelopes to @emptysock/engine's QueryChannel (packages/engine/src/bridge/QueryChannel.ts) running inside that live process and return its real answer — not a guess. No auth token: this is a localhost-only trust model, matching how this server already treats local file I/O. Only one live game is expected connected at a time; a newer connection replaces an older one rather than being queued.
What that means in practice, split by domain:
Physics, Scene, NavMesh, and Actor tools (
physics_*,scene_*,navmesh_*,actor_*) are all live: they relay real queries to the connected live game'sQueryChanneland return its real answer. With no live game connected, they return{ ok: false, error: { code: "no-live-instance", ... } }. Beyond that, each domain has its own honest "attached but not that system" error: physics queries against a scene with no initializedPhysicsSystemget"no-physics-world";actor_*against a scene with noActorSystemattached gets"no-actor-system";navmesh_*against a scene with no navmesh attached gets"no-navmesh". None of these ever collapse into a fabricated empty result — anavmesh_find_pathcall that finds no route reportspath: null, not an error, and is distinguishable from every one of the three "nothing to even ask" cases above.Save, GMS2 import, Story Graph export, and VisualScript validation are all real, working tools that operate on static project files on disk (save JSON,
.yyp/.yyfiles,.storyGraph.jsonfiles, aVisualScriptGraphpayload you hand it directly). No live game required, because none of these ever needed one.Battle damage estimation is a real, working, pure calculation — it reimplements
BattleSystem's default physical damage formula rather than driving a liveBattleSysteminstance, because that instance is a stateful turn machine meant to run inside a real game loop, not something this server has any business owning.emptysock_layer_infois reference documentation served as a tool response, not a stub — there's nothing to fake here, it's just handing back API docs.
One more note on terminology: the engine's recent module-package split moved VN, battle, and tilemap logic out of core @emptysock/engine into their own packages (@emptysock/vn, @emptysock/battle, @emptysock/tilemap). This server is a separate Node process and doesn't import the engine at all, so none of that split changes any code here — but the story_graph_export, battle_estimate_damage, and navmesh_* tools below are documented against the current package names (VNSystem now lives in @emptysock/vn, BattleSystem in @emptysock/battle, NavMeshSystem in @emptysock/tilemap) so you know which package's shape you're actually matching.
Security model, in plain terms
Path safety. Every filesystem path a caller supplies goes through a
SafeRelPathZod schema (no.., no leading/or\) and then gets resolved and checked again against the allowed base directory before any file touches disk. Two independent checks, because "the schema already rejected it" is a bad place to stop trusting yourself.Rate limiting. A sliding-window limiter sits in front of every tool call, keyed by tool name, tuned by
RATE_LIMIT_MAX/RATE_LIMIT_WINDOW_MS. It exists to stop a runaway agent loop from calling the same tool a thousand times in ten seconds, not to defend against a hostile network attacker (there is no network surface to attack).Audit logging. Every tool call gets logged as one line of JSON to stderr — tool name, status, duration. Never to stdout, because stdout is the actual MCP wire protocol and mixing log lines into it would corrupt every message after.
Input validation. Every tool argument is parsed with Zod before anything happens. Malformed input gets a clean
InvalidParamserror, not a stack trace or a half-executed side effect.No credentials, no secrets. There's no auth surface to configure because stdio has exactly one caller: the MCP host process that spawned this one.
Available tools
Status key: live — does real work (either standalone, or by relaying to a connected live game over the bridge). stub — validates input correctly but has no live-bridge query kind to call yet, so it always returns an honest ok: false error. docs — returns static reference information by design, not a stub.
NavMesh
Tool | Status | Parameters | Returns |
| live |
|
|
| live |
|
|
Relays to QueryChannel's navmeshFindPath/navmeshNearestNode query kinds, which reach an optional NavMeshQuerySource (@emptysock/tilemap's NavMeshSystem, wired in only if the live game's scene actually attached one) — a scene with no navmesh loaded is normal, not an error, and reports "no-navmesh" rather than "not-found".
Example call — find path:
{ "from": { "x": 0, "y": 0 }, "to": { "x": 100, "y": 50 }, "mapId": "level1" }Physics
Tool | Status | Parameters | Returns |
| live |
|
|
| live |
|
|
| live |
|
|
There's no physics_raycast_3d tool in this registry. 3D raycasting isn't offered at all right now, not even as a stub — implementing it needs a Rapier3D WASM build available server-side, which this server doesn't have (the engine's 3D physics runs in the browser's WASM context, not in Node). Don't call it expecting an error message with useful detail; you'll just get an unknown-tool error like any other made-up tool name.
Example call — overlap circle:
{ "center": { "x": 50, "y": 50 }, "radius": 20, "layerMask": 3 }Scene
Tool | Status | Parameters | Returns |
| live |
|
|
| live |
|
|
| live |
|
|
| live |
|
|
Example call — get component:
{ "sceneId": "gameplay", "entityId": "player-001", "componentType": "Transform" }Save
All save tools are sandboxed to SAVE_BASE_DIR. Path traversal (.., absolute paths) is rejected both at the schema layer and again when the path is resolved.
Slots are read and written using the engine's default GameSaveSlot shape ({ id, scene, data, timestamp, playtime }, from SaveSystem in @emptysock/engine). SaveSystem itself is generic over any Zod schema you construct it with, but these tools only speak the default shape — there's no way to carry an arbitrary Zod schema over MCP's JSON-RPC wire, so a game using a custom SaveSystem<TSlot> shape should treat these as opaque JSON storage rather than relying on the auto-filled id/timestamp/playtime convenience.
Tool | Status | Parameters | Returns |
| live |
| The |
| live |
|
|
| live |
|
|
| live |
|
|
Example call — write:
{ "slot": "autosave", "scene": "bridge-level", "data": { "level": 3, "score": 4200, "checkpoint": "bridge" } }Slot names are alphanumeric plus dashes/underscores only (slot1, autosave, new-game-plus).
Actor
Tool | Status | Parameters | Returns |
| live |
|
|
| live |
|
|
| live |
|
|
| live | none |
|
Relays to QueryChannel's actorSendMessage/actorBroadcast/actorInboxSize/actorList query kinds, which reach the live game's real ActorSystem directly. A scene can legitimately be live with zero actors running, so a missing ActorSystem is its own error, "no-actor-system", never "no-live-instance". The same ordering guarantee ActorSystem uses everywhere else applies: it drains every actor's inbox before calling update() on any actor, so a message sent during frame N is fully processed before frame N's update() logic runs.
Example call — send message:
{ "actorId": "enemy-spawner", "message": { "type": "SPAWN_WAVE", "payload": { "wave": 3 } } }GMS2
Tool | Status | Parameters | Returns |
| live |
|
|
| docs | none | A reference document for the |
gms2_inspect_project actually parses a real GameMaker Studio 2 project: it tolerates the trailing commas GameMaker's IDE always writes (not strict JSON) and reads the project's display name from the real "%Name" key rather than a "name" field that doesn't exist there.
Particles
Tool | Status | Parameters | Returns |
| stub |
| Reading returns a fixed default config merged with nothing real; writing returns |
Example call — read config:
{ "emitterId": "dust" }Example call — write config:
{
"emitterId": "dust",
"config": {
"emissionRate": 60,
"lifetime": { "min": 0.5, "max": 1.2 },
"shape": "circle",
"shapeRadius": 20,
"colorGradient": [16766464, 16711680]
}
}Story Graph
Tool | Status | Parameters | Returns |
| live |
| The parsed Story Graph — |
This is the Story Graph that @emptysock/vn's VNSystem runs at play time. Node types are "dialogue", "choice", and the VariableStore-driven "condition" type. Choice nodes carry their option labels in a plain options string array and, when any option is gated, a parallel optionWhens array of VariableCondition | null. A condition node carries a single condition. Branch destinations — a choice option's target, or a condition node's true/false targets — live on the graph's edges (keyed by fromPort), never on the node itself.
Example call:
{ "sceneId": "chapter1", "graphId": "intro" }Battle
Tool | Status | Parameters | Returns |
| live |
|
|
This genuinely reimplements @emptysock/battle's BattleSystem default physical damage formula (max(1, floor((effectiveAttack - effectiveDefense / 2) * power * (isCrit ? critMultiplier : 1)))) rather than driving a live BattleSystem turn machine — that machine's real job (turn order, status effects, subscriptions) is meant to run inside an actual game, and reimplementing all of that here risks drifting out of sync with the real thing. Use this one for sanity-checking combat balance numbers, not for simulating an actual fight.
Example call:
{ "effectiveAttack": 50, "effectiveDefense": 20, "power": 1.2 }VisualScript
Tool | Status | Parameters | Returns |
| live |
|
|
This runs a real, complete structural check against a VisualScriptGraph (the node graph VisualScriptComponent interprets): duplicate node ids, dangling next/connection targets, nodes unreachable from an onUpdate/onEvent entry node, and branch nodes missing a false branch. It does not execute the graph against a live VariableStore or ActorSystem — that's a different, much bigger job — but everything it does check, it checks for real.
Example call:
{
"graph": {
"nodes": [
{ "id": "update-1", "kind": "onUpdate", "next": ["branch-1"] },
{ "id": "branch-1", "kind": "branch", "next": ["set-1"], "variableIndex": 0, "comparator": "gt", "value": 5 },
{ "id": "set-1", "kind": "setSwitch", "next": [], "switchIndex": 0, "value": true }
],
"connections": [
{ "id": "c1", "from": "update-1", "to": "branch-1", "fromPort": 0 },
{ "id": "c2", "from": "branch-1", "to": "set-1", "fromPort": 0 }
]
}
}Development
npm run lint # TypeScript type-check (no emit)
npm test # run Vitest suite
npm run test:watch # watch modeTests live in src/tests/. They cover input validation, tool dispatch, and the security invariants that actually matter here (path traversal, unknown tool names, and confirming the stubs stay honest stubs).
Adding a tool
Create
src/tools/<domain>.ts— export atoolDefarray entry and ahandlerfunction.Register both in
src/tools/index.tsvia theregister()call inbuildRegistry().Add an entry to
api-reference.jsoninemptysock-engine.Add a skill file to
eleferrets/emptysock-ai-skills.Update this README's tools table. If it's a stub, say so plainly — don't let it look more finished than it is.
Shared helpers live in src/lib/:
parse(schema, raw)— Zod parse that throwsMcpError(InvalidParams)on failureSafeRelPath,SafeId,Vec2,Vec3,GameNum— reusable Zod schemastextResponse(data)— builds the standard MCP text content responsewrapError(err)— logs to stderr and re-throws asMcpError(InternalError)
Security model
Concern | Mitigation |
Malformed arguments | Zod |
Path traversal |
|
Shell injection | No |
Credential leakage | Secrets come from |
Oversized inputs | String lengths bounded on every schema field |
Unknown tools |
|
Available Tools
21 toolsactor_broadcastB
Broadcast a message to all registered actors in the current ActorSystem.
| Name | Required | Description | Default |
|---|---|---|---|
| message | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are present, so the description carries the full burden of behavioral disclosure. It only names the action and target; it does not reveal whether delivery is asynchronous or synchronous, how failures are handled, whether the message is copied per actor, or what side effects occur. This is a significant gap for a side-effecting operation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence with no filler. Every word contributes to identifying the operation and its scope, which is ideal for a tool of this simplicity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter tool with no annotations and no output schema, the description leaves important decisions to the agent: when to use broadcast versus targeted send, what message semantics are expected, and what the runtime behavior will be. The agent can infer the basic call shape but not the operational context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description adds no meaning beyond calling it 'a message'. It does not explain the required 'type' field, how 'payload' is used, or what message values are valid. With only one parameter, the description should compensate more for the schema's lack of detail.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action ('Broadcast'), the resource ('a message'), and the exact scope ('all registered actors in the current ActorSystem'). It is clear and implicitly contrasts with the sibling 'actor_send_message', which likely targets a single actor, though it does not explicitly name that alternative.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'all registered actors' implies the fan-out use case, and the sibling actor_send_message suggests a single-target alternative, but the description does not explicitly say when to choose this over targeted messaging or describe any prerequisites. Usage guidance is mostly left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
actor_inbox_sizeB
Return the number of pending messages in an actor's inbox.
| Name | Required | Description | Default |
|---|---|---|---|
| actorId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description must disclose behavioral traits. It states it returns a count but does not mention that it is a read-only operation, does not destroy or modify the inbox, or whether it has side effects. For a query tool, this is a minimal expectation. The description is too brief to disclose non-obvious behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, clear sentence with no verbosity issues. It is acceptably concise and front-loaded with the main action. However, it could use an additional sentence for usage guidance without becoming too long, so it earns a 4.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has a simple interface (one parameter, no output schema), but the description omits critical context: what the return value looks like (e.g., integer vs. formatted string), whether the inbox is cleared after reading, or any side effects like error conditions for invalid actorId. With no annotations or output schema, the description is insufficient for an agent to use the tool safely and correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has 0% coverage, meaning the description must compensate for the undocumented 'actorId' parameter. The description mentions 'an actor's inbox' which implies actorId identifies the actor, but it does not clarify the format or identifier type (e.g., numeric ID, string name). Baseline is 3 because the description provides some context beyond the schema, but not enough to fully clarify the parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states that the tool returns the count of pending messages in an actor's inbox. It uses a specific verb ('Return') and resource ('number of pending messages in an actor's inbox'), which is distinct from siblings like actor_send_message and actor_broadcast that send messages.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description does not explicitly state when to use this tool versus alternatives, nor does it mention any prerequisites (e.g., 'Use this to check if an actor has pending messages before processing'). It only implies usage by saying 'Return the count'.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
actor_send_messageA
Enqueue a message in a specific actor's inbox. The actor processes it on the next flush pass.
| Name | Required | Description | Default |
|---|---|---|---|
| actorId | Yes | ||
| message | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations to rely on, the description carries the transparency burden. It discloses two behavioral traits: the message is queued rather than synchronously handled, and processing occurs on the next flush pass. It stops short of describing error behavior or whether the message is delivered at most once, but the core side effect is clear.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, front-loaded with the action, with no filler or repeated schema information. Every word contributes meaning.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description is sufficient for a low-complexity side-effect tool: it states the action, target, and asynchronous processing behavior. There is no output schema to document, though a note on failure handling or message ordering would make it fully complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema parameter coverage is 0%, so the description must explain the parameters beyond their names. 'actorId' is loosely implied by 'specific actor' and 'message' by the verb 'enqueue', but the description adds no semantics about required message type, payload constraints, or behavior when the actor doesn't exist.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb ('Enqueue') and clearly identifies the resource ('a specific actor's inbox'), making the operation's intent unambiguous. The phrase 'specific actor' also implicitly differentiates it from broadcast-style messaging among the sibling tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It states that the message is enqueued and later processed on the next flush pass, which gives useful context for when this tool is appropriate. However, it does not directly contrast with actor_broadcast or other sibling tools, so the agent must infer the distinguishing conditions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
emptysock_layer_infoA
Return reference documentation for the EmptySock LayerSystem API: defineLayer, addToLayer, setVisible, setParallax, and RenderSystem integration.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are present, so the description carries the full burden. 'Return reference documentation' clearly indicates a read-only information-retrieval operation with no side effects weeds, and the function list further sets expectations about the content. It does not describe the format or size of the returned documentation, but the behavior itself is clear.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, front-loaded sentence states the verb ('Return'), the object ('reference documentation'), and the exact scope (EmptySock LayerSystem APIs). No redundant text or repetition of the tool name.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter, read-only documentation tool, the description tells an agent what it will get and which APIs are covered. It stops short of describing the result layout or how the reference material is presented, but nothing critical is missing for selecting and invoking such a simple tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters)Skip, so the schema already fully describes the input surface (none). The description adds no parameter-specific meaning, but none is needed. The baseline for zero-parameter tools is 4.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action ('Return reference documentation') and a precise subject (EmptySock LayerSystem API) with the exact list of covered functions (setVisible, setParallax, etc.). It is immediately distinguishable from all sibling tools, which perform scene/save/physics operations rather than returning API reference material.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The purpose implicitly tells an agent when this tool is useful — when EmptySock layer API reference documentation is needed — and the sibling context makes alternatives obvious. However, there is no explicit guidance about when to prefer this tool over, say, gms2_inspect_project or other documentation-capable siblings, and no exclusions are stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
gms2_inspect_projectA
Read a GameMaker Studio 2 .yyp project file and return a JSON summary: project name, asset counts by type, and lists of object and script names. Read-only — does not modify any files.
| Name | Required | Description | Default |
|---|---|---|---|
| yypPath | Yes | Absolute or relative path to the .yyp project file. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations are absent, so the description carries the full burden of disclosure; it explicitly states 'Read-only — does not modify any files,' which is the key trait an agent needs before invocation. It also discloses the JSON summary format of the return value. Missing finer error-case behavior (e.g., malformed or missing project files), but for a simple read tool the core behavioral profile is covered.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, zero filler. The opening verb is front-loaded, the return contract is stated in one clause, and the safety disclosure is appended concisely. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter, annotation-free, output-schema-free tool, the description covers the operation, the result shape, and the read-only guarantee. The only notable gap is error handling behavior for missing or invalid .yyp files, which is a minor omission for such a low-complexity tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% for the single yypPath parameter, so the schema already documents path semantics; the description restates 'project file' without adding new detail about allowed path forms or constraints. Baseline 3 applies because coverage is high and the description does not meaningfully extend parameter meaning.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb (Read), a specific resource (GameMaker Studio 2 .yyp project file), and enumerates the return payload (project name, asset counts by type, object/script names). No sibling tool targets GameMaker project inspection, so it is cleanly distinguished from the navmesh/physics/scene/save/actor siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied clearly: the agent should call this when it needs an overview of a GameMaker project file. However, the description never explicitly says when to use it versus alternatives or when not to use it, so there is no direct guidance beyond the purpose statement.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
particle_emitter_configA
Get or set a ParticleSystem emitter configuration by emitter ID. Omit config to read; provide config to write.
| Name | Required | Description | Default |
|---|---|---|---|
| config | No | ||
| emitterId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are present, so the description carries the full burden of behavioral disclosure. It does disclose the dual read/write behavior and the exact trigger for writing, which is useful. However, it does not describe the return value, whether writing merges or replaces the existing configuration, or how errors like unknown emitter IDs are handled.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is just two sentences with no filler, and the core purpose ('Get or set') is front-loaded. Every clause contributes essential guidance about how to invoke the tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers mode selection and resource identification, which is enough for a basic invocation. But with no output schema and no annotations, it omits important contextual details like what a read returns, what a write returns, and whether the write is a partial update or a full replacement, leaving some ambiguity about side effects.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It adds the key semantic that `config` is optional and controls read vs write mode, and the schema property names are largely self-explanatory (emitRate, maxParticles, etc.). Still, it leaves unspecified details like emitterId format, value ranges, units, and whether omitted config fields are preserved or reset on write.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses specific verbs ('get' or 'set') and a precise resource ('ParticleSystem emitter configuration' keyed by emitter ID). It clearly distinguishes both modes and is distinct from the sibling tools, which target other systems like physics, scenes, or save data.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly states when to read vs write: omit `config` to read, provide `config` to write. No overlapping sibling tool is mentioned, but the conditional instruction is clear enough for an agent to select the correct behavior without further context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
physics_body_stateA
Return the current position, velocity, and angular velocity of a physics body by entity ID.
| Name | Required | Description | Default |
|---|---|---|---|
| entityId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden of behavioral disclosure. It clearly indicates a read-style operation by saying 'Return', and it names the returned fields, but it does not mention failure behavior, coordinate units, or whether the entity must already have a physics body. This is adequate but incomplete.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, well-structured sentence with no filler. It front-loads the action and output, then states the input scope concisely.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple one-parameter getter with no output schema, the description covers the essential input and output. It lacks details about error cases, units, or coordinate systems, but the tool's complexity is low and the description is largely sufficient for an agent to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It only repeats the parameter name ('entityId') through the phrase 'by entity ID' and adds no detail about what IDs are valid, what entity types are expected, or how missing or invalid IDs are handled. The sole parameter is self-explanatory, but the description adds little semantic value.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Return') and a clear resource ('physics body'), and names the exact data returned: position, velocity, and angular velocity. It is sufficiently distinct from sibling tools like physics_raycast_2d/3d and navmesh tools, which serve different purposes.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance about when to choose this tool over alternatives, nor any mention of prerequisites or exclusions. The phrase 'by entity ID' implies an input requirement, but the description does not explain when this tool should be used versus scene_entity_info or other physics tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
physics_overlap_circleB
Return all entity IDs whose 2D colliders overlap a circle.
| Name | Required | Description | Default |
|---|---|---|---|
| center | Yes | ||
| radius | Yes | ||
| layerMask | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description makes the read-only selection behavior clear and notes the 2D collider scope. However, with no annotations, it omits details like whether touching colliders count, coordinate-space assumptions, or any side effects (likely none).
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One tightly worded, front-loaded sentence. No filler or restatement of the tool name; every word adds meaning.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
While the core behavior is clear, the description omits important context for correct use: layerMask semantics, coordinate space, whether touching colliders count, and what happens when nothing overlaps. With no annotations or output schema, an agent has limited guidance beyond the basic query shape.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema descriptions are 0%, so the description must carry semantic weight. It conveys that center/radius define the circle, but gives no meaning or usage guidance for the optional layerMask parameter and does not explain units or how center coordinates are interpreted.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Return'), a specific geometric query ('2D colliders overlap a circle'), and the output ('entity IDs'). It is immediately distinguishable from sibling raycast/pathfinding tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to choose this tool over alternatives like physics_raycast_2d, navmesh_find_path, or scene queries. It neither lists use cases nor exclusions, so the agent gets no routing help.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
physics_raycast_2dB
Cast a ray in 2D physics space and return the first hit entity, hit point, and normal.
| Name | Required | Description | Default |
|---|---|---|---|
| origin | Yes | ||
| direction | Yes | ||
| layerMask | No | ||
| maxDistance | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the behavioral burden. It clearly discloses the core outcome: cast a ray and return the first hit entity, hit point, and normal. However, it does not disclose behavior on no-hit cases, whether the ray interacts with all physics objects by default, or any coordinate-space assumptions, so transparency is partial.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single concise, front-loaded sentence with no wasted words. It states the action and the key returned values immediately. It loses one point because its brevity causes it to skip important parameter and edge-case information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given four parameters, no annotations, no output schema, and 0% schema description coverage, this description is not complete enough. It names the key return values but leaves out no-hit behavior, layer mask semantics, maxDistance interpretation, and whether the raycast is read-only, which an agent needs to call the tool confidently.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description adds no meaning for any of the four parameters. It does not mention origin, direction, maxDistance, or layerMask, leaving the agent to rely entirely on the bare schema with no explanation of semantics, normalization, units, or how the layer mask affects results.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb and resource: cast a ray in 2D physics space and return the first hit entity, point, and normal. The '2D' qualifier clearly distinguishes this from the sibling physics_raycast_3d, so an agent can separate it by name and description.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The context is inferred from '2D physics space', which suggests when to use it versus physics_raycast_3d or navmesh tools, but there is no explicit statement about when to prefer this tool or when not to use it. The usage guidance is implied rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
physics_raycast_3dC
Cast a ray in 3D physics space (Rapier3D) and return the first hit.
| Name | Required | Description | Default |
|---|---|---|---|
| origin | Yes | ||
| direction | Yes | ||
| layerMask | No | ||
| maxDistance | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description must disclose behavioral traits. It states the return is 'the first hit,' implying the closest intersection, but does not mention behavior on a miss (e.g., null result), whether the layer mask filters hits, or if the ray is infinitely thin. It lacks essential behavioral context for an agent to know what to expect.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, short sentence with no filler. It front-loads the verb and direct object, making it immediately clear. Perfectly concise and efficiently structured.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with 4 parameters, nested objects, no output schema, and no annotations, the description is severely incomplete. It does not explain the return value structure, error handling, coordinate system, or any usage nuances. An agent would not know how to interpret the result or handle edge cases, making the tool impractical to use reliably.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for parameter meaning. It does not mention any parameter, such as the format of origin/direction vectors, the meaning of maxDistance (units), or how layerMask is interpreted. The parameter names are self-explanatory but the description adds no semantics beyond them.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'Cast a ray in 3D physics space (Rapier3D) and return the first hit.' It clearly differentiates from sibling tools like physics_raycast_2d by specifying 3D and the physics engine. The purpose is unambiguous and distinct.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus alternatives such as physics_raycast_2d or physics_overlap_circle. There are no explicit when-to-use, when-not-to-use, or alternative suggestions. The agent must infer usage from the name and description alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
save_deleteA
Delete a save slot from disk.
| Name | Required | Description | Default |
|---|---|---|---|
| slot | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description must carry the behavioral burden. It does disclose that the operation removes a save slot from persistent storage, which is useful for a destructive call, but it does not mention whether deletion is reversible, what happens when the slot does not exist, or what is returned.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
At seven words, every word contributes and the core action is front-loaded. There is no filler or repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive tool with no annotations and no output schema, this one-liner is incomplete: it leaves the agent guessing about slot format, success/error signaling, and failure behavior. It is enough to understand the intent but not enough to invoke the tool with confidence in all cases.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema only gives 'slot' as a string with no description, and schema description coverage is 0%. The description's mention of 'a save slot' weakly links the parameter to the resource, but it does not explain expected slot format, valid values, or how the slot is selected.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States the operation explicitly: 'Delete' is a specific verb, 'save slot' is the resource, and 'from disk' adds location. This clearly differentiates it from sibling tools like save_read, save_write, and save_list, which perform non-deletion operations.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No explicit when-to-use guidance or alternatives are provided; the only signal is the verb 'Delete', so usage must be inferred from intent. There are no conditions, prerequisites, or exclusions stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
save_listA
List all available save slots.
| Name | Required | Description | Default |
|---|---|---|---|
| subdir | No | Optional sub-directory within save base (no traversal) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
There are no annotations, so the description carries the behavioral disclosure burden. 'List' clearly implies a read-only enumeration and no destructive side effects, but the description does not state what the returned items look like, whether subdir changes the scope, or how this relates to other save operations. This is adequate but minimal.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single front-loaded sentence with no filler. Every word—'List', 'all', 'available', 'save slots'—contributes to understanding the tool's core behavior.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's low complexity (one optional parameter, no output schema), the description conveys the high-level purpose well. However, it leaves the return value's structure unspecified and does not mention that subdir can restrict the listing, though that detail is present in the schema. This is a minor gap rather than a blocking one.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, with 'subdir' already documented as 'Optional sub-directory within save base (no traversal)'. The description adds no parameter-specific meaning, so the baseline score of 3 applies because the schema handles parameter documentation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('List') and a specific resource ('save slots'), making the tool's purpose immediately clear. It also distinguishes this tool from sibling save operations like save_read, save_write, and save_delete by framing it as an enumeration rather than an operation on a single save.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance about when to choose this tool over save_read, save_write, or save_delete. There is no mention of typical workflows such as scanning available slots before reading or writing, nor any explicit exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
save_readA
Read a save slot from disk and return its JSON data.
| Name | Required | Description | Default |
|---|---|---|---|
| slot | Yes | Alphanumeric slot name, e.g. "slot1" or "autosave" |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the behavioral burden. It does disclose that it reads from disk and returns JSON data, which implies a non-mutating operation. However, it does not mention error behavior for missing slots, whether the JSON is parsed or raw, or any side effects beyond reading.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single concise sentence, front-loads the verb and object, and includes the key output type (JSON data). No filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read tool with one documented parameterchers, the description covers the core behavior and output. It lacks notes on failure modes, whether the slot must already exist, or how the returned JSON is structured, which leaves some ambiguity for an agent calling this tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema already fully documents the slot parameter with an example, and the description references 'a save slot' without adding meaningful details beyond the schema. No additional parameter semantics are contributed.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Read') and resource ('a save slot from disk') and states the return value ('its JSON data'). This clearly distinguishes it from the sibling tools save_write, save_delete, and save_list.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is given about when to use this tool versus alternatives, nor any conditions or prerequisites. The description states what it does but does not explicitly frame it as the read counterpart to save_write/save_delete or mention when it is appropriate or inappropriate.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
save_writeA
Write JSON data to a save slot on disk.
| Name | Required | Description | Default |
|---|---|---|---|
| data | Yes | Arbitrary save data | |
| slot | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden for behavioral disclosure. It states the side effect of writing data to disk, but does not disclose whether the operation overwrites an existing save or creates a new one, whether it is atomic, or what failure modes exist (e.g., disk full, invalid slot). The description is minimally transparent about the outcome but lacks depth on error behavior and data durability.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single concise sentence that is front-loaded with the primary action and subject. It contains no wasted words and is efficient, though it could benefit from an additional sentence about slot semantics, but the brevity is not excessive.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a write operation with no annotations and no output schema, the description is minimalistic. It does not state whether the operation overwrites existing data, whether it creates a new save slot, or what the return value (if any) is. Given the tool's simplicity (2 parameters), the description is adequate for basic invocation but lacks details on overwrite behavior and return type, leaving some ambiguity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The 'data' parameter is described as 'Arbitrary save data', which adds a small layer of meaning, but the 'slot' parameter has no description in the schema, and the description does not clarify its format or constraints. However, since schema coverage is 50%, the description partially compensates by specifying the data type (JSON) and purpose (save slot), but it does not clarify slot semantics (e.g., numeric index vs string key). This is a moderate addition over the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a clear verb ('Write') and resource ('JSON data to a save slot on disk'), making its purpose unambiguous. The wording clearly identifies it as a save-write operation distinct from siblings like save_read, save_delete, and save_list.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description is implied to be the mechanism for writing JSON data to a save slot, but it does not explicitly state when to use it versus alternatives like save_update or save_overwrite if they existed. It does not mention prerequisites such as 'slot' format ('e.g., slot number or key') or whether the slot must exist. The context clarifies it is for writing, but the usage context is sparse.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
scene_create_entityC
Add a new entity to a scene. Returns the new entity's ID.
| Name | Required | Description | Default |
|---|---|---|---|
| tag | No | ||
| sceneId | Yes | ||
| components | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It discloses that the tool mutates a scene and returns the new entity's ID, but it does not state whether the operation is reversible, whether components must already exist, or what happens if the sceneId is invalid. For a mutation tool with zero annotation coverage, this is a significant gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single efficient sentence that front-loads the action and return value. It earns its place, though it could add a bit more context without becoming bloated.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with no annotations, no output schema, and 0% parameter coverage, the description is too thin. It does not explain the 'components' parameter, any constraints on 'tag', or error behavior. An agent would need to guess or inspect other tools to call this correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It only explains the return value, not the meaning of 'tag' or 'components'. The description adds no semantic detail beyond the schema's bare property names, leaving the agent to guess what values are valid.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Add') and resource ('a new entity to a scene'), and mentions the return value (entity ID). It is clear enough to distinguish from sibling tools like scene_list_entities or scene_entity_info, though it does not explicitly name them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this tool versus alternatives. The description does not mention prerequisites (e.g., scene must exist), nor does it contrast with scene_entity_info or scene_list_entities. Usage context is only implied by the verb 'Add'.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
scene_entity_infoC
Return the tag, active state, and component list for a specific entity.
| Name | Required | Description | Default |
|---|---|---|---|
| sceneId | Yes | ||
| entityId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are present, so the description carries the burden. The phrase 'Return...' implies a read-only operation, and the outputs are enumerated (tag, active state, component list). However, there is no disclosure of edge cases, failure behavior, or whether this reflects live state vs cached state. With no annotations, more behavioral context would be expected, but the tool's scope is at least made clear.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Single clear sentence; no filler. Front-loaded with what it returns.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only info tool with no annotations, no output schema, and no parameter descriptions, an agent needs more context: what the component list looks like, whether tag/state can be absent, ownership/visibility rules, and error behavior. The description is minimal but accurate.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% – neither parameter is explained in the input schema. The description implicitly ties the parameters to a scene and an entity ('specific entity', presumably within a scene), which helps, but it doesn't tell the agent how sceneId/entityId should be formatted, whether entityId is local to the scene, or what values are valid. Since coverage is 0%, the description needed to do more than infer the obvious parameter names.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names the exact resource (a specific entity) and the specific data returned: tag, active state, and component list. It clearly distinguishes the tool from raw entity listing (scene_list_entities) or component access (scene_get_component), though it doesn't explicitly name sibling alternatives.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this tool versus scene_get_component or scene_list_entities. The description says what it returns, but not the context where this is the right choice, nor exclusions/alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
scene_get_componentB
Retrieve the serialised state of a specific component on an entity.
| Name | Required | Description | Default |
|---|---|---|---|
| sceneId | Yes | ||
| entityId | Yes | ||
| componentType | Yes | PascalCase class name, e.g. Transform, PhysicsBody |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral burden. 'Retrieve' implies a read-only operation and 'serialised state' hints at the return format, but the description does not disclose error behavior, missing entity/component handling, or the exact serialization format.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence with no filler words. The core verb, resource, and scope are front-loaded, and every word earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple three-parameter getter, the description plus the componentType schema note is minimally viable, but the lack of an output schema makes the 'serialised state' return type vague, and the two identifier parameters remain underspecified. It is usable but leaves important retrieval details implicit.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is low at 33%, with only componentType described, and the description does not compensate for the undocumented sceneId and entityId parameters. It gives no guidance on where these identifiers come from or how they should be formatted.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Retrieve') and a specific resource ('serialised state of a specific component on an entity'), making the tool's purpose clear and distinguishing it from entity-level siblings such as scene_entity_info. It does not explicitly name an alternative it is not, so it stops short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use or when-not-to-use guidance is provided; the description only states what the tool does. An agent is not told how to choose this over scene_entity_info or scene_list_entities, nor are any prerequisites or context given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
scene_list_entitiesB
List all entity IDs currently active in a scene.
| Name | Required | Description | Default |
|---|---|---|---|
| sceneId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations present, the description carries the full burden. The phrase 'currently active' provides a useful time-scoped, read-only flavor, but the description is silent on output ordering, whether destroyed/hidden entities are excluded, or what an empty scene returns.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Seven words with zero filler. The term 'currently active' is meaningful and carries the behavioral nuance, and the sentence is front-loaded with the verb and resource.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Adequate for a simple one-parameter list tool, but with no output schema and no annotations, it does not state the return format (e.g., plain array of strings), ordering, or behavior for empty or invalid scenes.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the sceneId parameter has no documentation. The description only ties 'in a scene' loosely to the parameter, adding no real meaning beyond what the parameter name itself conveys (format, valid values, existence requirements).
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('List') and resource ('all entity IDs'), and scopes it to the active set in a scene. The word 'IDs' distinguishes this from sibling scene_entity_info, though it never names that alternative explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The use case is implied: call this when you need to enumerate the current ids in a scene. However, there is no explicit when-to-use/when-not-to-use guidance or mention of how it relates to sibling tools such as scene_entity_info or scene_create_entity.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
story_graph_exportA
Export the Story Graph (VNSystem) for a named scene as a JSON object containing nodes and edges.
| Name | Required | Description | Default |
|---|---|---|---|
| graphId | No | Graph identifier. Defaults to "default". | |
| sceneId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden. It does disclose the output shape (JSON object containing nodes and edges), which is useful behavioral context, but it does not explicitly state whether the export is read-only, handle missing scenes, or affect state. A stronger description would add a one-line guarantee of non-mutation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, focused sentence that front-loads the verb, resource, and expected output with no filler or repetition of the schema. Every part earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple 2-parameter export tool this is mostly complete: the agent knows what to call and what to get back. However it lacks any mention of the behavior of the optional graphId, any error or fallback behavior, and there is no output schema to give the agent more shape on the JSON structure. It covers the essentials but leaves some contextual gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema covers only graphId with a description; sceneId has none, giving 50% coverage. The description adds only a vague 'named scene' gloss for sceneId and says nothing about how graphId influences the export. Given the incomplete schema coverage, the description does not sufficiently compensate for the missing parameter context.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Export') with a specific resource (the Story Graph for a named scene) and tells what the result contains (JSON with nodes and edges). No sibling tool targets the story graph, so it is clearly distinguishable from alternatives.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The word 'Export' implies that this is the tool to call when story-graph data is needed, but there is no explicit 'when to use' statement, no exclusions, and no context about when other sibling tools would be more appropriate. The guidance is acceptable but only implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
21 tool updates
v0.1.0- First observed
actor_broadcast - First observed
actor_inbox_size - First observed
actor_send_message - First observed
emptysock_layer_info - First observed
gms2_inspect_project - First observed
navmesh_find_path - First observed
navmesh_nearest_node - First observed
particle_emitter_config - First observed
physics_body_state - First observed
physics_overlap_circle - First observed
physics_raycast_2d - First observed
physics_raycast_3d - First observed
save_delete - First observed
save_list - First observed
save_read - First observed
save_write - First observed
scene_create_entity - First observed
scene_entity_info - First observed
scene_get_component - First observed
scene_list_entities - First observed
story_graph_export
TDQS
Scored across 21 tools
Tools are clearly separated by domain (navmesh, physics, scene, save, actor, etc.) and each tool targets a distinct action or query. Even within physics, 2D and 3D raycasts are explicitly differentiated. No two tools appear to do the same thing.
All tools follow a consistent domain_verb_noun pattern (e.g., scene_create_entity, save_list, actor_send_message). Suffixes like _2d and _3d are uniform, and the one odd name (emptysock_layer_info) still adheres to the same style. Naming is predictable and coherent.
With 21 tools, the count falls in the 16-25 range, which is borderline heavy. However, the server covers many distinct subsystems, so each tool arguably earns its place, but the overall number feels slightly excessive for a single MCP server.
The set covers core read/query operations across multiple domains, but there are notable gaps: scene entities can be created and inspected but not updated or deleted, physics only offers queries (no manipulation), and particle control is limited to configuration. This leaves some lifecycle operations incomplete.
Maintenance
Related MCP Connectors
A comprehensive Model Context Protocol (MCP) server that enables AI assistants to control Unreal E…
Nifty's MCP server — exposes tasks, projects, messages, and files as tools for AI agents.
Your org's AI agents, tasks, runs, search, and brain files as MCP tools and resources.
Build, run and publish 3D games in the Zero engine from Claude Code, Cursor or Codex, over MCP.
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceEnables LLM-driven text game state management by exposing MCP tools for managing players, locations, items, entities, and abstract concepts.MIT
- AlicenseNot gradedqualityCmaintenanceAn MCP server enabling AI agents to author Unreal Engine 5 scenes directly, with tools for spawning actors, building PCG graphs, validating physics, generating terrain, and more through a single MCP connection.13MIT
- FlicenseNot gradedqualityDmaintenanceConnects Claude Code to the Unity Editor via MCP, enabling AI-driven control of scenes, assets, components, UI, animations, and more through 91 tools.2-
- AlicenseCqualityCmaintenanceEnables AI-driven game development by providing MCP tools to interact with the Godot editor, including scene editing, node manipulation, script attachment, and scene execution.2813 npmMIT