sts2-ironclad-agent
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@sts2-ironclad-agentWhat's my current hand and which moves are legal?"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Slay the Spire 2 Ironclad Agent
一个面向《杀戮尖塔 2》铁甲战士的 MCP 自动化与监控项目。它把游戏状态、合法动作、策略分流和动作校验放在本机控制器中;KEV 负责战斗战术,GPT-6 Luna 负责路线、奖励和商店等策略选择。
项目不会捆绑游戏文件、Steam 存档、Mod 二进制、模型权重、运行日志或 API 凭据。默认关闭游戏动作,默认示例不含 API Key。游戏数据以正在运行的游戏和 Mod 返回值为准。
状态与兼容性
当前维护目标是 Steam Main
v0.107.1。实际版本从安装目录的release_info.json读取;版本不匹配时应暂停并重新核验。游戏端需要 STS2MCP Mod 提供的本机 REST 桥;Python MCP 侧车本身不能读取游戏。
所选 STS2MCP 上游提交及本项目补丁记录在 桥接兼容说明。第三方仓库不会复制进本仓库;可用脚本按固定提交下载并应用补丁。
当前回合估算器只模拟受支持的精确卡牌效果。未知目标、意图、状态或卡牌效果需要保留不确定性;它不是完整战斗模拟器。
Related MCP server: sagaz-mcp
安装
需要 Windows、Python 3.11+、uv、Git,以及已安装的游戏和兼容的游戏端 Mod。若要编译 Mod,还需 .NET SDK 和本机游戏程序集。
git clone https://github.com/Mentat-Uran/sts2-ironclad-agent.git
Set-Location .\sts2-ironclad-agent
uv sync
Copy-Item config.example.toml config.toml在 config.toml 中填写本机游戏、KEV 和 OpenAI 兼容 Luna 服务地址。将 STS2_AGENT_CONFIG 指向本地 config.toml。该文件已被 Git 忽略。
先运行不会提交游戏动作的演示和检查:
uv run sts2-agent mock-demo
uv run sts2-agent preflight --config config.tomlpreflight 会检查服务和版本;只有桥接服务报告兼容的被动读取协议时,控制器才读取游戏状态。要在 Codex 中使用 MCP,可将 scripts/run-agent.ps1 配为项目 MCP stdio 命令。示例配置应使用你自己的工作区路径,不要提交个人 .codex 配置。
KEV 与 Luna
KEV 使用结构化判断接口
/v1/systemone,只可从当前合法动作候选中选择;不是聊天补全文本模型。Luna 使用 OpenAI Chat Completions 兼容接口。
base_url应包含一次/v1,客户端追加/chat/completions。模型 ID 在配置中指定。将 Luna 凭据放在启动进程的环境变量
STS2_LUNA_API_KEY中,或配置你自己的 SSH 密钥文件来源。SSH 方式只在你明确设置STS2_LUNA_KEY_SSH_TARGET与STS2_LUNA_KEY_REMOTE_PATH后启用;密钥只传给子进程,不写入项目文件或日志。run-agent.ps1会加载可选的本机忽略文件local.settings.ps1。直接运行autoplay_mcp.py时,应先在 PowerShell 中 dot-source 该本机文件,或自行设置上述环境变量。没有 Luna 凭据时,Luna 专属策略决策会安全暂停;不会自动换成其它模型。KEV 和 Luna 的上下文相互隔离。
项目样例默认使用本机回环地址。服务地址、模型可用性、认证方式和模型 ID 都需要由使用者按自己的服务配置。
只读监控面板
powershell.exe -NoProfile -ExecutionPolicy Bypass -File .\scripts\run-agent.ps1 -Mode monitor打开 http://127.0.0.1:8765。面板只绑定回环地址,显示当前游戏状态、合法动作、KEV/Luna 决策摘要和本机日志。日志可能包含牌组、候选项和游戏决策;默认只保存在 runtime/,该目录已从公开仓库排除。不要公开分享未经检查的运行日志。
游戏 Mod 桥接
先按 桥接兼容说明 下载固定上游提交并应用补丁:
powershell.exe -NoProfile -ExecutionPolicy Bypass -File .\scripts\bootstrap-sts2mcp.ps1
& .\vendor\STS2MCP\build.ps1 -GameDir $env:STS2_GAME_DIR安装/更新脚本要求显式提供 -GameDir,且只更新已识别的 STS2MCP 文件;它不会改动 Steam 存档或其它 Mod。编译、安装前请关闭游戏,并确认本机分支与版本兼容。
实时游戏动作
仓库及样例配置保持 game_actions_enabled = false。自动游玩会改变游戏存档;只有在检查好当前配置、存档和合法动作后,才显式启用动作门并使用 scripts/autoplay_mcp.py。不要用公开日志或模型输出绕过当前状态和合法动作校验。动作 POST 超时后控制器只重读状态,不重复提交。
项目结构
src/sts2_agent/:状态归一化、合法动作、控制器、MCP 服务、KEV/Luna/mock provider 与只读面板。.agents/skills/sts2-ironclad-agent/:铁甲战士通用 Codex Skill 和资料来源。.agents/skills/sts2-ironclad-luna/:仅供 Luna 使用的策略 Skill。docs/architecture.md:provider 分工、上下文边界与动作校验。docs/sources-and-versions.md:版本依据和中英文攻略来源。docs/known-issues.md:已知限制。patches/:针对固定 STS2MCP 上游提交的可审阅补丁。
许可证
本项目源代码采用 MIT 许可证。STS2MCP 补丁和其它第三方材料按其各自许可证使用;见 第三方声明。《杀戮尖塔 2》及其名称、图像和游戏内容归其权利人所有。本项目是非官方社区工具,与 Mega Crit 无隶属关系。
Available Tools
9 toolsautoplayB
Immediately continue Ironclad decisions and legal actions for up to 20 steps; stops on uncertainty or an optional Act 1 boss clear.
| Name | Required | Description | Default |
|---|---|---|---|
| batch_size | No | ||
| stop_after_act_one_boss | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description adds useful behavioral detail beyond the schema: 'Immediately', 'up to 20 steps', 'stops on uncertainty', and the optional Act 1 boss stop are meaningful traits. However, with no annotations provided, it does not disclose whether the tool mutates irreversible game state or what happens when it stops, which is a notable gap for an action-taking tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single dense sentence that front-loads the action and appends the relevant stopping conditions. Every clause adds useful information, with no filler or redundant restatement.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with two optional parameters, no required arguments, and an output schema, the description covers the core behavior and stopping conditions reasonably well. It omits important context such as mutation/side-effect expectations and explicit differentiation from step_once, which is more consequential given the absence of annotations.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. 'up to 20 steps' plausibly describes the batch_size boundation, and 'optional Act 1 boss clear' clearly maps to stop_after_act_one_boss. Still, batch_size is never named, and the relationship between the 20-step cap and the batch_size parameter is left implicit.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Immediately continue') and the resource ('Ironclad decisions and legal actions'), with concrete bounds ('up to 20 steps') and stopping conditions. It is not a tautology and is distinguishable from one-step tools by its multi-step scope, though it does not explicitly name or contrast any sibling.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit guidance on when to use autoplay versus alternatives like step_once, execute_proposal, or recommend_next_action. The description implies a batch/autonomous use case but does not state when not to use it or what preconditions must be met.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
execute_proposalC
Re-read the game and submit exactly one still-legal proposal. Dry-run is the default; no retries on timeout.
| Name | Required | Description | Default |
|---|---|---|---|
| proposal_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden. It discloses the dry-run default and the no-retry-on-timeout behavior, which are useful. However, it doesn't state whether the action is destructive, what happens on illegal proposals, or the nature of the response, leaving gaps.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise, two sentences, and front-loads the primary purpose. The dry-run and timeout details are secondary but relevant. No redundant information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with a single parameter and an output schema, the description is under-specified. It doesn't mention how to source proposal_id, what 'still-legal' means, or what the output represents, leaving an agent without enough context to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema provides no description for proposal_id (coverage 0%), and the tool description does not explain what proposal_id is, how to obtain it, or what format it should be in. This is a critical omission for a single-parameter tool.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action: re-read the game and submit a proposal, with a condition that it must be still-legal. This is clear and distinct from typical sibling tools like preflight or get_legal_actions, though it doesn't explicitly name an alternative.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is given on when to use this tool versus its siblings. The dry-run default and no-retry behavior are mentioned, but there is no statement about when this should be invoked (e.g., after planning) or when to prefer other tools like step_once or autoplay.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_game_snapshotA
Read state only from bridge schema 2, which advertises passive snapshots; refuse older bridges before requesting state.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden for behavioral disclosure. It does state that the tool is read-only ('Read state only') and implies a compatibility check/refusal for older bridges. However, it does not describe error behavior, return characteristics, or side effects beyond being read-only, leaving some ambiguity about what 'refuse' means in practice.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single compact sentence with no filler, and the core action is front-loaded. It loses a point only because the second clause is dense and uses domain-specific terms ('older bridges', 'advertises passive snapshots') without elaboration.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a no-parameter read-only tool with an output schema, the description covers the essential operational constraint: only use with bridge schema 2 and refuse older bridges. It does not explain the domain vocabulary, but the provided context signals and output schema likely fill in enough for an agent operating in this environment.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so there is nothing for the description to clarify beyond what the empty schema already shows. The baseline of 4 applies because no parameter documentation is needed.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool reads state from a specific source ('bridge schema 2') and that it provides passive snapshots, which aligns with the tool name. It does not explicitly compare itself to a sibling tool, but the 'state only' phrasing distinguishes it from action-oriented siblings like execute_proposal or step_once. Some jargon like 'bridge schema 2' is unexplained, which prevents a top score.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives a clear condition for use: only use with bridge schema 2 that advertises passive snapshots, and explicitly says to refuse older bridges before requesting state. This is strong when-to-use guidance, though it does not name alternative tools or explain what to do instead for older bridges.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_legal_actionsA
Read and validate the action candidates exposed by the current game screen.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It explicitly says 'Read', indicating a non-mutating operation, and adds 'validate' as extra behavioral contextches. However, it does not describe prerequisites, failure behavior, or whether any validation has side effects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, tight eleven-word sentence that front-loads the verb and resource. Every word earns its place, with no filler or repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a no-argument read-only helper with an output schema, this is largely complete. It could still more explicitly position the tool within the agent loop, such as calling it before execute_proposal, but the core context is present.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so there is no parameter semantics burden. The schema is fully covered and the description does not need to explain parameter meanings.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific operation ('Read and validate') and a specific resource ('action candidates exposed by the current game screen'). This clearly distinguishes it from siblings like get_game_snapshot, which returns broader game state, and execute_proposal, which performs a mutation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies the natural usage—querying available actions for the current screen—but it does not explicitly state when this tool should be preferred over get_game_snapshot, preflight, or recommend_next_action, nor does it list any exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
plan_map_routeA
Ask Luna to plan a full Ironclad route to the boss. Planning does not move the character.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. It discloses that planning does not move the character, which is a key side-effect, but it does not describe other potential effects such as whether the plan is stored, returned, or modifies game state. For a planning tool, this is a significant gap, but the description at least highlights the most important non-effect.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence that is concise, front-loaded with the purpose, and includes a crucial behavioral caveat. Every word earns its place, and there is no fluff. It is optimally structured for quick agent comprehension.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's simplicity (no parameters, simple action) and the presence of an output schema (which presumably describes return values), the description covers the essential aspects: what it does and its key side-effect. It lacks explicit context on prerequisites or when to use it, but the usage guidelines dimension already touches on that. Overall, it is reasonably complete for a tool of this simplicity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema is empty, with zero parameters. Schema coverage is 100% by definition. The description does not need to explain parameters because there are none. Per the calibration, a baseline of 4 is appropriate for tools with 0 parameters, and the description does not introduce any confusion.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states a specific verb 'plan' and a specific resource: 'a full Ironclad route to the boss'. It distinguishes itself from siblings by explicitly noting that planning does not move the character, which sets it apart from movement-oriented tools like step_once or execute_proposal. The purpose is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives a clear contextual hint by stating that planning does not move the character, implying that this tool is for planning only and not for actual movement. However, it does not explicitly name alternative tools or provide explicit when-not conditions. The exclusion of movement is implied but not fully explicit, so it falls short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
preflightA
Read-only check of the local game bridge, local game version, Kev model list and Luna key readiness.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description explicitly says 'Read-only check', which discloses the non-destructive nature of the tool. However, with no annotations provided, the description carries the full burden of behavioral disclosure. It does not mention what happens on failure, whether it returns detailed diagnostics, or any side effects (though read-only implies none). The read-only disclosure is valuable but incomplete.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single sentence that front-loads the read-only nature and lists the specific resources checked. No wasted words; every part of the sentence adds information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter read-only check with an output schema present, the description is largely complete. It names all the components being checked. It could be improved by stating when to run it (e.g., before starting a run), but the output schema likely covers return values. The main gap is usage timing, which is minor for a preflight tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has 0 parameters, so the schema provides no parameter information. The description compensates by listing the specific items being checked (game bridge, game version, Kev model list, Luna key readiness), which gives an agent a clear idea of what the tool inspects. With 0 params, baseline 4 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('check') and resource ('local game bridge, local game version, Kev model list and Luna key readiness'), which clearly distinguishes it from sibling tools that get snapshots, plan routes, or execute proposals. It lacks a title but the description is specific enough to identify the tool's purpose.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies this is a pre-flight readiness check to be run before other operations, but it does not explicitly state when to use it versus alternatives or when not to use it. The context signals show no annotations, so the description carries the burden, but it only implies usage rather than stating it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
recommend_next_actionC
Create one validated proposal: KEV for tactical choices, Luna for strategy, deterministic for forced or unique actions.
| Name | Required | Description | Default |
|---|---|---|---|
| route_id | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden of disclosing side effects, validation behavior, and persistence. 'Validated' hints at some checking logic, but the description does not say whether this creates state, requires authorization, or is safe to call repeatedly.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence that efficiently conveys the core action and mode-selection rules. It loses points only because the introduced terms KEV and Luna are unexplained, making the compactness cryptic rather than purely helpful.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
While an output schema exists, the definition still fails to explain the only parameter, the meaning of the strategy codes, or the validation semantics. For a non-trivial recommendation tool with no annotations, this leaves critical invocation details missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has one parameter, route_id, with 0% schema description coverage, so the description must compensate. It never mentions route_id, its nullability, or how it influences the proposal, leaving the agent unable to determine correct parameter usage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific deliverable ('one validated proposal') and contrasts with execution-oriented siblings like execute_proposal. The strategy qualifiers (KEV, Luna, deterministic) imply distinct proposal-generation modes, though the cryptic domain terms prevent full clarity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives internal selection guidance ('KEV for tactical choices, Luna for strategy, deterministic for forced or unique actions') but never states when to use this tool versus siblings such as get_legal_actions, plan_map_route, or execute_proposal. No exclusions or prerequisites are provided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
start_new_ironclad_run_stepB
Start a fresh Ironclad run with one advertised menu action per call; never switch profiles or overwrite a continue run.
| Name | Required | Description | Default |
|---|---|---|---|
| profile_id | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations, so description carries burden. It discloses that it starts fresh and doesn't overwrite a continue run, which is useful. However, it doesn't mention side effects, permissions, or the nature of 'advertised menu action'.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One sentence, no fluff. The key action is front-loaded. However, some terms like 'advertised menu action' are cryptic, but that's a clarity issue not conciseness.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given only one parameter and an output schema, the description still leaves significant gaps: what is a 'menu action', how does profile_id affect the run, what are the prerequisites? The negative constraints help but are insufficient for confident invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, and description does not explain profile_id parameter. The only hint is 'never switch profiles' which implies profile_id is used as-is, but the meaning is not clarified. The parameter has a default, but description adds no semantics.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States it starts a fresh Ironclad run, which is a clear verb+resource. The phrase 'one advertised menu action per call' adds specificity but is ambiguous without context. Does not explicitly differentiate from step_once or autoplay, but 'fresh' implies not continuing.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a negative constraint 'never switch profiles or overwrite a continue run' but does not explicitly state when to use this tool vs siblings. It implies starting a new run rather than continuing, but doesn't name alternatives or conditions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
step_onceC
Run one routed decision and, only when enabled in local config, submit at most one validated game action.
| Name | Required | Description | Default |
|---|---|---|---|
| route_id | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It does disclose a meaningful behavioral limit: it submits at most one validated game action, and only conditionally based on local config. But it does not clarify side effects, prerequisites, what 'validated' means, or whether the tool can return without acting.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single efficient sentence and front-loads the core action. It is not bloated, though its conciseness comes at the cost of important contextual detail.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool that can submit a game action)Skip? The description omits prerequisites, the role of local config, what 'routed decision' means, and how this tool fits with siblings like execute_proposal or autoplay. The output schema helps, but the tool context remains too sparse for confident invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the description never mentions route_id. The word 'routed' weakly hints that route_id selects the route, but the tool description does not explain how the parameter influences behavior or what null/default means.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific action ('Run one routed decision') and indicates it may submit a game action, so an agent gets a general sense of the tool. However, 'routed decision' is opaque and the description does not distinguish this from siblings like execute_proposal or autoplay, which also appear to advance or execute actions.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The only usage condition is 'when enabled in local config,' but there is no guidance about when to choose step_once over plan_map_route, recommend_next_action, execute_proposal, or autoplay. No alternatives or exclusions are named.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
9 tool updates
v0.1.0- First observed
autoplay - First observed
execute_proposal - First observed
get_game_snapshot - First observed
get_legal_actions - First observed
plan_map_route - First observed
preflight - First observed
recommend_next_action - First observed
start_new_ironclad_run_step - First observed
step_once
TDQS
Scored across 9 tools
Most tools have clear, distinct roles across the pipeline: readiness, snapshotting, legal actions, planning, recommendation, execution, stepping, autoplay, and new-run setup. The only mild boundary ambiguity is between execute_proposal and step_once, both of which can submit a single action, but their descriptions distinguish lower-level proposal submission from routed decision execution.
All names are readable snake_case, but they do not follow a single pattern. get_game_snapshot and get_legal_actions use get_, while plan_map_route, recommend_next_action, and execute_proposal use verb_noun, and preflight, step_once, and autoplay are one-off or compound names.
Nine tools is a well-scoped size for an autonomous game agent. Each tool supports a distinct stage of the play loop without excessive redundancy or unnecessary surface area.
The tool set covers the full gameplay automation loop: preflight checks, state reads, legal-action discovery, path planning, decision recommendation, proposal execution, single-step control, bounded autoplay, and starting a new run. Minor gaps exist around explicit stop/pause or richer configuration control, but the existing tools handle those through local config and autoplay stopping conditions.
Maintenance
Related MCP Connectors
MCP enforcement layer that intercepts AI agent actions and blocks rule violations before execution.
Pre-execution safety layer for autonomous agent wallets via MCP and x402.
Paid deterministic utilities and automation services for AI agents via MCP and x402.
Control plane for autonomous software labor. Agents claim objectives over MCP with audit trail.
Related MCP Servers
- FlicenseNot gradedqualityAmaintenanceEnables AI agents to play Slay the Spire 2 by exposing game state and actions through an MCP server, supporting combat, rewards, and run management.327-
- AlicenseNot gradedqualityBmaintenanceEnables transparent MCP proxying with a hash-chained effect ledger, classifying agent actions by reversibility, enforcing approval gates, and dry-run previews of sessions.MIT
- FlicenseBqualityAmaintenanceMCP server that enables AI agents to drive your real, logged-in browser via accessibility snapshots, with hardened credential brokering and local workflow memory.5330 npm-
- AlicenseNot gradedqualityBmaintenanceEnables an MCP agent to observe and command Slay the Spire 2 game state via HTTP, and optionally autopilot combat with a local deterministic policy, without LLM in the critical path.MIT