SatisfactoryMCP
SatisfactoryMCP is a local MCP server plus web map that reads your Satisfactory saves and game data to answer questions, plan factories, and visualize your world.
Query your world: progress, power (nameplate/measured), unlocked recipes, phase requirements, MAM research, collectibles, somersloops, power shards, storage, crates, and where the player is.
Plan factories: optimize with an LP over unlocked recipes and chosen resource sources; set objectives, exports, byproducts, Somersloop budgets, extractor clocks, and water extractors; produce layouts, bills of materials, recipe route comparisons, and unlock rankings.
Manage and analyze factories: propose, name, forget, and list factory labels; inspect factory health, uptime, issues, inputs/outputs, machines, recipes, power, nodes, and inter-factory links.
Trace material flow: walk save connections upstream or downstream from a machine, building type, or factory label.
Explore map and resources: list regions, describe locations, search resource nodes/conduits, rank build sites, and show targets on the local web map or satisfactory-calculator link.
Handle hard drives: list pending choices and rank options by counterfactual LP value across multiple objectives.
Visualize in browser: render terrain, factories, belts/pipes, power wiring, floors, crates, resource nodes, and region names on a local web map that follows saves live.
Stay read-only: only ever reads saves and game data; it never writes them.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@SatisfactoryMCPplan a factory for 10 turbo motors per minute using my save"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
SatisfactoryMCP
Ask Claude about your Satisfactory world — and get answers read straight from your own save files. This is an MCP server plus a local web map: the server plans factories with a real optimizer over the game's own recipe data and your actual progress, and the map renders your world in the browser — terrain, factories, belts and pipes, power wiring, floors, crates. Everything runs locally: your saves, the recipe data and the map artwork are all read off this machine, and nothing about your world is sent anywhere.

|
|
|
|
What you can ask
Phrased however you like — the model picks the tools:
"Plan a factory for 20 Modular Frames per minute using only recipes I've actually unlocked — what do I build, and how much power will it draw?"
"Which of my pending hard drives should I bank first, and why?"
"How healthy is my steel factory right now? Anything idle or starved?"
"Where am I standing, and what's the best spot near me for an aluminium setup?"
"What's still missing for Phase 3, and which factory is the bottleneck?"
"Trace my Reinforced Iron Plates upstream and tell me where the chain is thinnest."
"Compare the alternate recipes for Computers against what I'm running today."
"How much Quartz do I actually have, and which container is it in?"
"What was in the crate where I died?"
"Show the coal powerplant on the map." — answers with a link that opens the local web map above, zoomed to it. A satisfactory-calculator.com link comes second, for the vanilla world it knows; only the local one can draw what you built.
Plans balance every item honestly — a setup that would silently strand Heavy Oil Residue is reported infeasible instead of overstated — and every answer names the save file it read and how old it is. The server only ever reads your saves; it never writes them.
Area | Tools |
Game data |
|
Your world |
|
What you own |
|
Your factories |
|
Map |
|
Planning |
|
Hard drives |
|
Plus MCP resources (satisfactory://docs/summary, satisfactory://save/current,
satisfactory://map/regions) and three prompts that surface as slash commands:
design_factory, plan_power_plant, pick_hard_drive. The full surface, argument by
argument, is in docs/mcp-surface.md.
Related MCP server: minecraft-diagnostic-mcp
Getting started
You need:
Python ≥ 3.11 and uv
A local Satisfactory installation (Steam or Epic) — recipes and rates are read from the game's own data dump, so numbers stay correct when the game patches
Node.js — only to build the web map once; not needed for the MCP server alone
Windows is what it's developed and tested on; save and install auto-detection assume Windows paths, and both can be pointed elsewhere via environment variables (below)
git clone https://github.com/lukszi/SatisfactoryMCP.git
cd SatisfactoryMCP
uv syncConnect it to Claude
The MCP entry point is satisfactory-mcp (stdio). With Claude Code, register it at user scope
so it loads in any directory — you'll usually be asking about the game, not about this code:
claude mcp add --scope user satisfactory -- uv run --directory "/path/to/SatisfactoryMCP" satisfactory-mcpFor any other MCP client, the equivalent JSON configuration:
{
"mcpServers": {
"satisfactory": {
"command": "uv",
"args": ["run", "--directory", "/path/to/SatisfactoryMCP", "satisfactory-mcp"]
}
}
}The web map
Build the page once, then start the server:
cd src/satisfactory_mcp/interfaces/web/frontend
npm ci && npm run build
cd -
uv sync --extra web
uv run satisfactory-mcp-webThe map is at http://127.0.0.1:8712, and it follows your saves live as you play. It binds to
localhost on purpose: the API answers with the contents of your save directory and has no
authentication, so it is a local tool. To use another port, set SATISFACTORY_WEB_PORT for
both the web server and the MCP server, so the map links the tools print point at it.
Where your saves come from
Both locations are auto-detected:
Saves:
%LOCALAPPDATA%\FactoryGame\Saved\SaveGames— the game's own location. Override withSATISFACTORY_SAVES.Game data:
CommunityResources/Docs/en-US.jsonunder common Steam and Epic install paths. Override withSATISFACTORY_DOCS(pointing at theen-US.jsonfile itself).Your factory names and plans are kept in the platform's user data directory (
%LOCALAPPDATA%\satisfactory-mcp). Override withSATISFACTORY_USER_DATA. Pointing it at an empty folder lets you try naming without touching your real labels.satisfactory-mcp-web --port Nserves the map on another port.
Saves are grouped into worlds; within a world the newest save is used by default, and every answer says which file it read.
Licence
PolyForm Noncommercial 1.0.0. Free to use, modify and share for any noncommercial purpose, provided the required notice travels with copies:
Required Notice: Copyright Lukas Szimtenings (https://github.com/lukszi/SatisfactoryMCP)
Commercial use requires a separate licence from the owner — get in touch via GitHub (@lukszi). The licence covers this repository's code, tooling and extracted tables; it cannot and does not grant anyone rights over the game's content. All game-derived data describes Coffee Stain Studios' content — Coffee Stain retains all rights to Satisfactory and its assets, and this project is not affiliated with or endorsed by them.
The web map compiles Leaflet (BSD-2-Clause) into its bundle at build time from the npm package; the bundle is not committed, and every build carries Leaflet's licence text beside it.
Developing: the code layout, test suite, data generators, and the full data-provenance record live in docs/DEVELOPING.md. DESIGN.md is the design spine.
Available Tools
54 toolsadvise_hard_drive_pickA
Rank one pending drive's options by marginal value, via counterfactual LP.
Each option is solved for and against across several objectives, because a recipe can be worthless for power yet excellent for parts. Deltas are reported per objective and never collapsed into one score.
sources is plan_factory's selector list and means the same thing here, so the
baseline printed is the same quantity plan_factory reports for the same nodes.
| Name | Required | Description | Default |
|---|---|---|---|
| save | No | ||
| as_of | No | pin to one world state: a sav:… token from an earlier answer | |
| world | No | ||
| sources | No | ||
| hard_drive_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, and it does disclose meaningful behavior: each option is evaluated for and against across multiple objectives and deltas are reported per objective and never collapsed. It also clarifies that the printed baseline is the same quantity plan_factory reports. It stops short of stating safety/mutation status or cost, but the computational and reporting behavior is unusually well disclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action, then rationale and a semantics note. The 'recipe can be worthless for power yet excellent for parts' line earns its place by justifying the multi-objective design, though the jargon ('counterfactual LP') and the wording of the sources paragraph are slightly dense.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 5-parameter tool with no annotations and no output schema, the description helpfully characterizes the returned deltas but leaves most parameters and all behavioral/safety context undocumented. Adequate to understand intent, incomplete for confident invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 20% (as_of is the sole documented parameter), so the description must compensate. It does add real meaning for `sources` (it is plan_factory's selector list and yields the same baseline), but `save`, `world`, and `hard_drive_id` are left unexplained, so it covers at best a fraction of the gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'Rank one pending drive's options by marginal value, via counterfactual LP.' An agent can immediately distinguish this advising/ranking tool from the similarly-named sibling list_pending_hard_drive_choices, which enumerates rather than ranks.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied by the domain ('one pending drive') and the plan_factory comparison; there is no explicit statement of when to call this versus list_pending_hard_drive_choices or plan_factory, nor any prerequisite conditions. The plan_factory reference explains semantics rather than routing the agent between tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
alternates_for_itemB
Every automatable recipe that makes an item, alternates first.
When a save is readable, each row is marked HAVE or LOCKED, and a LOCKED one says which schematic would grant it -- a hard drive and a milestone are different work.
| Name | Required | Description | Default |
|---|---|---|---|
| item | Yes | ||
| save | No | ||
| as_of | No | pin to one world state: a sav:… token from an earlier answer | |
| world | No | ||
| include_locked | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does disclose useful output behavior: rows are marked HAVE or LOCKED and LOCKED rows name the granting schematic, with hard-drive vs milestone distinguished. It does not say whether 'save' is required for LOCKED data, what happens when no save is provided, or how 'world' interacts with 'save'.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences with no filler, and the core behavior (alternates first) is front-loaded. The second sentence is somewhat clipped but earns its place by explaining the HAVE/LOCKED output.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 5-parameter, no-annotation, no-output-schema tool, the description covers the return semantics well but leaves 'world', the relationship between 'world' and 'save', and the acceptable format for 'save' unaddressed. Adequate but with clear gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 20% (only 'as_of' is documented), so the description must compensate. It partially does: 'item' is the thing produced, 'save' governs HAVE/LOCKED marking, and 'include_locked' is implied by LOCKED rows. The 'world' parameter is never explained, leaving a real gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific resource and scope: 'every automatable recipe that makes an item, alternates first.' An agent can infer it returns recipe options for a given item, distinct from recipe_detail or search_recipes. It lacks a crisp verb and never names the sibling it is differentiated from, so it stops short of 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use, when-not, or alternative is stated. The description never mentions compare_recipe_options, recipe_detail, or search_recipes, so the agent must guess which sibling to call for a recipe-listing need. Only implicit context ('alternates first') hints at use.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
amend_factoryA
Add or drop individual machines on a label, without re-anchoring the rest of it.
name_factory re-anchors a label to whatever its selector picks, so correcting one
wrongly-included machine meant re-selecting the whole factory. This edits the
membership: every anchor add and drop do not name is left exactly as it was,
including ids this save no longer has.
add runs first, then drop, then prune_missing, which clears the anchors
list_factories reports gone. One machine is machine:<instance> on either side.
Dropping the last machine is refused -- deleting a label is forget_factory.
| Name | Required | Description | Default |
|---|---|---|---|
| add | No | selector terms for what to ADD. product:<item> | recipe:<name> | building:<class or name> | near:<place>@<radius_m> | base:<n> | line:<n> | slab:<n> | proposal:<n> | label:<name> | machine:<instance> | all. Terms are ANDed; comma-separated values inside one term are ORed; prefix a term with '-' to exclude it | |
| drop | No | selector terms for what to DROP, same grammar | |
| name | Yes | ||
| save | No | ||
| as_of | No | pin to one world state: a sav:… token from an earlier answer | |
| notes | No | ||
| world | No | ||
| dry_run | No | ||
| prune_missing | No | drop every anchor this save no longer has |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does much of it: it discloses operation ordering (add, then drop, then prune_missing), that unnamed anchors are left untouched even if their ids are gone, that dropping the last machine is refused, and what prune_missing clears. It stops short of stating whether the change is destructive/reversible, whether it needs write permissions, or what the result looks like.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The core purpose leads in the first sentence, followed by the rationale and the operation-ordering rules. Dense and mostly earned, though the second sentence's framing of the name_factory problem is slightly verbose for the information it delivers.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 9-parameter mutation tool with no annotations and no output schema, the description covers the central mechanics well but leaves several parameters (name, world, notes, dry_run, save) and the return/error surface unexplained. Adequate, with visible gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 44% across 9 parameters. The description explains the add/drop grammar, the machine:<instance> form, and prune_missing behavior, and implies the save semantics ('ids this save no longer has'), but name, world, notes, and dry_run receive no explanation in either the description or the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: add or drop individual machines on a label without re-anchoring the rest. It explicitly contrasts with name_factory (re-anchors) and forget_factory (deletes the label), so an agent can place it within the sibling set without opening schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Names the alternative name_factory and the condition that selects this tool instead (fixing one wrongly-included machine rather than re-selecting the whole factory), and routes label deletion to forget_factory. Both 'when to use' and 'when not to use' are covered.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
bomC
Flattened bill of materials: total raw and intermediate rates for qty/min of an item.
qty is a RATE, per minute. Solved by the LP, never by expanding the recipe
tree: Recycled Plastic and Recycled Rubber form a real 2-cycle, so an expansion
has no correct depth limit. Every row names the recipe chosen for that item,
because alternates change the totals materially.
| Name | Required | Description | Default |
|---|---|---|---|
| qty | No | ||
| item | Yes | ||
| save | No | ||
| as_of | No | pin to one world state: a sav:… token from an earlier answer | |
| limit | No | max rows (hard cap 25) | |
| world | No | ||
| offset | No | ||
| outlets | No | ||
| allow_sinks | No | ||
| only_recipes | No | ||
| exclude_recipes | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations the description carries the full burden, and it does disclose genuinely non-obvious behavior: the result is solved by an LP rather than recipe-tree expansion (because of a real 2-cycle), and each row names the recipe chosen. However it never states whether the call is read-only or has side effects, and the 'save'/'world' state behavior is left unexplained.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core meaning, and the parenthetical about the recycled-material 2-cycle earns its place by justifying why an LP is used instead of tree expansion. Slightly dense but no filler sentences.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an 11-parameter, solver-backed tool with no annotations and no output schema, the description leaves too much unsaid: most parameters, the meaning of save/world/as_of state, and the shape of the returned rows are never covered.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 18% across 11 parameters, so the description must compensate and largely does not. It usefully defines qty as a rate per minute, but item, save, world, outlets, allow_sinks, offset, only_recipes and exclude_recipes get no explanation anywhere.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: it produces a flattened bill of materials giving total raw and intermediate rates per qty/min of an item. That is concrete enough to distinguish it from recipe_detail or trace_upstream, though it does not name any sibling explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit when-to-use or when-not-to-use guidance and no named alternative. The reader can infer it is for totaling resource requirements, but nothing routes the agent between this and recipe_detail, trace_upstream, or alternates_for_item.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
collected_from_worldA
Map collectibles: how many exist, how many you took, what is left and what is closest.
Slugs, somersloops, Mercer spheres and their shrines, mushrooms, drop pods and the loot caches around them. Two sources, and neither is asked the other's question:
the map says what exists and where, read from the installed game's own cooked packages, so
placedis exact and every coordinate is exact;the save says what is gone. The world is not saved -- a save never mentions a slug still lying there -- so its destroyed-actor list is the collected list, and it is exact too.
remainingis the subtraction of the two.
Views: census (default) counts every category; collected and remaining list
individual placements with coordinates; nearest lists the remaining ones by distance
from near, defaulting to the player.
A placement in a cell no save has ever loaded is counted as remaining and reported as
never_streamed. It is never called present -- the map says where it is and nothing
on disk says whether it is still there.
| Name | Required | Description | Default |
|---|---|---|---|
| mode | No | retired -- write show= instead | |
| near | No | origin for show=nearest: 'x,y' in metres, 'me', or a factory name. Defaults to where the player is standing | |
| save | No | ||
| show | No | census | collected | remaining | nearest | census |
| as_of | No | pin to one world state: a sav:… token from an earlier answer | |
| group | No | one category, e.g. 'power_slug_blue'. Omit to see them all | |
| limit | No | max rows (hard cap 25) | |
| world | No | ||
| offset | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations the description carries the full burden, and it does disclose real behavior: the two data sources, that the save's destroyed-actor list *is* the collected list, that 'remaining' is a subtraction, and the important never_streamed caveat. It stops short of stating read-only status explicitly or any permission/rate context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded summary followed by a well-structured explanation of the two sources and the view modes using bullets. Most sentences carry information; the source/never_streamed prose is slightly long but each part earns its place for a tool this nuanced.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 9-parameter, annotation-free tool with no output schema, the description supplies the hard part: the conceptual model of where counts come from and what the views return. It is still incomplete on several parameters (save, world, offset) and does not indicate result shape for census.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 67%, so the schema documents most parameters. The description adds meaning for 'show' (what each view returns), 'near' (origin defaulting to the player) and 'as_of' (pinning a world state), but leaves 'save', 'world' and 'offset' entirely unexplained in both description and schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening line states a specific verb+resource and scope: a census of map collectibles covering slugs, somersloops, Mercer spheres, mushrooms and drop pods. It is clearly readable as 'count/list collectibles in the world'. It never differentiates itself from closely related siblings such as somersloops and power_shards, so it falls short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explains each view (census/collected/remaining/nearest) and when 'near' applies, which implies usage. However it never names an alternative tool or an explicit when-not-to-use condition, so the choice between this and the somersloops/power_shards siblings is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
commission_planA
In what order to switch a built plant on, without blowing the fuse.
This is a STARTUP order, not a build order, and the difference removes most of the problem. Building costs materials, not power -- a machine draws only when it runs -- so the whole plant can be constructed at leisure, drawing nothing, and then energised block by block. Nothing here tells you what to build first.
The constraint is one line, and it is hard: at every step, energised consumer draw must stay under the headroom plus generation from generators already burning fuel. Exceeding it in Satisfactory does not degrade gracefully -- the fuse blows and the whole grid stops until it is reset by hand, including the plant that was feeding it.
Generators are free to energise (0 MW draw, read from the dump), so a wave costs its consumers and refunds its generators, and that refund pays for the next wave.
Takes plan_factory's arguments, or recall a saved plan with plan=.
| Name | Required | Description | Default |
|---|---|---|---|
| plan | No | recall a saved plan by name | |
| save | No | ||
| as_of | No | pin to one world state: a sav:… token from an earlier answer | |
| limit | No | max rows (hard cap 25) | |
| world | No | ||
| clocks | No | ||
| offset | No | ||
| sloops | No | ||
| exports | No | ||
| sources | No | ||
| objective | No | max_mw | |
| allow_sinks | No | ||
| headroom_mw | No | grid power free for startup; default reads it from the save | |
| target_item | No | ||
| only_recipes | No | ||
| exclude_recipes | No | ||
| export_minimums | No | ||
| machine_cost_mw | No | ||
| only_free_nodes | No | ||
| extractor_clocks | No | ||
| water_extractors | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and provides strong behavioral context: the hard power constraint, the consequence of exceeding it (fuse blows and grid shuts down until manually reset), and the fact that generators are free to energise and refund power for later waves. It does not, however, disclose side effects of the save parameter, authentication needs, or return behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the core purpose and then unfolds the startup constraint and generator behavior in a logical sequence. It is somewhat long but the domain context is relevant, and no sentence feels purely redundant for such a specialised planning tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex 21-parameter tool with low schema description coverage and no output schema, the description does not provide enough invocation detail. It explains the conceptual model well but leaves most arguments and the expected return shape undocumented.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has 21 parameters with only 19% description coverage, and the description adds little: it mentions only plan= and that the tool takes plan_factory's arguments. Parameters such as save, world, clocks, objective, allow_sinks, headroom_mw, machine_cost_mw, and many recipe filters are not explained in the description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening line states a specific action and resource: determining the startup switch-on order for a built plant. It explicitly contrasts this with a build order and notes that it does not say what to build first, so an agent can distinguish it from planning tools like plan_factory.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context for when to use it: after construction, for startup energisation, not for deciding what to build. It also points to plan_factory arguments and the plan= recall option, but it does not explicitly exclude or compare against other sibling tools such as site_plan or plan_layout.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
compare_recipe_optionsA
Rank whole ROUTES to make an item by what each actually costs.
Not a recipe list -- alternates_for_item already does that. Each route is solved end to end with the LP, so the comparison is Crude -> Alt HOR -> Diluted Fuel against Crude -> Fuel, priced in raw resource per unit, whole buildings, net power, and byproducts needing an outlet.
| Name | Required | Description | Default |
|---|---|---|---|
| item | Yes | ||
| rate | No | ||
| save | No | ||
| as_of | No | pin to one world state: a sav:… token from an earlier answer | |
| limit | No | max rows (hard cap 25) | |
| world | No | ||
| outlets | No | ||
| allow_sinks | No | ||
| per_resource | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are present, so the description carries the full burden. It usefully discloses the computation mechanism ('solved end to end with the LP') and the result dimensions (raw resource per unit, whole buildings, net power, byproducts needing an outlet), but says nothing about read-only nature, permissions, or cost of invocation for a 9-parameter analysis call.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core purpose, then the differentiation, then the mechanics. The illustrative route example earns its space by making the LP comparison concrete. Two sentences plus example, little waste.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description appropriately explains what the comparison returns (pricing dimensions), which is a genuine plus. However, for a 9-parameter tool with 22% schema coverage and no annotations, seven undocumented parameters leave the definition materially incomplete for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 22% (2 of 9 params documented in-schema), so the description must compensate and largely does not. It only loosely hints at 'outlets' via 'byproducts needing an outlet'; rate, save, world, allow_sinks, and per_resource receive no explanation in either place.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Rank whole ROUTES to make an item by what each actually costs') and explicitly distinguishes itself from the sibling 'alternates_for_item'. The added contrast with a naive recipe list leaves no ambiguity about what this tool computes.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit when-not ('Not a recipe list -- alternates_for_item already does that') and routes the agent to the correct sibling for that case. It stops short of distinguishing from other planning siblings like plan_factory or propose_factories, so it is clear context without full alternatives coverage.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
cratesB
The crates lying on the ground: what is in each one and where to walk to get it.
A death crate is what you dropped when you died; a dismantle crate is the overflow from dismantling with a full inventory. Both are recoverable and NEITHER counts as spendable stock -- a crate deletes itself the moment it is emptied, so a plan that spent it would depend on somebody walking back there first.
Whose crate it is the save does not say: the crate's only saved property is its type, so there is no owner, no timestamp and no cause to report.
| Name | Required | Description | Default |
|---|---|---|---|
| save | No | ||
| as_of | No | pin to one world state: a sav:… token from an earlier answer | |
| limit | No | max rows (hard cap 25) | |
| world | No | ||
| offset | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden and does it well: it explains that both crate types are recoverable, are not spendable stock, and self-delete when emptied, and warns that a spent crate depends on someone walking back. It omits read-only/pagination behavior, but the domain semantics disclosed are substantive and non-obvious.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core purpose, then incremental caveats in three tight sentences with no filler. Slightly prose-heavy, but every sentence carries information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 5-parameter read tool with no output schema and no annotations, the description covers the crates' semantics but omits parameter behavior, pagination, and what the rows look like. Adequate on domain meaning, incomplete on invocation detail.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 40% (as_of and limit are documented; save, world, and offset are not), so the description needs to compensate and it does not mention parameters at all. An agent gets no guidance on how save, world, or offset scope the results, leaving half the parameters undocumented in both places.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The first sentence states a concrete verb and resource: it reports what is in each ground crate and where to walk to reach it. It also distinguishes two subtypes (death crate vs dismantle crate) that an agent would otherwise conflate. It does not name a sibling (e.g. stock/storage) to sharpen the boundary, so it falls short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied rather than stated: the warning that crates are 'NEITHER spendable stock' implicitly contrasts this tool with the stock/storage tools. There is no explicit when-to-call, when-not-to-call, or named alternative, so an agent must infer the trigger.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
describe_locationA
Name the region at a place, sample its elevation, and count what runs through.
at= takes 'x,y' in metres, or anything else this project prints an id for: me,
a named factory, slab:<n> from factory_map show=slabs -- including the bare
platforms nothing else would take -- or a chain:/pipe: run from search_conduits.
Returns 'off-map or ocean' rather than guessing the nearest land region.
Elevation is answered two ways and the two are never averaged. Where this machine
carries the extracted 1 m terrain field, terrain_m is one texel read at exactly
this coordinate, with the layer that answered, that layer's measured accuracy, and
the water surface and depth where water stands. Everything else is a SAMPLE
population reported with its count and spread: resource nodes rest on terrain and
are quoted as ground, foundations and buildings are quoted separately as built
elevation because a platform is wherever the player put it, and the gap between the
two is the fill already stacked there.
Belts and pipes are counted too, measured against the runs' drawn lines rather than
their corner points, so a conduit crossing mid-span is seen. With a readable save,
a zero here means nothing runs through -- absence in this output is absence in the
world. search_conduits lists the runs themselves.
| Name | Required | Description | Default |
|---|---|---|---|
| at | No | the place: 'x,y' in metres, 'me', a named factory, 'slab:<n>', or a run id like 'chain:7' | |
| save | No | ||
| as_of | No | pin to one world state: a sav:… token from an earlier answer | |
| world | No | ||
| radius_m | No | how far to look for known elevations, metres |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden and does so richly: it states that off-map points return 'off-map or ocean' rather than a guessed region, explains that elevation is never averaged across two methods, describes the terrain_m texel read with layer/accuracy/water details, distinguishes ground vs built elevation for nodes vs foundations, explains that conduits are measured against drawn lines not corner points, and notes that a zero count means true absence with a readable save.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the core purpose and organized logically from input formats to behavior to a sibling note. It is somewhat long and contains a few dense, stylized phrases, but for a tool with five parameters, no annotations, and no output schema, most sentences add useful context rather than filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given high complexity, no annotations, and no output schema, the description covers return behavior extensively—region, elevation modes, conduit counting, and absence semantics. It leaves `save`, `world`, and `as_of` largely unexplained, which is a meaningful gap for a tool that supports pinning to earlier world states, but the overall behavioral picture is strong.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 60%, and the description meaningfully extends the `at` parameter beyond the schema by specifying that it accepts 'x,y' in metres, `me`, a named factory, `slab:<n>` from `factory_map show=slabs`, or a `chain:`/`pipe:` run from `search_conduits`. It does not expand on `save`, `world`, `as_of`, or `radius_m`, so it only partially compensates for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The first sentence states a specific verb and three resources: name the region, sample elevation, count conduits. It later distinguishes this tool from `search_conduits` (which lists runs) and from `factory_map` (source of slab ids), so an agent can tell what makes this tool unique.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use the tool by describing what it returns, and it notes that `search_conduits` lists the runs themselves, but it never explicitly states when to choose this tool over alternatives like `whereami`, `show_on_map`, or `list_regions`. Usage is inferable, not spelled out.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
diff_vs_saveA
What to change to get from the factory you have to the one plan_factory plans.
Takes exactly plan_factory's arguments and re-solves, because the server keeps no state. Both tools print a plan id hashed over the arguments AND the save-derived solve inputs, so two responses carrying the same id are provably the same plan.
Machines are matched by IDENTITY, never by position: a manufacturer on (building, recipe), a generator on its building alone since its fuel is piped in rather than set on the machine, an extractor on the node it occupies. A Refinery running some other recipe is busy, not spare, so it never counts toward the plan.
Actions are ordered free-first -- UNPAUSE, then SETRECIPE on machines that produce nothing today, then BUILD. Stages follow the plan's own chain depth and the power arithmetic is INCREMENTAL, charging only the machines you have yet to place. Where a machine cannot be identified at all (Water Extractors have no recipe and no resolvable node) the answer is a RANGE, never a number.
Recall a stored plan with plan= and the diff is also grouped by STARTUP STAGE --
the same partition commission_plan emits -- so it answers "which stage am I in".
stage=<n> narrows to one stage's delta; stage=0 asks for the overview
without a stored plan, at the cost that the numbering moves when the arguments do.
Built and energised are DIFFERENT states and the save separates them in one direction only: a machine that produced in the last 300s window certainly had power, while one that did not may be unpowered, starved, blocked or idle. Grid membership is not persisted at all, so a stage is never reported as "unpowered" -- only as built with nothing proven running, which is exactly what a finished but not-yet-energised block looks like.
Saves are read-only: this never proposes writing one, and there is no dismantle action. Machines standing among the plan but not in it are listed for you to judge.
| Name | Required | Description | Default |
|---|---|---|---|
| plan | No | recall a saved plan by name | |
| save | No | ||
| as_of | No | pin to one world state: a sav:… token from an earlier answer | |
| limit | No | max rows (hard cap 25) | |
| stage | No | one startup stage's delta; 0 for the stage overview | |
| world | No | ||
| clocks | No | ||
| exports | No | ||
| factory | No | only count this factory's machines as already built | |
| sources | No | ||
| objective | No | max_mw | |
| show_cost | No | ||
| allow_sinks | No | ||
| target_item | No | ||
| only_recipes | No | ||
| exclude_recipes | No | ||
| export_minimums | No | ||
| machine_cost_mw | No | ||
| only_free_nodes | No | ||
| extractor_clocks | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and discharges it richly: identity-based matching rules, action ordering (UNPAUSE then SETRECIPE then BUILD), incremental power accounting, RANGE vs number for unidentifiable machines, the built-vs-energised asymmetry, non-persisted grid membership, and read-only saves with no dismantle action. This is exactly the kind of behavioral disclosure an agent needs.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
It is long, but the length is largely earned given the tool's conceptual complexity, and the first sentence leads with purpose. Paragraphing separates related concerns, though several sentences are dense enough that a skimming agent could lose the thread.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema and no annotations, yet the description explains the return shape well: plan-id hashing, stage-grouped diffs, RANGE answers, and provenance of sa: tokens. It is strong on concept but leaves many of the 20 parameters unexplained, which leaves the invocation surface incomplete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 25% across 20 parameters, so the description must compensate. It explains plan=, stage=, and factory= in prose and notes save handling, but leaves the majority of parameters (clocks, exports, sources, objective, show_cost, allow_sinks, target_item, recipe filters, machine_cost_mw, only_free_nodes, extractor_clocks, limit, world) unaddressed, so it only partially fills the gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening line states a specific purpose: computing what to change to turn the current factory into the one plan_factory plans. It distinguishes itself conceptually from plan_factory (it re-solves using that tool's arguments) and from commission_plan (whose stage partition it reuses), but the verb 'diff' is never stated plainly and the purpose is embedded in dense prose rather than front-loaded as a one-line statement.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives concrete usage cues: recall a stored plan with plan=, narrow with stage=<n>, use stage=0 for an overview without a stored plan. However it never frames when to reach for this tool versus plan_factory or commission_plan, and gives no exclusions or prerequisites beyond the read-only note.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
explain_byproductsB
Explain which byproducts stall a plan, and what can legally consume them.
Every item balance is an equality, so a byproduct with no consumer makes a plan INFEASIBLE rather than silently vanishing. This says WHICH item is stuck, whether it can be sunk (solids only -- a fluid must be consumed exactly or packaged first), and which recipes would absorb it, split into ones this world has unlocked and ones it does not.
Pass item to focus on one byproduct instead of the whole plan.
| Name | Required | Description | Default |
|---|---|---|---|
| item | No | ||
| save | No | ||
| as_of | No | pin to one world state: a sav:… token from an earlier answer | |
| limit | No | max rows (hard cap 25) | |
| world | No | ||
| exports | No | ||
| sources | No | ||
| objective | No | max_mw | |
| allow_sinks | No | ||
| target_item | No | ||
| exclude_recipes | No | ||
| export_minimums | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and delivers unusual depth: it explains that every item balance is an equality, that an unconsumeable byproduct makes the plan INFEASIBLE rather than being ignored, and that sinks apply to solids only since fluids must be consumed or packaged. It omits return format, pagination, and auth/cost behavior, so it is strong but not exhaustive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the core purpose, then explains the domain mechanics and the `item` scoping option in three compact paragraphs. Every sentence adds meaning with little filler, though the middle paragraph is dense.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 12-parameter tool with no annotations, no output schema, and 17% schema coverage, the description explains the domain semantics well but leaves the bulk of the parameters undocumented and never describes the response shape. An agent knows why to call it but not how to configure most of its inputs.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 17% of 12 parameters, so the description must compensate, yet it only clarifies the `item` parameter. The other ten parameters (save, world, exports, sources, objective, allow_sinks, target_item, exclude_recipes, export_minimums) are left with bare titles and no semantic explanation in either schema or description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Explain which byproducts stall a plan, and what can legally consume them') and describes the exact output focus (which item is stuck, sinkability, absorbing recipes). It does not name or contrast any sibling tool, but the purpose is distinctive enough to distinguish it from the surrounding factory/plan tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It states the scoping option ('Pass item to focus on one byproduct instead of the whole plan') and implicitly the use case (diagnosing infeasibility). However, there is no explicit when-to-use guidance, no prerequisites, and no comparison to alternatives such as factory_health or trace_upstream.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
factory_floorsA
How many decks a factory has and what stands on each.
Nothing in the save says "floor". Floors are recovered from geometry: a platform is a
4-connected run of 8 m foundation cells, and its storeys are the levels its tops
cluster at, with no assumed storey pitch. A band holding less than a quarter of the
platform's largest deck is marked minor -- a mezzanine or a machine plinth,
reported rather than merged away.
Whole-world by default: one row per platform, largest first. Narrow with platform=
(an index that is stable across calls on one save) or factory=, and a single
platform is answered floor by floor instead.
| Name | Required | Description | Default |
|---|---|---|---|
| save | No | ||
| as_of | No | pin to one world state: a sav:… token from an earlier answer | |
| limit | No | max rows (hard cap 25) | |
| world | No | ||
| offset | No | ||
| factory | No | a named factory, or any machine selector | |
| platform | No | one platform by the index this tool hands out |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden well: it discloses that floors are inferred from geometry, that no storey pitch is assumed, the threshold for marking a band 'minor' (under a quarter of the largest deck), that minors are reported rather than merged, and the default sort order. It omits whether this is strictly a read operation and how errors or ambiguous platforms are surfaced.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The answer is front-loaded in the first sentence and the remaining prose all adds derivational context. It is somewhat wordy across three paragraphs but no sentence is filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 7-parameter, no-output-schema, unannotated tool, the description covers the return shape (row per platform, or floor-by-floor), default ordering, and the core derivation rules. Remaining gaps are the unqualified save/world/as_of parameters and read-only confirmation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 57%. The description meaningfully augments platform= by noting the index is stable across calls on one save and explains the world-vs-narrowed behavior, but save, world, as_of, limit, and offset receive no added semantic detail beyond their schema hints.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening sentence states a specific resource and outcome: how many decks a factory has and what stands on each. It is clearly not a layout or map tool, and the geometry-derivation detail makes the scope concrete. It stops short of naming the closest siblings (factory_map, factory_query, factory_health) to draw an explicit boundary.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives clear operational context: whole-world by default with one row per platform, narrowable via platform= or factory=, and that a single-platform query switches to a floor-by-floor answer. There is no explicit when-not guidance or routing to an alternative tool, which keeps it below a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
factory_healthA
Measured uptime per machine, and WHY each stopped one is stopped.
The only measured numbers in this MCP. Every manufacturing building keeps a fixed 300-second productivity window; uptime is seconds-producing over that window.
States, worst first: paused, dead node (extractor bound to no resource --
a game update removed it), no recipe, blocked (output stack full),
starved (input empty), stalled (has input, output has room, still not running),
intermittent, saturated, unmonitored.
A stalled machine no generator can reach over the wires says so in its cause, and
every such machine is named in a note whatever state it is in -- separately for "no
wire at all" and "wired to a circuit no generator stands on", which are different
builds to finish. The save records no "has power" flag, so a machine a generator CAN
reach is never called unpowered here whatever the grid is doing: the wire is the only
electrical fact the file carries.
For a STARVED machine it also says what physically feeds the input it lacks: the run
that arrives, what stands at its far end and that feeder's own state, ONE hop back --
trace_upstream walks the rest. "No conduit of that medium arrives" and "one arrives
and the save joins its far end to nothing" are different rows and are never merged.
A missing FLUID is diagnosed on the plumbing manual's own ladder -- (1) connection,
(2) head lift, (3) flow rate -- and the cause names the FIRST rung that fires, so a
line that cannot climb to the machine is never answered with its supply rates. Solids
have no head-lift rung and read exactly as before.
Blocked counts as needing action, alongside dead node, no recipe, starved and
stalled: a full output box means nothing is taking what the machine makes. The sweep's
todo column counts those five.
The sweep over every factory also reports three plumbing faults that belong to no machine set: fluid buffers holding too little to output at their intake rate, pipeline pumps no wire reaches, and points where a line climbs above the head lift pushing it. All three are world-wide there, not scoped to a factory.
offset pages every table in the answer at once, worst first throughout.
| Name | Required | Description | Default |
|---|---|---|---|
| save | No | ||
| as_of | No | pin to one world state: a sav:… token from an earlier answer | |
| limit | No | max rows (hard cap 25) | |
| world | No | ||
| offset | No | ||
| factory | No | a label name, any selector, or 'all' for every named factory | all |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden and does so richly: it enumerates the state taxonomy worst-first, defines the `cause` semantics, explains that blocked/dead node/no recipe/starved/stalled feed the `todo` count, and describes how `offset` pages every table at once. It also discloses a genuine data limitation (the save records no 'has power' flag, so the wire is the only electrical fact) and that fluid diagnosis follows a fixed ladder. This is unusually complete behavioral disclosure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The key purpose sentence is correctly front-loaded, but the body is dense, jargon-heavy prose with long dependent clauses ('A stalled machine no generator can reach over the wires says so in its `cause`…') that is hard to parse quickly. Most content relates to real behavior, but the length-to-clarity ratio is poor for an agent scanning for a decision.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex diagnostic with no output schema, no annotations, and only 50% schema coverage, the description does describe the returned shape (per-machine states and causes, the sweep's `todo` column, world-wide plumbing faults). It is nearly complete, with the only real gap being the unexplained `save` and `world` inputs.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 50%, so the description must compensate for undocumented params. It does explain `offset` meaningfully ('pages every table in the answer at once, worst first'), which the schema leaves bare, and it implies the sweep semantics behind `factory`='all'. However, `save` and `world` receive no explanation, and `as_of`/`limit` are already covered by the schema. Partial compensation lands at the baseline 3.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening line states a specific thing measured (uptime per machine) and the diagnostic question it answers (why each stopped machine stopped), which is a clear verb+resource for an inspection tool. It also scopes itself against the sibling `trace_upstream` ('walks the rest'), so an agent can tell where this tool ends. It does not explicitly differentiate itself from adjacent reporting siblings such as `power_report` or `factory_query`, which keeps it at 4 rather than 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage context is clear: reach for this to find why machines stopped, and use `trace_upstream` to walk beyond the one hop this tool provides. It also explains that the world-wide sweep surfaces three plumbing faults not scoped to a factory. There is no explicit when-NOT-to-use statement or comparison to `power_report`/`factory_query`, so no exclusions are drawn.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
factory_mapA
Proposed factories, from power islands and belt topology, plus what is named.
Three independent signals are reported rather than one answer, because none is
right alone: power islands separate outposts but leave a grown-together base as one
476-machine blob; belt components shatter that blob into fragments; foundation slabs
are the sharpest of the three but say nothing about the ground-built parts of a
factory. Where they disagree, carve the difference with name_factory and a
product:, near: or slab: selector.
show=slabs also lists BARE platforms -- poured foundations carrying no machine yet -- with tile count, extent, bounding box and elevation, because a freshly built platform is a real place a build plan refers to. Pads under a stated tile threshold are summarised in one line.
| Name | Required | Description | Default |
|---|---|---|---|
| save | No | ||
| show | No | candidates | named | slabs | unlabelled | all | all |
| as_of | No | pin to one world state: a sav:… token from an earlier answer | |
| limit | No | max rows (hard cap 25) | |
| world | No | ||
| offset | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, and it does disclose real behavior: three signals are reported rather than one merged answer, each with known failure modes (476-machine blob, shattered fragments, no coverage of ground-built parts). It also states that show=slabs includes bare platforms with tile count, extent, bbox and elevation, and that small pads are collapsed to one line – useful return-shape context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose is front-loaded, which is correct, but the second sentence is a long unpacked clause chain ('because none is right alone: … blob; … fragments; … slab') that costs the reader effort for the information delivered. The show=slabs paragraph earns its place; the signal-disagreement sentence is bloated relative to its content.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and no annotations, the description has to describe both behavior and returns, and it partly does: the three-signal model, the slabs listing detail, and the bare-platform fields. It leaves the `save`/`world`/`offset` parameters and the overall response shape of the candidates/named modes unexplained, so it is solid but not fully complete for a 6-parameter tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 50%, with `show` and `as_of`/`limit` already documented in the schema. The description expands meaning for `show=slabs` beyond the bare enum list, which is real added value, but says nothing about `save`, `world`, or `offset`, leaving half of the parameter surface to the schema alone. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The first sentence names a specific resource – proposed factories – and the mechanism (power islands and belt topology) plus the naming overlay, which is more than a restatement of the name. It implicitly separates this from list_factories/propose_factories, but never names a sibling outright, so an agent must infer the boundary.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives one actionable route – when the three signals disagree, use `name_factory` with a product:/near:/slab: selector – which is genuine usage guidance. However, it never says when to reach for factory_map rather than factory_query, factory_sites, or list_factories, all of which sit adjacent in the sibling set.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
factory_queryA
Ask one thing about one factory: what it makes, needs, draws, or touches.
show accepts several at once, e.g. "balance,power,links". offset pages every
table in the answer at once, so asking for one aspect at a time is what you want
when a factory has more machines than fit.
summary size, position, top recipes, net power
balance per-item produced vs consumed vs net -- the sign is the point
outputs net surplus: it leaves the factory, or it backs up
inputs net deficit: it has to be fed in from outside
internal made and eaten inside the set -- the mark of a self-contained line
machines every machine with its building, recipe and clock
recipes / buildings counts
power draw vs generation, nameplate AND measured -- which factory is really burning the grid, rather than which could
nodes resource nodes its extractors sit on
links which other factories it exchanges material with
issues paused, recipe-less, or unresolved machines
Every rate is printed twice. NAMEPLATE is the machine's recipe rate at its saved clock, which a starved factory still reports in full. MEASURED is that rate scaled by the share of its own productivity window each machine spent producing -- the window that ended when the save was written, so a line idle at that moment measures 0 and is not broken. A machine keeping no monitor is left out of measured entirely and shown separately, because counting it in full there would invent output.
| Name | Required | Description | Default |
|---|---|---|---|
| of | No | retired -- write show= instead | |
| save | No | ||
| show | No | comma-separated: summary, machines, recipes, buildings, balance, inputs, outputs, internal, power, nodes, links, issues | summary |
| as_of | No | pin to one world state: a sav:… token from an earlier answer | |
| limit | No | max rows (hard cap 25) | |
| world | No | ||
| offset | No | ||
| factory | Yes | a label name, or any selector e.g. 'proposal:3' |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure and does so thoroughly. It explains the crucial NAMEPLATE vs MEASURED distinction, notes that idle machines measure 0 and are not broken, and warns that machines without monitors are excluded from measured to avoid inventing output. These are deep, non-obvious behavioral traits that an agent cannot infer from the schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the purpose, then organized into a bulleted list of show aspects followed by a measurement-semantics paragraph. It is dense and long, but each section earns its place by explaining how to interpret the returned data; there is little redundant filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given 8 parameters, no output schema, and moderate schema coverage, the description is largely complete for the main query workflow: it explains show aspects, pagination behavior, and rate semantics. It leaves save, world, and limit without description-level context, and does not say how to obtain a factory selector, but the central calling pattern is fully covered.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 63%, so the description must add meaning beyond the schema, and it does: it expands on what each show aspect returns (e.g., 'balance -- per-item produced vs consumed vs net -- the sign is the point', 'internal -- made and eaten inside the set'), clarifies that multiple show values are accepted, and explains that offset pages all tables together. It does not cover save, world, or limit semantics, but the core query parameters are well served.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening sentence states a specific verb and resource: 'Ask one thing about one factory: what it makes, needs, draws, or touches.' This clearly distinguishes a single-factory detail query from list-oriented siblings like list_factories. However, it stops short of explicitly naming alternative tools such as factory_health or factory_map for sibling differentiation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives concrete when-to-use guidance for aspects and pagination: 'offset pages every table in the answer at once, so asking for one aspect at a time is what you want when a factory has more machines than fit.' It does not explicitly compare the tool to siblings or state exclusions, but the context for selecting show combinations is clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
factory_sitesC
Built production buildings clustered into sites, largest first.
| Name | Required | Description | Default |
|---|---|---|---|
| save | No | ||
| as_of | No | pin to one world state: a sav:… token from an earlier answer | |
| limit | No | max rows (hard cap 25) | |
| world | No | ||
| offset | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are supplied, so the description carries the full disclosure burden. It reveals only two traits: results are ordered largest-first and only 'built' buildings are counted. It is silent on authentication/save requirements, pagination behavior, reversibility, and what the grouping actually yields, which is thin for a five-parameter read tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
It is a single short sentence, but its brevity is under-specification rather than efficient conciseness; it conveys almost no actionable information and there is no front-loaded verb or scope. The brevity does not earn its place because it omits everything an agent needs.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no annotations, no output schema, and five parameters at 40% coverage, the description would need to explain inputs, result shape, and applicability. Instead it offers one fragment, leaving the save/world scoping and pagination contract entirely unexplained.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 40% (as_of and limit are documented; save, world, offset are not), so the description is expected to compensate and does not mention a single parameter. Nothing explains what 'save' or 'world' select or how limit/offset paginate, leaving the majority of parameters ambiguous.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description describes the returned resource ('built production buildings clustered into sites, largest first') but supplies no verb and no scope, so it reads more like a return-value summary than a statement of what the tool does. It is not clearly distinguished from close siblings such as rank_build_sites, factory_map, or list_buildings, so an agent cannot tell which of these to pick from the text alone.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance, no mention of prerequisites (whether a save or world must be specified), and no routing to alternatives like rank_build_sites or factory_map. The only hint at behavior is the implicit 'largest first' ordering, which is not framed as selection criteria.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
forget_factoryB
Delete a factory label. The machines themselves are untouched.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | ||
| save | No | ||
| as_of | No | pin to one world state: a sav:… token from an earlier answer | |
| world | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, and it does disclose one important trait: the machines are not affected, so this is a label-only mutation. It says nothing about reversibility, whether the factory must be empty, or how the save/world/as_of parameters change behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences, with the destructive scope clarification front-loaded. No filler to trim.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with no annotations, no output schema, and three undocumented parameters, the description is too thin. The save/world/as_of semantics that likely govern what state gets mutated are unexplained, so an agent cannot call this correctly with confidence.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 25% (just `as_of`), and the description explains none of the four parameters. The meaning of `save` and `world` — which look like world-state selection inputs — is left entirely to guesswork.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('delete a factory label') and immediately scopes it ('machines themselves are untouched'), which distinguishes it from a teardown-style operation and from siblings like name_factory/rename_factory. It does not explicitly name a sibling alternative, but the resource is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no statement of when to use this versus alternatives such as forget_plan or rename_factory, and no prerequisites or preconditions are given. Usage is only implied by the verb 'delete'.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
forget_planA
Delete a saved plan. Nothing in the world is touched, and plan_log can undo it.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | ||
| save | No | ||
| as_of | No | pin to one world state: a sav:… token from an earlier answer | |
| world | No | ||
| base_rev | No | the plan version you read; needed to change an existing plan |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does well: it discloses the operation is destructive ("Delete"), scopes the blast radius ("Nothing in the world is touched"), and discloses recoverability ("plan_log can undo it"). It stops short of covering failure cases such as a non-existent plan name.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short, front-loaded sentences with zero filler; the verb and scope lead, and the reversibility note follows efficiently.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple delete with no output schema, the scope and reversibility disclosure is adequate, but the presence of four optional parameters with only 40% schema coverage means the description leaves meaningful invocation gaps unaddressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 40%, so the description must compensate and does not. The optional save, world, and base_rev parameters are undocumented in both places, and the description never explains that name identifies the plan or how as_of/base_rev affect the delete.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ("Delete a saved plan") that is immediately distinguishable from siblings like rename_plan, list_plans, and plan_log. However, it does not explicitly name the contrasting sibling or spell out the saved-plan vs. world-state distinction as a routing rule, keeping it just short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage context is implied rather than stated: mentioning that plan_log can undo it hints at recovery flow, but there is no explicit when-to-use or when-not-to-use guidance versus rename_plan or forgetting a factory. The agent must infer the intent.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_buildingsA
Buildings by kind: production, extractor, generator, logistics, foundation, ramp, wall, pillar, beam, architecture (all five families together), or all.
Rows are marked HAVE or LOCKED against the save when one can be read. That matters
most for logistics: a planner assuming a belt or pipe tier it has not unlocked
gets every line count wrong by a factor and nothing says so, which is the worst
failure mode a planner has.
Paged: all is 539 buildings and unpaged it ran to ~60k characters, which is not
an answer, it is a context eviction. The envelope says how many more there are and
which offset fetches them.
| Name | Required | Description | Default |
|---|---|---|---|
| kind | No | retired -- write building_kind= instead | |
| save | No | ||
| as_of | No | pin to one world state: a sav:… token from an earlier answer | |
| limit | No | max rows (hard cap 25) | |
| world | No | ||
| offset | No | ||
| building_kind | No | production | extractor | generator | logistics | foundation | ramp | wall | pillar | beam | architecture | all | production |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden and does a fair job: it discloses the HAVE/LOCKED save-relative marking, the hard framing of paging (envelope reports remaining count and offset), and the practical scale problem (539 rows, ~60k chars). It does not cover auth, mutation risk, or what the rows themselves contain, but the paging and save-coupling disclosure is genuinely beyond schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the kind list well, but the middle paragraph editorializes ('the worst failure mode a planner has', 'not an answer, it is a context eviction') at the cost of density. The justification for locked-state awareness is valuable, but it is delivered with more words than needed.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Seven parameters, no annotations, no output schema, and 57% schema coverage. The description covers paging and the locked/have semantic but leaves the save/as_of/world pinning parameters and the deprecated kind alias unexplained, so an agent must infer how to pin a world state or query a specific save.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 57%. The description enumerates the building_kind values and explains the paging parameters conceptually (envelope/offset), which partially compensates, but it never mentions save, as_of, world, or the retired 'kind' parameter that the schema flags as superseded.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Opens with a specific resource and enumeration of the kind values, so an agent knows this lists building definitions filtered by category. It is distinguishable from siblings like search_items or recipe_detail, though the verb is implicit in the noun phrase 'Buildings by kind' rather than stated outright.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It supplies useful context for when the HAVE/LOCKED marking matters (logistics tier assumptions) and explains the paging model, but it never states when to prefer this tool over siblings, nor any exclusions or prerequisites. Usage is implied rather than directed.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_factoriesC
Named factories for this world, with how much of each is still standing.
| Name | Required | Description | Default |
|---|---|---|---|
| save | No | ||
| as_of | No | pin to one world state: a sav:… token from an earlier answer | |
| world | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries full disclosure burden, yet it only adds that each entry includes a "still standing" fraction. It says nothing about save/as_of pinning semantics, world selection, or what the listing returns, leaving the mutability/safety picture unstated.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One sentence, front-loaded, with no wasted words. However, the brevity comes at the cost of specification rather than being tight-but-complete.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no annotations, no output schema, and three loosely documented parameters, the description is inadequate for correct invocation. It does not explain the save/as_of pinning model or defaults that determine which world state is listed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 33% (only as_of is documented as a sav: token), so the description must compensate for save and world. It mentions "this world" only obliquely, and never explains the save token or how save/world/as_of interact, so the parameter gap remains.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The verb is implicit ("Named factories for this world") and the resource (factories) is clear, but there is no differentiation from near siblings like factory_query, factory_map, or factory_health, which plausibly also enumerate factories. The agent can infer it is a listing tool, but the scope versus those siblings is ambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description offers no when-to-use, when-not-to-use, or alternative routing despite many factory-related siblings. Nothing tells the agent whether this is the right entry point versus factory_query or factory_health.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_pending_hard_drive_choicesC
The pending hard-drive choices stored in the save, with rerolls left.
| Name | Required | Description | Default |
|---|---|---|---|
| save | No | ||
| as_of | No | pin to one world state: a sav:… token from an earlier answer | |
| limit | No | max rows (hard cap 25) | |
| world | No | ||
| offset | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries full behavioral burden. It does not disclose read-only status, pagination behavior, side effects, or return format, leaving the agent to assume a safe read.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One terse sentence fragment, no wasted words, but its extreme brevity borders on under-specification.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with five parameters, no annotations, and no output schema, the description omits crucial usage and parameter context, leaving significant gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 40% (only as_of and limit have descriptions). The description adds no meaning to any of the five parameters, failing to compensate for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description is a noun phrase that restates the tool name, adding only 'with rerolls left' and 'stored in the save'. It does not state a verb or distinguish from the sibling tool advise_hard_drive_pick.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use guidance, prerequisites, or alternatives are provided. An agent must infer usage from the name alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_plansA
Plans saved for this world, their versions, and whether the world has moved under them.
name prints one plan in full instead: the arguments as they were stored, what its
source selectors resolved to when saved, and its whole siting. Nothing is solved, so
this answers "what did I ask for" -- plan_factory plan=<name> answers the other
question, what those arguments resolve to today, and pays an LP solve for it.
Every plan carries a version (v14). Pass it as base_rev to any tool that
changes the plan; plan_log lists the versions and undoes them.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | one plan's full stored request, siting and field, unsolved | |
| save | No | ||
| as_of | No | pin to one world state: a sav:… token from an earlier answer | |
| world | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It discloses important behavioral traits: nothing is solved, the output shows stored arguments and saved siting rather than current resolution, and plan versions are exposed via `base_rev`. It does not explicitly state read-only safety, auth needs, or pagination behavior, so it is strong but not complete.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the list behavior, then the `name` mode, then version semantics. Each paragraph earns its place, though some cross-tool guidance about `base_rev` and `plan_log` is slightly ancillary for a list tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema and no annotations, so the description must carry more context. It explains the main return semantics and the `name` mode well, but leaves `save` and `world` params, output ordering/shape, and read-only safety unaddressed. Adequate but with clear gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 50%. The description adds meaning for `name` and indirectly for `as_of` by discussing world-state pinning and versions, but `save` and `world` are left undocumented in both schema and description. The added value is real but incomplete against the low coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific resource and scope: saved plans for this world, their versions, and whether the world has moved under them. It also distinguishes the `name` mode as returning one plan's stored request rather than resolving it, and contrasts the tool with `plan_factory`. An agent can tell this apart from siblings without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly names alternatives and conditions: use this to answer "what did I ask for" without spending an LP solve, use `plan_factory plan=<name>` for current resolution, and use `plan_log` for version history and undo. The `name` parameter's behavior is also tied to a clear usage condition.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_regionsA
Named map regions, optionally only those containing a given resource.
Region names are ADVISORY: the boundaries are the game's own map areas, downsampled to a 256 m grid to publish and a 64 m one to look up in, so a name near a boundary can be one cell out. Use them to talk about places, not to compute with -- every node row also carries an exact grid cell.
anchor is a coordinate that provably lies in the region, which a centroid does not:
a concave region's mean lands on its neighbour's ground, and the map has drawn its
names at the anchor all along.
| Name | Required | Description | Default |
|---|---|---|---|
| resource | No | ||
| with_resource | No | retired -- write resource= instead |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and discharges it well: it discloses that names are advisory, that boundaries are downsampled to a 256 m grid to publish and 64 m to look up, that a boundary-adjacent name can be one cell out, and that `anchor` provably lies inside the region whereas a centroid may not. These are exactly the behavioral traits an agent needs to avoid misusing the output.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the purpose, then two tightly scoped paragraphs that each add non-redundant value (advisory-boundary caveat, anchor-vs-centroid rationale). No filler sentences or restated titles.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter, read-only listing tool with no output schema and no annotations, the description covers the semantics and caveats of the returned data well. It could still say a little more about the shape of a region record beyond `anchor`, but nothing needed to invoke it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 50% and the description compensates: it explains that `resource` restricts results to regions containing that resource, and effectively reinforces the schema's note that `with_resource` is retired in favor of `resource=`. The `anchor` discussion concerns output fields rather than inputs, so it does not leave an input parameter uninterpreted.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The first sentence states exactly what is returned (named map regions) and the optional filter (those containing a given resource), which is enough to distinguish it from siblings like whereami, describe_location, or search_resource_nodes. It is a noun phrase rather than a verb-led statement, but the resource and scope are unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly tells the agent when to use the results ('to talk about places, not to compute with') and points to the alternative for computation ('every node row also carries an exact grid cell'). It does not name which sibling to call instead when exact geometry is needed, so it falls short of full when/when-not/alternative routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_worldsA
List save games grouped by world, newest first.
Unsupported files are reported separately rather than failing the scan -- pre-1.0 saves cannot be parsed at all.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden for behavioral disclosure. It adds valuable context about unsupported files being reported separately and pre-1.0 saves being unparseable, which is beyond basic listing semantics. However, it does not mention return formatting or other edge cases.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is extremely concise: two sentences, front-loaded with the primary action, and the second sentence adds a meaningful caveat without wasting words. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter listing tool, the description covers purpose, sorting order, and error handling for unsupported files. It lacks explicit return structure details, but given the tool's simplicity and absent output schema, it is nearly complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so the baseline score is 4. There is nothing to add about parameter semantics; the description correctly focuses on behavior instead of documenting parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's function: 'List save games grouped by world, newest first.' This uses a specific verb and resource, and the grouping by world distinguishes it from sibling tools like list_factories.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for listing world save data but provides no explicit guidance on when to use this tool versus alternatives, nor any exclusions. It lacks direct sibling differentiation or conditional advice.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
mam_researchB
MAM research: what is left, what it costs, and what you can afford right now.
The MAM is where CAPABILITIES live, as opposed to recipes -- the Dimensional Depot, the Power Augmenter, and Production Amplifier, which is the one that lets a Somersloop go into a machine at all.
That last one has no flag in the save. BP_UnlockSubsystem_C records overclocking as
mIsBuildingOverclockUnlocked, but nothing anywhere in the file records production
amplification, so it is derived from the purchased-schematic set instead. Capability
rows are marked LOCKS so it is obvious which research gates a tool argument rather
than just adding a recipe.
| Name | Required | Description | Default |
|---|---|---|---|
| save | No | ||
| show | No | all | todo | affordable -- todo hides finished research | todo |
| as_of | No | pin to one world state: a sav:… token from an earlier answer | |
| limit | No | max rows (hard cap 25) | |
| query | No | filter by name, case-insensitive | |
| world | No | ||
| offset | No | ||
| search | No | retired -- write query= instead | |
| status | No | retired -- write show= instead |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the behavioral burden. It discloses meaningful internal behavior: MAM capabilities are distinct from recipes, production amplification is not directly recorded in the save and is derived from the purchased-schematic set, and capability rows are marked LOCKS to indicate tool-argument gating. It still does not state read-only safety, permissions, or return format, but the domain disclosure is strong.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The first sentence front-loads the purpose effectively, but the following paragraphs drift into detailed save-file implementation lore such as BP_UnlockSubsystem_C and mIsBuildingOverclockUnlocked. That context is relevant to trust the LOCKS marker, yet it is longer and more technical than needed for most invocation decisions.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given 9 parameters, no annotations, and no output schema, the description should explain more about invocation and expected results. It provides useful domain context but omits parameter usage, pagination behavior, and return shape, leaving the agent without guidance for several important call decisions.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description adds no parameter-specific meaning beyond the schema. With 9 parameters and only 67% schema description coverage, several parameters (save, world, offset) lack documentation in either place, and the description does not compensate or explain how to use query, show, as_of, or retired parameters. The schema does some work, but the description contributes nothing to parameter semantics.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific resource (MAM research) and the three questions it answers: what is left, what it costs, and what you can afford right now. It also distinguishes MAM capabilities from recipes, which helps separate it from recipe-oriented siblings. It does not explicitly name a sibling tool, so it stops short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by the opening question set and by the contrast between MAM capabilities and recipes. However, there is no explicit 'use this when' or 'do not use this when' guidance, and no named alternatives among the many sibling tools. The agent must infer the appropriate context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
milestonesA
HUB milestones: what is left, what each costs, and what you can afford right now.
The questions mam_research answers about the MAM tree, asked of the other ladder and
in the same words -- both walk one SchematicLadder priced against the same spendable
stock, so a status here means what it means there.
READY is about the BILL, not about access: tiers are opened by Space Elevator
deliveries, that gate is in no shipped data, and phase_requirements is where the
elevator stands.
| Name | Required | Description | Default |
|---|---|---|---|
| save | No | ||
| show | No | all | todo | affordable -- todo hides finished milestones | todo |
| tier | No | one HUB tier, 1-9. Omit for all of them | |
| as_of | No | pin to one world state: a sav:… token from an earlier answer | |
| limit | No | max rows (hard cap 25) | |
| query | No | filter by name, case-insensitive | |
| world | No | ||
| offset | No | ||
| search | No | retired -- write query= instead | |
| status | No | retired -- write show= instead |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, and it delivers a genuinely useful disclosure: READY reflects the bill, not access, and the access gate is not present in shipped data. That caveat prevents a real misinterpretation, though return format and pagination remain unstated.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The first line is well front-loaded, but the remaining prose is analogy-heavy and cryptic ('SchematicLadder', 'spendable stock', 'in the same words'), which costs the agent effort without proportional payoff.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 10-parameter, no-annotation, no-output-schema tool, the description explains the concept and the key caveat but leaves the parameter set and result shape largely to the schema. Adequate but with clear gaps for a tool of this complexity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 70%, so the schema already documents most parameters (show, tier, as_of, limit, query, retired params). The description hints at the show=affordable semantics but adds no syntax or format detail beyond the schema, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening line states the resource (HUB milestones) and the three things it returns: what's left, what each costs, and what you can afford now. It also distinguishes itself from siblings by explicitly framing itself as the same questions `mam_research` answers, applied to the other ladder. The purpose is clear despite the dense jargon.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It implicitly routes the agent by noting that `phase_requirements` is where the elevator gate stands and that `mam_research` covers the parallel ladder, but it never states an explicit 'use this when' condition. Usage must be inferred from the analogies rather than being spelled out.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
name_factoryA
Name a set of machines and persist it for this world.
The label stores the machine instance ids, which are stable across saves, so it
survives moving machines, adding to the factory, and autosave rotation. Calling
this again with the same name RE-ANCHORS it to the current selection, dropping every
machine the new selector misses -- amend_factory adds or drops a few without that,
and rename_factory changes the name without touching the membership.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | ||
| save | No | ||
| as_of | No | pin to one world state: a sav:… token from an earlier answer | |
| notes | No | ||
| split | No | keep only the largest spatial cluster | |
| world | No | ||
| expand | No | pull in everything belted to the result | |
| select | Yes | selector terms, ANDed. product:<item> | recipe:<name> | building:<class or name> | near:<place>@<radius_m> | base:<n> | line:<n> | slab:<n> | proposal:<n> | label:<name> | machine:<instance> | all. Terms are ANDed; comma-separated values inside one term are ORed; prefix a term with '-' to exclude it | |
| dry_run | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does substantial work: it discloses that membership is stored as stable instance ids surviving moves/autosave rotation, and warns that the same-name call is destructive to prior membership. It does not cover permissions, rate/latency, or what dry_run actually returns, which for a 9-param mutation tool is a real but smaller gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, front-loaded with the core action, then persistence rationale, then the destructive re-anchor warning and sibling contrasts. No filler; every sentence carries distinct information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 9-parameter, no-annotation, no-output-schema tool, the description covers the headline behavior and the destructive case well but omits dry_run, save/world scoping, and any hint of the return value (e.g. a label token). Adequate but with clear gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 44%, so the description must compensate, and it partially does by explaining the semantics of `name` (a label over stable instance ids) and the re-anchor behavior tied to `select`. But it says nothing about save, world, notes, or dry_run, leaving nearly half the parameters documented in neither place.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (name) and resource (a set of machines) plus the persistence effect ('persist it for this world'). It explicitly distinguishes itself from the two nearest siblings, amend_factory and rename_factory, so an agent can route without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Names both alternatives and the condition that selects each: re-calling with the same name re-anchors and drops non-matching machines, amend_factory adds/drops a few, rename_factory changes only the name. That is explicit when-to-use/when-to-use-something-else guidance rather than inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
phase_requirementsC
What the Space Elevator still wants, live record and deprecated record apart.
The per-phase item table in the save is DEPRECATED and frozen, so it is shown labelled rather than believed. Read the header line first.
| Name | Required | Description | Default |
|---|---|---|---|
| save | No | ||
| as_of | No | pin to one world state: a sav:… token from an earlier answer | |
| world | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, and it does disclose a genuine behavioral caveat: the per-phase item table is DEPRECATED and frozen, and is shown labelled rather than trusted. That is useful data-fidelity context. It still omits permissions, return shape, and error behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
It is short and front-loads the subject before the deprecation caveat. The awkward line breaks and lore-flavored phrasing cost it a point, but there is little filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, no annotations, and three loosely documented parameters, the description should work harder. It leaves the return format, the meaning of save/world, and the tool's relationship to sibling phase/milestone tools to inference.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 33% (only as_of is documented), and the description explains neither 'save' nor 'world', so it fails to compensate for the coverage gap. It does loosely gesture at the as_of concept ('pin to one world state') but that meaning already lives in the schema, adding nothing.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description does identify a specific resource (Space Elevator phase requirements) and distinguishes live vs. deprecated records, which no sibling covers. However the verb is oblique ('what the Space Elevator still wants') and it never states cleanly that it returns per-phase item requirements. An agent can guess the domain but not the exact output.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use or when-not-to-use guidance, and no sibling alternative is named despite many plausible overlaps (milestones, mam_research, unlocked_recipes). The only procedural hint is 'Read the header line first,' which concerns reading the output, not deciding to call the tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
plan_factoryA
Optimise a factory with an LP over this world's unlocked recipes.
sources says which resource nodes may feed the plan, as a list of selectors --
named regions, radii, grid cells, compass directions, or specific node ids::
["north"] everything in the northern half
["region:Northern Forest"] one named region
["near:0,-2000@900"] within 900 m of (0, -2000) metres
["node:BP_ResourceNode30_103"] one exact node (repeatable)
["grid:X3Y4", "grid:X3Y5"] specific grid cells
["north", "resource:Crude Oil"] narrow a location to one resourceOmit it and the whole map is in scope. Use search_resource_nodes to discover ids.
Machine counts are whole buildings at a derived clock: a 52.8 machine-equivalent result is reported as 53 machines at 99.6%. That is exact, always a clean ratio, and provably the power-optimal way to run that throughput, so ordinary ratio underclocking is automatic and needs no parameter.
extractor_clocks overclocks the SOURCE NODES only, e.g. [1.0, 1.5, 2.0, 2.5].
That is the usual play: a node set is fixed, so speed is the only way to get more
out of it, whereas overclocking production machines mostly burns power. Each
machine above 100% needs Power Shards, which nothing here counts.
clocks is only for asking a different question: passing [0.5, 1.0] lets the
solver SPREAD throughput over more machines to save power, which is real but not
free, so each machine is priced at machine_cost_mw (default 5 MW, just above
the 2.58 MW/machine that trade was measured to be worth). Overclock modes are not
offered by default because they consume Power Shards, which nothing here counts.
objective: max_mw | max_item | min_raw | min_machines | min_power. Every item is balanced as an EQUALITY, so a byproduct with no consumer makes the plan infeasible rather than silently vanishing.
exports is the whitelist of what may leave, and the single most load-bearing
argument here; default is power only, which is often infeasible for crude oil::
exports=["MW"] power out, plant must be self-powered
exports=["Plastic", "Rubber"] items out, NO power export
exports=["MW", "Plastic", "Rubber"] both -- MW must be listed explicitlyTwo things worth reading twice. The power token is MW (mw, power and
Power all work too), not the item name of anything. And exports
replaces the default rather than extending it: naming an item drops MW, which
is deliberate, because exporting MW also forbids drawing from the existing grid.
A token matching no item is refused by name rather than solved around.
sloops is a BUDGET, not a switch: it is how many Somersloops you will actually
commit, and the solver spends up to that many wherever they buy the most. Default 0
spends none, because only a fixed number exist on the whole map and a plan that
quietly assumed them would be unbuildable. Each one costs 4x power for 2x output on
its machine, so they are placed one at a time across many machines rather than
filling one -- output is linear in sloops and power is quadratic, so spreading wins.
required names recipes (exact name or class id) that must make their item: every
other recipe whose main product is that item is excluded. A locked or banned one is
refused by name.
logistics_items pins named items into the belt/pipe table however small their
flow, as rows ADDED to the limit biggest by volume. Without it, a two-item
question can fall off the bottom of a big plan's flow table.
save_as stores the request. Over an existing plan it needs base_rev, the
version you read (list_plans name=): edits to different settings since then merge,
the same setting changed by someone else is refused as outdated and nothing is saved.
site_at says where the plan will STAND. On its own it makes the plan's water
assumption MEASURED rather than assumed: the terrain at that pad is read and the note
quotes how much of it is under water, at what level, and how far below the dry ground.
It never changes a number the LP produced -- how many extractors a body of water holds
is placement geometry no data here carries. With save_as it is also recorded, with
yaw and footprint, so later calls can answer "does what stands there match it"
(diff_vs_save) and "show me" (show_on_map at='plan:'); a recalled plan that
was sited is measured at its own site without being told again. Use site_plan to set or
move the siting of an already-stored plan.
| Name | Required | Description | Default |
|---|---|---|---|
| plan | No | recall a saved plan by name | |
| save | No | ||
| as_of | No | pin to one world state: a sav:… token from an earlier answer | |
| limit | No | max rows (hard cap 25) | |
| world | No | ||
| clocks | No | ||
| sloops | No | Somersloops the plan may spend; 0 spends none | |
| exports | No | ||
| save_as | No | store this request under a name | |
| site_at | No | with save_as: record where this plan will STAND -- 'x,y[,z]' in metres, 'me', a factory name, 'slab:<n>' or a run id (the footprint's centre) | |
| sources | No | ||
| base_rev | No | the plan version you read; needed to save over an existing plan | |
| required | No | recipes that must make their item; others for it are excluded | |
| supplied | No | items another plan hands this one, {item: per-minute} | |
| objective | No | max_mw | |
| allow_sinks | No | ||
| for_factory | No | factory label this plan is for | |
| target_item | No | ||
| only_recipes | No | ||
| recycle_once | No | recipes that may run but must not feed each other, e.g. ['Recycled'] | |
| site_yaw_deg | No | site orientation: degrees about world Z, positive +X towards +Y | |
| site_footprint | No | site footprint 'WxD' in metres; blank = the layout's own square | |
| exclude_recipes | No | ||
| export_minimums | No | ||
| logistics_items | No | items whose belt/pipe rows to pin, whatever their volume | |
| machine_cost_mw | No | ||
| only_free_nodes | No | ||
| plan_notes_text | No | note stored with save_as | |
| extractor_clocks | No | ||
| water_extractors | No | how many Water Extractors your site can actually hold |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does so well: it explains the LP balances items as equalities (making byproducts infeasible rather than vanishing), reports machine counts as exact clocked ratios, discloses that unknown export tokens are refused by name, describes the base_rev merge/conflict semantics on save, and clarifies that site_at measures water but never alters LP numbers.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The length is largely justified for a 30-argument LP tool, and the purpose is front-loaded before the argument tour. Some asides ('Two things worth reading twice', the quadratic-vs-linear rationale) are informative rather than wasteful, but the block is dense enough that it could be tightened without losing meaning.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex tool with 30 params, no annotations, and no output schema, the description covers usage, infeasibility behavior, save/merge semantics, and related tools thoroughly. The remaining gap is the handful of unexplained arguments, so it is strong but not fully complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 53% across 30 params, and the prose compensates heavily by explaining sources selectors, clocks vs extractor_clocks, exports/MW token, sloops budget, objective, required, logistics_items, save_as/base_rev, and site_at. Still, several parameters (allow_sinks, only_free_nodes, only_recipes, exclude_recipes, export_minimums) receive no prose explanation at all.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The first sentence states a specific verb and resource ('Optimise a factory with an LP over this world's unlocked recipes'), and the body explicitly distinguishes the tool from siblings like search_resource_nodes (for discovery), list_plans (naming saved plans), site_plan (setting siting), diff_vs_save, and show_on_map. An agent can tell what this does and what it does not do.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Extensive guidance on when to use which argument: extractor_clocks is 'the usual play' vs clocks being 'only for asking a different question', exports replaces rather than extends the default, sloops is a budget not a switch. It also routes to search_resource_nodes for finding ids and site_plan for repositioning. However it never explicitly states when to choose plan_factory over plan_layout/commission_plan.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
plan_layoutA
Turn a plan into a buildable schematic: blocks, buses and floors.
Same arguments as plan_factory, plus show: "floors" (default, the stack),
"blocks" (every module with its size and rates), "buses" (item flows),
"trunks" (which resource nodes share each pipe or belt run into the site),
"materials" (what the whole thing costs to build, machines plus deck), or
"sites" (cut the plan into named modules and report what crosses between them).
This is a SCHEMATIC, not a blueprint. It gives modules, connections, floor assignment and a space budget. It deliberately does NOT give world coordinates or belt routing -- there is no terrain data here, so those would be invented.
Blocks are split by throughput: 46 Refineries needing 1380 m3/min of crude cannot share one manifold when a Mk2 pipe carries 600, so that is 3 blocks. Floors follow chain depth, with a logistics deck between each pair of production floors.
| Name | Required | Description | Default |
|---|---|---|---|
| plan | No | recall a saved plan by name | |
| save | No | ||
| show | No | floors | blocks | buses | trunks | materials | sites | floors |
| as_of | No | pin to one world state: a sav:… token from an earlier answer | |
| limit | No | max rows (hard cap 25) | |
| sites | No | show="sites": {"rig": ["Heavy Oil Residue", ...], "hall": ["MW"]} -- MW/power claims every generator | |
| world | No | ||
| clocks | No | ||
| detail | No | retired -- write show= instead | |
| sloops | No | Somersloops the plan may spend; 0 spends none | |
| exports | No | ||
| factory | No | fit the layout against this factory's existing platform | |
| sources | No | ||
| belt_tier | No | belt tier name; blank = the fastest you have unlocked | |
| objective | No | max_mw | |
| pipe_tier | No | pipe tier name; blank = the fastest you have unlocked | |
| allow_sinks | No | ||
| target_item | No | ||
| only_recipes | No | ||
| exclude_recipes | No | ||
| export_minimums | No | ||
| machine_cost_mw | No | ||
| only_free_nodes | No | ||
| order_floors_by | No | "chain" (build order) or "head" (minimise fluid lift) | chain |
| extractor_clocks | No | ||
| water_extractors | No | how many Water Extractors your site can actually hold | |
| max_floor_foundations | No | cap a deck at this many 8m foundations; 0 = one stage per deck |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does real work: it discloses the deliberate omissions (no world coordinates, no belt routing, because no terrain data) and explains the derivation rules (blocks split by throughput, floors follow chain depth with logistics decks between). It does not state whether the operation is read-only/side-effect-free or what cost/preconditions exist, so a 4 rather than 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core purpose, then the `show` modes, then the deliberate exclusions, then the derivation rationale. Sentences are dense but each earns its place. Slightly long, and the blocks/floors rationale paragraph could be trimmed, but structure is sound.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 27-parameter tool with no annotations, no output schema and 48% schema coverage, the description covers purpose and output modes well but leaves more than half the parameters and the entire side-effect/precondition surface unexplained. The 'same as plan_factory' pointer helps but shifts burden onto the sibling's definition.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Only 48% of 27 parameters are schema-documented, so the description must compensate. It richly explains `show` (the main control) and clarifies blocks/floors semantics, and defers the rest via 'Same arguments as plan_factory'. But many meaningful parameters (objective, allow_sinks, exports, sources, clocks, only_recipes, machine_cost_mw) get no explanation in either place, leaving the agent guessing.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening sentence states a concrete transformation (plan -> buildable schematic) and then enumerates the six output shapes via `show`. It explicitly positions itself against the sibling plan_factory ('Same arguments as plan_factory, plus show'), so the agent can distinguish it without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It tells the agent what the tool is for and names the alternative it extends (plan_factory), and it rules out a class of use cases ('schematic, not blueprint... no world coordinates or belt routing'), implicitly routing coordinate-level needs elsewhere. It does not give an explicit when-not-to-use clause or name the sibling that provides coordinates, so it stops short of 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
plan_logA
One plan's versions, newest first: who changed what, from chat or from the page.
undo=<v> writes a new version that reverses that one; restore=<v> writes a new
version equal to that one. Both take base_rev and merge like any other edit. A
forgotten plan is found too, so its forget can be undone.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | ||
| save | No | ||
| undo | No | undo this version (needs base_rev) | |
| as_of | No | pin to one world state: a sav:… token from an earlier answer | |
| limit | No | max rows (hard cap 25) | |
| since | No | list versions after this one | |
| world | No | ||
| restore | No | make the head equal this version (needs base_rev) | |
| base_rev | No | the plan version you read; needed to change an existing plan |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and largely succeeds: it discloses that undo/restore are non-destructive (they write a NEW version), that both require base_rev and merge like any other edit, and that forgotten plans remain retrievable so a forget can be undone. It omits permissions/auth requirements and rate limits, but the core behavioral traits are well covered.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The listing behavior is front-loaded in the first sentence, with the mutation semantics following as supporting detail. Every sentence conveys distinct information with no filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given 9 parameters, no annotations, and no output schema, the description is fairly complete: it describes the return ordering ('newest first'), the attribution ('who changed what'), and the mutation semantics. The main gaps are the undocumented save parameter and the lack of any note on how as_of/since interact with the returned rows.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
At 67% schema coverage, the schema already documents undo, restore, base_rev, as_of, limit, and since. The description meaningfully adds to the schema by clarifying that undo/restore create new versions rather than rewriting history and that they merge through the normal edit path. It leaves save and world undescribed in both places, so it does not fully compensate for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific resource and scope: 'One plan's versions, newest first: who changed what, from chat or from the page.' This clearly sets it apart from the sibling list_plans, which lists plans rather than the version history of a single plan. However, it never names list_plans or diff_vs_save as the alternatives, so the differentiation is implicit rather than explicit.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives good usage guidance for the mutation paths, explaining that undo=<v> and restore=<v> both require base_rev and merge like any other edit. But it offers no guidance on when to reach for plan_log versus list_plans, diff_vs_save, or site_plan, which is the primary routing decision an agent faces.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
power_reportB
Generation capacity vs machine draw, nameplate AND measured.
Nameplate is what everything built would draw running at once. Measured weights each machine by the 300 s productivity monitor the save already carries, which on a factory with idle blocks is a very different number -- and it is the one that says what is free right now. Both are shown because they answer different questions.
Generation is capacity on both figures, with one exception the answer names: a generator whose fuel or supplemental water has run dry AND whose own monitor read zero is listed as starved, because those MW will not arrive when the grid asks for them.
| Name | Required | Description | Default |
|---|---|---|---|
| save | No | ||
| as_of | No | pin to one world state: a sav:… token from an earlier answer | |
| world | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does unusually well conceptually: it defines nameplate vs the 300 s productivity-weighted measured figure, explains why both are returned, and documents the 'starved' labeling rule for dry-fuel generators whose monitor reads zero. It omits operational traits like permissions, cost, or latency, but the analytic semantics are genuinely disclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core comparison, then the explanation of the two figures and the starved exception. Slightly verbose in places ('Both are shown because they answer different questions'), but no sentence is empty.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema and no annotations, so the description must carry return-value and safety information; it conveys the report's contents (both figures, starved grouping) but never addresses the two undocumented input parameters or confirm this is a read-only query. Adequate for the conceptual core, incomplete around inputs.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Only as_of carries a schema description (33% coverage); save and world are undocumented in both schema and description. The description adds no parameter meaning at all, so it does nothing to compensate for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening line names the resource and scope precisely: generation capacity versus machine draw, in both nameplate and measured form. It is clear this returns a power analysis of a save, though it never explicitly differentiates itself from power-adjacent siblings like factory_health or world_summary.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no statement of when to call this versus alternatives, nor any prerequisite or context cue. The description explains what the numbers mean once you get them, but not why or when an agent should reach for this tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
power_shardsB
Power Shards held, committed and free, plus what an overclock plan would cost.
plan_machines machines at plan_clock costs shards per machine; the answer
says whether the free pool covers it.
| Name | Required | Description | Default |
|---|---|---|---|
| save | No | ||
| as_of | No | pin to one world state: a sav:… token from an earlier answer | |
| limit | No | max rows (hard cap 25) | |
| world | No | ||
| offset | No | ||
| plan_clock | No | ||
| plan_machines | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It discloses that the answer reports held, committed and free shards and whether the free pool covers a plan, but it does not state that this is a read-only query, nor does it describe pagination behavior or any side effects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two tight sentences with no filler. It front-loads the output content and then explains the overclock-plan computation, so every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with 7 parameters, no annotations, no output schema, and 29% schema coverage, the description is too thin. It covers the core shard accounting and plan cost, but omits parameter guidance, read-only status, and pagination or return-shape details an agent would need.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 29% (7 parameters, only as_of and limit have schema descriptions). The description adds meaning for plan_machines and plan_clock, but leaves save, world, offset, and the relationship between limit and offset unexplained, so it does not compensate for the low coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific resource (Power Shards) and what is reported: held, committed, free, and the cost of an overclock plan. It clearly tells an agent what the tool computes, though it does not explicitly differentiate itself from the sibling power_report tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by the explanation that plan_machines at plan_clock costs shards per machine and the answer says whether the free pool covers it. There is no explicit when-to-use statement, no when-not-to-use guidance, and no named alternative among the many sibling tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
propose_factoriesB
One coherence score over every signal, agglomerated into proposed factories.
Combines foundation slabs, proximity, belt connectivity, shared products and supply links. Validated leave-one-factory-out against the player's twelve hand-named factories: precision 1.000, recall 0.945, and precision was 1.000 on every fold -- it never merges two factories, it only ever splits one.
Use name_factory on what it proposes. unnamed_only=True answers "what have I
built and not named". The # column is the proposal:<n> selector every other tool
takes, and it counts over ALL proposals -- so it does not shift when you page.
| Name | Required | Description | Default |
|---|---|---|---|
| save | No | ||
| as_of | No | pin to one world state: a sav:… token from an earlier answer | |
| limit | No | max rows (hard cap 25) | |
| world | No | ||
| offset | No | ||
| max_span_m | No | cap on a proposal's diameter, metres | |
| unnamed_only | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and discloses several key traits: the signals combined, validation metrics, that it never merges factories but only splits them, and that the '#' column remains stable across pagination. It stops short of explicitly confirming read-only behavior or explaining whether any persistent state is modified.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The purpose is front-loaded in the first sentence, and the rest is mostly useful behavioral context. The validation statistics are somewhat verbose, but they support trust in the proposal quality and do not bury the main point.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 7-parameter tool with no annotations and no output schema, the description covers core behavior, follow-up action, and unnamed filtering adequately. It remains incomplete on several input parameters and does not fully describe output structure or side effects.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 43%, so the description needs to compensate but largely does not. It explains unnamed_only and pagination stability, but leaves save, world, as_of, offset, and max_span_m without added meaning beyond the sparse schema descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: it proposes factories by agglomerating a coherence score over signals. It is clearly not a general factory-list or query tool, but it does not explicitly differentiate itself from siblings like factory_query or list_factories.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives follow-up guidance ('Use name_factory on what it proposes') and a use case for unnamed_only=True ('what have I built and not named'). However, it does not state when to call this tool instead of alternatives such as factory_query or list_factories.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
rank_build_sitesA
Rank candidate fields for a new extraction site, best first.
Scores untapped REACHABLE capacity against spread, distance to your existing buildings, and purity mix. Every raw component is shown so you can re-weight: the single score is a starting point, not a verdict.
sources narrows the search area using the same selectors as
search_resource_nodes; omit it to search the whole map.
A ranking does not page: the rows below the cut score worse by construction, so raise
limit or narrow sources rather than looking for an offset.
| Name | Required | Description | Default |
|---|---|---|---|
| top | No | deprecated alias for limit | |
| save | No | ||
| as_of | No | pin to one world state: a sav:… token from an earlier answer | |
| limit | No | max rows (hard cap 25) | |
| world | No | ||
| sources | No | ||
| resource | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does meaningful work: it discloses the scoring inputs (reachable capacity, spread, distance to existing buildings, purity mix), that components are exposed for re-weighting, and that the result set is non-pageable by construction. It does not discuss permissions or whether it is purely read-only, but the ranking framing makes the operation's nature clear.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with purpose and scoring method, then usage and paging caveats in short paragraphs. Slightly padded by the double-backtick formatting and a couple of stylistic asides, but every sentence carries information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 7-parameter tool with no annotations and no output schema, the description covers ranking semantics and the non-paging contract well but omits the meaning of the required `resource` param and the `save`/`world`/`as_of` world-state controls, leaving gaps an agent must guess at.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 43% across 7 params, so the description must compensate. It adds real meaning for `sources` (same selectors as search_resource_nodes; omit to search whole map) and `limit` (raise rather than offset), but leaves `resource`, `save`, `world`, and `as_of` unexplained beyond sparse schema text.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Rank candidate fields for a new extraction site') and immediately differentiates from siblings by referencing search_resource_nodes selectors and contrasting with paging behavior. An agent can distinguish it from factory_sites, site_plan, or search_resource_nodes without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear usage context: omit `sources` to search the whole map, reuse search_resource_nodes selectors, and raise `limit` rather than seeking an offset. It stops short of naming when NOT to use it versus close siblings like site_plan or factory_sites, so no explicit exclusion guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
rank_unlocksA
What every locked alternate recipe would be worth to THIS plan.
One counterfactual per candidate: solve the plan, solve it again with the recipe added, report the difference. It answers "which unlock should I chase" with a number in the plan's own units instead of a tier list, because a recipe's worth depends entirely on what you already have.
A zero is an answer. Most candidates change nothing, and "you are not missing anything here" is a decision -- it is otherwise reached by walking the recipe tree by hand.
Deltas are an UPPER bound: a candidate needing a machine you have not built is judged as if you had it, and the machine is named. Alternates currently offered by a pending hard drive are flagged, which is the difference between "worth having" and "claimable now".
| Name | Required | Description | Default |
|---|---|---|---|
| plan | No | recall a saved plan by name | |
| save | No | ||
| as_of | No | pin to one world state: a sav:… token from an earlier answer | |
| limit | No | max rows (hard cap 25) | |
| query | No | only test alternates whose name matches | |
| world | No | ||
| clocks | No | ||
| search | No | retired -- write query= instead | |
| sloops | No | ||
| exports | No | ||
| sources | No | ||
| objective | No | max_mw | |
| allow_sinks | No | ||
| target_item | No | ||
| only_recipes | No | ||
| exclude_recipes | No | ||
| export_minimums | No | ||
| machine_cost_mw | No | ||
| only_free_nodes | No | ||
| extractor_clocks | No | ||
| water_extractors | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries real weight and does so: it explains the counterfactual method, that zeros are meaningful answers, that deltas are an UPPER bound assuming unbuilt machines, and that pending-hard-drive alternates are flagged. It omits cost/latency given it solves the plan once per candidate over up to 25 rows, so it is not fully transparent, but the behavioral disclosure is unusually rich.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core definition, then caveats that each earn their place (zero-as-answer, upper bound, hard-drive flag). Slightly essayistic in tone, but no sentence is padding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 21-parameter, no-output-schema tool, the description explains what the deltas mean but never sketches the shape of the result (one row per candidate? what columns?) or how a plan/world is identified among the many optional arguments. Adequate for intent, incomplete for invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 24% across 21 parameters, so the description is expected to compensate and does not. Terms like 'candidate' and 'plan' are used informally, but nothing explains plan, as_of, limit, query, world, or the many planning knobs (sloops, exports, objective, machine_cost_mw) — the agent is left to guess at most arguments.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a precise verb+resource (rank unlock candidates by marginal value to a given plan) and immediately differentiates from a tier list. An agent can distinguish it from siblings like alternates_for_item or compare_recipe_options without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The framing 'which unlock should I chase' with a plan-relative number implies the usage context, and the hard-drive flagging hints at the claimable-now case. But no sibling is named and there is no explicit when-not-to-use guidance (e.g., versus advise_hard_drive_pick or alternates_for_item), leaving the routing to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
recipe_detailA
Exact numbers for one recipe: rates, machine, power, unlock source.
Takes a class id OR a display name. Refusing the name cost a caller two round trips
to fetch an id this function could resolve itself, which is a poor trade for strictness
that buys nothing -- match_recipes already does exactly this resolution for
exclude_recipes.
| Name | Required | Description | Default |
|---|---|---|---|
| recipe_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description bears full responsibility. It does disclose a notable behavior: the tool accepts a display name and resolves it internally, which is nontrivial. Yet it omits other behavioral details such as error handling (e.g., what happens if the name is not found), case sensitivity, or the exact structure of the returned data. This is adequate but not thorough.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The first sentence is a model of conciseness, delivering the core purpose in under 10 words. The second sentence is longer and argumentative, defending the design choice to accept display names. While it adds useful context for an AI agent, it could be trimmed without losing essential guidance. Overall, it earns its place but is slightly verbose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple single-parameter lookup tool with no output schema, the description gives a good sense of the return fields (rates, machine, power, unlock source) and the accepted input. It does not mention error cases, but the tool is inherently simple and the description covers the key aspects. A short note on not-found behavior would make it complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema provides only a parameter name and type with no description (0% coverage). The description fully compensates by stating that the parameter accepts either a class ID or a display name, and even explains the trade-off. This is critical semantic information that the schema lacks, making the parameter's behavior completely clear.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: providing exact numbers (rates, machine, power, unlock source) for a single recipe. This distinguishes it from siblings like search_recipes (which searches) or compare_recipe_options (which compares multiple recipes). The scope is specific to one recipe's details.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The intended use case is implicit but clear: when you need exact numbers for a specific recipe. The description also explains that you can pass either a class ID or a display name, avoiding unnecessary round trips, and references match_recipes as an existing pattern for resolution. However, it does not explicitly say when to avoid this tool in favor of alternatives (e.g., for searching or comparison).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
rename_factoryB
Rename a factory label. The machines it holds are not touched and nothing re-anchors.
The label keeps its anchors, notes, centroid, signature and dates -- the name is the only thing here a player picked, and correcting one used to mean naming the whole selection again under a second name and forgetting the first.
Stored plans scoped to this factory follow the new name. Renaming onto a name this world already uses is refused and says which label holds it.
| Name | Required | Description | Default |
|---|---|---|---|
| to | Yes | the new name | |
| name | Yes | ||
| save | No | ||
| as_of | No | pin to one world state: a sav:… token from an earlier answer | |
| world | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden and does well: it discloses that machines are untouched, nothing re-anchors, the label keeps anchors/notes/centroid/signature/dates, stored plans follow the new name, and name collisions are refused with an explanatory message. It stops short of stating permissions or idempotency, keeping it from a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The core facts (rename label, machines untouched, plans follow, collision refused) are front-loaded, but the second paragraph's aside about 'correcting one used to mean naming the whole selection again' is narrative padding that doesn't help an agent decide or invoke.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Behavioral side effects are covered well for an unannotated tool with no output schema, but three of five parameters (save, as_of, world) are left undocumented by both schema and description, so the definition is not complete enough to invoke confidently.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 40% across 5 parameters. The description implies the 'name'/'to' pair but never names them, and says nothing about save, as_of, or world. It does not compensate for the documentation gap the schema leaves.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb+resource: 'Rename a factory label.' An agent immediately knows the operation. However, it never distinguishes itself from close siblings like name_factory or amend_factory, so the sibling boundary is left implicit.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit when-to-use or when-not-to-use guidance. The description mentions a collision case (renaming onto an existing name is refused) but does not tell the agent when to reach for rename_factory versus name_factory or amend_factory. Context is implied at best.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
rename_planA
Rename a saved plan. Nothing is re-solved and nothing else about it changes.
The plan keeps its key, its recorded field, its siting and its notes -- a name is the only thing here a player picked, and it was the only thing they could not correct without saving the plan again under a second name and forgetting the first.
| Name | Required | Description | Default |
|---|---|---|---|
| to | Yes | the new name | |
| name | Yes | ||
| save | No | ||
| as_of | No | pin to one world state: a sav:… token from an earlier answer | |
| world | No | ||
| base_rev | No | the plan version you read; needed to change an existing plan |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does well: it declares the mutation is narrowly scoped, that nothing is re-solved, and enumerates what is preserved (key, recorded field, siting, notes). It omits permissions/error behavior and the base_rev concurrency requirement, so it stops short of a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose is front-loaded in the first sentence, and the rationale that follows is short. The second sentence is slightly run-on and rhetorical ('forgetting the first'), but it earns most of its space by explaining the non-effects.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Covers the what-changes/what-doesn't question well for a mutation tool with no annotations and no output schema. But with 6 parameters at 50% coverage and no output schema, the definition leaves parameter usage and success/return behavior uncovered, so it is adequate rather than complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 50% and the description adds no parameter-level meaning at all. It never clarifies the crucial distinction between the 'name' parameter (which plan) and 'to' (the new name), nor the roles of save/world, leaving half the parameters undocumented in both places.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Rename a saved plan') and immediately scopes it by declaring what does not happen ('Nothing is re-solved and nothing else about it changes'). An agent can distinguish this from the factory-naming siblings (rename_factory, name_factory) purely from the resource noun.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The second sentence gives motivating context for when renaming is the right move ('a name is the only thing here a player picked, and it was the only thing they could not correct'), which implies usage. However, it never explicitly names an alternative (e.g. amend_factory, plan_factory) or states a when-not condition, so routing is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_conduitsA
Belt and pipe runs near a point or between two areas: ends, length, elevation.
The web map has drawn these all along; this is the text answer to "is there a pipe between those extractors and that platform, where does it run, how long is it". A run is one belt CHAIN (consecutive conveyor pieces, split at splitters, mergers and machines) or one placed pipeline piece. Longest first; each row carries both ends with what stands there where known, the drawn length, and the elevation span.
show="networks" answers the other size of question: one row per FLUID NETWORK in
the whole world, what each carries, how much pipe it is, where its middle is and
what it ends on. A network is one connected plumbing system, so that is the view
that tells you which system a run belongs to; radius_m and to do not narrow it,
and the distance column places each network relative to near.
near and to accept a coordinate in metres, me, a named factory, or one of this
tool's own run ids -- chain:7, pipe:333 -- which centres on that run's midpoint,
so the ids in the connects column can be followed one call at a time. With to
set, only runs passing within both radii are listed. Proximity is measured against
the runs' drawn lines, not their corner points, so a run crossing mid-span counts.
Long lists page with offset=, and the truncation line names the next offset --
a busy junction can carry hundreds of chains and the tail of that list is as real
as its head.
| Name | Required | Description | Default |
|---|---|---|---|
| to | No | second area: list only runs passing near BOTH, same forms as near | |
| kind | No | retired -- write conduit_kind= instead | |
| near | Yes | centre: 'x,y' in metres, 'me', a named factory, or a run id from this tool ('chain:7', 'pipe:333') | |
| save | No | ||
| show | No | runs | networks | runs |
| as_of | No | pin to one world state: a sav:… token from an earlier answer | |
| limit | No | max rows (hard cap 25) | |
| world | No | ||
| offset | No | ||
| radius_m | No | ||
| to_radius_m | No | radius around `to`, defaults to radius_m | |
| conduit_kind | No | belt | pipe | all |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, and it does so well: it discloses ordering ('longest first'), the columns returned (both ends, drawn length, elevation span, what it ends on), that proximity is measured against drawn lines not corner points, and that lists page via offset with a truncation line naming the next offset. It omits nothing critical behaviorally, though error/auth behavior is unaddressed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded and organized into one paragraph per concern, which is good structure for a 12-parameter tool. However it carries rhetorical padding ('the tail of that list is as real as its head', 'the web map has drawn these all along') that does not help an agent invoke it, so not every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 12-parameter tool with no output schema and no annotations, the description covers the core flow and the return shape adequately. But it never mentions `conduit_kind`, which is the central filter distinguishing belts from pipes, nor `save`/`world`, leaving real gaps an agent would have to infer.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 67%, but the description compensates for the key parameters: it details the accepted forms of `near`/`to` (metres coordinate, 'me', named factory, run id like 'chain:7'), explains that run ids centre on the run midpoint so `connects` ids can be followed call-by-call, and notes that `radius_m`/`to` do not narrow the networks view. It is silent on `conduit_kind`, `save`, and `world`.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource -- 'belt and pipe runs near a point or between two areas', plus a second mode ('show="networks"' returns fluid networks). An agent can tell exactly what class of object is returned. It never names a sibling tool (e.g. show_on_map, trace_upstream) to differentiate, which keeps it below a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear context for use -- it frames the tool as the text answer to 'is there a pipe between those extractors and that platform' -- and explains when each mode applies ('show="networks"' answers the other size of question). It also states the condition that selects the `to` filter. It lacks explicit when-NOT-to-use guidance and never compares to alternative tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_itemsC
Find items by name. Returns form, energy and sink points.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | max rows (hard cap 25) | |
| query | Yes | ||
| offset | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It states the return fields but does not mention pagination, sorting, matching behavior, result limits, or any side effects. For a search tool, this is minimal transparency.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two short sentences, front-loaded with the action and return fields. It contains no unnecessary words or repetition of schema details.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no output schema and no annotations, the description should explain the result format and pagination behavior. It lists return fields but not their structure (e.g., array vs object) or how limit/offset affect results. The description is too sparse for a search tool with 3 parameters.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 33%, with only 'limit' documented. The description clarifies that 'query' is for the item name but provides no meaning for 'offset' or additional details on limit/offset behavior. Thus, it does not compensate adequately for the low schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Find') with the resource 'items' and method 'by name,' clearly stating the tool's function. It mentions the return fields (form, energy, sink points), adding specificity, though it does not explicitly distinguish from sibling tools like search_recipes.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool compared to alternatives such as search_recipes, recipe_detail, or alternates_for_item. The description only states the action, leaving the agent to infer usage context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_recipesA
Search recipes by name, or by what they consume/produce. Marks HAVE/LOCKED.
consumes="Rubber" is the reverse lookup: every recipe that eats an item.
recipe_kind is "part" (default), "building" (build-gun costs), "manual" or
"all" -- and the header counts EVERY kind over the whole recipe table whatever it
is set to, so a part-only view still says how many buildings eat the item.
| Name | Required | Description | Default |
|---|---|---|---|
| kind | No | retired -- write recipe_kind= instead | |
| save | No | ||
| as_of | No | pin to one world state: a sav:… token from an earlier answer | |
| limit | No | max rows (hard cap 25) | |
| query | No | ||
| world | No | ||
| offset | No | ||
| consumes | No | ||
| produces | No | ||
| recipe_kind | No | part | building | manual | all | part |
| include_events | No | ||
| only_alternates | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does add real behavioral context: that results are marked HAVE/LOCKED and that header counts aggregate EVERY recipe kind regardless of the recipe_kind filter. It still omits return shape and pagination behavior, but the non-obvious counting quirk is exactly the kind of disclosure that matters.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded purpose statement followed by two tightly scoped clarifications with inline code examples. No filler sentences and nothing buried.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
A 12-parameter tool with 33% schema coverage, no annotations and no output schema needs substantially more than this. Undocumented booleans like include_events and only_alternates, plus save/world/as_of scoping and pagination via limit/offset, are left entirely to inference.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 33%, so the description must compensate and it partly does — clarifying consumes/produces semantics that the bare schema does not. But query, save, world, as_of, include_events, only_alternates, limit and offset get no explanation beyond their name/default, leaving a significant gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource (search recipes) and two distinct search modes: by name and by consume/produce relationships. It clearly is not search_items or recipe_detail, though it never names those siblings to sharpen the boundary.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explains one usage nuance well — that ``consumes="Rubber"`` is the reverse lookup for recipes that eat an item — and describes recipe_kind values. However it gives no guidance on when to use query vs consumes/produces, or when to prefer recipe_detail over a search.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_resource_nodesA
Resource nodes, in one of three views.
fields (default) clusters nodes within 200 m and ranks by yield -- "where is there a lot of iron".
nodes lists one row per node, ranked by yield, with ids reusable as selectors.
nearest lists one row per node ranked by DISTANCE from
near, with the distance shown -- "what is closest". Requiresnear.
sources is a list of selectors; locations union, filters intersect::
["north"] northern half of the map
["region:Northern Forest"] one named region
["near:0,-2000@800"] within 800 m of (0, -2000) metres
["grid:X3Y4"] one 1.024 km grid cell
["node:BP_ResourceNode26_99"] one specific node
["north", "resource:Crude Oil"] crude oil in the north
["bbox:-500,-2500,600,-1800"] a rectangle, metresnear accepts a coordinate in metres, me for the player, or the name of a
labelled factory -- "the nearest free coal to the coal powerplant" needs no
coordinates. Giving near in any view adds a distance column.
All three views page with offset=; the ranking is stable, so the tail of 127 iron
nodes is reachable 25 at a time.
Water is the exception to everything above. Open water carries no node, so asking for it returns only the fracking satellites; the bodies already being pumped, the pumps on each and the measured sea level are printed beside them instead.
| Name | Required | Description | Default |
|---|---|---|---|
| kind | No | node | well_sat | geyser | all | |
| mode | No | retired -- write show= instead | |
| near | No | origin for show=nearest: 'x,y' in metres, 'me', or a factory name | |
| save | No | ||
| show | No | fields | nodes | nearest | fields |
| as_of | No | pin to one world state: a sav:… token from an earlier answer | |
| group | No | retired -- write show= instead | |
| limit | No | max rows (hard cap 25) | |
| world | No | ||
| offset | No | ||
| purity | No | pure | normal | impure | all | |
| sources | No | ||
| resource | No | ||
| only_free | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations the description carries the full burden, and it discloses real behavior: stable ranking, offset paging, a hard 25-row cap, that `near` adds a distance column in any view, and the water exception where non-node fracking satellites plus sea level are returned instead. It does not state read-only-ness or auth/side-effect profile, which keeps it from a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the three-view framing, then groups selector syntax and paging compactly. It is long but nearly every block (selector grammar, near semantics, water exception) earns its place; only the water aside feels slightly tacked on.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 14-param, no-annotation tool with no output schema, the description covers query construction, mode selection, paging, and notable return quirks (distance column, printed sea level). It stops short of describing the shape of a `fields` cluster row or several lesser parameters, leaving some gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
It adds substantial meaning for `sources` (union of locations, intersection of filters) with syntax examples, and clarifies `near` and `offset`. But at 57% schema coverage across 14 parameters, several (save, world, as_of, purity, only_free, resource, kind) are left entirely to the schema, so the description only partially compensates.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific resource (resource nodes) and lays out its three modes precisely, including the question each answers ("where is there a lot of iron", "what is closest"). An agent knows exactly what it retrieves. It never names a sibling tool or draws an explicit boundary against search_items/search_conduits, so it stays short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives clear in-tool routing: fields for abundance, nodes for reusable ids, nearest for proximity, with the requirement that nearest needs `near`. It also gives worked query examples ("the nearest free coal to the coal powerplant"). It does not state when NOT to use this tool versus the many other search_* siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
select_machinesA
Preview which machines a selector picks, before naming them.
Worth running first on anything product-based: 17 machines make Concrete on the reference save, but 15 of them are a construction feed inside the steel site and only one is the player's "concrete setup".
slab:<n> answers "what stands on this platform", and answers it for an empty one
too: a poured platform with nothing on it yet is described rather than refused.
| Name | Required | Description | Default |
|---|---|---|---|
| save | No | ||
| as_of | No | pin to one world state: a sav:… token from an earlier answer | |
| split | No | keep only the largest spatial cluster | |
| world | No | ||
| expand | No | pull in everything belted to the result | |
| select | Yes | selector terms, ANDed. product:<item> | recipe:<name> | building:<class or name> | near:<place>@<radius_m> | base:<n> | line:<n> | slab:<n> | proposal:<n> | label:<name> | machine:<instance> | all. Terms are ANDed; comma-separated values inside one term are ORed; prefix a term with '-' to exclude it |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It usefully discloses two behaviors beyond expectation: product selectors over-match (construction feed inside the steel site) and `slab:<n>` describes an empty platform rather than refusing. However it never states the tool is read-only/preview-only in safety terms, nor how results are returned, leaving gaps for a 6-param tool with zero annotation coverage.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core purpose, then one grounded example and one edge-case note. Every sentence carries information; the Concrete example is a bit long but earns its place by teaching the over-match trap.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema and no annotations, so the description must carry the load. It covers purpose, the key selector gotcha, and slab behavior adequately for a preview tool, though it omits read-only framing and the meaning of save/world/as_of parameters, which an agent would still need.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 67%, and the description meaningfully supplements it: it explains the semantics of product-based selectors (over-matching and how they resolve) and the special `slab:<n>` term including its empty-platform behavior. It does not address save/world/as_of/split/expand, but it adds real semantic value beyond the schema's terse term list.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Preview which machines a selector picks') and adds the dry-run framing 'before naming them,' which signals it is the read-only preview preceding naming tools. It's clear, though it doesn't name the specific sibling (e.g. name_factory) it pairs with.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Worth running first on anything product-based' gives an explicit when-to-use condition and the concrete example (Concrete pulling 17 machines with only one intended) tells the agent exactly when this tool adds value. No explicit when-not or named alternative, so not a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
show_on_mapA
Map links centred on something: this project's own map, and the public one.
Two links for every place. The LOCAL one opens this project's web map, which draws the reader's own save -- their machines, their belts, their siting. The satisfactory-calculator.com one opens a third-party map of the vanilla world, which knows the terrain and the nodes and nothing the player built.
at is the same place vocabulary every other tool takes (see docs/selectors.md),
plus one kind of its own: resource:<name> centres on the centroid of EVERY node of
that resource and switches its overlays on, which is a viewport rather than a place
and is why no other tool accepts it.
Only the Crude Oil layer tokens are confirmed; the rest follow the same pattern and are flagged. A wrong token still opens the map in the right place, just without that overlay.
| Name | Required | Description | Default |
|---|---|---|---|
| at | Yes | any place -- 'x,y' in metres, 'me', a factory label, 'node:<id>', 'slab:<n>', 'chain:<n>'/'pipe:<n>', 'plan:<name>' -- or 'resource:Crude Oil' for every node of one resource | |
| save | No | ||
| zoom | No | ||
| as_of | No | pin to one world state: a sav:… token from an earlier answer | |
| world | No | ||
| layers | No | explicit sublayer tokens, overriding the guess |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden and does so well: it discloses that two links are returned, what each exposes (saved base vs vanilla world), that `resource:` switches overlays, and that only Crude Oil layer tokens are confirmed while a wrong token still opens the map at the right place. Missing only auth/side-effect framing, though this is plainly a read-only link generator.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the two-link purpose, then scoped to the `at` vocabulary and token caveats. It is prose-heavy and a touch long, but each paragraph adds decision-relevant detail rather than filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 6-parameter tool with no annotations and no output schema, the description adequately covers what the tool does, its return shape, and the key `at` semantics. The gaps are the undocumented save/zoom/world parameters, but the core invocation surface is well explained.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 50%. The description adds real meaning for `at` (the `resource:<name>` kind, docs pointer) and for layer tokens/overrides, but `save`, `zoom`, and `world` are undocumented in both the schema and the description, leaving half the surface unaddressed.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Opens with a specific verb+resource: it produces two map links (local project map and the public satisfactory-calculator.com one) centred on a place. It clearly distinguishes what each link shows and names the sibling vocabulary it reuses, so an agent immediately knows what the tool yields.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explains the place vocabulary shared with other tools (docs/selectors.md) and the one `resource:<name>` exception, which implies usage. However it never states when to pick this tool over siblings like describe_location, whereami, or factory_map, nor any when-not condition.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
site_planA
Record, update or clear WHERE a stored plan stands. Nothing is re-solved.
A plan stores what to build; this stores where -- origin (x, y and optionally z, in metres, at the footprint's centre), orientation (yaw about world Z, the same convention the save stores machine facing with), and footprint (width x depth, metres). The footprint defaults to the square plan_layout budgets for the plan's largest floor, and the record keeps track of whether it was measured or derived.
Once sited: plan recalls print the siting; diff_vs_save plan=<name> adds an
approximate what-stands-on-the-pad census; show_on_map at='plan:<name>'
centres a map link on the origin.
The siting is a RECORD of your decision, not a constraint on the solve -- re-running
the plan neither reads nor moves it, and save_as over the same name keeps it.
| Name | Required | Description | Default |
|---|---|---|---|
| at | No | site origin (the footprint's CENTRE): 'x,y[,z]' in metres, 'me', a factory name, 'slab:<n>' or a run id. Blank keeps the stored origin | |
| plan | Yes | ||
| save | No | ||
| as_of | No | pin to one world state: a sav:… token from an earlier answer | |
| clear | No | ||
| world | No | ||
| yaw_deg | No | degrees about world Z, positive +X towards +Y; omit to keep | |
| base_rev | No | the plan version you read; needed to change an existing plan | |
| footprint | No | 'WxD' in metres ('96' = square). Blank keeps the stored one, or derives the layout's own square if none is stored |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations exist, so the description carries full burden, and it delivers: it states the siting is a RECORD of a decision, that re-running the plan neither reads nor moves it, that save_as over the same name preserves it, and that the record tracks measured-vs-derived provenance. Missing permission/error or base_rev conflict behaviour keeps it from a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action, then layered into grouped rationale (what is stored, what happens once sited, the record-not-constraint caveat). Slightly long, but each paragraph carries distinct information rather than padding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 9-parameter mutation tool with no annotations and no output schema, the description supplies the domain model, defaults, provenance behaviour, and downstream effects an agent needs to invoke it correctly. It does not describe error or return behaviour, but nothing critical for correct invocation is absent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 56%, but the description compensates by defining the origin (footprint centre, metres), the yaw convention (about world Z, matching the save's machine-facing), and the footprint format ('WxD', defaults to the layout's square). It adds real meaning for at/yaw_deg/footprint beyond the schema, though save/as_of/world/base_rev semantics live only in the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Opens with a specific verb triple and resource: 'Record, update or clear WHERE a stored plan stands,' immediately scoping the tool to siting rather than solving. The 'Nothing is re-solved' line sharply distinguishes it from solve-side siblings like plan_layout and commission_plan.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explains the context of use (siting decisions), the default footprint derivation, and what becomes possible once sited (plan recalls, diff_vs_save census, show_on_map centring), naming concrete sibling tools. It stops short of explicit when-not-to-use guidance, but the 'record, not constraint' framing routes the agent well.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
somersloopsB
Somersloops held, slotted and owned -- the sibling of power_shards.
sloop_budget has existed since sloops became spendable and nothing exposed it, so
the only way to learn how many you had was to guess a sloops= budget and read the
shortfall warning: you had to guess the budget to discover the budget.
Free and committed are both exact. Slotted ones live in InventoryPotential, the same
component as Power Shards, so this counts slot contents rather than inverting a boost
multiplier.
| Name | Required | Description | Default |
|---|---|---|---|
| save | No | ||
| as_of | No | pin to one world state: a sav:… token from an earlier answer | |
| limit | No | max rows (hard cap 25) | |
| world | No | ||
| offset | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, and it does disclose real methodology: free and committed counts are exact, slotted ones come from `InventoryPotential` (the Power Shards component), and it counts slot contents rather than inverting a boost multiplier. However, it says nothing about read-only semantics, pagination behavior, or rate/scan cost.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The purpose is front-loaded, but the middle paragraph is a narrative detour about having to guess a budget to discover a budget that, while colorful, takes space without routing the agent. Three paragraphs is more than this definition needs.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema and no annotations mean the description must cover return shape, pagination (limit/offset), and the save/world scoping parameters, yet it covers none of them. For a 5-parameter listing tool, it leaves the agent under-equipped to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 40% (`save`, `world`, and `offset` are undocumented), and the description adds no parameter-level meaning at all, discussing only `sloop_budget`/`sloops=` from a different tool. It fails to compensate for the coverage gap on a 5-parameter tool.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description identifies the resource (somersloops) and its three states (held, slotted, owned), and explicitly names its sibling `power_shards`, letting an agent place it without opening the schema. The verb is only implicit ('counts slot contents'), but the reporting intent is clear enough to act on.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The second paragraph explains why the tool exists (nothing exposed `sloop_budget`; you previously had to guess a `sloops=` budget), which implies the usage context of checking somersloop holdings. It never states when to call this versus `power_shards` or other inventory tools, so usage is inferred rather than directed.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
stockA
What you own, by item: spendable stock apart from what merely exists.
Spendable is carried + storage + Dimensional Depot -- exactly the set every affordability check in this surface spends. Machine buffers and crate contents get their own columns and are never added in: buffer material is in transit, and a crate exists because something went wrong and deletes itself when emptied.
where=True answers "and where is it": one row per container or crate holding the
item, with the region it stands in and its coordinate. Carried and Depot stock has no
place, so it is reported on the summary line instead. Fluids are in m3.
| Name | Required | Description | Default |
|---|---|---|---|
| item | No | one item by name; omit for everything you own | |
| save | No | ||
| as_of | No | pin to one world state: a sav:… token from an earlier answer | |
| limit | No | max rows (hard cap 25) | |
| where | No | list the places holding it instead of the per-item totals | |
| world | No | ||
| offset | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden and does well: it defines exactly what is counted (carried + storage + Dimensional Depot), what is deliberately excluded (machine buffers, crate contents) and why, and notes fluids are in m3. It does not state the read-only nature, error behavior for unknown items, or pagination/row-cap semantics beyond what the schema already says.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core definition, and each subsequent sentence covers a distinct behavior (inclusions vs exclusions, where mode, units). Slightly dense and jargon-heavy ('Dimensional Depot', the affordability-check aside), and the stray ``where=True`` inline formatting is rough, but sentences generally earn their place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only query tool with no output schema, the description conveys enough about the shape of results (per-item totals, per-container rows with region/coordinates, summary line, fluid units). The unspecified parameters (save, world, offset) are the main remaining gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 57%. The description meaningfully expands the `where` parameter (one row per container/crate with region and coordinate, summary line for carried/Depot) and clarifies the scope of the item list, but `save`, `world`, and `offset` have no description anywhere, and `as_of` only relies on its schema text. This is close to the baseline 3 for partial coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening line names a specific resource and scope: 'What you own, by item: spendable stock apart from what merely exists.' It even carves out a boundary against the related 'crates' tool by stating crate contents are excluded, which helps distinguish it. It stops short of naming which sibling to use for other views (e.g. storage, crates, search_items), so differentiation is partial rather than complete.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implicit in the semantics: it explains that 'where=True answers "and where is it"', which tells the agent when to flip that flag. However there is no explicit when-to-use-this-vs-alternative guidance among the many inventory-adjacent siblings (storage, crates, search_items, collected_from_world).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
storageA
Which container holds what, where it stands, and how full it is.
Every storage container, Personal Storage Box, Depot uploader and fluid buffer the
player built -- not splitters and mergers, whose one to three items in transit are not
stock, and not machine buffers, which belong to their machine. Fullest first, or by
how much of item they hold when one is named.
fill is measured, not stated by the save: a container is its used slots over its
slots, stacking each item at its own stack size, and a buffer is its m3 over what the
class holds. It is - where either number is unknown.
| Name | Required | Description | Default |
|---|---|---|---|
| item | No | only containers holding this item, fullest first | |
| kind | No | retired -- write container_kind= instead | |
| near | No | centre: 'x,y' in metres, 'me', or a named factory | |
| save | No | ||
| as_of | No | pin to one world state: a sav:… token from an earlier answer | |
| empty | No | include containers with nothing in them | |
| limit | No | max rows (hard cap 25) | |
| world | No | ||
| offset | No | ||
| radius_m | No | ||
| container_kind | No | solid | fluid | all |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and discharges much of it: it defines exactly which container classes count as stock, explains that fill is computed rather than read from the save (used slots / slots, m3 over class capacity), and discloses the '-' sentinel for unknown values. It omits other behavioral facts an agent would want, such as which save/world is used by default and any auth or rate considerations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The one-line summary is front-loaded and the subsequent paragraphs are organized around inclusion, ordering, and metric definition. The prose is a little long and spends words on exclusions that could be tighter, but every sentence carries distinct information rather than restating the schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an 11-parameter query tool with no output schema and no annotations, the description explains the key derived field (fill), the row ordering, and what counts as a row, which is genuinely useful. However it never describes the shape of a returned row beyond fill, and leaves near/radius_m, save/world defaulting, and pagination (limit/offset) unexplained, so an agent cannot fully predict the response.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 64%, so several parameters (save, world, offset, radius_m) are documented nowhere. The description reinforces the item parameter and the ordering it triggers, but says nothing about near/radius_m pairing, container_kind, or how save/world are defaulted. Baseline 3 is appropriate given the schema does most of the work and the prose adds only marginal meaning.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening line gives a concrete triple of what it returns (which container holds what, where it stands, how full), and the second paragraph sharpens the scope by explicitly excluding splitters, mergers, and machine buffers. That is much clearer than a bare noun name. It loses a point because it never distinguishes itself from nearby siblings such as stock or crates, leaving the agent to infer the boundary.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied through scope and ordering: 'fullest first, or by how much of item they hold when one is named' suggests a survey/audit use case. There is no explicit when-to-use or when-not, no statement of prerequisites, and no named alternative (stock, crates, factory_query) for the cases this tool does not cover.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
trace_upstreamA
What feeds a machine, or what it feeds -- walked on the save's own connections.
factory_query answers this between LABELLED sets. This answers it for one machine or
one building type, which is the question a cutover actually asks: thirteen Oil
Extractors sit on the Spire nodes and twenty Fuel Generators are burning, and repiping
the wrong extractor first drops several GW.
Direction is READ, not guessed. Every material edge carries its connector role, and 92.5% of the connectors landing on a machine name their direction outright; the rest are all on extractors or generators, whose own nature settles them. Where even that fails the edge is walked BOTH ways -- over-reporting a feeder is recoverable, missing one is not.
Belts and pipes are walked THROUGH and left out of the table: a trace from the
generators touches 331 nodes at depth 72, nearly all of it conveyor. What the route
crossed is named instead in a note -- how many runs of each medium, and the ids
search_conduits takes for the ones that have them.
| Name | Required | Description | Default |
|---|---|---|---|
| save | No | ||
| seed | Yes | a machine instance, a factory label, or a building name | |
| as_of | No | pin to one world state: a sav:… token from an earlier answer | |
| limit | No | max rows (hard cap 25) | |
| world | No | ||
| direction | No | up (what feeds it) | down (what it feeds) | up |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does so richly: direction is 'READ, not guessed', 92.5% of connectors name their direction, and ambiguous edges on extractors/generators are walked BOTH ways with an explicit rationale ('over-reporting a feeder is recoverable, missing one is not'). It also discloses that belts/pipes are walked through and excluded from the table, with medium counts and search_conduits ids given in a note.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the core purpose in the first clause and the factory_query contrast immediately after. The prose is dense and stylized, and the 331-node/depth-72 example is somewhat illustrative rather than operational, but each sentence contributes behavioral detail rather than filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 6-param tool with no annotations and no output schema, the description supplies strong behavioral context, including that output is tabular ('left out of the table') and that a supplementary note carries conduit info. It leaves parameter-level meaning for save/world/as_of/limit uncovered, but the trace semantics themselves are well covered.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 67%, so several params (save, world, as_of, limit) are unspecified in both places. The description conveys the meaning of the direction choice ('what feeds a machine, or what it feeds') and the granularity of the seed ('one machine or one building type'), but adds no syntax or format detail for the remaining parameters. Baseline 3 fits.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific operation (walking the save's own material connections) and explicitly contrasts its scope with factory_query: 'factory_query answers this between LABELLED sets. This answers it for one machine or one building type.' That lets an agent route between the two. The only minor ambiguity is that the name says 'upstream' while the description covers both directions, resolved only by the direction param.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear context for when to pick this over factory_query (single machine or building type vs labelled sets) and frames the real-world question it answers (cutover ordering). It does not explicitly state when not to use it, but the alternative comparison is concrete enough to guide selection.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ui_contextB
What the web page has open, and what changed in plans since this session last looked.
Call it first when the user says "this", "here" or "what I have open": it names the page's view, plan and version, tab and selection, whether the page reads the same save as you, and every plan version and chat solve by someone else since your last look.
| Name | Required | Description | Default |
|---|---|---|---|
| save | No | ||
| world | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It discloses a rich set of returned context (page view, plan version, tab, selection, save comparison, others' plan versions and chat solves since last look), which is useful behavioral detail. However, it does not state that the operation is read-only, whether it requires an active page session, or any permissions or side effects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the tool's purpose and then follows with usage triggers. It is appropriately sized and largely free of filler, though the second paragraph is a long list-like sentence that could be marginally tighter.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema and no annotations, so the description must compensate. It describes returned context well, but it completely omits the meaning and usage of the two input parameters, leaving an agent unable to know when or how to supply 'save' or 'world'.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the description does not explain the two optional parameters 'save' and 'world'. It mentions 'save' only in the context of output comparison ('whether the page reads the same save as you'), not as an input parameter, and never mentions 'world' at all.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific resource – the web page's open state and plan changes since last look – and lists the exact context items it returns (view, plan, version, tab, selection, save comparison, others' activity). It is clear and specific, but it does not differentiate itself from any named sibling tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives explicit trigger conditions: call it first when the user says 'this', 'here', or 'what I have open'. That is clear contextual usage guidance, though it does not mention when not to use it or name any alternative tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
unlocked_recipesB
Which recipes this world has. Defaults to alternates, never all 872.
Sorted by name and paged with offset=, so the whole list is reachable.
| Name | Required | Description | Default |
|---|---|---|---|
| save | No | ||
| as_of | No | pin to one world state: a sav:… token from an earlier answer | |
| limit | No | max rows (hard cap 25) | |
| world | No | ||
| offset | No | ||
| only_alternates | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does disclose real behavior: the default only_alternates filter, name sorting, and offset-based paging that makes the full list reachable. It does not disclose auth/permission needs, whether results are cached or expensive, or any return shape, leaving meaningful gaps for a no-annotation tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences with zero filler, and the scope/default constraint is front-loaded before the paging note. Every clause carries information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 6-parameter query tool with no output schema and no annotations, the description covers the essential quirks (default filter, sorting, pagination) but omits the meaning of save, world, as_of, and limit, so an agent can call it correctly only for the default case.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 33%, so the description must compensate, and it does partially: it clarifies the only_alternates default and how offset= pages through results. It leaves save, world, as_of, and limit undocumented beyond the terse schema hints, so it falls short of full compensation without being empty.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific resource and scope: 'Which recipes this world has', and quantifies the default ('defaults to alternates, never all 872'). An agent can tell this is a recipe-listing tool, though it never names or contrasts with siblings like search_recipes, recipe_detail, or alternates_for_item.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage through the default-filter note but gives no explicit when-to-use or when-not-to-use guidance and no alternative routing against the ~50 sibling tools. An agent must infer whether this or search_recipes is the right call.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
whereamiB
Where the player is standing, and what is around them.
Position comes from the Char_Player_C pawn in the save, so it is wherever you
were when it was written -- an autosave can be several minutes stale. Use
near:me@<radius> as a source selector in the planning tools to scope work to
here.
| Name | Required | Description | Default |
|---|---|---|---|
| save | No | ||
| as_of | No | pin to one world state: a sav:… token from an earlier answer | |
| limit | No | max rows (hard cap 25) | |
| world | No | ||
| radius_m | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It usefully discloses the data source (Char_Player_C pawn) and a real behavioral caveat (autosave can be minutes stale), which is genuinely valuable. However, it says nothing about the shape or fields of what is returned, leaving the read behavior half-described.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences, front-loaded with the core purpose followed by caveat and cross-tool usage. Little waste, though the second sentence is slightly run-on.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No annotations and no output schema, with 5 parameters at 40% coverage, mean the description must do more. It covers staleness and the near:me idiom well but never describes return contents or the save/world/radius_m parameters, leaving an agent unable to predict the output.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 40%: as_of and limit are documented in the schema, but save, world, and radius_m are not. The description hints at a radius concept via 'near:me@<radius>' but does not explain save, world, or how radius_m maps to 'what is around them', so it fails to compensate for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb+resource: reporting the player's position and surroundings. It is clearly distinguishable from siblings like describe_location or world_summary, though it never explicitly names an alternative to disambiguate.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives clear context for use (staleness of the pawn-derived position) and even steers the agent toward 'near:me@<radius>' as a source selector in the planning tools. It stops short of explicit when-not-to-use guidance versus e.g. ui_context or describe_location.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
world_summaryC
Progress, power and problems for one world.
| Name | Required | Description | Default |
|---|---|---|---|
| save | No | ||
| as_of | No | pin to one world state: a sav:… token from an earlier answer | |
| world | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden, and it largely fails. It gestures at content categories (progress, power, problems) but says nothing about read-only nature, return shape, cost, or whether results are derived or stored. A mutation-free summary is implied only by the tool name.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
It is a single nine-word fragment with zero redundancy, so nothing wastes space. But it is under-specified rather than concise: there is no front-loaded statement of the action, so brevity here costs clarity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has 3 optional parameters, no annotations, and no output schema, so the description must explain both inputs and the shape of the returned summary. It does neither beyond naming topical categories, leaving the agent unable to predict what a 'summary' contains or how the three scoping inputs interact.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 33%: only 'as_of' is documented in the schema ('a sav:… token from an earlier answer'), while 'save' and 'world' are bare. The description's phrase 'for one world' loosely gestures at the 'world' parameter but adds no format, default, or interaction semantics (e.g., how save/world/as_of combine).
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The fragment 'Progress, power and problems for one world' names the subject matter (progress/power/problems) and scope ('one world'), which hints at a per-world summary read. However, it has no verb and never states what the tool actually does (retrieve? compute? compare?), so an agent must infer the operation from the name alone. It is more informative than a pure tautology but still vague.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance and no alternative named. Siblings such as list_worlds, power_report, factory_query, and describe_location clearly overlap, yet the description does nothing to route the agent between them. 'For one world' implies scoping but not selection criteria.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
51 tool updates
- Changed
advise_hard_drive_pick1 field changed- added
Input schema / properties / as_ofAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "description": "pin to one world state: a sav:… token from an earlier answer", + "title": "As Of" +}
- Changed
alternates_for_item2 fields changed- added
Input schema / properties / as_ofAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "description": "pin to one world state: a sav:… token from an earlier answer", + "title": "As Of" +} - added
Input schema / properties / worldAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "title": "World" +}
- Added
amend_factory - Changed
bom2 fields changed- added
Input schema / properties / as_ofAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "description": "pin to one world state: a sav:… token from an earlier answer", + "title": "As Of" +} - added
Input schema / properties / offsetAdded value: +{ + "default": 0, + "title": "Offset", + "type": "integer" +}
- Changed
collected_from_world8 fields changed- added
Input schema / properties / as_ofAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "description": "pin to one world state: a sav:… token from an earlier answer", + "title": "As Of" +} - added
Input schema / properties / mode / anyOfAdded value: +[ + { + "type": "string" + }, + { + "type": "null" + } +] - changed
Input schema / properties / mode / defaultPrevious value: -"census"New value: +null - changed
Input schema / properties / mode / descriptionPrevious value: -"census | collected | remaining | nearest"New value: +"retired -- write show= instead" - removed
Input schema / properties / mode / typeRemoved value: -"string" - changed
Input schema / properties / near / descriptionPrevious value: -"origin for mode=nearest: 'x,y' in metres, 'me', or a factory name. Defaults to where the player is standing"New value: +"origin for show=nearest: 'x,y' in metres, 'me', or a factory name. Defaults to where the player is standing" - added
Input schema / properties / offsetAdded value: +{ + "default": 0, + "title": "Offset", + "type": "integer" +} - added
Input schema / properties / showAdded value: +{ + "default": "census", + "description": "census | collected | remaining | nearest", + "title": "Show", + "type": "string" +}
- Changed
commission_plan3 fields changed- added
Input schema / properties / as_ofAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "description": "pin to one world state: a sav:… token from an earlier answer", + "title": "As Of" +} - changed
Input schema / properties / limit / defaultPrevious value: -60New value: +25 - added
Input schema / properties / offsetAdded value: +{ + "default": 0, + "title": "Offset", + "type": "integer" +}
- Changed
compare_recipe_options1 field changed- added
Input schema / properties / as_ofAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "description": "pin to one world state: a sav:… token from an earlier answer", + "title": "As Of" +}
- Added
crates - Changed
describe_location5 fields changed- added
Input schema / properties / as_ofAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "description": "pin to one world state: a sav:… token from an earlier answer", + "title": "As Of" +} - added
Input schema / properties / atAdded value: +{ + "default": "", + "description": "the place: 'x,y' in metres, 'me', a named factory, 'slab:<n>', or a run id like 'chain:7'", + "title": "At", + "type": "string" +} - removed
Input schema / properties / x_mRemoved value: -{ - "title": "X M", - "type": "number" -} - removed
Input schema / properties / y_mRemoved value: -{ - "title": "Y M", - "type": "number" -} - removed
Input schema / requiredRemoved value: -[ - "x_m", - "y_m" -]
- Changed
diff_vs_save1 field changed- added
Input schema / properties / as_ofAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "description": "pin to one world state: a sav:… token from an earlier answer", + "title": "As Of" +}
- Changed
explain_byproducts1 field changed- added
Input schema / properties / as_ofAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "description": "pin to one world state: a sav:… token from an earlier answer", + "title": "As Of" +}
- Added
factory_floors - Changed
factory_health2 fields changed- added
Input schema / properties / as_ofAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "description": "pin to one world state: a sav:… token from an earlier answer", + "title": "As Of" +} - added
Input schema / properties / offsetAdded value: +{ + "default": 0, + "title": "Offset", + "type": "integer" +}
- Changed
factory_map2 fields changed- added
Input schema / properties / as_ofAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "description": "pin to one world state: a sav:… token from an earlier answer", + "title": "As Of" +} - added
Input schema / properties / offsetAdded value: +{ + "default": 0, + "title": "Offset", + "type": "integer" +}
- Changed
factory_query7 fields changed- added
Input schema / properties / as_ofAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "description": "pin to one world state: a sav:… token from an earlier answer", + "title": "As Of" +} - added
Input schema / properties / of / anyOfAdded value: +[ + { + "type": "string" + }, + { + "type": "null" + } +] - changed
Input schema / properties / of / defaultPrevious value: -"summary"New value: +null - changed
Input schema / properties / of / descriptionPrevious value: -"comma-separated: summary, machines, recipes, buildings, balance, inputs, outputs, power, nodes, links, issues"New value: +"retired -- write show= instead" - removed
Input schema / properties / of / typeRemoved value: -"string" - added
Input schema / properties / offsetAdded value: +{ + "default": 0, + "title": "Offset", + "type": "integer" +} - added
Input schema / properties / showAdded value: +{ + "default": "summary", + "description": "comma-separated: summary, machines, recipes, buildings, balance, inputs, outputs, internal, power, nodes, links, issues", + "title": "Show", + "type": "string" +}
- Changed
factory_sites2 fields changed- added
Input schema / properties / as_ofAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "description": "pin to one world state: a sav:… token from an earlier answer", + "title": "As Of" +} - added
Input schema / properties / offsetAdded value: +{ + "default": 0, + "title": "Offset", + "type": "integer" +}
- Changed
forget_factory1 field changed- added
Input schema / properties / as_ofAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "description": "pin to one world state: a sav:… token from an earlier answer", + "title": "As Of" +}
- Changed
forget_plan2 fields changed- added
Input schema / properties / as_ofAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "description": "pin to one world state: a sav:… token from an earlier answer", + "title": "As Of" +} - added
Input schema / properties / base_revAdded value: +{ + "anyOf": [ + { + "type": "integer" + }, + { + "type": "null" + } + ], + "default": null, + "description": "the plan version you read; needed to change an existing plan", + "title": "Base Rev" +}
- Changed
list_buildings8 fields changed- added
Input schema / properties / as_ofAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "description": "pin to one world state: a sav:… token from an earlier answer", + "title": "As Of" +} - added
Input schema / properties / building_kindAdded value: +{ + "default": "production", + "description": "production | extractor | generator | logistics | foundation | ramp | wall | pillar | beam | architecture | all", + "title": "Building Kind", + "type": "string" +} - added
Input schema / properties / kind / anyOfAdded value: +[ + { + "type": "string" + }, + { + "type": "null" + } +] - changed
Input schema / properties / kind / defaultPrevious value: -"production"New value: +null - added
Input schema / properties / kind / descriptionAdded value: +"retired -- write building_kind= instead" - removed
Input schema / properties / kind / typeRemoved value: -"string" - added
Input schema / properties / limitAdded value: +{ + "default": 25, + "description": "max rows (hard cap 25)", + "maximum": 25, + "minimum": 1, + "title": "Limit", + "type": "integer" +} - added
Input schema / properties / offsetAdded value: +{ + "default": 0, + "title": "Offset", + "type": "integer" +}
- Changed
list_factories1 field changed- added
Input schema / properties / as_ofAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "description": "pin to one world state: a sav:… token from an earlier answer", + "title": "As Of" +}
- Changed
list_pending_hard_drive_choices3 fields changed- added
Input schema / properties / as_ofAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "description": "pin to one world state: a sav:… token from an earlier answer", + "title": "As Of" +} - added
Input schema / properties / limitAdded value: +{ + "default": 25, + "description": "max rows (hard cap 25)", + "maximum": 25, + "minimum": 1, + "title": "Limit", + "type": "integer" +} - added
Input schema / properties / offsetAdded value: +{ + "default": 0, + "title": "Offset", + "type": "integer" +}
- Changed
list_plans2 fields changed- added
Input schema / properties / as_ofAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "description": "pin to one world state: a sav:… token from an earlier answer", + "title": "As Of" +} - added
Input schema / properties / nameAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "description": "one plan's full stored request, siting and field, unsolved", + "title": "Name" +}
- Changed
list_regions2 fields changed- added
Input schema / properties / resourceAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "title": "Resource" +} - added
Input schema / properties / with_resource / descriptionAdded value: +"retired -- write resource= instead"
- Changed
mam_research9 fields changed- added
Input schema / properties / as_ofAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "description": "pin to one world state: a sav:… token from an earlier answer", + "title": "As Of" +} - added
Input schema / properties / offsetAdded value: +{ + "default": 0, + "title": "Offset", + "type": "integer" +} - added
Input schema / properties / queryAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "description": "filter by name, case-insensitive", + "title": "Query" +} - changed
Input schema / properties / search / descriptionPrevious value: -"filter by name, case-insensitive"New value: +"retired -- write query= instead" - added
Input schema / properties / showAdded value: +{ + "default": "todo", + "description": "all | todo | affordable -- todo hides finished research", + "title": "Show", + "type": "string" +} - added
Input schema / properties / status / anyOfAdded value: +[ + { + "type": "string" + }, + { + "type": "null" + } +] - changed
Input schema / properties / status / defaultPrevious value: -"todo"New value: +null - changed
Input schema / properties / status / descriptionPrevious value: -"all | todo | affordable -- todo hides finished research"New value: +"retired -- write show= instead" - removed
Input schema / properties / status / typeRemoved value: -"string"
- Added
milestones - Changed
name_factory2 fields changed- added
Input schema / properties / as_ofAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "description": "pin to one world state: a sav:… token from an earlier answer", + "title": "As Of" +} - changed
Input schema / properties / select / descriptionPrevious value: -"selector terms, ANDed. product:<item> | recipe:<name> | building:<class or name> | near:<x,y@radius_m or label@radius_m> | base:<n> | line:<n> | slab:<n> | proposal:<n> | label:<name> | all. Terms are ANDed; comma-separated values inside one term are ORed; prefix a term with '-' to exclude it"New value: +"selector terms, ANDed. product:<item> | recipe:<name> | building:<class or name> | near:<place>@<radius_m> | base:<n> | line:<n> | slab:<n> | proposal:<n> | label:<name> | machine:<instance> | all. Terms are ANDed; comma-separated values inside one term are ORed; prefix a term with '-' to exclude it"
- Changed
phase_requirements1 field changed- added
Input schema / properties / as_ofAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "description": "pin to one world state: a sav:… token from an earlier answer", + "title": "As Of" +}
- Changed
plan_factory6 fields changed- added
Input schema / properties / as_ofAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "description": "pin to one world state: a sav:… token from an earlier answer", + "title": "As Of" +} - added
Input schema / properties / base_revAdded value: +{ + "anyOf": [ + { + "type": "integer" + }, + { + "type": "null" + } + ], + "default": null, + "description": "the plan version you read; needed to save over an existing plan", + "title": "Base Rev" +} - added
Input schema / properties / requiredAdded value: +{ + "anyOf": [ + { + "items": { + "type": "string" + }, + "type": "array" + }, + { + "type": "null" + } + ], + "default": null, + "description": "recipes that must make their item; others for it are excluded", + "title": "Required" +} - added
Input schema / properties / site_atAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "description": "with save_as: record where this plan will STAND -- 'x,y[,z]' in metres, 'me', a factory name, 'slab:<n>' or a run id (the footprint's centre)", + "title": "Site At" +} - added
Input schema / properties / site_footprintAdded value: +{ + "default": "", + "description": "site footprint 'WxD' in metres; blank = the layout's own square", + "title": "Site Footprint", + "type": "string" +} - added
Input schema / properties / site_yaw_degAdded value: +{ + "default": 0, + "description": "site orientation: degrees about world Z, positive +X towards +Y", + "title": "Site Yaw Deg", + "type": "number" +}
- Changed
plan_layout7 fields changed- added
Input schema / properties / as_ofAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "description": "pin to one world state: a sav:… token from an earlier answer", + "title": "As Of" +} - added
Input schema / properties / detail / anyOfAdded value: +[ + { + "type": "string" + }, + { + "type": "null" + } +] - changed
Input schema / properties / detail / defaultPrevious value: -"floors"New value: +null - added
Input schema / properties / detail / descriptionAdded value: +"retired -- write show= instead" - removed
Input schema / properties / detail / typeRemoved value: -"string" - added
Input schema / properties / showAdded value: +{ + "default": "floors", + "description": "floors | blocks | buses | trunks | materials | sites", + "title": "Show", + "type": "string" +} - changed
Input schema / properties / sites / descriptionPrevious value: -"detail=\"sites\": {\"rig\": [\"Heavy Oil Residue\", ...], ...}"New value: +"show=\"sites\": {\"rig\": [\"Heavy Oil Residue\", ...], \"hall\": [\"MW\"]} -- MW/power claims every generator"
- Added
plan_log - Changed
power_report1 field changed- added
Input schema / properties / as_ofAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "description": "pin to one world state: a sav:… token from an earlier answer", + "title": "As Of" +}
- Changed
power_shards2 fields changed- added
Input schema / properties / as_ofAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "description": "pin to one world state: a sav:… token from an earlier answer", + "title": "As Of" +} - added
Input schema / properties / offsetAdded value: +{ + "default": 0, + "title": "Offset", + "type": "integer" +}
- Changed
propose_factories2 fields changed- added
Input schema / properties / as_ofAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "description": "pin to one world state: a sav:… token from an earlier answer", + "title": "As Of" +} - added
Input schema / properties / offsetAdded value: +{ + "default": 0, + "title": "Offset", + "type": "integer" +}
- Changed
rank_build_sites6 fields changed- added
Input schema / properties / as_ofAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "description": "pin to one world state: a sav:… token from an earlier answer", + "title": "As Of" +} - added
Input schema / properties / limitAdded value: +{ + "default": 5, + "description": "max rows (hard cap 25)", + "maximum": 25, + "minimum": 1, + "title": "Limit", + "type": "integer" +} - added
Input schema / properties / top / anyOfAdded value: +[ + { + "type": "integer" + }, + { + "type": "null" + } +] - changed
Input schema / properties / top / defaultPrevious value: -5New value: +null - added
Input schema / properties / top / descriptionAdded value: +"deprecated alias for limit" - removed
Input schema / properties / top / typeRemoved value: -"integer"
- Changed
rank_unlocks3 fields changed- added
Input schema / properties / as_ofAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "description": "pin to one world state: a sav:… token from an earlier answer", + "title": "As Of" +} - added
Input schema / properties / queryAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "description": "only test alternates whose name matches", + "title": "Query" +} - changed
Input schema / properties / search / descriptionPrevious value: -"only test alternates whose name matches"New value: +"retired -- write query= instead"
- Added
rename_factory - Added
rename_plan - Added
search_conduits - Changed
search_recipes7 fields changed- added
Input schema / properties / as_ofAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "description": "pin to one world state: a sav:… token from an earlier answer", + "title": "As Of" +} - added
Input schema / properties / kind / anyOfAdded value: +[ + { + "type": "string" + }, + { + "type": "null" + } +] - changed
Input schema / properties / kind / defaultPrevious value: -"part"New value: +null - added
Input schema / properties / kind / descriptionAdded value: +"retired -- write recipe_kind= instead" - removed
Input schema / properties / kind / typeRemoved value: -"string" - added
Input schema / properties / recipe_kindAdded value: +{ + "default": "part", + "description": "part | building | manual | all", + "title": "Recipe Kind", + "type": "string" +} - added
Input schema / properties / worldAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "title": "World" +}
- Changed
search_resource_nodes11 fields changed- added
Input schema / properties / as_ofAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "description": "pin to one world state: a sav:… token from an earlier answer", + "title": "As Of" +} - changed
Input schema / properties / group / descriptionPrevious value: -"deprecated alias for mode"New value: +"retired -- write show= instead" - added
Input schema / properties / kind / descriptionAdded value: +"node | well_sat | geyser | all" - added
Input schema / properties / mode / anyOfAdded value: +[ + { + "type": "string" + }, + { + "type": "null" + } +] - changed
Input schema / properties / mode / defaultPrevious value: -"fields"New value: +null - changed
Input schema / properties / mode / descriptionPrevious value: -"fields | nodes | nearest"New value: +"retired -- write show= instead" - removed
Input schema / properties / mode / typeRemoved value: -"string" - changed
Input schema / properties / near / descriptionPrevious value: -"origin for mode=nearest: 'x,y' in metres, 'me', or a factory name"New value: +"origin for show=nearest: 'x,y' in metres, 'me', or a factory name" - added
Input schema / properties / offsetAdded value: +{ + "default": 0, + "title": "Offset", + "type": "integer" +} - added
Input schema / properties / purity / descriptionAdded value: +"pure | normal | impure | all" - added
Input schema / properties / showAdded value: +{ + "default": "fields", + "description": "fields | nodes | nearest", + "title": "Show", + "type": "string" +}
- Changed
select_machines2 fields changed- added
Input schema / properties / as_ofAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "description": "pin to one world state: a sav:… token from an earlier answer", + "title": "As Of" +} - changed
Input schema / properties / select / descriptionPrevious value: -"selector terms, ANDed. product:<item> | recipe:<name> | building:<class or name> | near:<x,y@radius_m or label@radius_m> | base:<n> | line:<n> | slab:<n> | proposal:<n> | label:<name> | all. Terms are ANDed; comma-separated values inside one term are ORed; prefix a term with '-' to exclude it"New value: +"selector terms, ANDed. product:<item> | recipe:<name> | building:<class or name> | near:<place>@<radius_m> | base:<n> | line:<n> | slab:<n> | proposal:<n> | label:<name> | machine:<instance> | all. Terms are ANDed; comma-separated values inside one term are ORed; prefix a term with '-' to exclude it"
- Changed
show_on_map4 fields changed- added
Input schema / properties / as_ofAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "description": "pin to one world state: a sav:… token from an earlier answer", + "title": "As Of" +} - added
Input schema / properties / atAdded value: +{ + "description": "any place -- 'x,y' in metres, 'me', a factory label, 'node:<id>', 'slab:<n>', 'chain:<n>'/'pipe:<n>', 'plan:<name>' -- or 'resource:Crude Oil' for every node of one resource", + "title": "At", + "type": "string" +} - removed
Input schema / properties / targetRemoved value: -{ - "description": "'x,y' in metres, 'me', a factory label, a node id, or a resource name like 'Crude Oil'", - "title": "Target", - "type": "string" -} - changed
Input schema / requiredPrevious value: -[ - "target" -]New value: +[ + "at" +]
- Added
site_plan - Changed
somersloops3 fields changed- added
Input schema / properties / as_ofAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "description": "pin to one world state: a sav:… token from an earlier answer", + "title": "As Of" +} - added
Input schema / properties / limitAdded value: +{ + "default": 20, + "description": "max rows (hard cap 25)", + "maximum": 25, + "minimum": 1, + "title": "Limit", + "type": "integer" +} - added
Input schema / properties / offsetAdded value: +{ + "default": 0, + "title": "Offset", + "type": "integer" +}
- Added
stock - Added
storage - Changed
trace_upstream1 field changed- added
Input schema / properties / as_ofAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "description": "pin to one world state: a sav:… token from an earlier answer", + "title": "As Of" +}
- Added
ui_context - Changed
unlocked_recipes2 fields changed- added
Input schema / properties / as_ofAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "description": "pin to one world state: a sav:… token from an earlier answer", + "title": "As Of" +} - added
Input schema / properties / offsetAdded value: +{ + "default": 0, + "title": "Offset", + "type": "integer" +}
- Changed
whereami1 field changed- added
Input schema / properties / as_ofAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "description": "pin to one world state: a sav:… token from an earlier answer", + "title": "As Of" +}
- Changed
world_summary1 field changed- added
Input schema / properties / as_ofAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "description": "pin to one world state: a sav:… token from an earlier answer", + "title": "As Of" +}
42 tool updates
v0.1.0- First observed
advise_hard_drive_pick - First observed
alternates_for_item - First observed
bom - First observed
collected_from_world - First observed
commission_plan - First observed
compare_recipe_options - First observed
describe_location - First observed
diff_vs_save - First observed
explain_byproducts - First observed
factory_health - First observed
factory_map - First observed
factory_query - First observed
factory_sites - First observed
forget_factory - First observed
forget_plan - First observed
list_buildings - First observed
list_factories - First observed
list_pending_hard_drive_choices - First observed
list_plans - First observed
list_regions - First observed
list_worlds - First observed
mam_research - First observed
name_factory - First observed
phase_requirements - First observed
plan_factory - First observed
plan_layout - First observed
power_report - First observed
power_shards - First observed
propose_factories - First observed
rank_build_sites - First observed
rank_unlocks - First observed
recipe_detail - First observed
search_items - First observed
search_recipes - First observed
search_resource_nodes - First observed
select_machines - First observed
show_on_map - First observed
somersloops - First observed
trace_upstream - First observed
unlocked_recipes - First observed
whereami - First observed
world_summary
TDQS
Scored across 54 tools
Descriptions are unusually detailed and most tools have a specific niche, but overlapping factory_* and recipe/plan vocabularies create real selection risk. An agent could confuse factory_map, propose_factories, and factory_sites, or search_recipes, alternates_for_item, and compare_recipe_options without careful reading.
All names use snake_case, but the convention is mixed: some are verb_noun (search_items, plan_factory), while many are bare nouns (stock, bom, factory_map, world_summary). Pluralization and phrasing are uneven, though the set remains readable.
With 54 tools, the server is far beyond the 3-15 sweet spot and well past the 25+ 'too many' threshold. The domain is broad, but the surface feels like several toolsets merged, making discovery and context management difficult.
Coverage is strong for recipes, factories, plans, resources, power, research, collectibles, and map analysis, with lifecycle operations for plans and factory labels. Missing or thin areas include vehicle/train/drone logistics and blueprint editing, but core planning workflows are well supported.
Maintenance
Related MCP Connectors
A MCP server built for developers enabling Git based project management with project and personal…
MCP server for generating rough-draft project plans from natural-language prompts.
An MCP server for deep research or task groups
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceAn MCP server for managing recipes, meal plans, shopping lists, and more through a self-hosted Mealie instance.1MIT
- AlicenseNot gradedqualityDmaintenanceAn external MCP server for diagnosing Minecraft servers via backup analysis, local runtime, or Docker runtime, offering tools for plugin inspection, log analysis, configuration linting, and performance diagnostics.1MIT
- FlicenseNot gradedqualityCmaintenanceAn MCP server providing tools for YouTube search, AI trip planning, notes management, web search, and product price comparison, backed by a FastAPI backend.-
- AlicenseNot gradedqualityBmaintenanceMCP server for Minecraft that manages map points, provides recipes and guides, and enables inspection/download of files and worlds via the exaroton REST API.MIT
One platform up close — machines, storage, belts, pipes, wires
Floor view — pick a storey, see what stands on it
Terrain mode — hillshade from a 1 m heightfield read out of the game
A crate's contents, with icons extracted from the game's assets