sim-lab
Summary: Sim Lab is a local, offline MCP server that lets you discover simulation worlds, run and compare robotics experiments (swarms, formations, coverage, localisation, visual odometry), tune parameters, and manage floor plans and custom experiments.
Explore worlds —
cataloguelists every world: what it simulates and does not, each knob's unit and range, metrics (and which are better higher), plus example questions with ready variant sets.Browse experiments —
list_experimentsandget_experimentshow shipped and personal experiments with defaults, variants, metrics and YAML.Run single experiments —
run_experimentexecutes one run to completion (blocking), with parameter overrides or a named variant and an optional label; returns run id, status, headline metrics and report.Run campaigns —
run_campaignruns several named variants of one experiment in sequence and returns the comparison table.Tune parameters —
tunesearches a parameter space via grid, random, or Bayesian (GP + expected improvement) strategies for the best headline metric (budget capped at 40).Inspect past runs —
list_runsandget_runretrieve recent runs, or one run's manifest, metrics, report, log tail and folder contents.Compare runs —
compare_runsputs headline metrics side by side and names the best run per metric, direction-aware.Manage floor plans —
list_plansshows ASCII plans;save_planwrites your own (#wall,Ddoor,.free) to~/.simlab/plansaslayout: plan:<name>.Save custom experiments —
save_experimentvalidates and stores your own YAML (local command + params) under~/.simlab/experiments, with overwrite protection for shipped names.Safety and reproducibility — everything runs locally with no network or telemetry; params are range-checked and quoted without a shell, run folders are self-contained, and saved commands run with your rights.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@sim-labrun a swarm formation experiment with 12 drones and 30% message loss"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Sim Lab
A robotics simulation lab for Claude. Ask a question about a swarm, a formation, coverage, cooperative localisation or visual odometry, and Claude designs the experiment, runs it on your machine, compares the runs and answers with numbers you can reproduce.
"Does a formation of 12 drones still close when the radio drops 60 % of messages?" → one campaign, four runs, a compare table, an answer in metres, and a findings note with the run ids.
Everything runs locally. No account, no network, no telemetry. Python 3.10+, numpy and pyyaml.
What is in the box
Piece | What it does |
| 12 tools: catalogue, experiments, run, campaign, tune, runs, compare, floor plans. Stdio, stdlib JSON-RPC, no framework. |
| How Claude turns a question into a one-knob sweep, estimates cost, runs, compares, checks the seed and answers. |
| A one-page reproducible findings note: question, setup, table, reading, limits, reproduce. |
Swarm world | N agents in a 2-D arena with range-limited lossy, delayed radio; behaviours flock, formation, rendezvous, coverage, goto; belief filters (exponential, Kalman, unicycle EKF with range-bearing fusion and covariance intersection); obstacles and ASCII floor plans that block motion and radio; water and fog presets; point, unicycle, fixed-wing and quadrotor-lite vehicle models. |
Synthetic VO world | A camera path with typed noise (gaussian, drift, scale, outliers, mixed), dropout and latency, scored with TUM-style ATE and RPE. For testing evaluation chains and noise models. |
The lab | Experiments as YAML, runs as folders ( |
Related MCP server: AgentPrism Workflows
Install
claude plugin marketplace add najikay/claude-simlab
claude plugin install sim-lab@claude-simlab
python3 -m pip install numpy pyyamlOr from the Claude directory: install Sim Lab, then install the two packages once. The server starts as python3 and tells you (in the tool error) which interpreter it is running and what is missing.
Debian, Ubuntu, Homebrew refuse
pip installinto the system Python ("externally managed"). Use the distribution's packages (sudo apt install python3-numpy python3-yaml), orpython3 -m pip install --user --break-system-packages numpy pyyaml, or point Claude Code at a virtualenv's interpreter (below).Windows:
python3must resolve to a real interpreter. The Microsoft Store Python provides it; a python.org install provides onlypythonandpy. Either install Python from the Store, or register the server yourself with the interpreter you have, which also works for a virtualenv anywhere:claude mcp add sim-lab -- python "<path to the installed plugin>\servers\simlab_server.py" --home "<optional folder for runs>"(
claude plugin listshows where the plugin was installed.)
Runs are kept under ~/.simlab/runs/ (to move them, register the server yourself with --home <folder>, as in the Windows line above). Your own experiments and floor plans go in ~/.simlab/experiments/ and ~/.simlab/plans/; the tools save_experiment and save_plan write there.
A session
you: Does message loss above 30 % break a 12-drone ring formation?
claude: (reads the catalogue once, picks swarm-formation and the loss0 / loss30 / loss60 variants
plus a 0.9 run, says four runs will take about five seconds, runs them, compares)
No. At the defaults (12 agents, 20 m radio on a 40 m arena, 400 ticks, seed 7) the mean
slot error stays within 0.086–0.089 m from 0 % to 90 % loss. What loss does change is
collisions: 6 → 16 → 15 → 18, because agents learn about their neighbours later.
| run | msg_loss | formation_error_m | collisions | connected_pct |
| 20261002-220648-5229_swarm-formation | 0.0 | 0.086 | 6 | 100 |
| 20261002-220649-db2c_swarm-formation | 0.3 | 0.087 | 16 | 100 |
| 20261002-220650-afe4_swarm-formation | 0.6 | 0.089 | 15 | 100 |
| 20261002-220651-0fbd_swarm-formation | 0.9 | 0.088 | 18 | 100 |
A second seed agrees (0.074 m and 0.076 m at 0.6 and 0.9). The radio is dense enough that one
message in ten still carries consensus; what would break it is latency or range (the
delayed-comms and short-range variants), or walls blocking links (office-los). Want those?The numbers above are from real runs on this machine. Ask for "a findings note" and the second skill writes a page with the run ids as citations.
Tools
Tool | What it returns |
| Every world: what it simulates and what it does not, each knob with unit and range, the metrics, which are better when higher, example questions with ready variant sets. Claude reads this first. |
| The shipped experiments (and your own, flagged |
| One run to completion: run id, status, headline metrics, the report. Overrides by |
| Several variants of one experiment, then the compare table. |
| Grid, random or Bayesian (GP + expected improvement) search over a parameter space for the best headline metric. |
| Past runs with their metrics; one run with its manifest, report and folder. |
| A table across runs with the best run per metric, direction-aware. |
| ASCII floor plans ( |
| Your own experiment YAML, validated before it is written. Its |
Experiments shipped
Experiment | Question it is built for | Headline metric |
| Does the ring close under loss, latency, noise, walls, water, fog, a vehicle model? |
|
| How much of the arena (or the office) gets visited, and at what collision cost? |
|
| Odometry alone vs naive EKF fusion vs covariance intersection, with and without anchors. |
|
| How ATE and RPE grow with each noise model, dropout and latency. |
|
Each has 8 to 19 named variants (lossy-comms, office-los, water, fixed-wing, anchors2-ci, drift, …). docs/WORLDS.md explains the models; the catalogue is the authoritative list of knobs.
Reproducibility
A run is reproducible from its folder alone:
~/.simlab/runs/20261003-010101-a1b2_swarm-formation/
manifest.json the exact command, every parameter, the seed, timings, lab version
stdout.log
metrics.json {"headline": {"formation_error_m": 0.42, ...}}
report.md parameters, metrics, figures
trajectory.svg · metrics.svg · agents.jsonl · obstacles.jsonSeeds are explicit, the manifest records the lab, Python and NumPy versions and a hash of the experiment file and the floor plan, and a run that the lab process did not live to finish is marked interrupted rather than left running. The skill tells Claude to re-run a close call with another seed before believing it.
Data handling
Everything runs on your machine: the MCP server is a local process started by Claude Code; the worlds are Python scripts in this repository.
Nothing is sent anywhere. The plugin makes no network requests, has no telemetry and needs no account or key.
What is written: run folders under
~/.simlab(or the--homefolder), and the experiments and plans you ask Claude to save there. Delete the folder to delete everything.The server runs experiments only from the shipped YAML or from files in your lab folder. Commands are composed from the experiment's template and typed parameters, every value quoted as a single argument and run without a shell, and values are checked against the catalogue's ranges first.
The plugin reads no environment variables of its own. A world process receives only an allow-list (PATH, temp and home folders, Python's own), so keys and tokens in your environment are never passed to it.
save_experimentregisters a command of your own that the lab will run on later requests, with your rights. Claude is told to save only what you asked for, shipped names cannot be replaced by accident (an explicitoverwriteis needed), and every run's manifest records which file and plan it used, with their hashes.
Development
python3 -m pip install -r requirements-dev.txt # numpy, pyyaml, pytest, ruff, pillow (figures)
python3 -m pytest -q # world, dynamics, lab, server (fake stdio)
ruff check .
claude plugin validate --strict .
claude plugin eval . --runs 1 # the three skill evals under evals/
python3 docs/figures/make_figures.py # redraw the README figures from fresh runsRunning the lab without Claude:
from simlab.lab import Lab
lab = Lab()
r = lab.run("swarm-formation", {"msg_loss": 0.6}, label="loss60")
print(r["headline"])
print(lab.compare([r["run_id"]]))Evals
claude plugin eval . runs three cases with and without the plugin (ablation): designing a one-knob sweep from a question, reading a compare table honestly, and writing a findings note. Last run (Claude Code 2.1.288, three runs per case and arm, 2026-10-03):
case | with plugin | without | Δ |
design-sweep | 1.00 | 0.33 | +0.67 |
read-compare | 1.00 | 1.00 | 0.00 |
write-findings | 1.00 | 0.00 | +1.00 |
Reading a compare table is something Claude does well on its own; that case guards the skill's reading rules (direction of each metric, run ids as citations, a seed re-run before a conclusion) rather than adding capability. The cases do not start the MCP server (no sandbox-safe mock yet); they test what the skills teach Claude to do with the lab's output.
Limits (honest ones)
The worlds are kinematic and radio-based: no cameras, lidar or terrain; water and fog are presets (drag, a current, slow lossy links; shrunken sensing), not fluid or light models; obstacles stop agents and can block radio, doors do not open; no adversaries. The catalogue's not and missing entries say the same thing to Claude, so it will not promise what the lab cannot do. docs/ROADMAP.md lists what comes next.
Licence
PolyForm Noncommercial 1.0.0: free to use, change and share for any noncommercial purpose (personal projects, research, teaching, hobby and student work, charities and public institutions all count), with attribution. Commercial use needs a separate agreement; write to the author. See LICENSE and NOTICE.
Available Tools
12 toolscatalogueA
The worlds the lab can simulate: what each one is and is not, every knob with unit and range, the metrics and which are better when higher, and example questions with the variants that answer them. Read this first.
| Name | Required | Description | Default |
|---|---|---|---|
| world | No | one world id, else all |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden — and it does disclose the returned payload in detail: each world's capabilities and limits, knob units and ranges, metric directionality, and example questions with their answering variants. It does not state side-effect freedom or invocation cost, but for a read-only catalogue the content disclosure is the substance an agent needs.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The content is packed into one dense sentence plus a two-word imperative, with no filler. The ordering is slightly suboptimal — "Read this first" would land harder at the front than at the end — but nothing is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and no annotations, the description compensates by enumerating what the catalogue contains, so an agent knows what it will receive. It is complete for the task of orientation; only the read-only/side-effect status is left unstated.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the single `world` parameter's semantics ("one world id, else all") are fully documented in the schema. The description adds nothing about filtering or the id format, so the baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly identifies the resource: a catalogue of the simulated worlds, their configuration knobs, metrics, and example questions. It is readily distinguishable from siblings like list_experiments or run_experiment, which deal with experiments rather than the world definitions themselves, though it never names a sibling explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
"Read this first" gives an explicit sequencing instruction that tells the agent to consult this tool before acting on any sibling. It lacks any when-not guidance or explicit alternative routing, but the directive is strong and unambiguous.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
compare_runsB
Headline metrics side by side and the best run per metric (lower is better unless the catalogue says otherwise; a tie names no best). Unknown run ids are an error.
| Name | Required | Description | Default |
|---|---|---|---|
| run_ids | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It usefully discloses non-obvious behavior: the 'lower is better unless the catalogue says otherwise' metric rule, tie handling ('a tie names no best'), and error semantics for unknown run ids, but it omits return format, ordering, or permissions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two dense, front-loaded sentences with zero filler; the output description and the tie/error rules each earn their place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and no annotations, the description still conveys the conceptual return (metrics side by side, best per metric) and key error/tie behavior. For a low-complexity one-parameter tool this is nearly complete, though metric definition depends on the unexplained 'catalogue'.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% for the single run_ids parameter, so the description must compensate. It names run ids and clarifies that unknown ids are an error, and 'side by side' implies multiple ids, adding some meaning beyond the schema, but it doesn't spell out the array format or the minItems=2 constraint.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (compare) and resource (runs) and describes the output shape: headline metrics side by side plus the best run per metric. It is clearly distinguishable from siblings like get_run or list_runs, though it doesn't name them explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies comparison usage but gives no when-to-use conditions, no prerequisites, and no named alternatives versus siblings such as list_runs or get_run. An agent gets no explicit routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_experimentA
One experiment: description, every parameter with its default, the variants, the metrics it writes, and its YAML.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It does useful work by disclosing the return payload (parameters with defaults, variants, metrics, YAML), which is genuinely informative. However, it says nothing about read-only safety, behavior when the name does not exist, or permissions — and the response fields are described rather than the operation's behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence that enumerates the return contents with zero filler. Every clause conveys distinct information about what the agent will receive.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and no annotations, the description compensates well by listing the returned fields, which is exactly the gap an absent output schema leaves. It is thin only on input semantics and failure behavior for a simple one-parameter read tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the single required 'name' parameter is undocumented in both schema and description. The one identifier is self-evident (experiment name), which limits the damage, but the description's phrase 'every parameter with its default' refers to returned experiment parameters, not the input, and could momentarily confuse the agent.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names the resource (a single experiment) and enumerates exactly what it returns: description, parameters with defaults, variants, metrics, and YAML. The verb is implied by the name 'get' and the leading 'One experiment:' framing, and it is distinguishable from list_experiments by scope. It is clear but never states the operation explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit when-to-use or when-not-to-use guidance. The phrase 'One experiment' implicitly contrasts with list_experiments (bulk listing) versus fetching full detail for a single named experiment, so usage is inferable but not stated. No prerequisites, no error conditions, no sibling named.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_runB
One run in full: manifest (command, params, timings, versions, file hashes), metrics, report, the last lines of its stdout.log (the world's own output and errors) and the files in its folder.
| Name | Required | Description | Default |
|---|---|---|---|
| run_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral burden. It usefully discloses the payload shape, notably that stdout.log is only the last lines, which sets expectations about truncation. However, it omits auth/permission requirements, error behavior for unknown run_ids, and whether the read has side effects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single dense sentence with no filler; the payload enumeration is front-loaded and every clause conveys content. The parenthetical aside ('the world's own output and errors') is slightly editorial but still informative.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description must describe the return value, and it does so reasonably well. Gaps remain around how to source the run_id and what happens on failure, and with no annotations there is no safety profile to fall back on.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for the single run_id parameter, so the description must compensate and does not. It never states the run_id format, where to obtain a valid value, or whether it is an ID versus a name/path.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names the resource (one run, keyed by run_id) and enumerates its contents in detail, which lets an agent distinguish it from list_runs and compare_runs. The verb is implicit in the tool name rather than stated, but the payload enumeration is specific enough to convey purpose.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit when-to-use guidance, no statement of prerequisites (e.g., that run_id comes from list_runs), and no mention of when to prefer compare_runs or get_experiment instead. Usage is only inferable from the tool name.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_experimentsC
Experiments (shipped and the person's own) with their parameters and named variants.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations and no output schema, the description carries the full behavioral burden. It hints at the return shape ('parameters and named variants') but says nothing about ordering, pagination, whether it is a read-only safe operation, or any auth requirement. That is thin for a tool with zero structured behavioral coverage.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single short sentence fragment with no filler or redundancy. It is efficient, though being a fragment rather than a front-loaded sentence with a clear verb slightly weakens the opening.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter list tool with no annotations and no output schema, the description supplies only the rough content of the listing. It omits ordering, pagination, and how the result set differs from sibling list tools, leaving an agent to guess at runtime behavior.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has zero parameters, so there is nothing to document and the baseline of 4 applies. The description does not need to explain arguments and correctly does not attempt to.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description is a noun phrase ('Experiments ... with their parameters and named variants') rather than a verb+resource statement, though the name list_experiments implies enumeration. It does convey scope (shipped experiments plus the caller's own, including parameters and variants), which distinguishes it somewhat from get_experiment, but it never states that it retrieves a collection.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit guidance on when to call this versus get_experiment, catalogue, or list_runs, nor any mention of prerequisites or filters. The scope hint ('shipped and the person's own') implies a use case but leaves the agent to infer the alternative-tool routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_plansA
Floor plans for the swarm worlds (use as layout: plan:): size, walls, doors and the ASCII text.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It discloses the return payload shape (size, walls, doors, ASCII), which implies a safe read, but never explicitly states read-only behavior, permissions, or output format/pagination.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single compact sentence that front-loads the resource and then the payload. Dense and waste-free, though the parenthetical is slightly awkwardly embedded mid-sentence.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter list tool with no output schema, the description adequately covers what the tool returns and how the results are referenced. It could say more about whether all plans are always returned, but nothing essential to invoking it is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema declares zero parameters, so per the rubric the baseline is 4. There is nothing for the description to compensate for.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific resource (floor plans for the swarm worlds) and enumerates the returned content (size, walls, doors, ASCII text). The 'list' verb is implied by the name and the plural resource. It distinguishes from save_plan but not explicitly from catalogue or other list_* siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The parenthetical 'use as layout: plan:<name>' hints at how the returned names are consumed, which is genuine usage context. However, it never says when to call this versus save_plan or catalogue, so selection guidance is only implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_runsC
Recent runs, newest first, with status and headline metrics.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| experiment | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It conveys ordering ('newest first') and returned fields, but leaves 'recent' undefined, gives no default limit, no pagination or rate-limit notes, and no auth requirements for a read tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single tight sentence that front-loads the resource and ordering with no filler. It is efficient, though its brevity comes at the cost of missing detail rather than being genuinely complete.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 2-parameter list tool with no annotations and no output schema, the description should document parameter intent and default behavior. It covers return content and ordering but leaves both parameters and the meaning of 'recent' unexplained.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the description never mentions the 'limit' or 'experiment' parameters. The only implicit signal is that results are 'recent', vaguely suggesting a default time window, but neither parameter's meaning or bounds is explained.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific resource (runs) and the ordering and content (newest first, with status and headline metrics). It implicitly differs from siblings like get_run and list_experiments by being a list of runs, but it never names or contrasts those siblings explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance, no prerequisites, and no named alternative. The word 'recent' hints at a default scope but the description gives no condition that would route an agent to get_run or compare_runs instead.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
run_campaignB
Run several named variants of one experiment one after another and return the comparison table. Blocks until all of them end.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | ||
| params | No | parameter overrides, by name (see get_experiment) | |
| variants | No | variant names; default: all of them |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, and it does disclose one important trait: the call blocks until all variants finish. It omits failure semantics (what happens if one variant fails mid-sequence), whether results are persisted, and resource/duration implications of running several sequentially.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, no filler, with the core action and the blocking behavior both front-loaded. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
It correctly compensates for the absent output schema by naming the return value (comparison table) and covers blocking behavior for a nested-params tool. However, it leaves failure handling and result persistence unaddressed, which matters for a multi-run orchestration tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 67% — params and variants are documented in the schema, but the required name is bare. The description reinforces 'named variants of one experiment' and the default-to-all behavior is only in the schema, so it adds little beyond what the schema already states.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (run) and resource (several named variants of one experiment) plus the sequential nature and the returned comparison table. It implicitly separates itself from run_experiment by operating on multiple variants, though it never names that sibling.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit when-to-use or when-not-to-use guidance. An agent must infer that this is the multi-variant alternative to run_experiment, and nothing states prerequisites, cost, or when a single run or tune would be preferable.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
run_experimentA
Run an experiment to completion with optional parameter overrides or a named variant; returns the run id, status, headline metrics and the report (and the tail of its log when it failed). The call blocks until the run ends: a swarm run of 400 ticks takes a few seconds, 3000 ticks with 60 agents about a minute. Values outside the catalogue's ranges are refused.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | ||
| label | No | a short label stored on the run | |
| params | No | parameter overrides, by name (see get_experiment) | |
| variant | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are present, so the description carries the full burden and does well: it discloses that the call blocks until the run ends, gives concrete latency expectations (400 ticks ~seconds, 3000 ticks/60 agents ~a minute), states that out-of-range values are refused, and explains what is returned including failure logs. It is silent on permissions or persistence side effects, keeping it short of a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences, front-loaded with purpose and return values, then the blocking/latency behavior and validation rule. Every clause adds operational value with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description takes on return-value explanation and does so concretely (run id, status, headline metrics, report, failure log tail), plus runtime and validation behavior. Only the semantics of the required 'name' and 'label' inputs remain unaddressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 50%, so the description must compensate. It clarifies that params are overrides 'by name (see get_experiment)' and that variant is a named variant, but says nothing about the required 'name' or the 'label' parameter, leaving half the parameters unilluminated beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Run an experiment to completion') and adds the two input modes (parameter overrides or a named variant). This clearly separates it from read-oriented siblings like get_experiment, list_runs, and compare_runs.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear context for when this tool applies (synchronous execution of an experiment) and points to get_experiment for parameter names. It does not, however, address when to prefer sibling tools like run_campaign or tune, so the alternative-selection guidance is incomplete.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
save_experimentA
Save an experiment YAML of the person's own under ~/.simlab/experiments (validated: name, command with {run_dir}, params for every placeholder, variants that only set known params). The command is any local program the person wants the lab to run and score; it runs on their machine with their rights, so only save a command the person asked for. A shipped experiment's name is refused unless overwrite is true (then their file replaces it for every later run).
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | ||
| yaml | Yes | ||
| overwrite | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations, so description carries the burden. It discloses important behavior: validation rules, the command runs on the person's machine with their rights, and overwrite semantics for shipped experiments. Lacks details on return value or side effects beyond file creation, but covers key safety aspects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, front-loaded with core action, followed by validation rules and safety/overwrite details. Efficient, though slightly dense with parentheticals.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given 3 parameters, no annotations, and no output schema, the description provides sufficient context: validation, safety, overwrite behavior, and scope. It misses explicit mention of return value or pagination but that's acceptable. Overall, an agent would understand how to use it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, but the description explains semantic meaning for 'name' (refused for shipped unless overwrite true), 'yaml' (validated structure), and 'overwrite' (replaces file for later runs). However, it doesn't fully document the format or constraints of each parameter, leaving some ambiguity.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource: 'Save an experiment YAML'. It clarifies the scope ('of the person's own') and location ('~/.simlab/experiments'). This distinguishes it from siblings like save_plan or run_experiment, though it doesn't name them explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Implies usage when the person wants to save an experiment, and includes a safety caution ('only save a command the person asked for'). But it doesn't compare or contrast directly with sibling tools like save_plan or get_experiment; no explicit when-not-to-use guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
save_planA
Save an ASCII floor plan under the person's lab folder (~/.simlab/plans): # wall, D door, . free, at least 3x3, at least one free cell; it becomes layout plan:. A shipped plan's name is refused unless overwrite is true.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | ||
| text | Yes | ||
| overwrite | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does well: it discloses the storage path (~/.simlab/plans), the naming scheme (layout plan:<name>), the validation constraints (at least 3x3, at least one free cell), and the overwrite refusal behavior. It omits permission requirements and return values, but the mutation/refusal semantics are clearly stated.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single dense sentence that front-loads the core action and packs the format legend and constraints without filler. It is slightly overloaded, merging format spec, storage path, and overwrite rule into one clause chain.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 3-parameter mutation tool with no annotations and no output schema, the description covers storage location, naming, input format, validation, and overwrite behavior. Return-value details are not needed since there is no output schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for all three parameters and largely does: text is specified via the ASCII legend (# wall, D door, . free) plus size constraints, name via the layout naming scheme, and overwrite via the refusal-unless-true rule. Only minor format/edge details are left implicit.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource: saving an ASCII floor plan, plus the storage location and the resulting layout name. It distinguishes itself from read-oriented siblings like list_plans and from save_experiment, though it does not explicitly contrast them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The only usage condition given is implicit in the overwrite rule ('a shipped plan's name is refused unless overwrite is true'). There is no explicit when-to-use/when-not guidance or named alternatives, so the agent must infer the context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
tuneA
Search a parameter space for the best headline metric: strategy grid | random | bayes (Gaussian process + expected improvement). Each trial is a run and the call blocks until the budget is spent, so keep budget small (max 40). The direction comes from the catalogue unless minimize is given.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | ||
| space | Yes | {param: [values...]} | |
| budget | No | ||
| params | No | parameter overrides, by name (see get_experiment) | |
| minimize | No | ||
| strategy | No | ||
| objective | Yes | a headline metric name |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does well: it discloses that each trial is a run, that the call blocks until budget is spent, and that budget is capped at 40. It omits whether the run persists artifacts or requires specific auth, but the expensive/blocking nature is clearly surfaced.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three dense sentences, front-loaded with the core action, then strategy options, then blocking/budget warning. Efficient with little waste, though the parenthetical GP detail is optional.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 7-param, nested-object tool with no annotations and no output schema, the description covers behavior and key params but never states what the call returns (e.g., best params/metric) or how to inspect resulting trials via list_runs/get_run. Adequate but with a clear gap on return semantics.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 43%, so the description must compensate, and it does for the key params: strategy values (grid | random | bayes with GP+EI), budget (max 40), objective (headline metric), and minimize (overrides catalogue direction). Space and params are left to the schema; name is unexplained.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource: 'Search a parameter space for the best headline metric.' The tool's distinct behavior (tuning/optimization) is clear, though it does not explicitly contrast itself with siblings like run_campaign or run_experiment that also consume runs.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives useful context ('keep budget small (max 40)') and notes direction comes from the catalogue unless minimize is given, but offers no explicit when-to-use-this-vs-alternatives guidance relative to run_campaign/run_experiment. Usage is implied rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
12 tool updates
v0.1.0- First observed
catalogue - First observed
compare_runs - First observed
get_experiment - First observed
get_run - First observed
list_experiments - First observed
list_plans - First observed
list_runs - First observed
run_campaign - First observed
run_experiment - First observed
save_experiment - First observed
save_plan - First observed
tune
TDQS
Scored across 12 tools
Each tool maps to a distinct resource and action: catalogue for discovery, experiment CRUD/run (list/get/run/save), campaign batching, tune for search, run inspection (list/get/compare), and plan listing/saving. run_experiment, run_campaign, and tune all execute but are cleanly separated by scope (one, several variants, parameter search).
Strong verb_noun snake_case pattern throughout (list_experiments, get_experiment, run_campaign, compare_runs, save_plan). Two outliers—catalogue and tune—lack a noun and break the pattern, but the convention is otherwise predictable and readable.
12 tools is well-scoped for a simulation lab covering discovery, experiment lifecycle, campaign execution, tuning, run inspection, and plan authoring. Every tool earns its place with no redundant surface.
Core lifecycle is covered: discover, define, run, batch, tune, inspect, and compare. Minor gaps exist—no delete for experiments/plans, no single get_plan (only list_plans), and no run cancellation—but these are workable around for most workflows.
Maintenance
Related MCP Connectors
- alloyOAuthai.usealloy
Connect Claude, Cursor, Codex, and other AI tools to your robotics mission data.
- SimSenseOAuthai.simsense
Deploy sims to any screen. Control your displays with Claude.
Deploy sims to any screen. Control your displays with Claude.
Persistent cloud workspaces for AI agents: run commands, edit files, use git and a browser.
1
Related MCP Servers
- AlicenseNot gradedqualityBmaintenanceEnables ML researchers to manage experiments across local and remote AutoDL GPU instances, including experiment creation, training launch, run polling, and report writing via Claude Code.1MIT
- AlicenseNot gradedqualityAmaintenanceRun dynamic, multi-agent workflow scripts — agent(), parallel(), pipeline() — over real coding agents (Claude Code and OpenAI Codex), with deterministic journaling, resume, token budgets, and git-worktree isolation.7Apache 2.0
- AlicenseBqualityAmaintenanceLocal-first Agent OS that wraps Claude Code, Codex CLI, and other coding agents in a replayable Seed → Ledger → Runtime contract, driven by an interview → seed → execute → evaluate → evolve workflow loop.3416,680 PyPI6,172MIT
- FlicenseAqualityCmaintenanceLets you @mention a bot in a Slack channel or DM to send work to a local Claude Code session and watch it run, with live self-updating turn cards, threaded answers, progress checklists and a per-channel inbox with read cursors. It runs entirely on your machine over Slack Socket Mode—no tunnel or public IP—routes misdelivered messages to the right project, and masks credentials in anything it posts back.7-