dex-isaac-mcp
An MCP server that drives a live, persistent Isaac Sim session and launches Isaac Lab training runs — mostly for tuning and debugging articulated robots.
Run the sim: start/stop/reload a daemon (
sim_up,sim_down,sim_reload), check status, step physics, play/pause.Inspect: read joint positions/velocities, per-joint travel and tracking error, and a USD's articulation DOFs, loop-closure joints and roots.
Drive a robot: set normalized 0..1 joint targets, apply named poses (partially blended), wave every joint through its range, or run a range test that flags blocked joints.
Tune live: change stiffness/damping/effort, set software coupling ratios, and sweep one parameter over several values (restoring the original afterwards).
Build scenes: spawn cuboids, spheres, cylinders, capsules, cones or USD props (static or kinematic), list their world poses, remove them.
Capture: take screenshots, aim the viewport or a headless capture camera, auto-frame the robot, set a backdrop, and record captioned GIFs.
Train: launch headless Isaac Lab runs in their own containers, then poll status, tail logs, list checkpoints, read TensorBoard scalars, and stop runs.
Works with any articulated robot via a small JSON config (USD, driven/passive joints, couplings, gains, solver, poses), with Franka and Allegro examples included.
Drives a live, persistent NVIDIA Isaac Sim (Kit) session — starting/stopping the daemon, inspecting articulation joints and state, setting normalized joint targets and named poses, stepping/playing physics, capturing screenshots and tuning gains and couplings — and launches and monitors NVIDIA Isaac Lab reinforcement-learning training runs, including their status, logs, TensorBoard scalars and checkpoints.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@dex-isaac-mcpstart the sim with the Franka, move it to ready, and take a screenshot"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
dex-isaac-mcp

Recorded headless with the server's own tools (sim_frame_robot, sim_set_backdrop, sim_record_start), driving the Allegro hand example. Each caption is the request and the tool call it became.
An MCP server that lets an AI agent (Claude Code, or any MCP client) drive a live, persistent Isaac Sim session and launch Isaac Lab training runs.
Kit takes tens of seconds to boot. If every experiment is a fresh launch, most of your time goes to waiting. Here Kit starts once inside a daemon and stays up. Each tool call lands in that running session, so changing a gain, stepping physics or taking a screenshot costs a frame, not a relaunch.
MCP client ──stdio──▶ dex_isaac_mcp (host, plain python)
│ newline-delimited JSON over a Unix socket
▼
scripts/simd.py (Isaac Lab container, Kit stays up)
└─ one articulation, described by a robot JSONHow it differs from other Isaac Sim MCP servers: those mostly build scenes from language ("add a table and a Franka"). This one is a workbench for an agent that tunes and debugs a robot: it measures (range tests, tracking error, parameter sweeps that restore the original value) and runs training. Scene props are supported, but a scene builder is not the point.
Any articulated robot. Point it at a USD and a small JSON config. Franka and Allegro examples are included.
Normalized joint control. Targets are
0..1, where 0 is a joint's lower limit and 1 its upper limit, so agents don't need to know radians or meters.Named poses per robot (
home,fist, …), which can be blended part-way.Headless capture and GIF recording from an auto-framed camera, with captions. The clip above was made this way.
Props: spawn a table, a ball or a USD into the running scene and read back where they settle.
Measurements: joint state, per-joint travel and tracking error, a range test that finds blocked joints, and parameter sweeps inside one session.
Training control: each run is a detached
docker compose run. You can poll its status, TensorBoard scalars, checkpoints and logs.
Requirements
Linux with an NVIDIA GPU that Isaac Sim supports, plus the NVIDIA Container Toolkit
Docker with the Compose plugin
An NGC login to pull the Isaac Lab image:
docker login nvcr.io(user$oauthtoken, password: your NGC API key)Python ≥ 3.10 on the host, for the MCP server only
Related MCP server: omni-kit-mcp
Quick start
git clone <this repo> dex-isaac-mcp && cd dex-isaac-mcp
# 1. Build the image (Isaac Lab 2.3.2 base, pinned by digest)
cd docker && docker compose build && cd ..
# 2. Install the host-side server, editable so it finds this clone
pip install -e . # add [metrics] for train_metrics: pip install -e '.[metrics]'
# 3. Register it with Claude Code
claude mcp add isaac -- python -m dex_isaac_mcp
# 4. Allow the container to open windows (once per login; needed for screenshots)
xhost +local:dockerThen ask the agent something like "start the sim with the Franka, move it to ready and show me a screenshot." It will call sim_up, sim_set_pose, sim_step and sim_screenshot.
From PyPI instead: pip install dex-isaac-mcp. The daemon still runs from a clone (it needs docker/ and scripts/simd.py), so point the server at it:
claude mcp add isaac -e ISAAC_MCP_HOME=/abs/path/dex-isaac-mcp -- dex-isaac-mcpOther MCP clients can launch python -m dex_isaac_mcp (or the dex-isaac-mcp script) over stdio.
The first sim_up takes several minutes: Kit builds its shader cache and downloads Nucleus assets. Later starts are much faster.
Running the daemon by hand
sim_up deletes its container (--rm) when the daemon exits, so a crash on startup takes its traceback with it. To see the error, run the daemon in the foreground:
cd docker
docker compose run --rm simd scripts/simd.py --gui --robot examples/robots/franka.json
docker compose run --rm simd scripts/simd.py --headless --usd /path/in/container/robot.usd
docker compose run --rm simd scripts/simd.py --sliders --robot examples/robots/allegro_hand.json--sliders opens an omni.ui panel with one slider per driven joint. The socket stays live alongside it.
Tools
Session
Tool | What it does |
| Whether the daemon is up, plus its robot, driven joints, poses, couplings and gains |
| Start the daemon container ( |
| Stop the daemon |
| Restart with a different spawn property: robot, USD, |
Inspect
Tool | What it does |
| A USD's articulation DOFs, its loop-closure joints (excluded from the articulation) and its articulation roots. Reads the file, so it shows edits saved from the GUI |
| Positions and velocities, plus driven joints' normalized positions and limits |
| Per-joint travel and mean tracking error since the last reset |
| Viewport capture, returned as an image. Needs |
| Point the viewport camera |
Drive
Tool | What it does |
| Normalized targets, as a full vector or |
| Named poses from the robot config. |
| Advance N physics steps (default dt 1/120 s) |
| Run continuously, or pause |
| Sweep every driven joint through its range and return travel stats |
| Drive every joint from its lower limit toward a target and report the fraction of travel reached. Below 0.9 counts as blocked (self-collision, a binding linkage, too little effort) |
Scene
Tool | What it does |
| Add a prop to the live scene: |
| Every prop's current world pose |
| Delete a prop |
Capture and record
These work headless, with no GUI or viewport. They use a dedicated camera that is independent of the GUI view.
Tool | What it does |
| Aim the capture camera so the whole robot fills the frame, from a given direction. |
| Place the capture camera by hand |
| A plain colored panel behind the robot. Use it with |
| One frame, returned as an image |
| Record every Nth physics step, with a caption drawn on each frame, to an animated GIF under |
Tune
Tool | What it does |
| Live |
| Live software-mimic ratios, keyed by follower joint |
| Try several values of one live parameter ( |
Train
Tool | What it does |
| Launch a headless training run in its own container and return at once. |
| Running and recent runs, and log directories holding checkpoints |
| Container state, checkpoints and latest scalars |
| Tail a running container's output |
| List TensorBoard tags, or get a downsampled series for one tag |
| Checkpoints with step and size |
| Stop a run |
Training runs do not depend on the daemon or on the MCP session. They keep going after the client disconnects.
Robot config
A robot is one JSON file. Only usd is required. Unknown keys are rejected, so a typo fails loudly instead of quietly falling back to a default.
{
"name": "franka",
"usd": "{ISAACLAB_NUCLEUS_DIR}/Robots/FrankaEmika/panda_instanceable.usd",
"fix_root_link": true,
"spawn_pos": [0, 0, 0],
"init_joint_pos": {"panda_joint4": -2.81, "panda_joint6": 3.04, ".*": 0.0},
"driven_joints": ["panda_joint[1-7]", "panda_finger_joint.*"],
"passive_joints": [],
"couplings": [{"leader": "joint_a", "follower": "joint_b", "ratio": 1.0}],
"actuator": {"stiffness": 400, "damping": 40, "effort": 87, "velocity": null},
"solver": {"pos_iters": 32, "vel_iters": 4, "self_collisions": true},
"camera": {"eye": [1.8, 1.8, 1.4], "target": [0, 0, 0.4]},
"poses": {"ready": {"panda_joint[1357]": 0.5, "panda_finger_joint.*": 1.0}}
}Key | Meaning |
| Local path (relative paths resolve from the JSON's own directory), a URL, or a path using |
| Spawn pose in the joints' own units (rad / m), keyed by regex. It must lie inside every joint's limits or spawning fails. The default is all zeros, which is out of range for e.g. Franka's joint 4 |
| Regexes (full match) for joints that take commands. Default |
| Joints whose angle is owned by a constraint, such as a closed-chain linkage. They get a zero-stiffness drive, because a live PD drive fights the constraint and the mechanism jitters |
| Software mimic joints: follower target = |
| Implicit PD gains and effort/velocity limits for the driven joints. Live-tunable |
| Spawn properties. Changing them needs |
|
|
Using your own robot without forking
Keep the robot config and assets in your own repo, and add them to the container with a compose override that you list in COMPOSE_FILE. Set it in the MCP server's environment, using absolute paths:
# my-robot/mcp-compose.yaml
services:
simd:
volumes:
- /abs/path/my-robot:/workspace/my-robotclaude mcp add isaac \
-e COMPOSE_FILE=/abs/path/dex-isaac-mcp/docker/docker-compose.yaml:/abs/path/my-robot/mcp-compose.yaml \
-e ISAAC_MCP_ROBOT=/workspace/my-robot/robot.json \
-- python -m dex_isaac_mcpTraining defaults
Out of the box, train_start runs Isaac Lab's stock skrl script inside the isaac-lab service. It tags the run name onto the log directory (logs/skrl/<experiment>/<timestamp>_ppo_torch_<run_name>/), which the other train_* tools use to find the run. To use your own launcher, set these in the environment the MCP server starts in:
Variable | Default |
| the clone this package was installed from (editable install); required for a PyPI install |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| robot config |
|
|
The script must accept --task, --headless and, when given, --num_envs, --seed, --max_iterations and --checkpoint. Tasks from your own extension need to be importable inside the container, either installed into the image or mounted.
Design notes
These are the constraints the code is built around. Most were learned by breaking them.
Every Kit call happens on the main thread. Kit, PhysX and USD are not thread-safe. Socket threads only parse JSON and queue requests, and the main loop executes them between physics steps. Answering from a reader thread appears to work, then corrupts the stage under load.
Spawn properties are frozen. Replacing a spawned articulation needs
SimulationContext.stop(), which blocks on a timeline event that only advances while the Kit loop pumps. A command runs on that loop, so the call never returns.omni.usdnew_stage()has the same trap. So the USD, solver iterations and self-collision need a restart (sim_reload), and gains stay live.A Unix socket, not TCP. The repo is bind-mounted and the container runs as the host uid, so the host sees the socket file directly, with no port mapping. Paths are capped at 107 bytes (
AF_UNIX). If your checkout is deep, setISAAC_MCP_SOCKET.The host side imports no Isaac code.
protocol.py,robot.pyandtraining.pyare stdlib-only. The MCP server adds onlymcp. Nothing on the host needs isaaclab, torch or a GPU.Cache directories are committed with
.gitkeep. If Docker auto-creates a bind-mount source, it is root-owned, and Kit then dies withregistry cache path is not setbefore any script runs.The base image is pinned by digest. A re-pulled tag once shipped
/isaac-simas mode 750, and every non-root container lost its Python.Recorded GIFs are stabilized. The renderer's denoiser shimmers: between two frames of a motionless scene, about 9% of background pixels change slightly, and a GIF re-encodes every one of them. Holding sub-threshold changes and using one shared palette took a 9-second clip from 15 MB to 1.2 MB.
The daemon always renders, even headless (
enable_cameras). Without rendering, PhysX never registers a prop spawned at runtime. Prop poses are read from fabric, because the USD transform and the PhysX CPU query both stay at the spawn pose, and creating a PhysX tensor view mid-simulation crashes CUDA.No floor for range tests (
ground=False) on anything whose links can reach the ground. Otherwise the test measures the floor, not the robot.
Development
python -m unittest discover tests # host-side tests: no Isaac, no GPU
ruff check .dex_isaac_mcp/protocol.Client is a handy debugging client:
from dex_isaac_mcp.protocol import Client
with Client() as c:
print(c.call("status"))
c.call("set_pose", name="ready"); c.call("step", n=240)Status
Tested against Isaac Lab 2.3.2 (Isaac Sim 5.x) and mcp 2.3, over the stdio protocol, with the GUI on:
Franka and Allegro examples, plus a custom closed-linkage hand through a compose override:
sim_up, poses, screenshots,sim_range_test,sim_sweep(restores the original value),sim_down.Headless daemon: gains, frozen-parameter rejection, wave.
Headless capture and recording, Allegro and Franka: auto-framing, backdrop, captions, GIF output (the clip at the top).
Props, GUI and headless: a sphere and a cylinder dropped onto a static table settle at exactly table height plus their radius and half-height.
Training, against Isaac Lab's stock skrl script:
Isaac-Cartpole-v0launched, polled, logged, checkpointed and read back through everytrain_*tool, plus a run stopped mid-training. skrl'swrite_interval: autowrites no TensorBoard scalars on a very short run (5 iterations), sotrain_metricscomes back empty there; 50 iterations gives 18 tags.
Issues and PRs are welcome.
License
MIT, see LICENSE.
Available Tools
26 toolssim_downA
Stop the sim daemon; its container removes itself.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It does disclose one genuinely non-obvious trait — the container removes itself, i.e. this is irreversible teardown — but says nothing about in-flight state, permissions, or what happens to any running training/sim data.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One clause plus one semicolon-linked consequence, front-loaded with the action. No padding, no restatement of the name.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained, and a zero-param tool needs no argument documentation. The only shortfall is that the shutdown's effect on existing sim state is left implicit.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so the baseline is 4; there is nothing for the description to clarify beyond the empty schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Stop the sim daemon'), which cleanly contrasts with the sibling sim_up. It does not explicitly differentiate from sim_reload or sim_status, so it falls short of full sibling disambiguation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by the action itself: an agent can infer this is the shutdown counterpart to sim_up. There is no explicit when-to-use guidance, no stated preconditions, and no routing between sim_down and sim_reload.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
sim_get_joint_stateB
Joint positions (rad/m) and velocities, plus driven joints' normalized positions and limits.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full disclosure burden. It implies a read ('get') and lists the returned quantities, but says nothing about side effects, whether it blocks until a simulation step, or whether it requires a running/paused sim — and much of what it does say overlaps with the output schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single dense sentence that front-loads the primary return values (positions, velocities) and appends the secondary driven-joint data. No wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, the description need not explain return values, and for a zero-argument getter this is largely sufficient. The one real gap is absence of any routing hint versus sim_inspect_joints.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so the baseline is 4. The description correctly implies a global, unfiltered read of all joint state with no arguments needed.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific resource and enumerates exactly what state is returned: joint positions in rad/m, velocities, and driven joints' normalized positions and limits. It is clear what the tool yields. It does not, however, distinguish itself from the sibling sim_inspect_joints, leaving the reader unsure which to pick.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance, no prerequisites, and no mention of the closely related sim_inspect_joints or sim_get_stats siblings. The description only describes the payload, so the agent must infer call context on its own.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
sim_get_statsA
Per-joint travel and driven-joint mean tracking error since the last reset.
Travel on a passive joint proves a linkage transmits: with zero stiffness it moves only if its constraint moves it.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations, so the description carries the full burden. It discloses that values are accumulated 'since the last reset,' which is critical behavioral context for interpreting the stats. It does not mention reset behavior itself or return format, but the output schema likely covers the latter.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core purpose in one efficient sentence. The second paragraph is somewhat tangential—useful as a domain hint but not strictly necessary for invocation; still, it's brief.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no parameters, a rich output schema, and no annotations, the description provides the essential scope ('since the last reset') and domain context for interpretation. It is nearly complete, missing only explicit routing to sibling tools.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Zero parameters, so baseline is 4. The description appropriately avoids discussing nonexistent parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Clearly states the resource (per-joint travel and driven-joint mean tracking error) and the scope (since last reset), which distinguishes it from siblings like sim_get_joint_state by returning accumulated statistics rather than instantaneous state. However, it doesn't explicitly contrast with sim_status or sim_inspect_joints.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this tool versus sim_get_joint_state, sim_status, or others. The second paragraph explains a concept (travel on a passive joint) but offers no operational context or exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
sim_inspect_jointsA
List a USD's articulation DOFs, joints excluded from the articulation (loop closures), and roots.
Reads the file, not the live scene, so it shows edits saved from the Isaac GUI. Defaults to the loaded USD.
| Name | Required | Description | Default |
|---|---|---|---|
| usd | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations, so the description carries the burden, and it discloses the critical trait that results come from the file rather than the live scene and default to the loaded USD. It does not cover error behavior or the usd format, but the key data-source caveat is clearly surfaced.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with what is listed, and every clause (loop closures, file-vs-live, default) earns its place. Slightly dense but no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists so return values need not be described, and for a single-optional-parameter read tool the description covers scope, source, and default. Only the usd parameter format is left unspecified.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The single usd parameter has 0% schema coverage, so the description must compensate; it only explains the default ('loaded USD') and gives no hint of expected format (path, name, prim). Partial compensation over an otherwise undocumented parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb (List) plus precise resources: articulation DOFs, loop-closure joints, and roots. The 'Reads the file, not the live scene' clause distinguishes it from the live-state sibling sim_get_joint_state, so an agent can tell them apart without opening schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The file-vs-live-scene note implies when this tool is appropriate (inspecting saved GUI edits) but never explicitly states when to use it instead of sim_get_joint_state or names that alternative. Usage must be inferred rather than read.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
sim_list_posesB
Named poses from the robot config, as {pose: {joint_pattern: 0..1}}.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It discloses the return shape ({pose: {joint_pattern: 0..1}}), which is useful behavioral context, but it never confirms this is a read-only operation with no side effects — an inference an agent can reasonably make but that is not stated.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single compact sentence with no wasted words, front-loading the resource and source. It is a fragment rather than a sentence, which slightly weakens readability but costs no clarity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained, and with zero parameters the input side is trivially covered. The main omission is guidance on when this list is the right call versus the other inspection tools.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so the baseline of 4 applies. The description correctly adds nothing about arguments because there are none.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description identifies a specific resource (named poses) and its source (the robot config), which cleanly separates it from siblings like sim_set_pose and sim_inspect_joints. It omits an explicit verb such as 'list', leaving that to be inferred from the tool name.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no statement of when to call this tool versus sim_inspect_joints, sim_get_joint_state, or sim_set_pose, and no prerequisites or exclusions. The agent must infer usage entirely from the name and description.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
sim_playA
Run the sim continuously in real time (True) or pause it (False).
| Name | Required | Description | Default |
|---|---|---|---|
| playing | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It discloses the key semantic of the mutation (continuous real-time execution vs. paused), but says nothing about prerequisites (e.g., sim must be up via sim_up), idempotency, or state side effects of toggling playback.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One sentence, front-loaded with the action, and the boolean mapping is given inline with zero filler. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need no explanation, and the tool is a simple binary toggle. The description covers purpose and the parameter fully; the only gap is the unstated relationship to siblings like sim_step and sim_status and any required preconditions.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the sole boolean parameter 'playing' has no schema documentation. The description compensates by fully defining the semantics of both values (True = run continuously in real time, False = pause), which is exactly what is needed for a single-parameter tool.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific action (run continuously in real time / pause) on a named resource (the sim), so an agent knows exactly what the tool does. It does not explicitly distinguish itself from the closely related sibling sim_step, which is the main thing keeping it from a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The parenthetical explains what each argument value does, which implies when to call it (to resume or pause continuous simulation). However, there is no explicit when-to-use framing and no mention of the alternative sim_step for stepping the simulation instead of running it continuously.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
sim_range_testA
Drive every driven joint from its lower limit toward target, report the fraction reached.
A joint below 0.9 is blocked: self-collision, a binding linkage, or too little effort for the load. reset=True first puts every joint at its lower limit and verifies it. Start the daemon with ground=False if the robot can reach the floor, or the test measures the floor.
| Name | Required | Description | Default |
|---|---|---|---|
| reset | No | ||
| steps | No | ||
| target | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden, and it does disclose meaningful behavior: the drive direction, the 0.9 blocked threshold with its three possible causes, and what reset does to joint state. It stops short of saying what state the robot is left in after the test or what effort/permissions the daemon needs.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four short sentences with the core action front-loaded, followed by interpretation, then configuration caveats. Dense but every sentence carries information; no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return format need not be spelled out, and the description still explains how to read the reported fraction. The main gaps are the unexplained `steps` parameter and the absence of any statement about the robot's post-test state, which matters for a tool that physically drives joints.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It explains `target` (drive toward it) and `reset` (re-home and verify first), but `steps` is never mentioned, leaving a third of the parameters undocumented in both schema and prose.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: driving every driven joint from its lower limit toward `target` and reporting the fraction reached. This is a distinctive diagnostic action that an agent can separate from sim_inspect_joints or sim_set_targets, though no sibling is named explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives concrete operating conditions: reset=True first re-homes and verifies joints, and the daemon must be started with ground=False if the robot can reach the floor, otherwise the test measures the floor instead of joint range. It does not, however, say when to prefer this over sim_sweep or sim_inspect_joints.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
sim_reloadA
Restart the daemon to change a spawn property: robot, USD, solver iterations, self-collision.
One Kit start rather than one per experiment; gains stay live via sim_set_params. Omitted robot/usd fall back to the defaults, not to what was loaded before.
| Name | Required | Description | Default |
|---|---|---|---|
| gui | No | ||
| usd | No | ||
| robot | No | ||
| ground | No | ||
| pos_iters | No | ||
| self_collisions | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It does disclose a genuinely important trait: omitted robot/usd fall back to defaults rather than the previously loaded values, which is a destructive-feeling reset. It says nothing about state loss, downtime, or permissions, so it is only partially transparent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two dense sentences, front-loaded with the action and scope, with the alternative and the default-fallback caveat following. Nothing is wasted, though the second sentence is terse enough to be slightly cryptic.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need no explanation, and the description covers the core restart semantics and default fallback. However, for a 6-param mutation with zero annotation coverage, it leaves gui/ground unexplained and omits the disruption profile an agent would want before restarting a daemon.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% and there are 6 params, so the description must compensate. It names robot, usd, solver iterations (pos_iters) and self-collision (self_collisions), and clarifies default-fallback semantics for robot/usd, but it never mentions gui or ground and gives no format details for usd.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (restart) plus resource (the daemon) and enumerates the spawn properties it can change (robot, USD, solver iterations, self-collision). This clearly separates it from sim_up/sim_down and from sim_set_params, which it names as the tool for live gains. It stops short of 5 only because 'daemon' is unexplained jargon for the Kit process.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives the operative condition (change a spawn property) and contrasts it with the alternative for live values (sim_set_params), plus a note on why one restart beats one per experiment. There is no explicit statement about when NOT to reload versus sim_up/sim_down, so it is context rather than full routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
sim_screenshotA
Capture the viewport as an image. Requires the daemon started with gui=True.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It usefully discloses the gui=True prerequisite, which is a genuine behavioral constraint. But it omits return characteristics (image format, size, viewport vs full scene), latency/cost, and failure behavior when gui is disabled.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, zero filler, with the action stated first and the prerequisite immediately after. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple zero-parameter tool this is nearly complete — purpose plus the one gating precondition. The only gap is the absence of any indication of what the returned image looks like (format/dimensions), which matters since there is no output schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so there is nothing for the description to clarify; baseline for a parameterless tool is 4. The empty schema is self-consistent with the description implying a no-argument capture.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource: 'Capture the viewport as an image.' An agent immediately understands it returns a rendered screenshot of the simulation viewport, which is clearly distinct from siblings like sim_set_camera or sim_get_stats. It stops short of explicitly contrasting itself with any sibling, so 4 rather than 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a real precondition — the daemon must be started with gui=True — which tells the agent when this tool will actually work. However, it offers no guidance on when to prefer this over alternatives (e.g. sim_set_camera + capture workflow) or what to do if the daemon wasn't started with gui=True.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
sim_set_cameraB
Point the viewport camera, e.g. eye=[1.2, 1.2, 1.0] target=[0, 0, 0.3].
| Name | Required | Description | Default |
|---|---|---|---|
| eye | Yes | ||
| target | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden, and it discloses almost nothing: not whether this mutates persistent scene state or is a transient viewport-only change, not whether coordinates are world-space or robot-base-relative, and not whether a render must occur afterward. For a camera-control tool with zero annotation coverage this is a substantial gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence that states the action and immediately grounds it with a concrete example. Nothing is padded or repeated from the schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be described, and with only two required numeric-vector params the surface is small. However, missing coordinate-frame conventions and persistence semantics leave an agent guessing about how the values will be interpreted.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% — the schema gives only titles ('Eye', 'Target') with no semantic text. The example (eye=[1.2, 1.2, 1.0], target=[0, 0, 0.3]) does imply 3-element numeric vectors and that 'eye' is the camera position while 'target' is the look-at point, which partially compensates, but the coordinate frame and units are never stated.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb ('point') and resource ('viewport camera'), which cleanly separates it from siblings like sim_set_pose and sim_set_targets that manipulate robot/object targets rather than the viewport. It does not explicitly call out that distinction, so it falls short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no statement of when to use this tool, when not to, or which sibling to prefer (e.g., sim_screenshot or sim_set_pose). The inline example hints at typical values but conveys no usage conditions or prerequisites.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
sim_set_couplingA
Change software-coupling ratios live, keyed by follower joint. 0 releases a follower.
Couplings are declared in the robot config; a follower is driven to ratio x its leader's target.
| Name | Required | Description | Default |
|---|---|---|---|
| ratios | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, and it does add real behavioral meaning: values are keyed by follower joint, 0 releases a follower, and couplings must be declared in the robot config. It still omits error behavior, persistence across steps/reloads, and what happens to unspecified couplings, leaving meaningful gaps for a mutation tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the core action and immediately followed by the one non-obvious value rule ('0 releases a follower'). No filler; every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists so return values need no explanation, and the single parameter is well covered semantically. The main residual gap for an annotation-free mutation tool is failure/edge behavior and whether the change persists, which keeps it from a 5.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% and the parameter is an untyped key->number map, so the description must compensate. It does: the keys are follower joints, the values are the ratio multiplier applied to the leader's target, and 0 has the special meaning of releasing the follower. Value limits/units are not given, preventing a 5.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb+resource ('Change software-coupling ratios') and clarifies the mechanism ('a follower is driven to ratio x its leader's target'), which is distinguishable from siblings like sim_set_targets or sim_set_params. It never explicitly names or contrasts with those siblings, so it stops short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The word 'live' hints that this is a runtime change, but there is no explicit when-to-use guidance, no when-not, and no naming of alternatives such as sim_set_targets or sim_set_pose. Usage must be inferred from the coupling semantics alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
sim_set_paramsC
Change the driven joints' PD gains and effort limit on the live articulation.
| Name | Required | Description | Default |
|---|---|---|---|
| effort | No | ||
| damping | No | ||
| stiffness | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It discloses that the change applies to the 'live' articulation (implying runtime mutation), but says nothing about prerequisites (e.g. sim_up), reversibility, units, or what happens to parameters left null.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single tight sentence with no filler, front-loading the verb and the affected resource. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Output schema exists, so return values needn't be explained, but for an unannotated mutation tool with 0% parameter coverage the description is too thin. Units, null semantics, and the requirement that the articulation be live are all left unstated.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% and none of the three parameters (effort, damping, stiffness) are named or documented. 'PD gains' loosely maps to stiffness/damping and 'effort limit' to effort, but this is inference the description never makes explicit, and null/default behavior is unexplained.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (change), resource (PD gains and effort limit), and scope (driven joints on the live articulation). This distinguishes it from sim_set_targets and sim_set_pose, which mutate different properties. Lacks explicit naming of siblings but the resource is specific enough to differentiate.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use guidance, no prerequisites, and no alternatives named. The phrase 'on the live articulation' hints that the sim must be running, but the agent gets no routing help against sim_set_targets, sim_set_pose, or sim_set_coupling.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
sim_set_poseA
Command a named pose. amount blends toward it: 1 = as authored, 0.5 = halfway.
Halfway from the lower limits by default, or from the current targets with from_current=True. Follow with sim_step, then sim_screenshot to verify.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | ||
| amount | No | ||
| from_current | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does well: it explains the blending semantics of amount (1 = as authored, 0.5 = halfway) and the default reference frame (lower limits vs current targets). It omits whether the command is immediate, requires stepping, or affects persisted state.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences, front-loaded with the core action and then the parameter semantics and follow-up workflow. No wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need no explanation. The description covers the two non-obvious parameters and the verification workflow, leaving only minor gaps around the name argument and state-mutation effects.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate. It clearly defines amount's scale and from_current's two modes, but leaves the required name parameter's source (presumably sim_list_poses) unexplained.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
"Command a named pose" states a specific verb and resource, and the blending sentences clarify the effect. It is distinguishable from sibling list/mutation tools like sim_list_poses and sim_set_targets, though it never names those siblings explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives a clear operational workflow ("Follow with sim_step, then sim_screenshot to verify"), which tells the agent how to sequence this call. It stops short of naming alternatives or stating when NOT to use this tool versus sim_set_targets.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
sim_set_targetsA
Command the driven joints, normalized 0..1 (lower..upper limit).
Pass unit as a full vector in driven order (see sim_status), or joints
to set individual joints by name. Targets persist; advance with sim_step.
| Name | Required | Description | Default |
|---|---|---|---|
| unit | No | ||
| joints | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It usefully discloses that targets persist until sim_step advances the sim, but says nothing about permissions, error behavior when vectors are wrong length, or blocking characteristics. One genuine behavioral trait is disclosed, so a mid score is appropriate.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short, front-loaded sentences: the core action and normalization first, then the parameter routing, then the persistence/next-step note. No filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained, and the description covers argument modes, units, and the step coupling. For a two-optional-parameter command tool it is nearly complete, with only error/edge-case behavior left implicit.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% (bare array/object with no descriptions), so the description must compensate. It does so well: it explains that `unit` is a full vector in driven order and `joints` maps joint names to values, plus the 0..1-to-limit normalization semantics that appear nowhere in the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Command the driven joints') with the normalization scheme (0..1 mapped to lower..upper limit), so the agent understands this sets joint targets rather than poses or parameters. It does not explicitly differentiate itself from siblings like sim_set_pose or sim_set_params, which keeps it from a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear either/or guidance for the two parameters: pass `unit` as a full driven-order vector, or `joints` to address individual joints by name, and points to sim_status for the driven order. It also tells the agent the next step ('advance with sim_step'), but offers no explicit when-not or comparison against alternative setters.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
sim_statusA
Report whether the sim daemon is up, its robot, driven joints, poses and parameters.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. 'Report' implies a read-only operation, which is a minimal behavioral hint, but there is no explicit statement about safety, side effects, permissions, or behavior when the daemon is down.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no redundant or filler content. It lists the reported items efficiently.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The output schema exists, so return values need not be explained. However, with no annotations and many siblings, the description is minimally adequate: it states purpose but omits usage context and behavioral traits that would help an agent choose correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so the baseline is 4. The description correctly adds no parameter details because none are needed.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Report') and resource (sim daemon status, robot, driven joints, poses, parameters). The listed contents clearly distinguish it from narrow siblings like sim_get_joint_state, sim_list_poses, and sim_get_stats.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use guidance or alternatives are given. It only says what the tool reports, leaving the agent to infer that this is a general status check rather than a targeted query like sim_get_stats or sim_inspect_joints.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
sim_stepA
Advance the simulation n physics steps (dt from sim_status, default 1/120 s).
| Name | Required | Description | Default |
|---|---|---|---|
| n | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It usefully discloses that dt is taken from sim_status with a 1/120 s default, which is real behavioral context, but says nothing about blocking behavior, whether stepping mutates shared state, or performance for large n.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with the core action first and the timing detail in a parenthetical. Nothing is wasted and nothing is buried.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter, zero-required tool with an output schema, the description covers the action, the parameter, and the timestep source, which is enough to call it correctly. Only the when-to-use-vs-siblings gap remains.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, and it does: it defines n as the number of physics steps and clarifies that the timestep is not a parameter but is derived from sim_status. The only omission is the schema's default of 60 for n.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Advance the simulation') and quantifies the unit of work as 'n physics steps', so the agent knows exactly what it does. It does not differentiate itself from siblings like sim_play or sim_up, which is the only thing keeping it from a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies the use case (incremental stepping of the sim) but never says when to choose this over sim_play (continuous run) or sim_up/sim_down. No preconditions or exclusions are given, so usage must be inferred from the verb alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
sim_sweepA
Compare several values of one live parameter inside the single session.
param: stiffness, damping, effort, or "coupling:". For each value
it applies the parameter and runs test ("wave" or "range") for steps.
The original value is restored afterwards, even on error. A diverging
solver is recorded as a result, not raised.
| Name | Required | Description | Default |
|---|---|---|---|
| test | No | wave | |
| param | Yes | ||
| steps | No | ||
| values | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does well: it discloses that the original value is restored afterwards even on error, that a diverging solver is recorded as a result rather than raised, and that everything runs in a single session. It omits permission/auth or performance/timeout caveats, but the key state and error-handling behaviors are disclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose is front-loaded in the first sentence, followed by tightly packed behavioral and parameter detail. No filler sentences, though the enum lists make it read a touch list-like rather than prose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need no explanation. Given the 4-param schema at 0% coverage and no annotations, the description covers the essential behavior and parameter meaning adequately, though it never states the result shape for a sweep (per-value outcomes) which would help an agent interpret output.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, and it largely does: it enumerates the valid param choices (stiffness, damping, effort, or coupling:<follower>), the test options ("wave" or "range"), and explains steps. The values array is only implied as the set to sweep, leaving its semantics slightly thin.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource: comparing several values of one live parameter within a single session. The sweep scope functionally distinguishes it from sim_set_params (single set) and sim_wave/sim_range_test (single run). It does not explicitly name any sibling, so it stops short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The mechanics (apply each value, run the test, restore) imply the use case of parameter comparison, but there is no explicit when-to-use, when-not, or routing to alternatives like sim_set_params or sim_range_test. Usage is implied rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
sim_upA
Start the daemon container and wait until it answers. Idempotent.
robot: a robot config JSON, path relative to the repo root
(e.g. examples/robots/franka.json). usd: a USD path instead of / overriding
the config's (every joint driven if no config). gui=True is needed for
sim_screenshot; the host must have run xhost +local:docker once.
ground=False spawns without a floor (use for range tests).
extra_args go to scripts/simd.py verbatim (e.g. ["--pos-iters", "64"]).
The first start downloads Nucleus assets and builds shader caches: minutes.
| Name | Required | Description | Default |
|---|---|---|---|
| gui | No | ||
| usd | No | ||
| robot | No | ||
| ground | No | ||
| timeout | No | ||
| extra_args | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does well: idempotency, a blocking wait, a host prerequisite (xhost +local:docker), and a significant latency warning on first start (Nucleus asset download and shader cache builds take minutes). Missing only what happens on failure/timeout.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded purpose sentence, then a compact per-parameter block with zero filler. Slightly terse formatting, but every line adds information an agent needs.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Output schema exists so return values need not be described. The description covers prerequisites, latency, idempotency, and most parameter semantics for a 6-param tool, leaving only the timeout parameter's meaning unaddressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate and largely does: it explains robot (repo-relative path with example), usd (overrides config), gui (screenshot dependency), ground (floor removal for range tests), and extra_args (passed verbatim to scripts/simd.py with example). Only timeout is left unexplained.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Start the daemon container and wait until it answers'), plus a scoping trait ('Idempotent'). Combined with the sim_* sibling family, an agent can distinguish this from sim_down/sim_reload without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives conditional guidance tied to other tools: gui=True is required for sim_screenshot, and ground=False is for range tests. It doesn't explicitly name sim_down as the inverse or state when NOT to call it (e.g. already running), but the operational context is clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
sim_waveB
Sweep every driven joint through its range for n steps; return travel stats.
| Name | Required | Description | Default |
|---|---|---|---|
| n | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. It states that joints are swept and stats are returned, but does not clarify whether simulation state is modified, whether the motion is reversible, or what permissions or constraints apply.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, front-loaded sentence with no redundant or filler content. It delivers the core action and scope immediately.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained further. However, with no annotations and several similar siblings, the description leaves key context unaddressed: side effects on simulation state and why an agent should pick this over sim_sweep or sim_range_test.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema provides no description for the sole parameter 'n', so the description must compensate. It adds meaning by saying 'for n steps,' but gives no range, format, or constraints beyond the default already in the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Sweep') and a clear resource scope ('every driven joint through its range for n steps'). This distinguishes it from single-joint or step tools, but it does not explicitly differentiate from close siblings like sim_sweep or sim_range_test.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is given on when to choose this over sim_sweep, sim_range_test, or sim_step. The implied usage is clear from the verb, but there are no conditions or alternatives mentioned.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
train_checkpointsB
List a run's saved checkpoints with step and size.
| Name | Required | Description | Default |
|---|---|---|---|
| run_name | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the burden, but the 'List' verb clearly signals a non-mutating read operation, which is the key behavioral trait here. It adds nothing about pagination, ordering, or whether checkpoints can be deleted or restored, so it is only minimally transparent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One short sentence, front-loaded with the verb and resource, with no filler. Nothing can be trimmed without losing information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return shape need not be explained, and this is a simple one-parameter read. However, the description never links run_name to the run identifiers produced by sibling tools like train_list, leaving a small but real gap for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The single parameter run_name has 0% schema description coverage and the description adds no meaning beyond the schema title 'Run Name'. It does not say whether the value must match a run returned by train_list or what format is expected.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (List) and resource (a run's saved checkpoints) and even previews the returned fields (step and size). It is distinguishable from siblings like train_list or train_logs by resource, though the description never explicitly contrasts them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use guidance and no alternatives named. An agent must infer from the name alone that this is for inspecting a run's checkpoint artifacts rather than listing runs (train_list) or reading logs (train_logs).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
train_listA
List training runs: live or recent containers, and log directories holding checkpoints.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It does disclose the sources it enumerates (live or recent containers plus checkpoint log directories), which tells the agent this is a broad discovery call rather than a filtered one, but it says nothing about read-only safety, staleness/recency semantics of 'recent', or whether it touches remote hosts.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with the verb and resource first, followed by a compact colon-style enumeration of what is listed. No filler, no restatement of the title.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return-value explanation is unnecessary, and with zero parameters the schema side is fully covered. The description is adequate for a simple enumeration tool, though it leaves the boundary against train_status and train_checkpoints undefined.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so there is nothing for the description to clarify; baseline 4 applies. The description correctly avoids inventing filter semantics that do not exist in the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('List') and resource ('training runs') and adds scope detail about what a run comprises (live/recent containers, log directories holding checkpoints). It is clear on its own, but it does not differentiate itself from siblings like train_status or train_checkpoints, and the mention of 'log directories holding checkpoints' partially overlaps with train_checkpoints.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance, no exclusions, and no named alternative. An agent must infer whether to call train_list vs train_status vs train_checkpoints, especially given the checkpoint overlap. Nothing tells the agent under what conditions this is the right call.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
train_logsA
Tail a running training container's stdout/stderr.
| Name | Required | Description | Default |
|---|---|---|---|
| tail | No | ||
| run_name | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. 'Tail' and 'stdout/stderr' disclose that this is a read of log output and hint at streaming/last-N-lines behavior, but it never states whether the call blocks or streams, whether it is read-only, or how it behaves for a stopped container.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single tight sentence with the key constraint ('running') front-loaded. No filler, no redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be described. The description covers the core action but leaves gaps for a log-retrieval tool: streaming vs snapshot behavior, handling of a non-running container, and the meaning of the tail count are all unstated.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate. The verb 'tail' loosely maps to the integer 'tail' parameter (default 200, implying number of lines), and 'training container' implies run_name selects the target, but neither parameter is explicitly explained in terms of format or semantics.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb (tail) and resource (a running training container's stdout/stderr), which is far more informative than the bare name train_logs. It implicitly distinguishes itself from siblings like train_metrics and train_status by scoping to raw container output, though it never explicitly contrasts them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'a running training container' implies the tool is only applicable while training is active, which is useful context. However, there is no explicit guidance on when to use this versus train_metrics or train_status, and no statement of what happens if the run has finished.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
train_metricsA
TensorBoard scalars for a run: no tag lists tag names; a tag returns its series, downsampled.
| Name | Required | Description | Default |
|---|---|---|---|
| tag | No | ||
| run_name | Yes | ||
| max_points | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It usefully discloses the dual-mode behavior and that the returned series is downsampled, which is real behavioral information. It omits any note on auth/permissions, error behavior, or the effect of max_points on downsampling, leaving notable gaps for an annotation-free tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler; the resource is stated first and the mode logic follows. Efficient and readable, though the colon-separated clauses are slightly dense.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be described, and the key tag-mode behavior is covered. run_name is self-evident and max_points has a schema default. The description is nearly complete, missing only explicit max_points semantics.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It explains the tag parameter's two modes well (null lists names, a value returns the series), but says nothing explicit about run_name or max_points beyond the passing hint 'downsampled'. Partial compensation, not full.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific resource and scope: 'TensorBoard scalars for a run', with a clear verb-implied retrieval action. It is not a tautology and the two operating modes (tag absent vs present) are stated. It does not, however, explicitly differentiate itself from siblings like train_logs, so it falls short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'no tag lists tag names; a tag returns its series' tells the agent how to switch between the two modes, which is genuinely useful selection guidance within the tool. But there is no explicit when-to-use guidance versus alternatives such as train_logs or train_checkpoints, leaving sibling routing to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
train_startA
Launch a headless Isaac Lab training run in its own container; returns at once.
task is a registered gym id (e.g. Isaac-Cartpole-v0). run_name defaults to '_'; keep it to find the run later. extra_args go to the training script verbatim (e.g. Hydra overrides). device pins one GPU index. Independent of the daemon and of this MCP session.
| Name | Required | Description | Default |
|---|---|---|---|
| seed | No | ||
| task | Yes | ||
| device | No | ||
| num_envs | No | ||
| run_name | No | ||
| checkpoint | No | ||
| extra_args | No | ||
| max_iterations | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does add real behavioral context: it is asynchronous ('returns at once'), container-isolated, and explicitly independent of the daemon and the MCP session. It still omits failure behavior, resource/GPU requirements, and any concurrency limits.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded purpose followed by tight per-parameter notes; each sentence adds information. The trailing 'Independent of the daemon...' line is slightly disconnected but still earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists so return values need not be explained, and the description covers the async launch semantics and the most ambiguous parameters. The gaps are the four undocumented params and the absence of any sibling routing, which for an 8-param async launcher is a modest shortfall rather than a serious one.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% across 8 params, so the description must compensate and it only does so for half: task (gym id + example), run_name (default pattern + retention purpose), extra_args (verbatim + Hydra example), device (GPU index). seed, num_envs, checkpoint, and max_iterations are left to their self-explanatory names.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Launch a headless Isaac Lab training run in its own container') with the key scope fact that it returns at once. This is clearly distinguishable from the sim_* control tools and from the train_status/train_logs readers.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by 'keep it [run_name] to find the run later', which gestures at train_status/train_logs, but no sibling is named and there is no explicit when-to-use or when-not-to-use guidance. An agent can infer intent but is not routed.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
train_statusB
Container state, checkpoints, and latest TensorBoard scalar values for one run.
| Name | Required | Description | Default |
|---|---|---|---|
| run_name | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full disclosure burden. It does add one behavioral trait — only the *latest* scalar values are returned, implying history lives elsewhere — but says nothing about permissions, whether the tool blocks on a running container, or how a missing run is handled.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler; the most important information (what comes back, for how many runs) leads the sentence.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be spelled out. However, with five overlapping train_* siblings and no annotations, the description should clarify how this differs from train_metrics, train_checkpoints, and train_logs — that routing information is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% and the single parameter run_name has no schema description. The phrase 'for one run' hints that run_name selects a single run, which is weak but non-zero compensation for the undocumented parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific resource and its payload (container state, checkpoints, latest TensorBoard scalars) scoped to one run, so an agent knows it is a read/status tool. It lacks an explicit verb and does not distinguish itself from the overlapping siblings train_checkpoints, train_metrics, and train_logs, which cover much of the same ground.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance, no prerequisites, and no mention of any alternative. Given five sibling train_* tools with overlapping semantics, the absence of routing guidance is a real gap.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
train_stopC
Stop a training run's container.
| Name | Required | Description | Default |
|---|---|---|---|
| run_name | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. It says the tool stops a container but does not describe whether the stop is graceful or forced, whether it is irreversible, what happens to the training run's state, or any authentication or side-effect considerations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence with zero wasted words. It states the action and target immediately, which is appropriately concise for a simple tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Although an output schema exists and return values need not be explained, the description is incomplete for a mutation tool with no annotations. It omits usage context, behavioral implications, and parameter details, leaving significant gaps for an agent to invoke the tool with confidence.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the single parameter, run_name, is undocumented in both the schema and the description. The description implies that the tool operates on a training run, but it adds no format, syntax, or meaning beyond the parameter's name.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb, 'Stop', and a specific resource, 'a training run's container', which clearly distinguishes it from sibling tools like train_start and train_status. However, it does not explicitly differentiate from those siblings by naming them or describing conditions, so it falls short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives such as train_start or train_list. It implies usage but offers no context, prerequisites, or exclusions, leaving the agent to infer when stopping is appropriate.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
26 tool updates
v0.1.0- First observed
sim_down - First observed
sim_get_joint_state - First observed
sim_get_stats - First observed
sim_inspect_joints - First observed
sim_list_poses - First observed
sim_play - First observed
sim_range_test - First observed
sim_reload - First observed
sim_screenshot - First observed
sim_set_camera - First observed
sim_set_coupling - First observed
sim_set_params - First observed
sim_set_pose - First observed
sim_set_targets - First observed
sim_status - First observed
sim_step - First observed
sim_sweep - First observed
sim_up - First observed
sim_wave - First observed
train_checkpoints - First observed
train_list - First observed
train_logs - First observed
train_metrics - First observed
train_start - First observed
train_status - First observed
train_stop
TDQS
Scored across 26 tools
Tools are cleanly partitioned into sim_* and train_* domains, and within sim the verbs target distinct actions (status, step, play, set_targets vs set_pose vs set_coupling). Minor overlap between sim_wave, sim_range_test, and sim_sweep is resolved by their descriptions.
All tool names use snake_case with consistent sim_ or train_ prefixes and predictable verb_noun structure, e.g. sim_set_targets, sim_get_joint_state, train_start, train_checkpoints.
26 tools is slightly above the ideal single-domain range, but the server covers two substantial subsystems: live Isaac Sim control and headless Isaac Lab training orchestration. The count is reasonable for that dual scope.
The surface covers sim lifecycle (up/down/reload, status, joint state, poses, params, camera, screenshot, experiments) and training lifecycle (start/list/status/logs/stop/checkpoints/metrics). Minor gaps include no explicit generic sim reset and no training-run cleanup/delete operation.
Maintenance
Related MCP Connectors
Drive real devices from your AI Coding tool. Embed a client SDK (Unity, Godot, Flutter, iOS/macOS, Android, React Native, Web) in your app, then capture screenshots, traverse the UI tree, inject taps and key events, and run automated test tasks on the physical device over a secure relay.
Control Unreal Engine to browse assets, import content, and manage levels and sequences. Automate…
LLM chat, text tools, image generation, editing, batch image jobs, and asynchronous video generation
- naturaliOAuthai.naturali
Configure AI agents, give them knowledge and tools, and read back every generation they run.
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceEnables AI tools like Windsurf and Claude to control NVIDIA Isaac Sim and Isaac Lab through natural language, providing tools for scene inspection, prim management, physics simulation, and robot spawning.8Apache 2.0
- AlicenseAqualityBmaintenanceEnables driving Omniverse Kit apps (Isaac Sim, Isaac Lab) over MCP, allowing agents to control simulations, run Python, and call namespace-scoped tools via a single bridge.12MIT
- AlicenseCqualityBmaintenanceEnables AI assistants to control NVIDIA Isaac Sim by building scenes, loading assets, operating robots and humans, reading sensors, managing simulation, and creating Action Graphs.1291MIT
- AlicenseAqualityCmaintenanceEnables an AI agent to drive Roblox Studio, including running Luau code, syncing Rojo projects, starting solo or multiplayer play tests, and retrieving logs, screenshots, and GUI information.11MIT