Skip to main content
Glama
Milokucia

dex-isaac-mcp

by Milokucia

dex-isaac-mcp

An agent driving the Allegro hand through dex-isaac-mcp

Recorded headless with the server's own tools (sim_frame_robot, sim_set_backdrop, sim_record_start), driving the Allegro hand example. Each caption is the request and the tool call it became.

An MCP server that lets an AI agent (Claude Code, or any MCP client) drive a live, persistent Isaac Sim session and launch Isaac Lab training runs.

Kit takes tens of seconds to boot. If every experiment is a fresh launch, most of your time goes to waiting. Here Kit starts once inside a daemon and stays up. Each tool call lands in that running session, so changing a gain, stepping physics or taking a screenshot costs a frame, not a relaunch.

 MCP client ──stdio──▶ dex_isaac_mcp (host, plain python)
                          │  newline-delimited JSON over a Unix socket
                          ▼
                       scripts/simd.py (Isaac Lab container, Kit stays up)
                          └─ one articulation, described by a robot JSON

How it differs from other Isaac Sim MCP servers: those mostly build scenes from language ("add a table and a Franka"). This one is a workbench for an agent that tunes and debugs a robot: it measures (range tests, tracking error, parameter sweeps that restore the original value) and runs training. Scene props are supported, but a scene builder is not the point.

  • Any articulated robot. Point it at a USD and a small JSON config. Franka and Allegro examples are included.

  • Normalized joint control. Targets are 0..1, where 0 is a joint's lower limit and 1 its upper limit, so agents don't need to know radians or meters.

  • Named poses per robot (home, fist, …), which can be blended part-way.

  • Headless capture and GIF recording from an auto-framed camera, with captions. The clip above was made this way.

  • Props: spawn a table, a ball or a USD into the running scene and read back where they settle.

  • Measurements: joint state, per-joint travel and tracking error, a range test that finds blocked joints, and parameter sweeps inside one session.

  • Training control: each run is a detached docker compose run. You can poll its status, TensorBoard scalars, checkpoints and logs.

Requirements

  • Linux with an NVIDIA GPU that Isaac Sim supports, plus the NVIDIA Container Toolkit

  • Docker with the Compose plugin

  • An NGC login to pull the Isaac Lab image: docker login nvcr.io (user $oauthtoken, password: your NGC API key)

  • Python ≥ 3.10 on the host, for the MCP server only

Related MCP server: omni-kit-mcp

Quick start

git clone <this repo> dex-isaac-mcp && cd dex-isaac-mcp

# 1. Build the image (Isaac Lab 2.3.2 base, pinned by digest)
cd docker && docker compose build && cd ..

# 2. Install the host-side server, editable so it finds this clone
pip install -e .            # add [metrics] for train_metrics: pip install -e '.[metrics]'

# 3. Register it with Claude Code
claude mcp add isaac -- python -m dex_isaac_mcp

# 4. Allow the container to open windows (once per login; needed for screenshots)
xhost +local:docker

Then ask the agent something like "start the sim with the Franka, move it to ready and show me a screenshot." It will call sim_up, sim_set_pose, sim_step and sim_screenshot.

From PyPI instead: pip install dex-isaac-mcp. The daemon still runs from a clone (it needs docker/ and scripts/simd.py), so point the server at it:

claude mcp add isaac -e ISAAC_MCP_HOME=/abs/path/dex-isaac-mcp -- dex-isaac-mcp

Other MCP clients can launch python -m dex_isaac_mcp (or the dex-isaac-mcp script) over stdio.

The first sim_up takes several minutes: Kit builds its shader cache and downloads Nucleus assets. Later starts are much faster.

Running the daemon by hand

sim_up deletes its container (--rm) when the daemon exits, so a crash on startup takes its traceback with it. To see the error, run the daemon in the foreground:

cd docker
docker compose run --rm simd scripts/simd.py --gui --robot examples/robots/franka.json
docker compose run --rm simd scripts/simd.py --headless --usd /path/in/container/robot.usd
docker compose run --rm simd scripts/simd.py --sliders --robot examples/robots/allegro_hand.json

--sliders opens an omni.ui panel with one slider per driven joint. The socket stays live alongside it.

Tools

Session

Tool

What it does

sim_status

Whether the daemon is up, plus its robot, driven joints, poses, couplings and gains

sim_up

Start the daemon container (robot, usd, gui, ground, extra_args) and wait until it answers. Does nothing if it is already up

sim_down

Stop the daemon

sim_reload

Restart with a different spawn property: robot, USD, pos_iters, self-collision

Inspect

Tool

What it does

sim_inspect_joints

A USD's articulation DOFs, its loop-closure joints (excluded from the articulation) and its articulation roots. Reads the file, so it shows edits saved from the GUI

sim_get_joint_state

Positions and velocities, plus driven joints' normalized positions and limits

sim_get_stats

Per-joint travel and mean tracking error since the last reset

sim_screenshot

Viewport capture, returned as an image. Needs gui=True

sim_set_camera

Point the viewport camera

Drive

Tool

What it does

sim_set_targets

Normalized targets, as a full vector or {joint: value}

sim_list_poses / sim_set_pose

Named poses from the robot config. amount blends toward a pose, starting from the lower limits or from the current targets

sim_step

Advance N physics steps (default dt 1/120 s)

sim_play

Run continuously, or pause

sim_wave

Sweep every driven joint through its range and return travel stats

sim_range_test

Drive every joint from its lower limit toward a target and report the fraction of travel reached. Below 0.9 counts as blocked (self-collision, a binding linkage, too little effort)

Scene

Tool

What it does

sim_spawn_object

Add a prop to the live scene: cuboid, sphere, cylinder, capsule, cone or a USD file, with collision. static for a fixed table or wall, kinematic for a body contact cannot move

sim_list_objects

Every prop's current world pose

sim_remove_object

Delete a prop

Capture and record

These work headless, with no GUI or viewport. They use a dedicated camera that is independent of the GUI view.

Tool

What it does

sim_frame_robot

Aim the capture camera so the whole robot fills the frame, from a given direction. raise_frac leaves room for captions

sim_set_capture_camera

Place the capture camera by hand

sim_set_backdrop

A plain colored panel behind the robot. Use it with ground=False for clean footage

sim_capture

One frame, returned as an image

sim_record_start / sim_record_caption / sim_record_stop

Record every Nth physics step, with a caption drawn on each frame, to an animated GIF under .cache/recordings/. Anything that steps the sim is recorded

Tune

Tool

What it does

sim_set_params

Live stiffness / damping / effort on the driven joints

sim_set_coupling

Live software-mimic ratios, keyed by follower joint

sim_sweep

Try several values of one live parameter (stiffness, damping, effort, coupling:<follower>), running wave or range for each. Restores the original value afterwards. A diverging solver is recorded as a result rather than raised

Train

Tool

What it does

train_start

Launch a headless training run in its own container and return at once. extra_args go to the script verbatim (e.g. Hydra overrides). device pins a GPU

train_list

Running and recent runs, and log directories holding checkpoints

train_status

Container state, checkpoints and latest scalars

train_logs

Tail a running container's output

train_metrics

List TensorBoard tags, or get a downsampled series for one tag

train_checkpoints

Checkpoints with step and size

train_stop

Stop a run

Training runs do not depend on the daemon or on the MCP session. They keep going after the client disconnects.

Robot config

A robot is one JSON file. Only usd is required. Unknown keys are rejected, so a typo fails loudly instead of quietly falling back to a default.

{
  "name": "franka",
  "usd": "{ISAACLAB_NUCLEUS_DIR}/Robots/FrankaEmika/panda_instanceable.usd",
  "fix_root_link": true,
  "spawn_pos": [0, 0, 0],
  "init_joint_pos": {"panda_joint4": -2.81, "panda_joint6": 3.04, ".*": 0.0},
  "driven_joints": ["panda_joint[1-7]", "panda_finger_joint.*"],
  "passive_joints": [],
  "couplings": [{"leader": "joint_a", "follower": "joint_b", "ratio": 1.0}],
  "actuator": {"stiffness": 400, "damping": 40, "effort": 87, "velocity": null},
  "solver": {"pos_iters": 32, "vel_iters": 4, "self_collisions": true},
  "camera": {"eye": [1.8, 1.8, 1.4], "target": [0, 0, 0.4]},
  "poses": {"ready": {"panda_joint[1357]": 0.5, "panda_finger_joint.*": 1.0}}
}

Key

Meaning

usd

Local path (relative paths resolve from the JSON's own directory), a URL, or a path using {ISAAC_NUCLEUS_DIR} / {ISAACLAB_NUCLEUS_DIR}

init_joint_pos

Spawn pose in the joints' own units (rad / m), keyed by regex. It must lie inside every joint's limits or spawning fails. The default is all zeros, which is out of range for e.g. Franka's joint 4

driven_joints

Regexes (full match) for joints that take commands. Default .*

passive_joints

Joints whose angle is owned by a constraint, such as a closed-chain linkage. They get a zero-stiffness drive, because a live PD drive fights the constraint and the mechanism jitters

couplings

Software mimic joints: follower target = ratio × leader target. They are applied as drive targets rather than PhysX mimic constraints, so a large ratio cannot blow up the solver

actuator

Implicit PD gains and effort/velocity limits for the driven joints. Live-tunable

solver

Spawn properties. Changing them needs sim_reload

poses

{name: {joint_regex: 0..1}}. Later patterns win, so {".*": 0, "thumb.*": 1} works. A pattern that matches no driven joint is an error

Using your own robot without forking

Keep the robot config and assets in your own repo, and add them to the container with a compose override that you list in COMPOSE_FILE. Set it in the MCP server's environment, using absolute paths:

# my-robot/mcp-compose.yaml
services:
  simd:
    volumes:
      - /abs/path/my-robot:/workspace/my-robot
claude mcp add isaac \
  -e COMPOSE_FILE=/abs/path/dex-isaac-mcp/docker/docker-compose.yaml:/abs/path/my-robot/mcp-compose.yaml \
  -e ISAAC_MCP_ROBOT=/workspace/my-robot/robot.json \
  -- python -m dex_isaac_mcp

Training defaults

Out of the box, train_start runs Isaac Lab's stock skrl script inside the isaac-lab service. It tags the run name onto the log directory (logs/skrl/<experiment>/<timestamp>_ppo_torch_<run_name>/), which the other train_* tools use to find the run. To use your own launcher, set these in the environment the MCP server starts in:

Variable

Default

ISAAC_MCP_HOME

the clone this package was installed from (editable install); required for a PyPI install

ISAAC_MCP_COMPOSE_DIR

<repo>/docker

ISAAC_MCP_TRAIN_SERVICE

isaac-lab

ISAAC_MCP_TRAIN_WORKDIR

/workspace/isaaclab

ISAAC_MCP_TRAIN_SCRIPT

scripts/reinforcement_learning/skrl/train.py

ISAAC_MCP_LOGS_DIR

<repo>/logs (mounted at /workspace/isaaclab/logs)

ISAAC_MCP_RUN_NAME_ARG

agent.agent.experiment.experiment_name={run_name} (empty = don't pass one)

ISAAC_MCP_TRAIN_ARGS

hydra.run.dir=/tmp/hydra hydra.output_subdir=null: appended to every run. The stock script otherwise writes Hydra's outputs/ into the root-owned /workspace/isaaclab and dies. Set it empty for a non-Hydra script

ISAAC_MCP_SIMD_SERVICE

simd

ISAAC_MCP_ROBOT

robot config sim_up loads when none is given (unset: Franka example)

ISAAC_MCP_SOCKET

<repo>/.cache/simd.sock

The script must accept --task, --headless and, when given, --num_envs, --seed, --max_iterations and --checkpoint. Tasks from your own extension need to be importable inside the container, either installed into the image or mounted.

Design notes

These are the constraints the code is built around. Most were learned by breaking them.

  • Every Kit call happens on the main thread. Kit, PhysX and USD are not thread-safe. Socket threads only parse JSON and queue requests, and the main loop executes them between physics steps. Answering from a reader thread appears to work, then corrupts the stage under load.

  • Spawn properties are frozen. Replacing a spawned articulation needs SimulationContext.stop(), which blocks on a timeline event that only advances while the Kit loop pumps. A command runs on that loop, so the call never returns. omni.usd new_stage() has the same trap. So the USD, solver iterations and self-collision need a restart (sim_reload), and gains stay live.

  • A Unix socket, not TCP. The repo is bind-mounted and the container runs as the host uid, so the host sees the socket file directly, with no port mapping. Paths are capped at 107 bytes (AF_UNIX). If your checkout is deep, set ISAAC_MCP_SOCKET.

  • The host side imports no Isaac code. protocol.py, robot.py and training.py are stdlib-only. The MCP server adds only mcp. Nothing on the host needs isaaclab, torch or a GPU.

  • Cache directories are committed with .gitkeep. If Docker auto-creates a bind-mount source, it is root-owned, and Kit then dies with registry cache path is not set before any script runs.

  • The base image is pinned by digest. A re-pulled tag once shipped /isaac-sim as mode 750, and every non-root container lost its Python.

  • Recorded GIFs are stabilized. The renderer's denoiser shimmers: between two frames of a motionless scene, about 9% of background pixels change slightly, and a GIF re-encodes every one of them. Holding sub-threshold changes and using one shared palette took a 9-second clip from 15 MB to 1.2 MB.

  • The daemon always renders, even headless (enable_cameras). Without rendering, PhysX never registers a prop spawned at runtime. Prop poses are read from fabric, because the USD transform and the PhysX CPU query both stay at the spawn pose, and creating a PhysX tensor view mid-simulation crashes CUDA.

  • No floor for range tests (ground=False) on anything whose links can reach the ground. Otherwise the test measures the floor, not the robot.

Development

python -m unittest discover tests   # host-side tests: no Isaac, no GPU
ruff check .

dex_isaac_mcp/protocol.Client is a handy debugging client:

from dex_isaac_mcp.protocol import Client
with Client() as c:
    print(c.call("status"))
    c.call("set_pose", name="ready"); c.call("step", n=240)

Status

Tested against Isaac Lab 2.3.2 (Isaac Sim 5.x) and mcp 2.3, over the stdio protocol, with the GUI on:

  • Franka and Allegro examples, plus a custom closed-linkage hand through a compose override: sim_up, poses, screenshots, sim_range_test, sim_sweep (restores the original value), sim_down.

  • Headless daemon: gains, frozen-parameter rejection, wave.

  • Headless capture and recording, Allegro and Franka: auto-framing, backdrop, captions, GIF output (the clip at the top).

  • Props, GUI and headless: a sphere and a cylinder dropped onto a static table settle at exactly table height plus their radius and half-height.

  • Training, against Isaac Lab's stock skrl script: Isaac-Cartpole-v0 launched, polled, logged, checkpointed and read back through every train_* tool, plus a run stopped mid-training. skrl's write_interval: auto writes no TensorBoard scalars on a very short run (5 iterations), so train_metrics comes back empty there; 50 iterations gives 18 tags.

Issues and PRs are welcome.

License

MIT, see LICENSE.

Available Tools

26 tools
sim_downA

Stop the sim daemon; its container removes itself.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden. It does disclose one genuinely non-obvious trait — the container removes itself, i.e. this is irreversible teardown — but says nothing about in-flight state, permissions, or what happens to any running training/sim data.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One clause plus one semicolon-linked consequence, front-loaded with the action. No padding, no restatement of the name.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained, and a zero-param tool needs no argument documentation. The only shortfall is that the shutdown's effect on existing sim state is left implicit.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so the baseline is 4; there is nothing for the description to clarify beyond the empty schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Stop the sim daemon'), which cleanly contrasts with the sibling sim_up. It does not explicitly differentiate from sim_reload or sim_status, so it falls short of full sibling disambiguation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied by the action itself: an agent can infer this is the shutdown counterpart to sim_up. There is no explicit when-to-use guidance, no stated preconditions, and no routing between sim_down and sim_reload.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

sim_get_joint_stateB

Joint positions (rad/m) and velocities, plus driven joints' normalized positions and limits.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full disclosure burden. It implies a read ('get') and lists the returned quantities, but says nothing about side effects, whether it blocks until a simulation step, or whether it requires a running/paused sim — and much of what it does say overlaps with the output schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single dense sentence that front-loads the primary return values (positions, velocities) and appends the secondary driven-joint data. No wasted words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, the description need not explain return values, and for a zero-argument getter this is largely sufficient. The one real gap is absence of any routing hint versus sim_inspect_joints.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so the baseline is 4. The description correctly implies a global, unfiltered read of all joint state with no arguments needed.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific resource and enumerates exactly what state is returned: joint positions in rad/m, velocities, and driven joints' normalized positions and limits. It is clear what the tool yields. It does not, however, distinguish itself from the sibling sim_inspect_joints, leaving the reader unsure which to pick.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no prerequisites, and no mention of the closely related sim_inspect_joints or sim_get_stats siblings. The description only describes the payload, so the agent must infer call context on its own.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

sim_get_statsA

Per-joint travel and driven-joint mean tracking error since the last reset.

Travel on a passive joint proves a linkage transmits: with zero stiffness it moves only if its constraint moves it.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations, so the description carries the full burden. It discloses that values are accumulated 'since the last reset,' which is critical behavioral context for interpreting the stats. It does not mention reset behavior itself or return format, but the output schema likely covers the latter.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core purpose in one efficient sentence. The second paragraph is somewhat tangential—useful as a domain hint but not strictly necessary for invocation; still, it's brief.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no parameters, a rich output schema, and no annotations, the description provides the essential scope ('since the last reset') and domain context for interpretation. It is nearly complete, missing only explicit routing to sibling tools.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Zero parameters, so baseline is 4. The description appropriately avoids discussing nonexistent parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Clearly states the resource (per-joint travel and driven-joint mean tracking error) and the scope (since last reset), which distinguishes it from siblings like sim_get_joint_state by returning accumulated statistics rather than instantaneous state. However, it doesn't explicitly contrast with sim_status or sim_inspect_joints.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool versus sim_get_joint_state, sim_status, or others. The second paragraph explains a concept (travel on a passive joint) but offers no operational context or exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

sim_inspect_jointsA

List a USD's articulation DOFs, joints excluded from the articulation (loop closures), and roots.

Reads the file, not the live scene, so it shows edits saved from the Isaac GUI. Defaults to the loaded USD.

ParametersJSON Schema
NameRequiredDescriptionDefault
usdNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations, so the description carries the burden, and it discloses the critical trait that results come from the file rather than the live scene and default to the loaded USD. It does not cover error behavior or the usd format, but the key data-source caveat is clearly surfaced.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with what is listed, and every clause (loop closures, file-vs-live, default) earns its place. Slightly dense but no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return values need not be described, and for a single-optional-parameter read tool the description covers scope, source, and default. Only the usd parameter format is left unspecified.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The single usd parameter has 0% schema coverage, so the description must compensate; it only explains the default ('loaded USD') and gives no hint of expected format (path, name, prim). Partial compensation over an otherwise undocumented parameter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Specific verb (List) plus precise resources: articulation DOFs, loop-closure joints, and roots. The 'Reads the file, not the live scene' clause distinguishes it from the live-state sibling sim_get_joint_state, so an agent can tell them apart without opening schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The file-vs-live-scene note implies when this tool is appropriate (inspecting saved GUI edits) but never explicitly states when to use it instead of sim_get_joint_state or names that alternative. Usage must be inferred rather than read.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

sim_list_posesB

Named poses from the robot config, as {pose: {joint_pattern: 0..1}}.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It discloses the return shape ({pose: {joint_pattern: 0..1}}), which is useful behavioral context, but it never confirms this is a read-only operation with no side effects — an inference an agent can reasonably make but that is not stated.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single compact sentence with no wasted words, front-loading the resource and source. It is a fragment rather than a sentence, which slightly weakens readability but costs no clarity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained, and with zero parameters the input side is trivially covered. The main omission is guidance on when this list is the right call versus the other inspection tools.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so the baseline of 4 applies. The description correctly adds nothing about arguments because there are none.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description identifies a specific resource (named poses) and its source (the robot config), which cleanly separates it from siblings like sim_set_pose and sim_inspect_joints. It omits an explicit verb such as 'list', leaving that to be inferred from the tool name.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no statement of when to call this tool versus sim_inspect_joints, sim_get_joint_state, or sim_set_pose, and no prerequisites or exclusions. The agent must infer usage entirely from the name and description.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

sim_playA

Run the sim continuously in real time (True) or pause it (False).

ParametersJSON Schema
NameRequiredDescriptionDefault
playingNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It discloses the key semantic of the mutation (continuous real-time execution vs. paused), but says nothing about prerequisites (e.g., sim must be up via sim_up), idempotency, or state side effects of toggling playback.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One sentence, front-loaded with the action, and the boolean mapping is given inline with zero filler. Every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need no explanation, and the tool is a simple binary toggle. The description covers purpose and the parameter fully; the only gap is the unstated relationship to siblings like sim_step and sim_status and any required preconditions.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the sole boolean parameter 'playing' has no schema documentation. The description compensates by fully defining the semantics of both values (True = run continuously in real time, False = pause), which is exactly what is needed for a single-parameter tool.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific action (run continuously in real time / pause) on a named resource (the sim), so an agent knows exactly what the tool does. It does not explicitly distinguish itself from the closely related sibling sim_step, which is the main thing keeping it from a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The parenthetical explains what each argument value does, which implies when to call it (to resume or pause continuous simulation). However, there is no explicit when-to-use framing and no mention of the alternative sim_step for stepping the simulation instead of running it continuously.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

sim_range_testA

Drive every driven joint from its lower limit toward target, report the fraction reached.

A joint below 0.9 is blocked: self-collision, a binding linkage, or too little effort for the load. reset=True first puts every joint at its lower limit and verifies it. Start the daemon with ground=False if the robot can reach the floor, or the test measures the floor.

ParametersJSON Schema
NameRequiredDescriptionDefault
resetNo
stepsNo
targetNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden, and it does disclose meaningful behavior: the drive direction, the 0.9 blocked threshold with its three possible causes, and what reset does to joint state. It stops short of saying what state the robot is left in after the test or what effort/permissions the daemon needs.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four short sentences with the core action front-loaded, followed by interpretation, then configuration caveats. Dense but every sentence carries information; no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return format need not be spelled out, and the description still explains how to read the reported fraction. The main gaps are the unexplained `steps` parameter and the absence of any statement about the robot's post-test state, which matters for a tool that physically drives joints.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It explains `target` (drive toward it) and `reset` (re-home and verify first), but `steps` is never mentioned, leaving a third of the parameters undocumented in both schema and prose.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: driving every driven joint from its lower limit toward `target` and reporting the fraction reached. This is a distinctive diagnostic action that an agent can separate from sim_inspect_joints or sim_set_targets, though no sibling is named explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives concrete operating conditions: reset=True first re-homes and verifies joints, and the daemon must be started with ground=False if the robot can reach the floor, otherwise the test measures the floor instead of joint range. It does not, however, say when to prefer this over sim_sweep or sim_inspect_joints.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

sim_reloadA

Restart the daemon to change a spawn property: robot, USD, solver iterations, self-collision.

One Kit start rather than one per experiment; gains stay live via sim_set_params. Omitted robot/usd fall back to the defaults, not to what was loaded before.

ParametersJSON Schema
NameRequiredDescriptionDefault
guiNo
usdNo
robotNo
groundNo
pos_itersNo
self_collisionsNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It does disclose a genuinely important trait: omitted robot/usd fall back to defaults rather than the previously loaded values, which is a destructive-feeling reset. It says nothing about state loss, downtime, or permissions, so it is only partially transparent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two dense sentences, front-loaded with the action and scope, with the alternative and the default-fallback caveat following. Nothing is wasted, though the second sentence is terse enough to be slightly cryptic.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need no explanation, and the description covers the core restart semantics and default fallback. However, for a 6-param mutation with zero annotation coverage, it leaves gui/ground unexplained and omits the disruption profile an agent would want before restarting a daemon.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% and there are 6 params, so the description must compensate. It names robot, usd, solver iterations (pos_iters) and self-collision (self_collisions), and clarifies default-fallback semantics for robot/usd, but it never mentions gui or ground and gives no format details for usd.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (restart) plus resource (the daemon) and enumerates the spawn properties it can change (robot, USD, solver iterations, self-collision). This clearly separates it from sim_up/sim_down and from sim_set_params, which it names as the tool for live gains. It stops short of 5 only because 'daemon' is unexplained jargon for the Kit process.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives the operative condition (change a spawn property) and contrasts it with the alternative for live values (sim_set_params), plus a note on why one restart beats one per experiment. There is no explicit statement about when NOT to reload versus sim_up/sim_down, so it is context rather than full routing guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

sim_screenshotA

Capture the viewport as an image. Requires the daemon started with gui=True.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A3.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It usefully discloses the gui=True prerequisite, which is a genuine behavioral constraint. But it omits return characteristics (image format, size, viewport vs full scene), latency/cost, and failure behavior when gui is disabled.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, zero filler, with the action stated first and the prerequisite immediately after. Every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple zero-parameter tool this is nearly complete — purpose plus the one gating precondition. The only gap is the absence of any indication of what the returned image looks like (format/dimensions), which matters since there is no output schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so there is nothing for the description to clarify; baseline for a parameterless tool is 4. The empty schema is self-consistent with the description implying a no-argument capture.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource: 'Capture the viewport as an image.' An agent immediately understands it returns a rendered screenshot of the simulation viewport, which is clearly distinct from siblings like sim_set_camera or sim_get_stats. It stops short of explicitly contrasting itself with any sibling, so 4 rather than 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives a real precondition — the daemon must be started with gui=True — which tells the agent when this tool will actually work. However, it offers no guidance on when to prefer this over alternatives (e.g. sim_set_camera + capture workflow) or what to do if the daemon wasn't started with gui=True.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

sim_set_cameraB

Point the viewport camera, e.g. eye=[1.2, 1.2, 1.0] target=[0, 0, 0.3].

ParametersJSON Schema
NameRequiredDescriptionDefault
eyeYes
targetYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden, and it discloses almost nothing: not whether this mutates persistent scene state or is a transient viewport-only change, not whether coordinates are world-space or robot-base-relative, and not whether a render must occur afterward. For a camera-control tool with zero annotation coverage this is a substantial gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence that states the action and immediately grounds it with a concrete example. Nothing is padded or repeated from the schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be described, and with only two required numeric-vector params the surface is small. However, missing coordinate-frame conventions and persistence semantics leave an agent guessing about how the values will be interpreted.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% — the schema gives only titles ('Eye', 'Target') with no semantic text. The example (eye=[1.2, 1.2, 1.0], target=[0, 0, 0.3]) does imply 3-element numeric vectors and that 'eye' is the camera position while 'target' is the look-at point, which partially compensates, but the coordinate frame and units are never stated.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb ('point') and resource ('viewport camera'), which cleanly separates it from siblings like sim_set_pose and sim_set_targets that manipulate robot/object targets rather than the viewport. It does not explicitly call out that distinction, so it falls short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no statement of when to use this tool, when not to, or which sibling to prefer (e.g., sim_screenshot or sim_set_pose). The inline example hints at typical values but conveys no usage conditions or prerequisites.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

sim_set_couplingA

Change software-coupling ratios live, keyed by follower joint. 0 releases a follower.

Couplings are declared in the robot config; a follower is driven to ratio x its leader's target.

ParametersJSON Schema
NameRequiredDescriptionDefault
ratiosYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden, and it does add real behavioral meaning: values are keyed by follower joint, 0 releases a follower, and couplings must be declared in the robot config. It still omits error behavior, persistence across steps/reloads, and what happens to unspecified couplings, leaving meaningful gaps for a mutation tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, front-loaded with the core action and immediately followed by the one non-obvious value rule ('0 releases a follower'). No filler; every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return values need no explanation, and the single parameter is well covered semantically. The main residual gap for an annotation-free mutation tool is failure/edge behavior and whether the change persists, which keeps it from a 5.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% and the parameter is an untyped key->number map, so the description must compensate. It does: the keys are follower joints, the values are the ratio multiplier applied to the leader's target, and 0 has the special meaning of releasing the follower. Value limits/units are not given, preventing a 5.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a specific verb+resource ('Change software-coupling ratios') and clarifies the mechanism ('a follower is driven to ratio x its leader's target'), which is distinguishable from siblings like sim_set_targets or sim_set_params. It never explicitly names or contrasts with those siblings, so it stops short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The word 'live' hints that this is a runtime change, but there is no explicit when-to-use guidance, no when-not, and no naming of alternatives such as sim_set_targets or sim_set_pose. Usage must be inferred from the coupling semantics alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

sim_set_paramsC

Change the driven joints' PD gains and effort limit on the live articulation.

ParametersJSON Schema
NameRequiredDescriptionDefault
effortNo
dampingNo
stiffnessNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It discloses that the change applies to the 'live' articulation (implying runtime mutation), but says nothing about prerequisites (e.g. sim_up), reversibility, units, or what happens to parameters left null.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single tight sentence with no filler, front-loading the verb and the affected resource. Every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Output schema exists, so return values needn't be explained, but for an unannotated mutation tool with 0% parameter coverage the description is too thin. Units, null semantics, and the requirement that the articulation be live are all left unstated.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% and none of the three parameters (effort, damping, stiffness) are named or documented. 'PD gains' loosely maps to stiffness/damping and 'effort limit' to effort, but this is inference the description never makes explicit, and null/default behavior is unexplained.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (change), resource (PD gains and effort limit), and scope (driven joints on the live articulation). This distinguishes it from sim_set_targets and sim_set_pose, which mutate different properties. Lacks explicit naming of siblings but the resource is specific enough to differentiate.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance, no prerequisites, and no alternatives named. The phrase 'on the live articulation' hints that the sim must be running, but the agent gets no routing help against sim_set_targets, sim_set_pose, or sim_set_coupling.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

sim_set_poseA

Command a named pose. amount blends toward it: 1 = as authored, 0.5 = halfway.

Halfway from the lower limits by default, or from the current targets with from_current=True. Follow with sim_step, then sim_screenshot to verify.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYes
amountNo
from_currentNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does well: it explains the blending semantics of amount (1 = as authored, 0.5 = halfway) and the default reference frame (lower limits vs current targets). It omits whether the command is immediate, requires stepping, or affects persisted state.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences, front-loaded with the core action and then the parameter semantics and follow-up workflow. No wasted words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need no explanation. The description covers the two non-obvious parameters and the verification workflow, leaving only minor gaps around the name argument and state-mutation effects.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate. It clearly defines amount's scale and from_current's two modes, but leaves the required name parameter's source (presumably sim_list_poses) unexplained.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

"Command a named pose" states a specific verb and resource, and the blending sentences clarify the effect. It is distinguishable from sibling list/mutation tools like sim_list_poses and sim_set_targets, though it never names those siblings explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives a clear operational workflow ("Follow with sim_step, then sim_screenshot to verify"), which tells the agent how to sequence this call. It stops short of naming alternatives or stating when NOT to use this tool versus sim_set_targets.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

sim_set_targetsA

Command the driven joints, normalized 0..1 (lower..upper limit).

Pass unit as a full vector in driven order (see sim_status), or joints to set individual joints by name. Targets persist; advance with sim_step.

ParametersJSON Schema
NameRequiredDescriptionDefault
unitNo
jointsNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden. It usefully discloses that targets persist until sim_step advances the sim, but says nothing about permissions, error behavior when vectors are wrong length, or blocking characteristics. One genuine behavioral trait is disclosed, so a mid score is appropriate.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short, front-loaded sentences: the core action and normalization first, then the parameter routing, then the persistence/next-step note. No filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained, and the description covers argument modes, units, and the step coupling. For a two-optional-parameter command tool it is nearly complete, with only error/edge-case behavior left implicit.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% (bare array/object with no descriptions), so the description must compensate. It does so well: it explains that `unit` is a full vector in driven order and `joints` maps joint names to values, plus the 0..1-to-limit normalization semantics that appear nowhere in the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Command the driven joints') with the normalization scheme (0..1 mapped to lower..upper limit), so the agent understands this sets joint targets rather than poses or parameters. It does not explicitly differentiate itself from siblings like sim_set_pose or sim_set_params, which keeps it from a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear either/or guidance for the two parameters: pass `unit` as a full driven-order vector, or `joints` to address individual joints by name, and points to sim_status for the driven order. It also tells the agent the next step ('advance with sim_step'), but offers no explicit when-not or comparison against alternative setters.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

sim_statusA

Report whether the sim daemon is up, its robot, driven joints, poses and parameters.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.5/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. 'Report' implies a read-only operation, which is a minimal behavioral hint, but there is no explicit statement about safety, side effects, permissions, or behavior when the daemon is down.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no redundant or filler content. It lists the reported items efficiently.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The output schema exists, so return values need not be explained. However, with no annotations and many siblings, the description is minimally adequate: it states purpose but omits usage context and behavioral traits that would help an agent choose correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so the baseline is 4. The description correctly adds no parameter details because none are needed.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Report') and resource (sim daemon status, robot, driven joints, poses, parameters). The listed contents clearly distinguish it from narrow siblings like sim_get_joint_state, sim_list_poses, and sim_get_stats.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance or alternatives are given. It only says what the tool reports, leaving the agent to infer that this is a general status check rather than a targeted query like sim_get_stats or sim_inspect_joints.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

sim_stepA

Advance the simulation n physics steps (dt from sim_status, default 1/120 s).

ParametersJSON Schema
NameRequiredDescriptionDefault
nNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It usefully discloses that dt is taken from sim_status with a 1/120 s default, which is real behavioral context, but says nothing about blocking behavior, whether stepping mutates shared state, or performance for large n.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with the core action first and the timing detail in a parenthetical. Nothing is wasted and nothing is buried.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-parameter, zero-required tool with an output schema, the description covers the action, the parameter, and the timestep source, which is enough to call it correctly. Only the when-to-use-vs-siblings gap remains.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, and it does: it defines n as the number of physics steps and clarifies that the timestep is not a parameter but is derived from sim_status. The only omission is the schema's default of 60 for n.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Advance the simulation') and quantifies the unit of work as 'n physics steps', so the agent knows exactly what it does. It does not differentiate itself from siblings like sim_play or sim_up, which is the only thing keeping it from a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies the use case (incremental stepping of the sim) but never says when to choose this over sim_play (continuous run) or sim_up/sim_down. No preconditions or exclusions are given, so usage must be inferred from the verb alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

sim_sweepA

Compare several values of one live parameter inside the single session.

param: stiffness, damping, effort, or "coupling:". For each value it applies the parameter and runs test ("wave" or "range") for steps. The original value is restored afterwards, even on error. A diverging solver is recorded as a result, not raised.

ParametersJSON Schema
NameRequiredDescriptionDefault
testNowave
paramYes
stepsNo
valuesYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does well: it discloses that the original value is restored afterwards even on error, that a diverging solver is recorded as a result rather than raised, and that everything runs in a single session. It omits permission/auth or performance/timeout caveats, but the key state and error-handling behaviors are disclosed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose is front-loaded in the first sentence, followed by tightly packed behavioral and parameter detail. No filler sentences, though the enum lists make it read a touch list-like rather than prose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need no explanation. Given the 4-param schema at 0% coverage and no annotations, the description covers the essential behavior and parameter meaning adequately, though it never states the result shape for a sweep (per-value outcomes) which would help an agent interpret output.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, and it largely does: it enumerates the valid param choices (stiffness, damping, effort, or coupling:<follower>), the test options ("wave" or "range"), and explains steps. The values array is only implied as the set to sweep, leaving its semantics slightly thin.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource: comparing several values of one live parameter within a single session. The sweep scope functionally distinguishes it from sim_set_params (single set) and sim_wave/sim_range_test (single run). It does not explicitly name any sibling, so it stops short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The mechanics (apply each value, run the test, restore) imply the use case of parameter comparison, but there is no explicit when-to-use, when-not, or routing to alternatives like sim_set_params or sim_range_test. Usage is implied rather than stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

sim_upA

Start the daemon container and wait until it answers. Idempotent.

robot: a robot config JSON, path relative to the repo root (e.g. examples/robots/franka.json). usd: a USD path instead of / overriding the config's (every joint driven if no config). gui=True is needed for sim_screenshot; the host must have run xhost +local:docker once. ground=False spawns without a floor (use for range tests). extra_args go to scripts/simd.py verbatim (e.g. ["--pos-iters", "64"]). The first start downloads Nucleus assets and builds shader caches: minutes.

ParametersJSON Schema
NameRequiredDescriptionDefault
guiNo
usdNo
robotNo
groundNo
timeoutNo
extra_argsNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does well: idempotency, a blocking wait, a host prerequisite (xhost +local:docker), and a significant latency warning on first start (Nucleus asset download and shader cache builds take minutes). Missing only what happens on failure/timeout.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded purpose sentence, then a compact per-parameter block with zero filler. Slightly terse formatting, but every line adds information an agent needs.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Output schema exists so return values need not be described. The description covers prerequisites, latency, idempotency, and most parameter semantics for a 6-param tool, leaving only the timeout parameter's meaning unaddressed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate and largely does: it explains robot (repo-relative path with example), usd (overrides config), gui (screenshot dependency), ground (floor removal for range tests), and extra_args (passed verbatim to scripts/simd.py with example). Only timeout is left unexplained.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Start the daemon container and wait until it answers'), plus a scoping trait ('Idempotent'). Combined with the sim_* sibling family, an agent can distinguish this from sim_down/sim_reload without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives conditional guidance tied to other tools: gui=True is required for sim_screenshot, and ground=False is for range tests. It doesn't explicitly name sim_down as the inverse or state when NOT to call it (e.g. already running), but the operational context is clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

sim_waveB

Sweep every driven joint through its range for n steps; return travel stats.

ParametersJSON Schema
NameRequiredDescriptionDefault
nNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. It states that joints are swept and stats are returned, but does not clarify whether simulation state is modified, whether the motion is reversible, or what permissions or constraints apply.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single, front-loaded sentence with no redundant or filler content. It delivers the core action and scope immediately.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained further. However, with no annotations and several similar siblings, the description leaves key context unaddressed: side effects on simulation state and why an agent should pick this over sim_sweep or sim_range_test.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema provides no description for the sole parameter 'n', so the description must compensate. It adds meaning by saying 'for n steps,' but gives no range, format, or constraints beyond the default already in the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Sweep') and a clear resource scope ('every driven joint through its range for n steps'). This distinguishes it from single-joint or step tools, but it does not explicitly differentiate from close siblings like sim_sweep or sim_range_test.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is given on when to choose this over sim_sweep, sim_range_test, or sim_step. The implied usage is clear from the verb, but there are no conditions or alternatives mentioned.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

train_checkpointsB

List a run's saved checkpoints with step and size.

ParametersJSON Schema
NameRequiredDescriptionDefault
run_nameYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the burden, but the 'List' verb clearly signals a non-mutating read operation, which is the key behavioral trait here. It adds nothing about pagination, ordering, or whether checkpoints can be deleted or restored, so it is only minimally transparent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One short sentence, front-loaded with the verb and resource, with no filler. Nothing can be trimmed without losing information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return shape need not be explained, and this is a simple one-parameter read. However, the description never links run_name to the run identifiers produced by sibling tools like train_list, leaving a small but real gap for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The single parameter run_name has 0% schema description coverage and the description adds no meaning beyond the schema title 'Run Name'. It does not say whether the value must match a run returned by train_list or what format is expected.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (List) and resource (a run's saved checkpoints) and even previews the returned fields (step and size). It is distinguishable from siblings like train_list or train_logs by resource, though the description never explicitly contrasts them.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance and no alternatives named. An agent must infer from the name alone that this is for inspecting a run's checkpoint artifacts rather than listing runs (train_list) or reading logs (train_logs).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

train_listA

List training runs: live or recent containers, and log directories holding checkpoints.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden. It does disclose the sources it enumerates (live or recent containers plus checkpoint log directories), which tells the agent this is a broad discovery call rather than a filtered one, but it says nothing about read-only safety, staleness/recency semantics of 'recent', or whether it touches remote hosts.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with the verb and resource first, followed by a compact colon-style enumeration of what is listed. No filler, no restatement of the title.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return-value explanation is unnecessary, and with zero parameters the schema side is fully covered. The description is adequate for a simple enumeration tool, though it leaves the boundary against train_status and train_checkpoints undefined.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so there is nothing for the description to clarify; baseline 4 applies. The description correctly avoids inventing filter semantics that do not exist in the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('List') and resource ('training runs') and adds scope detail about what a run comprises (live/recent containers, log directories holding checkpoints). It is clear on its own, but it does not differentiate itself from siblings like train_status or train_checkpoints, and the mention of 'log directories holding checkpoints' partially overlaps with train_checkpoints.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no exclusions, and no named alternative. An agent must infer whether to call train_list vs train_status vs train_checkpoints, especially given the checkpoint overlap. Nothing tells the agent under what conditions this is the right call.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

train_logsA

Tail a running training container's stdout/stderr.

ParametersJSON Schema
NameRequiredDescriptionDefault
tailNo
run_nameYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. 'Tail' and 'stdout/stderr' disclose that this is a read of log output and hint at streaming/last-N-lines behavior, but it never states whether the call blocks or streams, whether it is read-only, or how it behaves for a stopped container.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single tight sentence with the key constraint ('running') front-loaded. No filler, no redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be described. The description covers the core action but leaves gaps for a log-retrieval tool: streaming vs snapshot behavior, handling of a non-running container, and the meaning of the tail count are all unstated.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate. The verb 'tail' loosely maps to the integer 'tail' parameter (default 200, implying number of lines), and 'training container' implies run_name selects the target, but neither parameter is explicitly explained in terms of format or semantics.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb (tail) and resource (a running training container's stdout/stderr), which is far more informative than the bare name train_logs. It implicitly distinguishes itself from siblings like train_metrics and train_status by scoping to raw container output, though it never explicitly contrasts them.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'a running training container' implies the tool is only applicable while training is active, which is useful context. However, there is no explicit guidance on when to use this versus train_metrics or train_status, and no statement of what happens if the run has finished.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

train_metricsA

TensorBoard scalars for a run: no tag lists tag names; a tag returns its series, downsampled.

ParametersJSON Schema
NameRequiredDescriptionDefault
tagNo
run_nameYes
max_pointsNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It usefully discloses the dual-mode behavior and that the returned series is downsampled, which is real behavioral information. It omits any note on auth/permissions, error behavior, or the effect of max_points on downsampling, leaving notable gaps for an annotation-free tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler; the resource is stated first and the mode logic follows. Efficient and readable, though the colon-separated clauses are slightly dense.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be described, and the key tag-mode behavior is covered. run_name is self-evident and max_points has a schema default. The description is nearly complete, missing only explicit max_points semantics.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It explains the tag parameter's two modes well (null lists names, a value returns the series), but says nothing explicit about run_name or max_points beyond the passing hint 'downsampled'. Partial compensation, not full.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific resource and scope: 'TensorBoard scalars for a run', with a clear verb-implied retrieval action. It is not a tautology and the two operating modes (tag absent vs present) are stated. It does not, however, explicitly differentiate itself from siblings like train_logs, so it falls short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'no tag lists tag names; a tag returns its series' tells the agent how to switch between the two modes, which is genuinely useful selection guidance within the tool. But there is no explicit when-to-use guidance versus alternatives such as train_logs or train_checkpoints, leaving sibling routing to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

train_startA

Launch a headless Isaac Lab training run in its own container; returns at once.

task is a registered gym id (e.g. Isaac-Cartpole-v0). run_name defaults to '_'; keep it to find the run later. extra_args go to the training script verbatim (e.g. Hydra overrides). device pins one GPU index. Independent of the daemon and of this MCP session.

ParametersJSON Schema
NameRequiredDescriptionDefault
seedNo
taskYes
deviceNo
num_envsNo
run_nameNo
checkpointNo
extra_argsNo
max_iterationsNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does add real behavioral context: it is asynchronous ('returns at once'), container-isolated, and explicitly independent of the daemon and the MCP session. It still omits failure behavior, resource/GPU requirements, and any concurrency limits.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded purpose followed by tight per-parameter notes; each sentence adds information. The trailing 'Independent of the daemon...' line is slightly disconnected but still earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return values need not be explained, and the description covers the async launch semantics and the most ambiguous parameters. The gaps are the four undocumented params and the absence of any sibling routing, which for an 8-param async launcher is a modest shortfall rather than a serious one.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% across 8 params, so the description must compensate and it only does so for half: task (gym id + example), run_name (default pattern + retention purpose), extra_args (verbatim + Hydra example), device (GPU index). seed, num_envs, checkpoint, and max_iterations are left to their self-explanatory names.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Launch a headless Isaac Lab training run in its own container') with the key scope fact that it returns at once. This is clearly distinguishable from the sim_* control tools and from the train_status/train_logs readers.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied by 'keep it [run_name] to find the run later', which gestures at train_status/train_logs, but no sibling is named and there is no explicit when-to-use or when-not-to-use guidance. An agent can infer intent but is not routed.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

train_statusB

Container state, checkpoints, and latest TensorBoard scalar values for one run.

ParametersJSON Schema
NameRequiredDescriptionDefault
run_nameYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full disclosure burden. It does add one behavioral trait — only the *latest* scalar values are returned, implying history lives elsewhere — but says nothing about permissions, whether the tool blocks on a running container, or how a missing run is handled.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler; the most important information (what comes back, for how many runs) leads the sentence.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be spelled out. However, with five overlapping train_* siblings and no annotations, the description should clarify how this differs from train_metrics, train_checkpoints, and train_logs — that routing information is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% and the single parameter run_name has no schema description. The phrase 'for one run' hints that run_name selects a single run, which is weak but non-zero compensation for the undocumented parameter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific resource and its payload (container state, checkpoints, latest TensorBoard scalars) scoped to one run, so an agent knows it is a read/status tool. It lacks an explicit verb and does not distinguish itself from the overlapping siblings train_checkpoints, train_metrics, and train_logs, which cover much of the same ground.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no prerequisites, and no mention of any alternative. Given five sibling train_* tools with overlapping semantics, the absence of routing guidance is a real gap.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

train_stopC

Stop a training run's container.

ParametersJSON Schema
NameRequiredDescriptionDefault
run_nameYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. It says the tool stops a container but does not describe whether the stop is graceful or forced, whether it is irreversible, what happens to the training run's state, or any authentication or side-effect considerations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence with zero wasted words. It states the action and target immediately, which is appropriately concise for a simple tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Although an output schema exists and return values need not be explained, the description is incomplete for a mutation tool with no annotations. It omits usage context, behavioral implications, and parameter details, leaving significant gaps for an agent to invoke the tool with confidence.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the single parameter, run_name, is undocumented in both the schema and the description. The description implies that the tool operates on a training run, but it adds no format, syntax, or meaning beyond the parameter's name.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb, 'Stop', and a specific resource, 'a training run's container', which clearly distinguishes it from sibling tools like train_start and train_status. However, it does not explicitly differentiate from those siblings by naming them or describing conditions, so it falls short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives such as train_start or train_list. It implies usage but offers no context, prerequisites, or exclusions, leaving the agent to infer when stopping is appropriate.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 26 tool updatesv0.1.0
    • First observedsim_down
    • First observedsim_get_joint_state
    • First observedsim_get_stats
    • First observedsim_inspect_joints
    • First observedsim_list_poses
    • First observedsim_play
    • First observedsim_range_test
    • First observedsim_reload
    • First observedsim_screenshot
    • First observedsim_set_camera
    • First observedsim_set_coupling
    • First observedsim_set_params
    • First observedsim_set_pose
    • First observedsim_set_targets
    • First observedsim_status
    • First observedsim_step
    • First observedsim_sweep
    • First observedsim_up
    • First observedsim_wave
    • First observedtrain_checkpoints
    • First observedtrain_list
    • First observedtrain_logs
    • First observedtrain_metrics
    • First observedtrain_start
    • First observedtrain_status
    • First observedtrain_stop

TDQS

A3.6/5.0

Scored across 26 tools

Disambiguation5/5

Tools are cleanly partitioned into sim_* and train_* domains, and within sim the verbs target distinct actions (status, step, play, set_targets vs set_pose vs set_coupling). Minor overlap between sim_wave, sim_range_test, and sim_sweep is resolved by their descriptions.

Naming Consistency5/5

All tool names use snake_case with consistent sim_ or train_ prefixes and predictable verb_noun structure, e.g. sim_set_targets, sim_get_joint_state, train_start, train_checkpoints.

Tool Count4/5

26 tools is slightly above the ideal single-domain range, but the server covers two substantial subsystems: live Isaac Sim control and headless Isaac Lab training orchestration. The count is reasonable for that dual scope.

Completeness4/5

The surface covers sim lifecycle (up/down/reload, status, joint state, poses, params, camera, screenshot, experiments) and training lifecycle (start/list/status/logs/stop/checkpoints/metrics). Minor gaps include no explicit generic sim reset and no training-run cleanup/delete operation.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    D
    maintenance
    Enables AI tools like Windsurf and Claude to control NVIDIA Isaac Sim and Isaac Lab through natural language, providing tools for scene inspection, prim management, physics simulation, and robot spawning.
    8
    Apache 2.0
  • A
    license
    A
    quality
    B
    maintenance
    Enables driving Omniverse Kit apps (Isaac Sim, Isaac Lab) over MCP, allowing agents to control simulations, run Python, and call namespace-scoped tools via a single bridge.
    1
    2
    MIT
  • A
    license
    C
    quality
    B
    maintenance
    Enables AI assistants to control NVIDIA Isaac Sim by building scenes, loading assets, operating robots and humans, reading sensors, managing simulation, and creating Action Graphs.
    129
    1
    MIT