embodify-mcp
This MCP server gives an agent embodied robot control: inspect tasks, run episodes, observe cameras/state, move arms, and operate grippers in simulated or remote backends.
Browse task catalogues with
list_tasks(suites, tasks, init layouts) without affecting a running episode.Start/reset episodes with
reset_task, optionally choosing suite, task index, and initial layout, and get first camera images plus robot state.Read camera images and end-effector/gripper state without moving or spending budget via
observe.Move end effectors relative by translation and optional rotation with
move_relative, receiving stop reasons, achieved displacement, and remaining distance.Open or close grippers with
set_gripper, checking opening width and stop conditions like settled or blocked.Inspect current run status, task, step budget, and robot state with
get_session_info.End the current episode and save frames, event logs, and summary artifacts with
stop_episode.Work across backends: FakeSim, LIBERO, RoboDojo, and remote simulation over SSH/TCP; RoboTwin and SO-101 are coming soon.
Support fair evaluation with hidden task success, configurable budgets, and live monitor/replay (as described in the README).
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@embodify-mcplist the LIBERO tasks and start an episode"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
First release (0.1.0a1). Supports LIBERO and RoboDojo today; support for RoboTwin and the LeRobot SO-101 arm is coming soon.
Why Embodify
Most embodied agents are built from scratch: a dedicated harness wraps a model, hands it a fixed set of robot actions, and nothing else. The agent you use every day already has what those harnesses lack:
Context and memory. It manages long sessions and remembers across them.
It knows you. Your preferences, your projects, your lab setup.
It talks with you in the terminal, IDE, desktop app or chat you already use. You can ask how it is going, step in, correct it, or teach it mid-task.
It has a computer. It writes and runs code, uses its tools, searches the web and reads papers.
Embodify keeps all of that and adds a body. Install the plugin and your agent gets robot observation and control tools over MCP — a protocol it already uses for everything else — plus skills that teach it how to operate robots and how to achieve recursive self-improvement (RSI) from its own experience.
Why "best"
A general agent with a body is stronger than a harness that can only move a robot:
It can think with tools, not only act. When a task needs geometry it can write a script; when it needs a fact it can search; when it needs perception it can run a model.
It keeps track. Long-horizon manipulation fails when the agent forgets what it already tried. Mature agents manage context and memory well.
It works with you. It asks you when a task is ambiguous, takes your guidance when it gets stuck, and remembers your corrections next time.
It controls robots in its native language. Robot control arrives as ordinary MCP tool calls, the same shape as every other tool the agent uses.
It builds on frontier embodied models. Your agent runs on models with frontier embodied manipulation skills, such as GPT-6 Astra and Claude Opus 5.5, and every model upgrade makes your robot better, with no retraining.
It improves itself recursively. It turns every episode into lessons, lessons into rules and rules into new skills, and gets better the more it works.
Zero-shot, few-shot and in-context learning come easily. A general agent takes on new tasks without training: describe a task in plain language (zero-shot), show it a few examples (few-shot), or put instructions, demonstrations and past experience in its context (in-context learning, ICL).
Related MCP server: Mobile Device MCP
What's inside
Embodify is an agent plugin with two parts, and installing the plugin gives your agent both:
Part | What it gives your agent |
MCP server ( | Tools to list tasks, start an episode, observe cameras and robot state, move end effectors, open and close grippers, and coordinate two arms. |
Skills | Know-how: running a careful observe–act loop, keeping a profile of the robot body and cameras, and learning from past episodes. More perception skills are coming soon. |
flowchart LR
A["Your agent<br/>Claude Code · Codex · …"] -- "MCP (stdio)" --> B["embodify-mcp<br/>episodes · budgets · logs"]
K["Embodify skills"] -. "loaded by" .-> A
B --> I["Backend interface"]
I --> L["LIBERO"]
I --> R["RoboDojo"]
I --> T["RoboTwin (WIP)"]
I --> H["SO-101 and other real robots (WIP)"]
I -. "SSH / TCP" .-> G["Remote server"]The MCP server is the front end. Each simulator, benchmark or robot is a backend behind one small interface, so adding a new one does not change what the agent sees. Highlights:
Works with any MCP host. No LLM calls inside; your agent keeps its own model, memory, skills and tools.
One call, one motion. The control loop runs next to the simulator. Each action returns the actual displacement, remaining error, a stop reason and fresh camera images.
Remote simulation. Run the simulator on your lab's GPU server over SSH while the agent stays on your laptop. Heartbeats keep high-latency links alive; if the link drops, the episode aborts cleanly and
reset_taskreconnects.Fair evaluation. When you use Embodify to evaluate an agent's manipulation skills, task success is recorded for humans and never exposed to the agent.
Live monitor and replay. Watch running episodes live in a browser, or replay past ones frame by frame.
Backends
Backend | Robot | Status |
| Kinematic diagnostic, one or two arms | ✅ Included, no simulator needed |
| Franka Panda, 130 tasks in 5 suites | ✅ Included |
| Dual ARX X5, all 54 simulation tasks | ✅ Included |
RoboTwin | Dual-arm manipulation benchmark | 🚧 Coming soon |
LeRobot SO-101 | Real robot: 5-DoF arm with gripper | 🚧 Coming soon |
| Any backend above, on another machine | ✅ SSH or trusted TCP |
Coming soon: robot arms such as xArm, mobile manipulators, humanoids and dexterous hands. If there is a simulator or robot you would like us to support, open a GitHub issue with details about it. We welcome contributions of new backends.
For research
Embodify doubles as a testbed for the embodied manipulation abilities of language models and agent frameworks. Run GPT, Claude, Gemini, Qwen and other models, through Claude Code, Codex or your own agent framework, on the same simulated benchmarks with the same tools, and compare them directly.
Standard benchmarks: LIBERO (130 tasks in 5 suites) and RoboDojo (54 tasks), with RoboTwin coming soon.
Fair by design: task success is recorded for evaluators and never shown to the agent, and
--lock-taskfixes the scene.Controlled budgets: episode budgets and per-call step limits are configurable and logged.
Complete records: tool calls, robot states and camera frames are saved for replay and analysis.
Typical studies include zero-shot robot manipulation by frontier models, comparisons of agent frameworks on the same model, the effect of skills and memory, and recursive self-improvement across episodes. See citation for how to cite Embodify and what to report.
Quick start
One-line setup
Paste this into Claude Code, Codex or any other agent you use:
Install Embodify (https://github.com/YidaYang/embodify) for yourself, then set up a LIBERO environment or connect to my existing simulator.That's it. Your agent installs the Embodify plugin, sets up the simulator and tells you when to restart the session. (Agents: follow the agent setup guide.)
Manual install
Embodify starts its MCP server with uv, straight
from PyPI. If you do not have uv yet, install it with python -m pip install uv
or the official installer.
Claude Code
claude plugin marketplace add YidaYang/embodify
claude plugin install embodify@embodifyCodex
codex plugin marketplace add YidaYang/embodify
codex plugin add embodify@embodifyThe plugin brings both the MCP server, registered as embodify, and the skills.
The server starts on the Fake backend, so you can try the tools right away;
connect a real simulator next. It reads its
settings from ~/.embodify/config.json (see backend setup).
To check the package on its own, without an agent:
uvx --from embodify-mcp embodify-mcp-smoke # end-to-end check: no simulator, GPU or model neededOther agents
Register the server in your agent's MCP configuration using examples/mcp.json.
If your agent supports Agent Skills (
SKILL.mdfolders), copy or link the folders in plugin/skills/ into its skills directory.
Restart the session, then ask your agent something like "Reset the task, describe what the cameras show, then raise the gripper 5 cm."
Connect a real simulator
Let your agent do this step. Paste:
Connect Embodify to a real simulator for me: set up LIBERO on this machine, or connect to my existing simulator. Follow https://github.com/YidaYang/embodify/blob/main/docs/agent-setup.md.Your agent installs the simulator or connects to yours over SSH, writes the
server settings to ~/.embodify/config.json and asks you to restart the
session. To do it by hand, see backend setup.
Watch robots live and replay episodes
Add "monitor-port": 8765 to ~/.embodify/config.json and open
http://127.0.0.1:8765 to watch the robot work live, with every camera view,
the robot state and each tool call your agent makes, or to replay any past
episode frame by frame. To browse saved runs without a running server:
uvx --from embodify-mcp embodify-mcp-monitor # reads ~/.embodify/runsThe monitor listens on localhost only.
MCP tools
Tool | Purpose |
| Current run, task, step budget and robot state; no images |
| Browse the backend's task catalogue |
| Start an episode and return the first camera images |
| Camera images and robot state, without moving |
| Move one end effector by a translation and optional rotation |
| Open or close a gripper |
| Move several arms together with common progress (two-arm backends) |
| End the episode and write its record |
Translations are in meters, rotations in radians, quaternions in xyzw order. Frames and stop reasons are defined in the action contract.
Skills
The plugin contains these skills:
Skill | Status | What it does |
v0 | The observe–act loop: small moves, reading stop reasons, verifying grasps, releasing in a separate call | |
v0 | Keep a profile of the robot: arms, cameras, frames, calibration and measured behavior | |
v0 | Write a lesson after every episode, read relevant lessons before the next, promote repeated lessons to rules | |
Coming soon | Open-vocabulary segmentation of camera images | |
Coming soon | Pixel to 3D position from depth and calibration |
Skills are grouped into three families that will keep growing:
Embodiment knowledge: how this particular robot is built, where its cameras are and how it actually moves.
Manipulation tools: perception, measurement and planning tools, including ideas from work such as Code as Policies.
Recursive self-improvement: turning experience into lessons, lessons into rules and eventually into new skills, improving recursively.
Roadmap
Backends: RoboTwin; LeRobot SO-101 as the first real robot, with workspace limits, human-judged success and an emergency stop; more real robots and simulators.
Embodiments: mobile manipulators, humanoids, dexterous hands.
Observations: depth images, camera calibration and more through the tools.
Skills: segmentation, depth ranging and other perception tools; code-as-policy style tools; a stronger recursive self-improvement loop, and more.
Development
python -m pip install -e ".[dev]"
python -m pytest -q
python tools/check_release.pySee contributing, writing a backend, backend setup, action contract and security.
License and citation
Original code is licensed under Apache-2.0. Bundled third-party code keeps its original notices. See citation for how to cite Embodify.
Available Tools
7 toolsget_session_infoA
Read the state of the current run: task context, instruction, step count and robot state. Renders no images and uses no step budget; callable at any time, including before the first reset_task. This only says where you are now. To see which tasks are available, use list_tasks. Returns run (short run ID), status (idle/running/stopped), step (simulation steps used), max_steps (the episode limit), steps_left (simulation steps left), task (suite, task_index, name, instruction, init_state_index), task_locked (true means the server fixed this run's scene, and reset_task accepts no scene arguments), and state (eef_pos_m is the XYZ position in meters, eef_rpy_rad the RPY orientation in radians, gripper holds command/opening_m). While status is idle there is no task or state; instead it returns default_scene (the scene a reset_task without arguments starts).
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden and does so well: it discloses side-effect-free behavior ('Renders no images and uses no step budget'), availability at any lifecycle point, and the idle-vs-running state differences (no task/state while idle, default_scene returned instead). It also explains task_locked semantics, which an agent could not infer elsewhere.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with purpose then constraints, and every sentence adds usable information. The dense return-field enumeration is long but each element (eef_pos_m units, gripper fields) earns its place for an agent with no output schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, so the description must carry the return contract, and it fully enumerates run, status, step, max_steps, steps_left, task, task_locked, and state including units. Combined with the idle-state caveat, an agent has everything needed to call and interpret this tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so the baseline is 4. The description instead documents the rich return shape, which is useful but not a parameter concern.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Read the state of the current run') and immediately enumerates the scope (task context, instruction, step count, robot state). It explicitly distinguishes itself from the sibling list_tasks, so an agent can route correctly without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit when-to-use conditions ('callable at any time, including before the first reset_task') and names the alternative for a different need ('To see which tasks are available, use list_tasks'). This is unambiguous routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_tasksA
List the tasks FakeSim offers. Independent of the current run: callable at any time, renders no images, uses no step budget and does not change a running episode. Without arguments: returns suites, the name and task count of every suite, plus the total task count. Start here, pick a suite, then look inside it. With suite: returns task_index, instruction and n_init_states (the number of initial layouts of the same task) for every task in that suite. Pass the chosen suite and task_index to reset_task to start.
| Name | Required | Description | Default |
|---|---|---|---|
| suite | No | The suite whose tasks to list; omit to get only the suite overview. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations present, the description carries the full burden and does so: it discloses that the call is independent of the current run, renders no images, consumes no step budget, and does not alter a running episode. These are exactly the side-effect and cost traits an agent needs before calling in an interactive simulation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the most important constraint (independent of the current run) and then organizes the rest by mode, with no filler sentences. Slightly dense and could be tightened, but every sentence carries information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description correctly documents return contents for both modes (suites and counts; task_index, instruction, n_init_states) and explains n_init_states inline. An agent has everything needed to call it and interpret the result.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description goes further by spelling out what omitting 'suite' yields (suite names, per-suite task counts, total) versus what passing it yields (task_index, instruction, n_init_states). That is meaningful added semantics, though it does not cover type/format constraints.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('List the tasks FakeSim offers') and immediately distinguishes the two invocation modes (no argument = suite overview; with suite = per-task detail). It is clearly separable from siblings like reset_task and observe.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit workflow: 'Start here, pick a suite, then look inside it' and 'Pass the chosen suite and task_index to reset_task to start.' It names the downstream tool and the condition that selects it, leaving nothing to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
move_relativeA
Move the end effector by delta_xyz meters in the FakeSim world frame (z up, table surface at z=0.80 m), optionally rotating by delta_rpy about the robot base X/Y/Z axes (radians). The server approaches the target step by step with P control and returns only when a definite stop condition triggers, not after a fixed number of steps. action.stop_reason says why the call stopped, one of four:
reached: Within the target tolerance (8 mm in position, 0.02 rad in orientation); the motion is complete.
stalled: Barely moved for 3 consecutive steps (less than 1 mm in total) without reaching the target, which usually means something is blocking it: the table, an object, or a joint limit. Repeating the same command has no effect; change direction, or lift or back off first. This is the main signal of contact.
step_cap: Hit the per-call internal step limit while the end effector was still moving toward the target. Call again with the remaining displacement in the same direction to continue.
budget: The episode's total step budget ran out during the motion; only stop_episode can be called afterwards. Motion is a best-effort approach and may fall short of the requested delta: judge progress by action.achieved_delta_xyz (actual displacement of this call, meters) and action.remaining_distance_m (meters still missing to the requested target), and do not assume the request was achieved. Returns images (two separate PNGs), state (eef_pos_m is the end-effector XYZ position in meters, eef_rpy_rad the RPY orientation in radians, gripper holds command/opening_m), action (the requested delta_xyz/delta_rpy, executed_steps, stop_reason, achieved_delta_xyz, remaining_distance_m), step (simulation steps used after the call), steps_left (simulation steps left) and status (running). When the budget is already 0, returns a step_limit error instead of silently doing nothing.
| Name | Required | Description | Default |
|---|---|---|---|
| delta_rpy | No | [droll, dpitch, dyaw], increments about the robot base X/Y/Z axes in radians; omit to keep the current orientation. | |
| delta_xyz | Yes | [dx, dy, dz] in meters. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does so: it discloses best-effort motion, P-control stepping, the four stop conditions with their meanings, the budget exhaustion path (only stop_episode afterwards), and the step_limit error when budget is 0. These are exactly the non-obvious behaviors an agent needs before calling.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long but information-dense and front-loaded: the core motion semantics come first, then the stop reasons, then the return payload. The stop-reason list is the one section that could be tightened, but every item carries decision-relevant content.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 2-parameter motion tool with no output schema and no annotations, the description covers frame conventions, units, failure modes, continuation strategy, and the full return payload (images, state, action fields, step, steps_left, status). Nothing an agent needs to call it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% so the baseline is 3, but the description adds real meaning beyond the schema: the world frame with z up, the table surface at z=0.80 m, that delta_rpy rotates about the robot base X/Y/Z axes in radians, and that omitted delta_rpy keeps the current orientation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: move the end effector by delta_xyz meters in a named frame, with optional delta_rpy rotation. This is clearly distinct from siblings like set_gripper, observe, or stop_episode, so an agent can select it without ambiguity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives strong operational guidance: how to interpret each stop_reason, that 'stalled' means repeating the same command has no effect and the agent should change direction or lift first, and that 'step_cap' should be followed by another call with the remaining displacement. It doesn't explicitly frame when to prefer another tool, but the conditions are clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
observeA
Read the current RGB images from two cameras and the end-effector/gripper state, without advancing the simulation or using the step budget. Returns images (agentview and robot0_eye_in_hand: two separate PNGs), state (eef_pos_m is the end-effector XYZ position in meters, eef_rpy_rad the orientation about the base XYZ axes in radians, gripper holds the target command and opening_m in meters), step (simulation steps used), steps_left (simulation steps left) and status (running).
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It discloses the key behavioral traits: it does not advance the simulation and does not consume the step budget, and it enumerates the exact return fields with units. It does not cover error conditions or permissions, but for a non-mutating observation tool this is largely adequate.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
It is a single dense sentence but front-loads purpose before listing returns; every clause provides useful information. It could be slightly more scannable, but there is no filler or repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and no annotations, the description fully documents the return shape (two labeled PNG images, state fields with units, step counters, status) and the side-effect profile, leaving no significant gap for an agent invoking this zero-parameter read tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so the baseline is 4. There are no parameters to clarify and the description appropriately focuses on outputs instead.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Read) and resource (current RGB images and end-effector/gripper state) and adds the crucial scope constraint 'without advancing the simulation or using the step budget,' which sharply distinguishes it from the actuation siblings like move_relative and set_gripper.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The 'without advancing the simulation' clause implies when to use it (inspection without stepping), but it never names an alternative or says when not to use it. Sibling get_session_info is not mentioned as a differentiating option.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
reset_taskA
Start a new FakeSim episode. All three arguments are optional: omit them all to use the server's default scene (usually task 0 of the suite with initial layout 0), which is the most common use. To switch tasks, use list_tasks first: without arguments it lists the suites, with suite it gives the task_index and instruction of every task in that suite; then pass the chosen suite / task_index. init_state_index picks another object layout of the same task. If the server was started with --lock-task (task_locked is true in get_session_info), passing any of these arguments returns a task_locked error. If an episode is already running, returns already_running; call stop_episode first. Returns run (short run ID), task (suite, task_index, name, instruction, init_state_index), images (agentview and robot0_eye_in_hand: two separate PNGs), state (eef_pos_m is the end-effector XYZ position in meters; eef_rpy_rad.roll_x, pitch_y and yaw_z are the orientation about the robot base X/Y/Z axes in radians; gripper.command is the open/close target and gripper.opening_m the gripper opening in meters), step (simulation steps used) and steps_left (simulation steps left in this episode).
| Name | Required | Description | Default |
|---|---|---|---|
| suite | No | Switch to another suite; omit to keep the current one. See list_tasks (without arguments) for the suites. | |
| task_index | No | Which task of the suite; omit to use the server's default scene (task 0 when switching suites). See list_tasks with suite for the options. | |
| init_state_index | No | Which initial layout of the task; 0 if omitted. Must be below the task's n_init_states. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations exist, so the description carries the full burden and does so: it discloses two distinct failure modes (task_locked, already_running), the precondition that stop_episode must precede it, default-argument behavior, and the full shape of the return payload. This is far beyond what a bare 'reset' would convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The most important fact — that all three arguments are optional and omitting them is the common path — is bolded and front-loaded, which is well judged. The trailing enumeration of the return payload is long and reads as a run-on, but it is justified because no output schema exists.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no annotations and no output schema, the description must supply safety, preconditions, error behavior, and return semantics on its own — and it does all four, including unit-bearing field explanations (eef_pos_m, eef_rpy_rad, gripper.opening_m). Nothing an agent needs to call this correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents suite, task_index, and init_state_index including their defaults and the n_init_states bound. The description largely restates that ('init_state_index picks another object layout', task 0 default) rather than adding new syntax or constraints, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Opens with a specific verb and resource ('Start a new FakeSim episode') and immediately clarifies the scope of the operation (fresh episode, not a continuation). An agent can distinguish it from stop_episode, observe, and get_session_info without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit routing: omit all args for the default scene, call list_tasks first to switch tasks (and it even explains what list_tasks returns with and without args), and call stop_episode first if an episode is already running. The --lock-task precondition is stated as a when-not condition.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
set_gripperA
Open or close the gripper. The server keeps sending the command and watches the opening width, and returns only when a definite stop condition triggers, not after a fixed number of steps. action.stop_reason says why the call stopped, one of four:
settled: After at least 10 internal steps of the same command, the opening varied by less than 0.5 mm over the last 2 steps. This only means the opening is momentarily stable, possibly because something blocks it; it does not guarantee a fully open gripper, a firm grasp or a released object. Check state.gripper.opening_m and the images.
step_cap: Hit the per-call internal step limit before the opening was confirmed stable; the gripper may still be starting or moving. Repeat the same command to continue.
budget: The episode's total step budget ran out during the action; only stop_episode can be called afterwards.
no_op: The request matches the current gripper target and no gripper call is pending; nothing was executed and no budget was used. It does not mean an object was released. Returns images (two separate PNGs), state (eef_pos_m is the end-effector XYZ position in meters, eef_rpy_rad the RPY orientation in radians, gripper.command the gripper target and gripper.opening_m the opening in meters), action (the requested gripper target, executed_steps, stop_reason, opening_m before and after), step (simulation steps used after the call), steps_left (simulation steps left) and status (running). When the budget is already 0, returns a step_limit error and the gripper target stays unchanged.
| Name | Required | Description | Default |
|---|---|---|---|
| state | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full load and does so exceptionally: it explains the adaptive stop algorithm, that 'settled' does not imply a firm grasp or released object, the no_op case, the budget-exhaustion lockout, and the step_limit error leaving the target unchanged. This is exactly the operational context an agent needs for a stateful actuator.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The first sentence front-loads purpose plus the key behavioral mechanic, then the four stop reasons are cleanly bulleted. Length is justified because it simultaneously serves as the missing output documentation (no output schema exists); every clause is load-bearing.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No annotations and no output schema mean the description must cover both behavior and returns, and it does: it enumerates the returned fields (images, state, action, step, steps_left, status) with units and meanings, plus the error path. Nothing an agent needs to call or interpret this tool correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% and there is one parameter, but its enum values open/close are self-explanatory. The description never names or explains the 'state' parameter directly, and its references to 'the requested gripper target' do not map cleanly onto a parameter called state, so compensation for the coverage gap is only partial.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
"Open or close the gripper" is a precise verb+resource that an agent can act on immediately, and the gripper is a domain no sibling tool (move_relative, observe, stop_episode) touches. It stops short of explicitly positioning itself against those siblings, but the resource is unmistakable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The stop_reason breakdown effectively tells the agent what to do next: repeat the same command on step_cap, and note that only stop_episode is callable once budget is exhausted. That is concrete downstream guidance, though it never states the motivating use case (e.g. grasp vs release) or a when-not-to-use condition.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
stop_episodeA
End the current simulation and save its frames, event log and summary. Returns run (short run ID), status (stopped) and artifacts (events: path of events.jsonl, frames: PNG frame directory, summary: path of summary.json).
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and actually discloses meaningful side effects: frames, event log and summary are persisted, and it enumerates the returned run ID, status and artifact paths. It does not state preconditions (must an episode be active?) or error behavior, which a terminal state-changing tool should ideally cover.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tightly written sentences: the action is front-loaded and the second sentence enumerates return fields with parenthetical formats (events.jsonl path, PNG frame directory, summary.json path). Every clause carries information, with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description usefully documents the return shape itself (run, status, artifacts with paths), which makes it reasonably self-sufficient. It falls short only on lifecycle preconditions and what happens if no episode is running, which matter for a terminal operation in a stateful simulation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so the schema imposes no semantic burden and the baseline of 4 applies. Nothing in the description contradicts or obscures the empty parameter set.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a concrete verb and resource ('End the current simulation and save its frames, event log and summary'), so the agent knows exactly what happens. It is distinguishable from siblings like observe or move_relative, though it never explicitly contrasts with reset_task, which is the closest ambiguous alternative.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied: 'End the current simulation' suggests calling it to wrap up an episode, but there is no explicit when-to-use, no when-not-to-use, and no guidance on how it relates to reset_task in the sibling set. An agent must infer the workflow position.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
7 tool updates
v0.1.0- First observed
get_session_info - First observed
list_tasks - First observed
move_relative - First observed
observe - First observed
reset_task - First observed
set_gripper - First observed
stop_episode
TDQS
Scored across 7 tools
The tools map cleanly to distinct phases of the simulation loop: start (reset_task), observe (observe), act (move_relative, set_gripper), end (stop_episode), and introspect (get_session_info, list_tasks). There is minor overlap in that observe and get_session_info both return robot state, but their descriptions clearly separate their roles (images during a run vs. run metadata callable anytime).
Six of seven tools follow a verb_noun pattern (reset_task, move_relative, set_gripper, stop_episode, get_session_info, list_tasks). observe is a bare verb without a noun, a minor deviation, but overall the naming is predictable and readable.
Seven tools exactly cover the full episode lifecycle without redundancy or bloat. Every tool earns its place: reset, observe, two action primitives, stop, session info, and task listing are all necessary for the stated purpose.
The surface covers task discovery, episode reset, observation, action, episode termination, and session introspection, which is the core loop. A minor gap is the lack of an explicit reward/success query, though stop_episode artifacts (summary, events) likely provide results for agents able to read them.
Maintenance
Related MCP Connectors
- DazbenchOAuthapp.dazbench
Task management your AI agents can actually run. One line becomes a context-ready task over MCP.
Remote MCP learning coach for coding agents.
Agent Replay Debugger MCP — record every agent step + deterministic replay. Step-debugger for
MCP server for Clipkit — gives AI agents a video toolbox via the Clipkit schema.
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceA robot-control MCP server enabling LLMs to control Unitree robots and DJI Tello drones with mock and hardware backends, built-in routines, and sequence execution.12MIT
- AlicenseNot gradedqualityAmaintenanceAn MCP server that lets AI agents control iOS and Android devices (tap, scroll, type, take screenshots, read UI trees, and run code). Works with multiple devices at the same time.102 npm47MIT
- AlicenseBqualityAmaintenanceGive any LLM agent a real Android or iPhone. 62 MCP tools: tap, swipe, type, screenshot, screen-tree reading, app launch, camera, TTS, crash reports, batched execution. Android via ADB, iPhone via WebDriverAgent, on-device inference, Docker+KVM emulators. Works with Claude Code, Cursor, LangChain, LlamaIndex, and any MCP client. MIT.6694 PyPI370MIT
- FlicenseNot gradedqualityCmaintenanceMCP server for controlling a simulated robot arm with vision-based pick-and-place, driven by LLM or manual control.1-