abc_sim
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@abc_simstart a trial to put plastic bottles in the bin and show me the camera views"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
ABC Inspect Playground
Control a simulated robot in Chrome or through a TUI agent using MCP. Runs locally with ABC Sim, MuJoCo, and Inspect Robots—no robot hardware or policy weights needed.
Roadmap · Integration measurements · AGPL-3.0-or-later

Quick start
Requires Git, uv, Chrome, and a graphics context for MuJoCo cameras. Tested on
macOS with an Apple M3 and 16 GiB RAM; other platforms are unverified.
git clone https://github.com/kimjune01/abc-inspect-playground.git
cd abc-inspect-playground
bash scripts/bootstrap.sh # downloads pinned dependencies and assets
bash scripts/serve.shOpen http://127.0.0.1:8876/play and click Start. If the server is already running, reuse it.
Control | Action |
W / S | Forward / back |
A / D | Left / right |
↑ / ↓ | Up / down |
← / → | Yaw left / right |
Space | Open / close gripper |
R / arm selector | Swap left and right arms |
Evaluate | End the attempt and score the result |
Reset | Start a fresh scene |
Escape | Stop sending commands |
Hold keys to move, or use the on-screen buttons. Releasing keys or leaving the window stops new commands; an in-flight command completes. Cameras refresh after each command, and physics pauses between commands. Translation uses world axes; the wrist retains its tilt. Pitch/roll controls are not implemented.
The status shows the remaining step budget: 1,000 ticks for box tasks (about 34 simulated seconds), 5,000 for letter blocks (about 170 seconds). Trials end with a “Step limit reached” message. Reset to continue; restart the server after 256 starts. The eight most recent trials remain available through the API; older logs stay on disk. After restarting, refresh the page and click Start.
Choose a scene and task in the navigation bar. Switching saves the current attempt and starts a fresh one.
Scene | Tasks | Success criterion |
Opaque box | One, two, or three objects | Exactly that many objects inside, through the flap |
Letter blocks | CAT, DOG, FISH | Correct letters in a tight left-to-right row on the table, faces up |
Evaluate ends the attempt without advancing physics. Grey means unscored, green Pass · 1/1, red Fail · 0/1, amber unavailable. Let objects finish falling before evaluating. Scores appear only after the trial ends.
These are ABC's documented task variants and native evaluators. The pinned ABC randomizer overrides fixed counting goals, so our adapter restores the declared one/two/three-object directive through its evaluator configuration hook. Geometry and scoring logic are unchanged.
Related MCP server: VELMA
Agent control
MCP endpoint: http://127.0.0.1:8876/mcp. Browser and agents share one active trial, so coordinate control.
.codex/config.toml registers abc_sim for a trusted project session. Native TUI
discovery remains unverified; a subagent has driven the simulator using this
protocol helper:
uv run python -m abc_inspect.client start_trial \
'{"request_id":"trial-1","seed":7,"max_steps":100,"task":"count_into_opaque_box"}'The response includes metadata and paths to three camera PNGs. View them, then use the returned session ID and sequence:
uv run python -m abc_inspect.client jog_arm \
'{"session_id":"SESSION_ID","request_id":"move-1","expected_sequence":0,"translation":[0.01,0,0]}'
uv run python -m abc_inspect.client finish_trial \
'{"session_id":"SESSION_ID","request_id":"finish-1","expected_sequence":1,"reason":"done"}'Tool | Purpose |
| Create a trial and return state and cameras |
| Read state without advancing physics |
| Cartesian translation, yaw, and gripper control through IK |
| Apply 14 absolute joint/gripper targets |
| Finish without moving; return metrics and log path |
Tasks:
count_one_into_opaque_box,count_two_into_opaque_box,count_three_into_opaque_box,spell_cat,spell_dog,spell_fish; alsocount_into_opaque_box(sampled goal) andput_plastic_bottles_in_bin(MCP default).Observations: cameras, robot joints and grasp-site poses, limits, sequence, and time. No live object poses or task scores.
Jog limits: translation ≤0.026 m, yaw ±0.12 rad. Gripper:
0closed,1open,nullretains its command. Joint commands allow changes ≤0.2 rad and ≤0.25 gripper range.Timing: 1–30 ticks per action at about 29.41 simulated Hz; 15-minute idle timeout.
Retries: use fresh request IDs and the latest sequence. After a timeout, retry the same ID and arguments. Cached responses can be old; observe before the next decision.
Lifecycle: reconnects preserve trials; server restarts invalidate them. Retry images live on disk, and finished workers release their resources.
See server.py for tool signatures and defaults.
Recording and evaluation
Inspect owns every reset and physics step. Commands drive real MuJoCo actuators;
Mink supplies inverse kinematics. Trials save native JSON, camera frames, and tool
transcripts under outputs/trials/. Full agent conversations are not captured.
eval_status: success means the run completed; abc_success is the task score.
This is a development sandbox: local agents still have shell access. A controlled
human–model comparison also needs matched inputs, budgets, and held-out trials.
Scripted pick-and-place
uv run python -m abc_inspect.pick_place --output outputs/pick-place/run
uv run python scripts/render_replay.py outputs/pick-place/runThe renderer requires ffmpeg. Open outputs/demo/index.html; use a fresh run
directory each time. This is a recording, separate from live play.
The script uses known object geometry and Mink IK to lift a cube into the box.
pick_place=1 verifies a lift over 12 cm, an open gripper, and the cube inside the
box. ABC's separate counting objective may still score zero. One fixed seed is
tested; this does not establish vision-agent competence or reliable bottle handling.
Development
uv run pytest -q
uv run ruff check src tests
uv run mypy src/abc_inspectThe 37 tests cover physics, cameras, resets, scoring, retries, concurrency,
transports, scene/task switching, keyboard-control endpoints, and resource cleanup. Source lives in
src/abc_inspect/; local artifacts in outputs/ are ignored.
The UI follows June’s design notes: explicit states, feedback at the point of action, grouped controls, and semantic color.
Ideas and attribution
This playground connects existing robotics and evaluation work:
ABC / ABC Sim — the bimanual robot environment, cameras, objects, task definitions, and success evaluators. All six browser tasks come from the pinned ABC task catalog, using its counting and spelling scorers.
Inspect Robots — separating policy, embodiment, task, and scorer; the actual rollout loop and evaluation records used here. The Evaluate button ends that rollout and invokes its scorer.
RoboDojo — inspiration for a shared suite of manipulation tasks and reproducible policy comparisons. This playground uses ABC tasks; it does not run or reproduce RoboDojo's official benchmark.
MuJoCo and Mink — physical simulation and inverse kinematics for Cartesian keyboard/agent control.
Model Context Protocol — the tool interface connecting TUI agents to a persistent simulator. Related implementations are credited in DERISK.md.
Our contribution is the ABC–Inspect adapter, queued external policy, session/retry
protocol, and shared browser/MCP controls. ABC and Inspect are pinned in
scripts/bootstrap.sh and pyproject.toml; Python dependencies are locked in uv.lock.
The letter-block artwork is by Cherryvania, under CC BY 4.0, distributed through ABC. See ABC's asset credits for other upstream models and licenses.
License
Copyright (C) 2026 June and contributors. Original code is licensed under AGPL-3.0-or-later. Dependencies and rendered third-party assets retain their respective licenses; downloaded upstream sources and assets are not bundled.
This server cannot be deployed
Maintenance
Related MCP Connectors
Read-only MCP server for the OPERANT AI operating-agent calibration benchmark.
Hosted MCP endpoint with realistic fake data for prototyping agents. 12 tools, no setup.
- FensoryOAuthcom.fensory
Trading MCP server for AI agents, with live market data, account reads and controlled execution.
Related MCP Servers
- AlicenseNot gradedqualityBmaintenanceMCP server for AI agents to drive Gazebo / gz-sim simulation, with offline mock mode for CI/demos.4MIT
- FlicenseNot gradedqualityBmaintenanceMCP server for controlling a simulated robot arm with vision-based pick-and-place, driven by LLM or manual control.1-
- AlicenseNot gradedqualityBmaintenanceAn MCP server that enables LLM agents to control a two-joint MuJoCo robot arm through tools for moving, reading state, and resetting. It deliberately omits inverse kinematics, so the agent must infer joint angles from observations.MIT
- FlicenseNot gradedqualityCmaintenanceEnables agents to safely drive a simulated Franka Panda pick-and-place cell through MCP tools, with gated motion, emergency stop, structured errors, and audit logging.-