Skip to main content
Glama

ABC Inspect Playground

Control a simulated robot in Chrome or through a TUI agent using MCP. Runs locally with ABC Sim, MuJoCo, and Inspect Robots—no robot hardware or policy weights needed.

Roadmap · Integration measurements · AGPL-3.0-or-later

ABC Sim playground with three camera views and keyboard controls.

Quick start

Requires Git, uv, Chrome, and a graphics context for MuJoCo cameras. Tested on macOS with an Apple M3 and 16 GiB RAM; other platforms are unverified.

git clone https://github.com/kimjune01/abc-inspect-playground.git
cd abc-inspect-playground
bash scripts/bootstrap.sh  # downloads pinned dependencies and assets
bash scripts/serve.sh

Open http://127.0.0.1:8876/play and click Start. If the server is already running, reuse it.

Control

Action

W / S

Forward / back

A / D

Left / right

↑ / ↓

Up / down

← / →

Yaw left / right

Space

Open / close gripper

R / arm selector

Swap left and right arms

Evaluate

End the attempt and score the result

Reset

Start a fresh scene

Escape

Stop sending commands

Hold keys to move, or use the on-screen buttons. Releasing keys or leaving the window stops new commands; an in-flight command completes. Cameras refresh after each command, and physics pauses between commands. Translation uses world axes; the wrist retains its tilt. Pitch/roll controls are not implemented.

The status shows the remaining step budget: 1,000 ticks for box tasks (about 34 simulated seconds), 5,000 for letter blocks (about 170 seconds). Trials end with a “Step limit reached” message. Reset to continue; restart the server after 256 starts. The eight most recent trials remain available through the API; older logs stay on disk. After restarting, refresh the page and click Start.

Choose a scene and task in the navigation bar. Switching saves the current attempt and starts a fresh one.

Scene

Tasks

Success criterion

Opaque box

One, two, or three objects

Exactly that many objects inside, through the flap

Letter blocks

CAT, DOG, FISH

Correct letters in a tight left-to-right row on the table, faces up

Evaluate ends the attempt without advancing physics. Grey means unscored, green Pass · 1/1, red Fail · 0/1, amber unavailable. Let objects finish falling before evaluating. Scores appear only after the trial ends.

These are ABC's documented task variants and native evaluators. The pinned ABC randomizer overrides fixed counting goals, so our adapter restores the declared one/two/three-object directive through its evaluator configuration hook. Geometry and scoring logic are unchanged.

Related MCP server: VELMA

Agent control

MCP endpoint: http://127.0.0.1:8876/mcp. Browser and agents share one active trial, so coordinate control.

.codex/config.toml registers abc_sim for a trusted project session. Native TUI discovery remains unverified; a subagent has driven the simulator using this protocol helper:

uv run python -m abc_inspect.client start_trial \
  '{"request_id":"trial-1","seed":7,"max_steps":100,"task":"count_into_opaque_box"}'

The response includes metadata and paths to three camera PNGs. View them, then use the returned session ID and sequence:

uv run python -m abc_inspect.client jog_arm \
  '{"session_id":"SESSION_ID","request_id":"move-1","expected_sequence":0,"translation":[0.01,0,0]}'

uv run python -m abc_inspect.client finish_trial \
  '{"session_id":"SESSION_ID","request_id":"finish-1","expected_sequence":1,"reason":"done"}'

Tool

Purpose

start_trial

Create a trial and return state and cameras

observe

Read state without advancing physics

jog_arm

Cartesian translation, yaw, and gripper control through IK

move_joints

Apply 14 absolute joint/gripper targets

finish_trial

Finish without moving; return metrics and log path

  • Tasks: count_one_into_opaque_box, count_two_into_opaque_box, count_three_into_opaque_box, spell_cat, spell_dog, spell_fish; also count_into_opaque_box (sampled goal) and put_plastic_bottles_in_bin (MCP default).

  • Observations: cameras, robot joints and grasp-site poses, limits, sequence, and time. No live object poses or task scores.

  • Jog limits: translation ≤0.026 m, yaw ±0.12 rad. Gripper: 0 closed, 1 open, null retains its command. Joint commands allow changes ≤0.2 rad and ≤0.25 gripper range.

  • Timing: 1–30 ticks per action at about 29.41 simulated Hz; 15-minute idle timeout.

  • Retries: use fresh request IDs and the latest sequence. After a timeout, retry the same ID and arguments. Cached responses can be old; observe before the next decision.

  • Lifecycle: reconnects preserve trials; server restarts invalidate them. Retry images live on disk, and finished workers release their resources.

See server.py for tool signatures and defaults.

Recording and evaluation

Inspect owns every reset and physics step. Commands drive real MuJoCo actuators; Mink supplies inverse kinematics. Trials save native JSON, camera frames, and tool transcripts under outputs/trials/. Full agent conversations are not captured.

eval_status: success means the run completed; abc_success is the task score. This is a development sandbox: local agents still have shell access. A controlled human–model comparison also needs matched inputs, budgets, and held-out trials.

Scripted pick-and-place

uv run python -m abc_inspect.pick_place --output outputs/pick-place/run
uv run python scripts/render_replay.py outputs/pick-place/run

The renderer requires ffmpeg. Open outputs/demo/index.html; use a fresh run directory each time. This is a recording, separate from live play.

The script uses known object geometry and Mink IK to lift a cube into the box. pick_place=1 verifies a lift over 12 cm, an open gripper, and the cube inside the box. ABC's separate counting objective may still score zero. One fixed seed is tested; this does not establish vision-agent competence or reliable bottle handling.

Development

uv run pytest -q
uv run ruff check src tests
uv run mypy src/abc_inspect

The 37 tests cover physics, cameras, resets, scoring, retries, concurrency, transports, scene/task switching, keyboard-control endpoints, and resource cleanup. Source lives in src/abc_inspect/; local artifacts in outputs/ are ignored.

The UI follows June’s design notes: explicit states, feedback at the point of action, grouped controls, and semantic color.

Ideas and attribution

This playground connects existing robotics and evaluation work:

  • ABC / ABC Sim — the bimanual robot environment, cameras, objects, task definitions, and success evaluators. All six browser tasks come from the pinned ABC task catalog, using its counting and spelling scorers.

  • Inspect Robots — separating policy, embodiment, task, and scorer; the actual rollout loop and evaluation records used here. The Evaluate button ends that rollout and invokes its scorer.

  • RoboDojo — inspiration for a shared suite of manipulation tasks and reproducible policy comparisons. This playground uses ABC tasks; it does not run or reproduce RoboDojo's official benchmark.

  • MuJoCo and Mink — physical simulation and inverse kinematics for Cartesian keyboard/agent control.

  • Model Context Protocol — the tool interface connecting TUI agents to a persistent simulator. Related implementations are credited in DERISK.md.

Our contribution is the ABC–Inspect adapter, queued external policy, session/retry protocol, and shared browser/MCP controls. ABC and Inspect are pinned in scripts/bootstrap.sh and pyproject.toml; Python dependencies are locked in uv.lock.

The letter-block artwork is by Cherryvania, under CC BY 4.0, distributed through ABC. See ABC's asset credits for other upstream models and licenses.

License

Copyright (C) 2026 June and contributors. Original code is licensed under AGPL-3.0-or-later. Dependencies and rendered third-party assets retain their respective licenses; downloaded upstream sources and assets are not bundled.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    B
    maintenance
    An MCP server that enables LLM agents to control a two-joint MuJoCo robot arm through tools for moving, reading state, and resetting. It deliberately omits inverse kinematics, so the agent must infer joint angles from observations.
    MIT
  • F
    license
    Not graded
    quality
    C
    maintenance
    Enables agents to safely drive a simulated Franka Panda pick-and-place cell through MCP tools, with gated motion, emergency stop, structured errors, and audit logging.
    -