Skip to main content
Glama
README.md
# agent-runner-mcp

> An MCP server that exposes the three-platform sandboxed runner protocol to **any MCP client** (Claude Code, Codex, and others): run tasks in an evidence-protected sandbox, read the EXIT protocol, and get an autopsy report. Zero dependencies — the MCP layer is hand-rolled. Every claim carries an experiment number.

中文版见 [README.zh-CN.md](./README.zh-CN.md)。

[![license](https://img.shields.io/badge/license-Apache--2.0-blue.svg)](./LICENSE)
[![ci](https://github.com/Wang-Lin-Chang/agent-runner-mcp/actions/workflows/ci.yml/badge.svg)](https://github.com/Wang-Lin-Chang/agent-runner-mcp/actions/workflows/ci.yml)

## Why this exists

Agent frameworks each go their own way, but "run tasks reliably in a sandbox with evidence left behind" is a common hard need. This server wraps [dsh-witness](https://github.com/Wang-Lin-Chang/dsh-witness)'s runner protocol (lock=pid:startSec, EXIT:<code>, the autopsy taxonomy) into MCP tools — **Claude Code measured ✓ Connected**; Codex and other MCP clients can connect over the same protocol.

## Tools

| Tool | Semantics |
|---|---|
| `task_run` | Sandboxed command execution (Windows ACL / Linux bwrap / macOS sandbox-exec chosen per platform), returns the task directory |
| `task_wait` | Wait for the terminal state (EXIT:<code> written) |
| `task_output` | Incrementally read out.log by byte offset |
| `task_autopsy` | Generate the autopsy report (autopsy-spec format: manner/evidence/verdict/D-01~D-09) |
| `task_kill` | Kill by lock pid (crash experiments) |
| `task_adopt` | Three-evidence adoption adjudication (lock parsing + process liveness + exit protocol) |

## Quick start

```sh
# Install (git source)
dsh plugin --profile <name> add "github:Wang-Lin-Chang/agent-runner-mcp#v0.1.1"

# Register with Claude Code
claude mcp add agent-runner -- node dist/server.js

# Or run the tests directly (15-assertion protocol measurement)
npm test
```

## Acceptance evidence

- EXP-1: MCP protocol layer measured 15/15 (handshake / six tools / EXIT:0→D-01 / crash→D-08 / EXIT:1→D-02)
- EXP-2: Claude Code 2.1.92 real client `✓ Connected` (initialize + tools/list handshake)
- EXP-3: Windows ACL persistence verdict (observed artifacts separated from the evidence zone)
- EXP-4: Death-semantics matrix aligned with the autopsy-spec taxonomy

## Honest boundaries

- Claude Code in-session tool calls need a login state — this machine is not logged in: the protocol handshake is measured, session calls are unmeasured, not claimed.
- This server is the **runner protocol layer**: full registry-level adoption/event-sourcing/caching lives in dsh-witness.
- Under the Windows CI admin environment the ACL sandbox does not apply → the runner's fail-closed (EXIT:-998) is the protocol-correct response; sandbox capabilities are measured in non-admin environments.
- Offline applicability: architecturally no network dependency (local processes + file protocol); multi-day offline runs are unmeasured.

## License

Apache-2.0

TDQS

A3.9/5.0

Scored across 6 tools

Disambiguation5/5

Each tool has a clearly distinct purpose: run starts a task, wait blocks for completion, output reads logs, autopsy analyzes results, kill terminates, and adopt verifies state. No two tools overlap in function; even wait and adopt differ in blocking vs. check semantics.

Naming Consistency5/5

All tools follow the exact same pattern: task_ followed by a single verb (run, wait, output, autopsy, kill, adopt). This is perfectly consistent and predictable.

Tool Count5/5

With 6 tools, the server is well-scoped for task lifecycle management. Each tool earns its place, covering the core operations without redundancy.

Completeness5/5

The tool set covers the full task lifecycle: create/run, wait for completion, read output, kill, analyze, and adopt. No obvious gaps for the stated purpose of running commands in a sandbox and managing their results.

Maintenance

ActivitySlowing
ResponsivenessNo issues