Skip to main content
Glama
wesleyzhangwq

tiny-agent-sandbox

README.md
# Tiny Agent Sandbox

A small, inspectable Docker-backed sandbox that agents can call as an
[MCP](https://modelcontextprotocol.io/) tool.

It runs Python, JavaScript, or shell snippets in a new container with no network, no host mounts, a
read-only root filesystem, a non-root user, dropped capabilities, and hard resource/output limits.
Results are returned as structured data for an agent to reason about.

> [!WARNING]
> This is a learning and local-development project, not a production multi-tenant security boundary.
> Docker containers share a kernel. Read [SECURITY.md](SECURITY.md) before using untrusted code.

## What the agent sees

The MCP server exposes two tools:

- `sandbox_status()` reports Docker readiness and missing runtime images.
- `run_code(language, code, timeout_seconds)` returns `stdout`, `stderr`, exit status, timeout and
  truncation flags, duration, and the runtime image.

The agent cannot choose an image, command, mount, environment variable, network policy, or Docker
flag. Those remain operator-controlled policy.

## Quick start

Prerequisites: Docker, Python 3.11+, and [uv](https://docs.astral.sh/uv/).

```bash
git clone https://github.com/wesleyzhangwq/tiny-agent-sandbox.git
cd tiny-agent-sandbox

docker pull python:3.12-alpine
docker pull node:22-alpine
docker pull alpine:3.22

uv sync --no-editable
uv run --no-editable tas doctor
uv run --no-editable tas run python -c 'print(sum(range(10)))'
```

Start the MCP server over stdio:

```bash
uv run --no-editable tiny-agent-sandbox
```

## Connect an MCP client

For clients that accept the common `mcpServers` JSON shape:

```json
{
  "mcpServers": {
    "tiny-agent-sandbox": {
      "command": "uvx",
      "args": [
        "--from",
        "git+https://github.com/wesleyzhangwq/tiny-agent-sandbox",
        "tiny-agent-sandbox"
      ]
    }
  }
}
```

Pull the three runtime images before starting the client. Tool calls use `--pull=never`, so an agent
cannot cause an image download.

Example tool input:

```json
{
  "language": "python",
  "code": "import statistics\nprint(statistics.mean([2, 4, 9]))",
  "timeout_seconds": 5
}
```

Example structured result:

```json
{
  "language": "python",
  "image": "python:3.12-alpine",
  "exit_code": 0,
  "stdout": "5\n",
  "stderr": "",
  "timed_out": false,
  "output_truncated": false,
  "duration_ms": 180
}
```

## Default policy

| Control | Default |
| --- | --- |
| Network | disabled |
| Root filesystem | read-only |
| Writable storage | 64 MiB ephemeral `/tmp` (`noexec`, `nosuid`) |
| User | UID/GID 65534 |
| Linux capabilities | all dropped |
| Privilege escalation | disabled |
| Memory / swap | 256 MiB / 256 MiB |
| CPU | 1 core |
| Processes | 64 |
| File descriptors | 64 |
| Wall time | 5 seconds default, 30 seconds maximum |
| Input / combined output | 64 KiB / 128 KiB |

The fixed policy is assembled in
[`src/tiny_agent_sandbox/runner.py`](src/tiny_agent_sandbox/runner.py). Timeout and output overflow
remove the named container rather than merely terminating the Docker client.

## Development

```bash
uv sync --no-editable --extra dev
uv run --no-editable ruff check .
uv run --no-editable pytest -m "not integration"

docker pull python:3.12-alpine node:22-alpine alpine:3.22
uv run --no-editable pytest -m integration
```

The preliminary ecosystem notes and design tradeoffs are in
[`docs/research.md`](docs/research.md).

## Roadmap

- Optional gVisor (`runsc`) backend with explicit runtime detection.
- Digest-pinned, operator-configurable runtime images.
- Per-session workspaces with strict size and lifetime limits.
- Egress proxy with destination allowlists and audit logs.
- Concurrency quotas and OpenTelemetry execution traces.

## License

Apache-2.0

TDQS

A3.8/5.0

Scored across 2 tools

Disambiguation4/5

The two tools serve clearly distinct purposes: run_code executes code while sandbox_status reports environment readiness. There is no meaningful overlap between them, so an agent can easily distinguish which to use.

Naming Consistency5/5

Both tools follow a consistent verb_noun pattern: run_code and sandbox_status. The naming style is uniform snake_case with clear, descriptive verbs and nouns.

Tool Count3/5

With only 2 tools, the surface feels thin for a sandbox server that presumably supports code execution. A richer surface might include file operations, image management, or environment configuration, but for a narrowly-scoped execution sandbox two tools can be reasonable.

Completeness3/5

The core execute-and-check lifecycle is covered: run code and check readiness. However, there are notable gaps such as no image listing, no stop/cleanup of sandboxes, and no way to provision or configure runtimes beyond status reporting.

Maintenance

ActivitySlowing
ResponsivenessNo issues