tiny-agent-sandbox
# Tiny Agent Sandbox
A small, inspectable Docker-backed sandbox that agents can call as an
[MCP](https://modelcontextprotocol.io/) tool.
It runs Python, JavaScript, or shell snippets in a new container with no network, no host mounts, a
read-only root filesystem, a non-root user, dropped capabilities, and hard resource/output limits.
Results are returned as structured data for an agent to reason about.
> [!WARNING]
> This is a learning and local-development project, not a production multi-tenant security boundary.
> Docker containers share a kernel. Read [SECURITY.md](SECURITY.md) before using untrusted code.
## What the agent sees
The MCP server exposes two tools:
- `sandbox_status()` reports Docker readiness and missing runtime images.
- `run_code(language, code, timeout_seconds)` returns `stdout`, `stderr`, exit status, timeout and
truncation flags, duration, and the runtime image.
The agent cannot choose an image, command, mount, environment variable, network policy, or Docker
flag. Those remain operator-controlled policy.
## Quick start
Prerequisites: Docker, Python 3.11+, and [uv](https://docs.astral.sh/uv/).
```bash
git clone https://github.com/wesleyzhangwq/tiny-agent-sandbox.git
cd tiny-agent-sandbox
docker pull python:3.12-alpine
docker pull node:22-alpine
docker pull alpine:3.22
uv sync --no-editable
uv run --no-editable tas doctor
uv run --no-editable tas run python -c 'print(sum(range(10)))'
```
Start the MCP server over stdio:
```bash
uv run --no-editable tiny-agent-sandbox
```
## Connect an MCP client
For clients that accept the common `mcpServers` JSON shape:
```json
{
"mcpServers": {
"tiny-agent-sandbox": {
"command": "uvx",
"args": [
"--from",
"git+https://github.com/wesleyzhangwq/tiny-agent-sandbox",
"tiny-agent-sandbox"
]
}
}
}
```
Pull the three runtime images before starting the client. Tool calls use `--pull=never`, so an agent
cannot cause an image download.
Example tool input:
```json
{
"language": "python",
"code": "import statistics\nprint(statistics.mean([2, 4, 9]))",
"timeout_seconds": 5
}
```
Example structured result:
```json
{
"language": "python",
"image": "python:3.12-alpine",
"exit_code": 0,
"stdout": "5\n",
"stderr": "",
"timed_out": false,
"output_truncated": false,
"duration_ms": 180
}
```
## Default policy
| Control | Default |
| --- | --- |
| Network | disabled |
| Root filesystem | read-only |
| Writable storage | 64 MiB ephemeral `/tmp` (`noexec`, `nosuid`) |
| User | UID/GID 65534 |
| Linux capabilities | all dropped |
| Privilege escalation | disabled |
| Memory / swap | 256 MiB / 256 MiB |
| CPU | 1 core |
| Processes | 64 |
| File descriptors | 64 |
| Wall time | 5 seconds default, 30 seconds maximum |
| Input / combined output | 64 KiB / 128 KiB |
The fixed policy is assembled in
[`src/tiny_agent_sandbox/runner.py`](src/tiny_agent_sandbox/runner.py). Timeout and output overflow
remove the named container rather than merely terminating the Docker client.
## Development
```bash
uv sync --no-editable --extra dev
uv run --no-editable ruff check .
uv run --no-editable pytest -m "not integration"
docker pull python:3.12-alpine node:22-alpine alpine:3.22
uv run --no-editable pytest -m integration
```
The preliminary ecosystem notes and design tradeoffs are in
[`docs/research.md`](docs/research.md).
## Roadmap
- Optional gVisor (`runsc`) backend with explicit runtime detection.
- Digest-pinned, operator-configurable runtime images.
- Per-session workspaces with strict size and lifetime limits.
- Egress proxy with destination allowlists and audit logs.
- Concurrency quotas and OpenTelemetry execution traces.
## License
Apache-2.0
TDQS
Scored across 2 tools
The two tools serve clearly distinct purposes: run_code executes code while sandbox_status reports environment readiness. There is no meaningful overlap between them, so an agent can easily distinguish which to use.
Both tools follow a consistent verb_noun pattern: run_code and sandbox_status. The naming style is uniform snake_case with clear, descriptive verbs and nouns.
With only 2 tools, the surface feels thin for a sandbox server that presumably supports code execution. A richer surface might include file operations, image management, or environment configuration, but for a narrowly-scoped execution sandbox two tools can be reasonable.
The core execute-and-check lifecycle is covered: run code and check readiness. However, there are notable gaps such as no image listing, no stop/cleanup of sandboxes, and no way to provision or configure runtimes beyond status reporting.