bunkervm
Provides a LangChain-compatible tool that enables LangChain agents to run code in isolated microVMs with automatic recording, snapshot, and replay.
Allows LangGraph agents to run code in BunkerVM sandboxes with recording, rewinding, and diffing capabilities.
Provides an OpenAI Agents SDK tool that lets OpenAI agents execute code in isolated microVMs with time-travel debugging and state restoration.
Click on "Install Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@bunkervmdiff the last two sandbox sessions"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
That's a real run: three commands mutate x to 1100, one line rewinds the sandbox to step 2, and x is 11 again — actual VM memory restored, not the script re-executed. That's what this repo does.
Is this for you?
If you're building with LangChain, LangGraph, the OpenAI Agents SDK, or an MCP client like Claude Desktop or VS Code Copilot, and you've ever asked "wait, what did the agent actually do right before it broke?" — yes. BunkerVM gives every sandboxed run a rewind button and a diff tool, on your own machine, for free.
If you need managed infrastructure for thousands of concurrent sandboxes, this isn't that — see Why not E2B / Daytona / Modal? below.
Related MCP server: declaw-mcp-server
The problem
AI agents execute code on your machine. When something goes wrong — and it will — you have no way to see what the agent actually did, rewind to the moment before it broke, or compare why one agent succeeded and another failed.
Containers share your kernel (escapes are real).
Cloud sandboxes send your data to someone else's server.
Neither gives you observability into agent behaviour.
BunkerVM solves all three: isolation, observability, and time-travel.
What it does
Each sandbox is a Firecracker microVM — the same technology behind AWS Lambda. Own kernel, own filesystem, hardware-level (KVM) isolation. Not a container.
On top of that, BunkerVM adds capabilities that no other sandbox provides:
Record every execution
from bunkervm import Sandbox
with Sandbox(record=True) as sb:
sb.run("import pandas as pd")
sb.run("df = pd.read_csv('/data/input.csv')")
sb.run("df['total'] = df.price * df.qty")
sb.run("df.to_csv('/output/result.csv')")
# Every step recorded: command, output, filesystem changes, VM snapshotRewind to any point
sb.restore(step=2) # VM state rewinds to after read_csv
sb.run("df.describe()") # explore from that exact pointThe VM's memory, CPU registers, filesystem — everything reverts to exactly what it was after step 2. Not a re-run. An actual restore from a Firecracker snapshot.
See what changed
for cp in sb.history():
print(f"step {cp['step']}: {cp['command']}")
if cp['trace']:
for f in cp['trace']['files_created']:
print(f" + {f['path']} ({f['size']} bytes)")step 1: import pandas as pd
step 2: df = pd.read_csv('/data/input.csv')
~ /data/input.csv (read)
step 3: df['total'] = df.price * df.qty
step 4: df.to_csv('/output/result.csv')
+ /output/result.csv (1247 bytes)Compare two agents
Every recorded session gets an auto-generated ID (printed when the sandbox exits, or via sb.session_id). Run the same task through two agents, then:
bunkervm diff d0c13cb74d85 f29a61bb02e7Agent Diff
Session A: d0c13cb74d85 (12 steps, 3400ms)
Session B: f29a61bb02e7 (8 steps, 1200ms)
Files only in A: /tmp/debug.log, /tmp/retry_3.py
Files only in B: /output/result.csv
step 1 [same] import pandas as pd
step 2 [same] df = pd.read_csv('/data/input.csv')
step 3 [diff]
A: df = df.dropna()
B: df = df.fillna(0)
step 4 [diff]
A: # crashed — KeyError: 'total'
B: df['total'] = df.price * df.qty ← OKAgent A dropped rows and lost a required column. Agent B filled missing values and succeeded. Without diff, you'd never know why.
Quick start
pip install bunkervmfrom bunkervm import run_code
result = run_code("print('Hello from a microVM!')")
print(result) # Hello from a microVM!VM boots, code runs, VM dies. Your host was never touched.
How it works
AI Agent
│
▼
bunkervm (host) ──vsock──▶ Firecracker MicroVM
│ ┌────────────────────┐
│ record=True │ Alpine Linux │
│ ─────────▶ │ Own kernel │
│ snapshot() │ exec_agent.py │
│ trace() │ (filesystem trace) │
│ restore() └────────────────────┘
│ KVM hardware isolation
▼
~/.bunkervm/sessions/ ~/.bunkervm/snapshots/
d0c13cb74d85.json d0c13cb74d85-step1/ vmstate + memory
d0c13cb74d85-step2/ vmstate + memoryFirecracker provides the isolation. BunkerVM adds the instrumentation layer:
Layer | What it does |
exec_agent (inside VM) | Traces filesystem changes per command — files created, modified, deleted, bytes written |
Firecracker API (host→VM) | Pauses VM, snapshots CPU + memory state to disk, resumes — all via Firecracker's built-in snapshot API |
Snapshot manager (host) | Stores and indexes snapshots at |
Session recorder (host) | Chains commands → traces → snapshots into a replayable session JSON |
No custom kernel modules. No eBPF. No ptrace. The VM is the isolation boundary; the API socket is the control plane. Pure Python, stdlib-only transport.
Named checkpoints & replaying a session
restore(step=N) rewinds to an auto-recorded step. For a checkpoint you want to name and return to deliberately — e.g. right after a slow setup step — use checkpoint():
with Sandbox() as sb:
sb.run("import torch; model = torch.load('bert.pt')")
sb.checkpoint("model-loaded") # snapshot: 45ms
sb.run("output = model(bad_input)") # crashes
sb.restore(step=1) # restore: <100ms
sb.run("output = model(good_input)") # worksEvery record=True session is saved to ~/.bunkervm/sessions/<id>.json on exit and can be replayed from the CLI, independent of the process that created it:
bunkervm replay d0c13cb74d85 --traceSession: d0c13cb74d85
Steps: 5
Recorded: 2026-03-29 23:15
step 1 [ok] 34ms x = 42
step 2 [ok] 23ms print(x * 2)
step 3 [ok] 22ms import os; os.makedirs('/tmp/output', exist_ok=True)
step 4 [ok] 21ms open('/tmp/output/result.txt', 'w').write(str(x))
step 5 [ok] 21ms print(open('/tmp/output/result.txt').read())Why not E2B / Daytona / Modal?
Those are hosted sandbox platforms — good at giving your agent a place to run. BunkerVM is a local, self-hosted debugger for whatever sandbox your agent already runs in. As of writing, none of the major hosted sandboxes ship automatic action recording, mid-session VM snapshot/restore, and cross-run diffing together:
BunkerVM | E2B / Daytona / Modal | |
Isolation | Firecracker microVM (hardware/KVM) | Firecracker or container, depending on provider |
Hosting | Local, self-hosted — nothing leaves your machine | Cloud-hosted |
Auto-records every command | ✅ | ❌ (manual snapshot primitives at best) |
Mid-session restore | ✅ full VM state (memory + fs) | Fork-from-snapshot, not automatic rewind |
Diff two agent runs | ✅ | ❌ |
Cost | Free, open source | Usage-billed |
Trade-off: you run it on your own machine (needs /dev/kvm or WSL2), and it won't scale to thousands of concurrent sandboxes the way a hosted platform will. If you need managed multi-tenant infra, use one of those. If you need to see exactly what your agent did and rewind to before it broke, that's what this is for.
Integrations
MCP (Claude Desktop, VS Code Copilot, any MCP client)
bunkervm vscode-setup # generates .vscode/mcp.json, works on Windows WSL2
bunkervm server # stdio for Claude Desktop
bunkervm server --transport sse # SSE for web8 MCP tools: sandbox_exec, sandbox_write_file, sandbox_read_file, sandbox_list_dir, sandbox_upload_file, sandbox_download_file, sandbox_status, sandbox_reset.
Any agent framework
secure_agent() wraps a single-tool adapter around whatever you already have, no BunkerVM-specific toolkit required:
from bunkervm import secure_agent
runtime = secure_agent()
tool = runtime.as_tool() # LangChain-compatible tool (requires langchain-core)
tool = runtime.as_openai_tool() # OpenAI Agents SDK tool (requires openai-agents)Install
pip install bunkervmRequirements: Linux with /dev/kvm, or Windows WSL2 (enable nested virtualization). Python 3.10+.
The Firecracker binary + kernel + rootfs (~100MB) auto-download on first run. Or download from Releases.
Add to %USERPROFILE%\.wslconfig:
[wsl2]
nestedVirtualization=trueThen: wsl --shutdown
Problem | Fix |
|
|
Permission denied |
|
Bundle download fails | Manual download from Releases → |
VM won't start |
|
git clone https://github.com/ashishgituser/bunkervm.git
cd bunkervm
sudo bash build/setup-firecracker.sh
sudo bash build/build-sandbox-rootfs.sh
pip install -e ".[dev]"
pytest tests/CLI
bunkervm demo # see it in action
bunkervm run script.py # run a script in a sandbox
bunkervm run -c "print(42)" # inline code
bunkervm replay <session-id> --trace # replay recorded session
bunkervm diff <session-a> <session-b> # compare two agent runs
bunkervm snapshot list # list VM snapshots
bunkervm snapshot delete <name> # delete a snapshot
bunkervm server --transport sse # MCP server
bunkervm info # system readiness checkContributing
See CONTRIBUTING.md.
Security
See SECURITY.md.
License
MIT
This server cannot be installed
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Servers
- AlicenseAqualityBmaintenanceConnects AI agents to microsandbox for creating lightweight sandboxes, executing code, managing files, and monitoring resources.1918312Apache 2.0

declaw-mcp-serverofficial
Alicense-qualityBmaintenanceRuns AI-generated code in secure Firecracker microVMs with opt-in network policy enforcement, PII scanning, prompt injection defense, and audit logging. Exposes MCP tools for running commands, managing files, and the full sandbox lifecycle.301Apache 2.0- Alicense-qualityBmaintenanceEnables AI agents to create, manage, and execute code in isolated Firecracker microVM sandboxes via the MCP protocol, with support for sandbox lifecycle and file operations.2Apache 2.0
- Alicense-qualityAmaintenancePersistent, secure LXC sandbox environments for AI agents with native MCP support.269Apache 2.0
Related MCP Connectors
Agent Replay Debugger MCP — record every agent step + deterministic replay. Step-debugger for
Live browser debugging for AI assistants — DOM, console, network via MCP.
OCR, transcription, file extraction, and image generation for AI agents via MCP.
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/ashishgituser/bunkervm'
If you have feedback or need assistance with the MCP directory API, please join our Discord server