agent-sandbox
agent-sandbox
AIコーディングエージェントが、実際のKubernetesクラスターに対して実際のインフラストラクチャコマンドを実行できるようにするためにこれを構築しました。ただし、常時有効な認証情報を保持することはなく、監視なしで何かを破壊することもできません。
3つのMCPツール。各呼び出しは、gVisorでサンドボックス化されたKubernetes Job内で実行され、その単一のアクションのためにVaultが発行する短命で狭いスコープの認証情報を使用します。破壊的な操作はすべて、人間の承認ゲートで停止します。
AI Agent (Claude Desktop / Cursor)
│ MCP protocol (stdio)
▼
┌──────────────────────────────────────────┐
│ MCP Server src/agent_sandbox │
│ 3 tools -> guardrails -> broker -> │
│ sandbox -> audit │
└───────┬──────────────────────┬───────────┘
│ │
▼ ▼
┌────────────────┐ ┌──────────────────────────┐
│ Credential │ │ Sandbox Runner │
│ Broker (Vault) │ │ K8s Job + gVisor │
│ 10-min leases │ │ restricted PSS │
│ per-action │ │ default-deny NetworkPolicy│
│ scope │ │ cpu/mem limits, deadline │
└────────────────┘ └──────────────────────────┘
│ │
└──────────┬───────────┘
▼
┌──────────────────────┐
│ Guardrails + Approval│
│ policy.yaml, SQLite, │
│ agent-sandbox CLI │
└──────────────────────┘なぜこれを構築したのか
私が使用していたAIツールが、本番インフラストラクチャに対して、稼働中のリソースの置き換えを強制するTerraformの変更を提案したことがありました。そのプランは一見普通に見えました。失敗の原因はモデルが間違っていたことではなく、もっともらしいプランと破壊的なapplyの間に何も介在していなかったことでした。
このプロジェクトを、欠けていたレイヤーとして、動作するコードで構築しました:
エージェントは再利用可能な認証情報を決して保持しない
すべてがホストに害を及ぼせない場所で実行される
破壊的な変更は停止して人間を待つ
すべてのアクションが記録に残る
make demo は、私が遭遇したまさにそのシナリオを再現します。1行のラベル編集で実行中のDeploymentの置き換えが強制され、ゲートがそれを検出します。
Related MCP server: safe-runbook-mcp
クイックスタート
Docker、kind、kubectl、vault、terraform、Python 3.11+ が必要です。
brew install kind kubectl hashicorp/tap/vault terraform
make up # ~5 minutes from cold: cluster, CNI, gVisor, Vault, image, verify
make demo # the forces-replacement guardrail demo
make down # tear it all downmake up は冪等です。最後に make verify を実行して終了します。これは分離の主張を 証明 するもので、単に主張するだけではありません(下記参照)。
エージェントをそれに向ける
cp examples/claude_desktop_config.json \
~/Library/Application\ Support/Claude/claude_desktop_config.jsonCursor: examples/cursor_mcp.json を .cursor/mcp.json にコピーします。次に、エージェントに "demo-app のポッドステータスを確認" または "k8s-demo terraform をプラン" と依頼します。
3つのツール
ツール | リスク | 動作 |
| 低 | 即座に実行。認証情報は1つのネームスペース内の |
| 低 | 即座に実行。プランを保存し、後続のapplyがレビューされた差分を 正確に 実行するようにします。 |
| 高 |
|
4つのコンポーネント
1. サンドボックス実行 — src/agent_sandbox/sandbox.py
ツール呼び出しごとに使い捨てのJobを1つ作成します。すべての制御には特定の理由があります:
制御 | 防止するもの |
| システムコールはgVisorセントリーに到達し、ホストカーネルには到達しない |
PSS | root、権限昇格、ケーパビリティ、書き込み可能なrootfs |
| サンドボックス内の任意の周囲のクラスターID |
デフォルト拒否の | インターネットへの送信、横移動、メタデータエンドポイント |
| 暴走したジョブがノードを枯渇させたり、永久にハングしたりするのを防ぐ |
| 失敗した破壊的なアクションが静かに再試行されるのを防ぐ |
認証情報は ファイル としてマウントされ、環境変数としては決してマウントされません。環境変数は kubectl describe、/proc、クラッシュダンプを通じて漏洩します。
2. 認証情報ブローカー — src/agent_sandbox/broker.py
cred = broker.issue_scoped_credential("k8s_get_pod_status")
# -> Vault mints a ServiceAccount + Role + RoleBinding, 10-minute lease
# -> revoked immediately after the Job finishesエージェントは自分でスコープを選択できません。 スコープはアクションから導出されます。
デフォルトで拒否。 マッピングされたスコープがないアクションには認証情報が付与されません。
爆発半径が強制されます。 ターゲット以外のネームスペースを要求すると拒否されます。
トークンはモジュールから決して出ません。
Credential.__repr__はtoken=<redacted>を出力するため、偶発的なログでも漏洩しません。
これを手動で検証しました:pod-reader トークンは demo-app 内のポッドを一覧表示でき、kube-system では 拒否 され、シークレットでは 拒否 され、リースが失効した瞬間に動作を停止し、ServiceAccountを残しません。
3. ガードレール — policy/policy.yaml、src/agent_sandbox/guardrails.py
デフォルトで拒否:MCPツールを登録するだけでは、呼び出し可能にはなりません。 ポリシーに存在しないツールは拒否されるため、機能を追加するには意図的なリスク層の決定が必要です。
承認は明らかな攻撃に対して強化されています:
単回使用 — 1つのSQLiteトランザクション内で消費されるため、2つの同時applyが同じ承認を使用できません
パラメータバインド — 正確なツール+パラメータのハッシュにバインドされるため、
k8s-demoの承認をprod-clusterに対して再生できません期限付き — デフォルトで30分
帯域外 — 別のCLIプロセスを介して付与されます。承認するMCPツールはありません。エージェントには自分の要求を承認するコードパスがありません。
4. MCPサーバー — src/agent_sandbox/server.py
公式Python SDK(mcp 2.0、MCPServer)に基づいています。トランスポート層は意図的に薄く、それ自体には権限がありません。そこにバグがあっても、エージェントができることを広げることはできません。ポリシーとAPIサーバーのPod Securityアドミッションが実際の制御だからです。
監査ログ
すべての呼び出しは、相関されたイベントトレイルを var/audit.jsonl に出力します:
tool.request -> guardrail.decision -> credential.issued -> sandbox.started
-> sandbox.completed -> credential.revoked -> tool.resultmake audit
./.venv/bin/agent-sandbox audit --request-id req-4239b8bb5459 --json認証情報の値は書き込み前に再帰的にスクラブされます。スコープ、リースID、TTLは保持されます。テストでは、JWT形式の文字列がログに到達しないことをアサートします。
検証済み、想定ではない
このプロジェクトには、主張 するのは簡単で、実際には持っていないことが2つあります。そのため、それらを信頼しませんでした。make verify は両方をライブクラスターに対してテストします:
== 1. gVisor kernel check ==
kernel reported: Linux version 4.19.0-gvisor
PASS: sandbox runs on the gVisor sentry kernel
== 2. NetworkPolicy egress enforcement check ==
PASS: baseline connectivity works (got PONG)
PASS: default-deny egress enforced (traffic blocked)これは構築中に実際の問題を検出しました。 kindのデフォルトCNI(kindnet)は NetworkPolicy オブジェクトを受け入れ、静かに無視します。デフォルト拒否の送信ポリシーを適用しましたが、ポッド間トラフィックは依然として通過しました。サンドボックスはロックダウンされているように見えながら、完全なネットワークアクセスを持っていたでしょう。kindnetを無効にしてCalicoをインストールすることで修正しました。Calicoは実際に強制します。scripts/install-calico.sh を参照してください。
APIサーバーの許可リストに関連する罠にも遭遇しました:ClusterIP は機能しません。kube-proxyがCalicoが送信を評価する前に実際のエンドポイントにDNATするためです。症状は、ポリシー拒否イベントなしにサンドボックスがハングするだけでした。scripts/apply-sandbox-policy.sh に文書化されています。
正直な制限
gVisorは実行されますが、これは依然としてkindです。 kindノード(Docker DesktopのLinux VM内のコンテナ)に
runscをインストールし、アクティブであることを確認しました。これは本物のgVisorサンドボックスですが、本番用に強化されたノードではありません。AWS/STSパスは条件付きです。
scripts/vault-setup.shは、実際のAWS認証情報が存在する場合にのみVaultのAWSシークレットエンジンを設定します。存在しない場合はスキップされ、その旨が表示されます。デモを完全に見せるためにそのパスを偽装したくありませんでした。ライブで実証可能な 認証情報パスはKubernetesのもので、完全に本物です:動的ServiceAccount、実際のRBAC、実際のリース、実際の失効。Vaultは開発モードで実行されます — インメモリ、ルートトークン
root、シールなし。ローカルプロジェクトには問題ありませんが、そのままデプロイするものではありません。破壊的シグナル検出はプラン出力の文字列マッチングです。 これは人間のための 表面化 支援であり、セキュリティ境界ではありません。
terraform_applyはすでに高リスク層であり、スキャンが何を見つけてもゲートされます。シングルノードクラスター のため、Terraform状態を保持するPVCは1つのノード上の
ReadWriteOnceです。
レイアウト
cluster/ kind config, RuntimeClass, namespaces, RBAC, network policy
images/ sandbox runner image (terraform + kubectl, providers vendored)
policy/ guardrail policy: risk tiers and destructive signals
scripts/ up/down, gVisor + Calico install, verification, demo
src/ the package: broker, sandbox, guardrails, approvals, audit, MCP
terraform/ demo module managed by the agent
tests/ 56 unit tests + a real-stdio MCP integration checkテスト
make test # 56 unit tests, no cluster required
make test-mcp # drives the server over real MCP stdio (needs the stack up)
make verify # proves gVisor + NetworkPolicy enforcement on the live clusterAvailable Tools
3 toolsk8s_get_pod_statusGet pod statusA
Read-only. Lists pods and their phase in the target namespace, executed inside a gVisor sandbox with a credential scoped to get/list/watch pods in that one namespace. Runs immediately; no approval needed.
| Name | Required | Description | Default |
|---|---|---|---|
| namespace | No | demo-app |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full disclosure burden and does it well: it states read-only semantics, gVisor sandbox isolation, the exact credential scope (get/list/watch on one namespace), and that no approval gate exists. It stops short of pagination or error behavior, but the safety and permission profile is unusually well covered for an annotation-free tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the safety-critical 'Read-only.' qualifier followed by scope, environment, and approval status. Every sentence adds information; no filler or repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return-format explanation is unnecessary, and the description covers read-only status, execution environment, credential scope, and approval behavior for a single-parameter tool. The one remaining gap is that the namespace default and format are never mentioned in either the description or the schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for the lone parameter. It only gestures at 'the target namespace' without naming the parameter, giving the default value ('demo-app'), or indicating accepted namespace formats. The agent must open the schema to learn anything actionable about the input.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb and resource ('Lists pods and their phase') plus an explicit scope ('in the target namespace'), so the agent knows exactly what is returned. It does not differentiate from siblings, but the siblings (terraform_plan/terraform_apply) are an unrelated domain, so no routing confusion exists. Minor tension: the name says 'pod status' (singular) while the description lists pods (plural).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Read-only' and 'Runs immediately; no approval needed' imply when this is safe to call, and the namespace scoping implies the context. There is no explicit when-to-use/when-not statement or named alternative, so usage is only implied rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
terraform_applyTerraform applyA
HIGH RISK -- mutates real infrastructure. Calling without approval_id does NOT apply anything: it computes the plan, records a pending approval, and returns an approval id for a human to review out-of-band. Once a human has approved, call again with that approval_id to execute the reviewed plan. Approvals are single-use, expiring, and bound to these exact parameters.
| Name | Required | Description | Default |
|---|---|---|---|
| approval_id | No | ||
| working_dir | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden and does so well: it flags HIGH RISK mutation, explains that the no-approval path is non-destructive (computes plan, records pending approval), and discloses approval lifecycle traits (single-use, expiring, parameter-bound). This is far beyond what a bare 'apply' would convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four tight sentences with the risk warning and the two-phase contract front-loaded; no filler. Dense but readable, and every clause adds operational information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Output schema exists so return-value documentation is unnecessary, and the description still names the key return (approval id). Combined with the risk warning and approval lifecycle, an agent has everything needed to call this correctly and safely.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate; it fully explains approval_id semantics (optional, single-use, expiring, must match exact parameters, absence means no apply). working_dir is never explained, which is the one remaining gap, but the high-risk parameter is covered.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('mutates real infrastructure', Terraform apply) and immediately clarifies its two-phase nature, which separates it from terraform_plan's pure-preview role. An agent can tell what this does and when it would fire without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit conditional workflow: call without approval_id to compute a plan and register a pending approval, then call again with the returned approval_id to execute. It also states the human review happens out-of-band, so the agent knows not to wait on this call for the mutation to occur.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
terraform_planTerraform planA
Computes a Terraform diff for a module under the terraform/ root. Runs immediately in the sandbox and persists the plan so a subsequent apply can execute exactly the reviewed diff. Output is annotated when the plan contains destructive changes such as forced replacement.
| Name | Required | Description | Default |
|---|---|---|---|
| working_dir | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does well: it discloses immediate sandbox execution (no confirmation gate), that the plan is persisted as durable state, and that output is annotated on destructive changes like forced replacement. It omits auth/permission requirements and failure behavior, which keeps it out of 5 territory.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences with no filler, front-loaded with what the tool does before the sandbox/persistence and destructive-change details.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained, and the description covers execution model, persistence, and destructive-change signaling. The main residual gap is parameter semantics for working_dir, which neither schema nor description pins down.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for the single working_dir parameter, so the schema adds nothing. The description hints that the module lives 'under the terraform/ root', which weakly constrains the path, but it never states whether working_dir is absolute, relative, or relative to that root.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Computes a Terraform diff for a module'), and scopes it to the terraform/ root. It also implicitly distinguishes itself from the sibling terraform_apply by framing itself as the review step that a 'subsequent apply' consumes.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear context for use: it runs immediately and produces a persisted plan that a later apply executes, which tells the agent this belongs before terraform_apply. It stops short of explicit when-not-to-use guidance or naming the alternative tool directly.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
3 tool updates
v0.1.0- First observed
k8s_get_pod_status - First observed
terraform_apply - First observed
terraform_plan
TDQS
Scored across 3 tools
Each tool has a clearly distinct purpose: read-only Kubernetes pod status, a Terraform plan diff, and a gated Terraform apply. The plan/apply pair is explicitly differentiated by the approval workflow, so an agent can reliably choose the right tool.
Names use consistent snake_case with a domain prefix and a verb (k8s_get_pod_status, terraform_plan, terraform_apply). The k8s tool adds an object noun while the Terraform tools do not, a minor deviation but still predictable.
Three tools across two distinct domains (Kubernetes and Terraform) is thin for an agent sandbox. Each tool earns its place, but the surface feels minimal relative to the apparent infrastructure-management scope.
Kubernetes coverage is limited to reading pod phase with no logs, events, describe, or any mutating operation, and Terraform lacks init/validate, destroy, and state inspection. The core plan-then-apply lifecycle is well covered, but notable gaps remain for a sandbox meant to work with real infrastructure.
Maintenance
Related MCP Connectors
Fail-closed policy guardrails for AI agents running kubectl, terraform, helm, and argocd.
Security gateway for AI agents: policy, approval, and audited execution, no secrets shared.
Security reviews for coding agents: diffs checked against your org policy and live infrastructure.
- emisarOAuthdev.emisar
Let AI operate servers without SSH. Choose actions, approve risky changes, and audit every step.
Related MCP Servers
- FlicenseNot gradedqualityBmaintenanceGive AI agents Zero-Trust access to production infrastructure without the risks of granting them shell access. Actions are bounded by policy and an on-host runner.337-
- AlicenseAqualityBmaintenanceEnables AI agents to safely inspect and execute version-controlled operational runbooks with policy checks, dry-run planning, and out-of-band approvals.3MIT
- AlicenseAqualityBmaintenanceEnables AI coding agents to plan, policy-check, and cost-estimate Terraform changes while requiring human approval through Slack or CLI before any terraform apply can execute.98 npmMIT
- AlicenseAqualityBmaintenanceEnables AI coding agents to safely execute commands, run tests, and modify project files inside disposable, policy-enforced Docker sandboxes that are isolated from the host machine and its credentials.15MIT