Skip to main content
Glama

agent-sandbox

Я создал это, чтобы ИИ-агент для написания кода мог выполнять реальные команды инфраструктуры против реального кластера Kubernetes — не имея при этом постоянных учетных данных и не имея возможности что-либо разрушить без присмотра.

Три MCP-инструмента. Каждый вызов выполняется внутри Kubernetes Job в песочнице gVisor с кратковременными, узко ограниченными учетными данными, которые Vault выпускает для этого единственного действия. Любые разрушительные действия останавливаются на этапе одобрения человеком.

AI Agent (Claude Desktop / Cursor)
        │  MCP protocol (stdio)
        ▼
┌──────────────────────────────────────────┐
│  MCP Server            src/agent_sandbox │
│    3 tools -> guardrails -> broker ->    │
│    sandbox -> audit                      │
└───────┬──────────────────────┬───────────┘
        │                      │
        ▼                      ▼
┌────────────────┐   ┌──────────────────────────┐
│ Credential     │   │ Sandbox Runner           │
│ Broker (Vault) │   │ K8s Job + gVisor         │
│ 10-min leases  │   │ restricted PSS           │
│ per-action     │   │ default-deny NetworkPolicy│
│ scope          │   │ cpu/mem limits, deadline │
└────────────────┘   └──────────────────────────┘
        │                      │
        └──────────┬───────────┘
                   ▼
        ┌──────────────────────┐
        │ Guardrails + Approval│
        │ policy.yaml, SQLite, │
        │ agent-sandbox CLI    │
        └──────────────────────┘

Зачем я это создал

Один ИИ-инструмент, которым я пользовался, однажды предложил изменение Terraform для производственной инфраструктуры, которое привело бы к принудительной замене живого ресурса. План выглядел обычным. Причиной сбоя была не ошибка модели — а то, что между правдоподобным планом и разрушительным применением не было ничего.

Я создал этот проект как недостающий слой, в рабочем коде:

  • агент никогда не имеет учетных данных, которые можно использовать повторно

  • всё выполняется там, где не может навредить хосту

  • разрушительные изменения останавливаются и ждут человека

  • каждое действие фиксируется

make demo воспроизводит именно тот сценарий, с которым я столкнулся. Изменение метки в одну строку приводит к принудительной замене работающего Deployment, и шлюз это ловит.

Related MCP server: safe-runbook-mcp

Быстрый старт

Требуются Docker, kind, kubectl, vault, terraform и Python 3.11+.

brew install kind kubectl hashicorp/tap/vault terraform
make up      # ~5 minutes from cold: cluster, CNI, gVisor, Vault, image, verify
make demo    # the forces-replacement guardrail demo
make down    # tear it all down

make up идемпотентен. Он завершается запуском make verify, который доказывает утверждения об изоляции, а не просто заявляет их (см. ниже).

Направьте на него агента

cp examples/claude_desktop_config.json \
   ~/Library/Application\ Support/Claude/claude_desktop_config.json

Cursor: скопируйте examples/cursor_mcp.json в .cursor/mcp.json. Затем попросите агента "проверить статус подов в demo-app" или "спланировать terraform для k8s-demo".

Три инструмента

Инструмент

Риск

Поведение

k8s_get_pod_status(namespace)

низкий

Выполняется немедленно. Учетные данные ограничены get/list/watch pods в одном namespace.

terraform_plan(working_dir)

низкий

Выполняется немедленно. Сохраняет план, чтобы последующее применение выполняло точно проверенный diff.

terraform_apply(working_dir, approval_id?)

высокий

Без approval_id: вычисляет план, записывает ожидающее одобрение, ничего не применяет. С ним: использует одобрение и применяет сохраненный план.

Четыре компонента

1. Изолированное выполнение — src/agent_sandbox/sandbox.py

Одноразовый Job на каждый вызов инструмента. Каждый контроль здесь имеет конкретную причину:

Контроль

Предотвращает

runtimeClassName: gvisor

Системные вызовы попадают в sentry gVisor, а не в ядро хоста

PSS restricted, применяемый API-сервером

root, повышение привилегий, capabilities, записываемая rootfs

automountServiceAccountToken: false

Любая фоновая идентичность кластера внутри песочницы

NetworkPolicy с запретом по умолчанию + allowlist API-сервера

Исходящий интернет-трафик, боковое перемещение, метаданные-эндпоинты

resources.limits, activeDeadlineSeconds

Вышедший из-под контроля Job, который голодает узел или зависает навсегда

backoffLimit: 0

Неудачное разрушительное действие, которое молча повторяется

Учетные данные монтируются как файл, а не переменная окружения — переменные окружения утекают через kubectl describe, /proc и дампы при сбоях.

2. Брокер учетных данных — src/agent_sandbox/broker.py

cred = broker.issue_scoped_credential("k8s_get_pod_status")
# -> Vault mints a ServiceAccount + Role + RoleBinding, 10-minute lease
# -> revoked immediately after the Job finishes
  • Агент никогда не выбирает свою область действия. Область выводится из действия.

  • Запрет по умолчанию. Действие без сопоставленной области не получает учетных данных.

  • Радиус поражения ограничен. Запрос любого namespace, кроме целевого, отклоняется.

  • Токен никогда не покидает модуль. Credential.__repr__ выводит token=<redacted>, так что даже случайный лог не может его раскрыть.

Я проверил это вручную: токен pod-reader перечисляет поды в demo-app, отклоняется в kube-system, отклоняется для секретов и перестает работать в момент отзыва его аренды — не оставляя за собой ServiceAccount.

3. Защитные механизмы — policy/policy.yaml, src/agent_sandbox/guardrails.py

Запрет по умолчанию: регистрация MCP-инструмента недостаточна, чтобы сделать его вызываемым. Инструмент, отсутствующий в политике, отклоняется, поэтому добавление возможности требует осознанного решения о категории риска.

Одобрения защищены от очевидных атак:

  • одноразовые — расходуются в одной SQLite-транзакции, так что два одновременных применения не могут использовать одно одобрение

  • привязаны к параметрам — привязаны к хэшу точного инструмента + параметров, так что одобрение для k8s-demo нельзя воспроизвести против prod-cluster

  • с истечением срока — 30 минут по умолчанию

  • внеполосные — выдаются через отдельный CLI-процесс. Нет MCP-инструмента для одобрения чего-либо; у агента нет пути кода для одобрения собственного запроса.

4. MCP-сервер — src/agent_sandbox/server.py

Построен на официальном Python SDK (mcp 2.0, MCPServer). Транспортный слой намеренно тонкий и не дает собственных полномочий — ошибка там не может расширить возможности агента, потому что политика и Pod Security admission API-сервера являются фактическими контролами.

Журнал аудита

Каждый вызов создает коррелированную цепочку событий в var/audit.jsonl:

tool.request -> guardrail.decision -> credential.issued -> sandbox.started
   -> sandbox.completed -> credential.revoked -> tool.result
make audit
./.venv/bin/agent-sandbox audit --request-id req-4239b8bb5459 --json

Значения учетных данных рекурсивно очищаются перед записью; область, идентификатор аренды и TTL сохраняются. Тест проверяет, что ни одна строка в форме JWT никогда не попадает в журнал.

Проверено, а не предположено

Две вещи в этом проекте легко заявить и тихо не иметь, поэтому я не принимал их на веру. make verify проверяет обе на живом кластере:

== 1. gVisor kernel check ==
     kernel reported: Linux version 4.19.0-gvisor
  PASS: sandbox runs on the gVisor sentry kernel
== 2. NetworkPolicy egress enforcement check ==
  PASS: baseline connectivity works (got PONG)
  PASS: default-deny egress enforced (traffic blocked)

Это выявило реальную проблему, пока я его создавал. Стандартный CNI kind (kindnet) принимает объекты NetworkPolicy и молча игнорирует их — я применил политику запрета исходящего трафика по умолчанию, и трафик между подами всё равно проходил. Песочница выглядела бы изолированной, имея при этом полный сетевой доступ. Я исправил это, отключив kindnet и установив Calico, который применяет политики по-настоящему. См. scripts/install-calico.sh.

Я столкнулся с похожей ловушкой при добавлении API-сервера в allowlist: ClusterIP не работает, потому что kube-proxy делает DNAT к реальному эндпоинту до того, как Calico оценивает исходящий трафик. Симптомом была песочница, которая просто зависала без события об отказе политики, объясняющего это. Документировано в scripts/apply-sandbox-policy.sh.

Честные ограничения

  • gVisor работает, но это всё еще kind. Я установил runsc внутри узла kind (контейнер в Linux VM Docker Desktop) и проверил, что он активен. Это настоящая песочница gVisor, а не узел, закаленный для продакшена.

  • Путь AWS/STS условен. scripts/vault-setup.sh настраивает AWS secrets engine Vault только при наличии реальных учетных данных AWS; без них он пропускается и сообщает об этом. Я не хотел имитировать этот путь только для того, чтобы демо выглядело полным. Живой, демонстрируемый путь учетных данных — это Kubernetes, который полностью реален: динамические ServiceAccounts, реальный RBAC, реальные аренды, реальный отзыв.

  • Vault работает в dev-режиме — в памяти, root-токен root, без печати. Подходит для локального проекта, но не для развертывания как есть.

  • Обнаружение разрушительных сигналов — это сопоставление строк в выводе плана. Это вспомогательное средство для человека, а не граница безопасности — terraform_apply уже относится к высокому уровню риска и блокируется независимо от того, что найдет сканирование.

  • Одноузловой кластер, поэтому PVC, хранящий состояние Terraform, имеет ReadWriteOnce на одном узле.

Структура

cluster/      kind config, RuntimeClass, namespaces, RBAC, network policy
images/       sandbox runner image (terraform + kubectl, providers vendored)
policy/       guardrail policy: risk tiers and destructive signals
scripts/      up/down, gVisor + Calico install, verification, demo
src/          the package: broker, sandbox, guardrails, approvals, audit, MCP
terraform/    demo module managed by the agent
tests/        56 unit tests + a real-stdio MCP integration check

Тестирование

make test       # 56 unit tests, no cluster required
make test-mcp   # drives the server over real MCP stdio (needs the stack up)
make verify     # proves gVisor + NetworkPolicy enforcement on the live cluster

Available Tools

3 tools
k8s_get_pod_statusGet pod statusA

Read-only. Lists pods and their phase in the target namespace, executed inside a gVisor sandbox with a credential scoped to get/list/watch pods in that one namespace. Runs immediately; no approval needed.

ParametersJSON Schema
NameRequiredDescriptionDefault
namespaceNodemo-app

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full disclosure burden and does it well: it states read-only semantics, gVisor sandbox isolation, the exact credential scope (get/list/watch on one namespace), and that no approval gate exists. It stops short of pagination or error behavior, but the safety and permission profile is unusually well covered for an annotation-free tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, front-loaded with the safety-critical 'Read-only.' qualifier followed by scope, environment, and approval status. Every sentence adds information; no filler or repetition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return-format explanation is unnecessary, and the description covers read-only status, execution environment, credential scope, and approval behavior for a single-parameter tool. The one remaining gap is that the namespace default and format are never mentioned in either the description or the schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate for the lone parameter. It only gestures at 'the target namespace' without naming the parameter, giving the default value ('demo-app'), or indicating accepted namespace formats. The agent must open the schema to learn anything actionable about the input.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a specific verb and resource ('Lists pods and their phase') plus an explicit scope ('in the target namespace'), so the agent knows exactly what is returned. It does not differentiate from siblings, but the siblings (terraform_plan/terraform_apply) are an unrelated domain, so no routing confusion exists. Minor tension: the name says 'pod status' (singular) while the description lists pods (plural).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'Read-only' and 'Runs immediately; no approval needed' imply when this is safe to call, and the namespace scoping implies the context. There is no explicit when-to-use/when-not statement or named alternative, so usage is only implied rather than stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

terraform_applyTerraform applyA

HIGH RISK -- mutates real infrastructure. Calling without approval_id does NOT apply anything: it computes the plan, records a pending approval, and returns an approval id for a human to review out-of-band. Once a human has approved, call again with that approval_id to execute the reviewed plan. Approvals are single-use, expiring, and bound to these exact parameters.

ParametersJSON Schema
NameRequiredDescriptionDefault
approval_idNo
working_dirYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden and does so well: it flags HIGH RISK mutation, explains that the no-approval path is non-destructive (computes plan, records pending approval), and discloses approval lifecycle traits (single-use, expiring, parameter-bound). This is far beyond what a bare 'apply' would convey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four tight sentences with the risk warning and the two-phase contract front-loaded; no filler. Dense but readable, and every clause adds operational information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Output schema exists so return-value documentation is unnecessary, and the description still names the key return (approval id). Combined with the risk warning and approval lifecycle, an agent has everything needed to call this correctly and safely.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate; it fully explains approval_id semantics (optional, single-use, expiring, must match exact parameters, absence means no apply). working_dir is never explained, which is the one remaining gap, but the high-risk parameter is covered.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('mutates real infrastructure', Terraform apply) and immediately clarifies its two-phase nature, which separates it from terraform_plan's pure-preview role. An agent can tell what this does and when it would fire without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit conditional workflow: call without approval_id to compute a plan and register a pending approval, then call again with the returned approval_id to execute. It also states the human review happens out-of-band, so the agent knows not to wait on this call for the mutation to occur.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

terraform_planTerraform planA

Computes a Terraform diff for a module under the terraform/ root. Runs immediately in the sandbox and persists the plan so a subsequent apply can execute exactly the reviewed diff. Output is annotated when the plan contains destructive changes such as forced replacement.

ParametersJSON Schema
NameRequiredDescriptionDefault
working_dirYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does well: it discloses immediate sandbox execution (no confirmation gate), that the plan is persisted as durable state, and that output is annotated on destructive changes like forced replacement. It omits auth/permission requirements and failure behavior, which keeps it out of 5 territory.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences with no filler, front-loaded with what the tool does before the sandbox/persistence and destructive-change details.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained, and the description covers execution model, persistence, and destructive-change signaling. The main residual gap is parameter semantics for working_dir, which neither schema nor description pins down.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% for the single working_dir parameter, so the schema adds nothing. The description hints that the module lives 'under the terraform/ root', which weakly constrains the path, but it never states whether working_dir is absolute, relative, or relative to that root.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Computes a Terraform diff for a module'), and scopes it to the terraform/ root. It also implicitly distinguishes itself from the sibling terraform_apply by framing itself as the review step that a 'subsequent apply' consumes.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear context for use: it runs immediately and produces a persisted plan that a later apply executes, which tells the agent this belongs before terraform_apply. It stops short of explicit when-not-to-use guidance or naming the alternative tool directly.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 3 tool updatesv0.1.0
    • First observedk8s_get_pod_status
    • First observedterraform_apply
    • First observedterraform_plan

TDQS

A3.9/5.0

Scored across 3 tools

Disambiguation5/5

Each tool has a clearly distinct purpose: read-only Kubernetes pod status, a Terraform plan diff, and a gated Terraform apply. The plan/apply pair is explicitly differentiated by the approval workflow, so an agent can reliably choose the right tool.

Naming Consistency4/5

Names use consistent snake_case with a domain prefix and a verb (k8s_get_pod_status, terraform_plan, terraform_apply). The k8s tool adds an object noun while the Terraform tools do not, a minor deviation but still predictable.

Tool Count3/5

Three tools across two distinct domains (Kubernetes and Terraform) is thin for an agent sandbox. Each tool earns its place, but the surface feels minimal relative to the apparent infrastructure-management scope.

Completeness3/5

Kubernetes coverage is limited to reading pod phase with no logs, events, describe, or any mutating operation, and Terraform lacks init/validate, destroy, and state inspection. The core plan-then-apply lifecycle is well covered, but notable gaps remain for a sandbox meant to work with real infrastructure.

Maintenance

ActivitySlowing
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • F
    license
    Not graded
    quality
    B
    maintenance
    Give AI agents Zero-Trust access to production infrastructure without the risks of granting them shell access. Actions are bounded by policy and an on-host runner.
    337
    -
  • A
    license
    A
    quality
    B
    maintenance
    Enables AI coding agents to safely execute commands, run tests, and modify project files inside disposable, policy-enforced Docker sandboxes that are isolated from the host machine and its credentials.
    15
    MIT