Capsule Bash Server
OfficialCapsule Bash MCP
AI 에이전트가 안전하고 지속적인 샌드박스 환경에서 bash 명령을 실행할 수 있도록 하는 MCP 서버입니다.
작동 방식
각 세션은 WebAssembly 샌드박스 내부에서 실행됩니다. 샌드박스는 다음을 제공합니다:
지속적인 상태: 세션 내의 명령 간에 cwd, 환경 변수 및 파일 시스템 변경 사항이 유지됩니다.
파일 시스템 차이(diff): 모든
run응답에는 디스크에서 변경된 내용의 차이가 포함됩니다.격리된 메모리: 각 세션은 고유한 주소 공간을 가지며 세션 간 데이터 유출이 없습니다.
호스트 접근 불가: 샌드박스는 호스트 파일 시스템이나 네트워크에 접근할 수 없습니다.
Capsule Bash에 대해 더 알아보세요.
Related MCP server: execkit-mcp
도구
도구 | 설명 |
| 샌드박스 세션에서 bash 명령을 실행합니다. stdout, stderr, 종료 코드, 파일 시스템 차이 및 현재 상태(cwd + env)를 반환합니다. |
| 세션의 파일 시스템 및 상태(cwd, 환경 변수)를 초기 값으로 재설정합니다. |
| 모든 활성 세션을 나열합니다. |
세션
동일한 session_id 내의 명령은 호출 간에 cwd, 환경 변수 및 파일 시스템 상태를 공유합니다.
예시
AI 에이전트에게 다음과 같이 요청하세요:
"숫자 목록의 평균을 구하는 Python 스크립트를 작성해 줘."
에이전트는 순차적으로 run을 호출합니다:
{ "command": "mkdir -p /data && cd /data", "session_id": "custom_session" }
{ "command": "echo 'nums = [x for x in [1, 2, 3, []] if isinstance(x, int)]\nprint(sum(nums) / len(nums))' > avg.py", "session_id": "custom_session" }
{ "command": "python3 avg.py", "session_id": "custom_session" }각 호출은 stdout, stderr, exitCode, 파일 시스템 diff 및 업데이트된 state를 반환하여 컨텍스트를 풍부하게 하고 대화 기록을 추적합니다.
설정
MCP 클라이언트 구성(예: Claude Desktop, Cursor)에 추가하세요:
{
"mcpServers": {
"bash": {
"command": "npx",
"args": ["-y", "@capsule-run/bash-mcp"]
}
}
}제한 사항
모든 bash 명령 및 옵션이 구현된 것은 아닙니다. 명령이 누락되었거나 예상치 않게 작동하는 경우 이슈를 열어 알려주세요.
Available Tools
3 toolsresetA
Reset a session's filesystem and shell state (cwd, env vars) to their initial values. Useful to start fresh without creating a new session.
| Name | Required | Description | Default |
|---|---|---|---|
| session_id | No | The session to reset. Defaults to 'default'. | default |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It clearly explains that the tool resets filesystem and shell state to initial values, which is specific and transparent about the behavior. No side effects or additional information is provided, but the action is straightforward.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description consists of two sentences, both front-loaded with essential information: the action and scope in the first sentence, usage context in the second. No wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's simplicity (one parameter, no output schema), the description provides all necessary information: what it does, what it affects, and when to use it. No gaps remain.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 100% description coverage, and the description adds no additional meaning beyond what the schema already provides for the single parameter 'session_id'. Baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb 'reset' and the resource 'session's filesystem and shell state', specifying what is reset (cwd, env vars) and distinguishing from siblings 'run' and 'sessions' which have different purposes.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly says 'useful to start fresh without creating a new session', giving a clear usage scenario and implying an alternative (creating a new session). However, it does not explicitly state when not to use it or list other alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
runB
Execute a Bash command in the sandboxed environment
| Name | Required | Description | Default |
|---|---|---|---|
| command | Yes | The bash command to execute. | |
| session_id | No | Identifier for the shell session. Commands within the same session share cwd, env, and filesystem state. Defaults to 'default'. | default |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description bears full burden for behavioral disclosure. It only notes the sandboxed environment but fails to describe side effects (e.g., file system changes, output handling, error behavior) or any permissions needed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, focused sentence with no wasted words. It is appropriately sized for a simple tool and front-loades key information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's simplicity (2 parameters, no output schema, no annotations), the description is adequate but not thorough. It lacks information about return values (e.g., stdout/stderr) and potential side effects, which could help an agent understand the tool's full behavior.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, with clear explanations for both parameters (command and session_id). The tool description adds no extra parameter information. Baseline of 3 is appropriate as the schema already documents parameter semantics.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description 'Execute a Bash command in the sandboxed environment' clearly states the action (execute) and the resource (Bash command in sandbox). It distinguishes from siblings 'reset' and 'sessions' by implying this tool is for running commands, not managing sessions.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus alternatives like 'reset' or 'sessions'. The description does not mention prerequisites, scenarios, or exclusions, leaving the agent without context for tool selection.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
sessionsA
List all active session IDs and their current state.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Lacking annotations, the description states the tool lists active sessions, conveying a non-destructive, read-only behavior. But it omits details like whether state includes sub-attributes or if sessions are global.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, front-loaded sentence with no wasted words. Efficiently captures the core functionality.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no output schema, the description could be more complete by hinting at return structure or example use. It covers the basic purpose but lacks actionable output details.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
No parameters; schema coverage is 100% by default. The description does not need to add param info. Baseline 4 for 0 parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Describes a specific action (list) and resource (active session IDs and state). Clearly distinguishes from siblings 'reset' and 'run' which imply modification actions.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies a read-only listing use case, and with siblings being action-oriented, the context is clear. However, no explicit when-not or alternative guidance is provided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
3 tool updates
v0.1.0- First observed
reset - First observed
run - First observed
sessions
TDQS
Scored across 3 tools
Each tool has a unique, clear purpose: reset reverts state, run executes commands, and sessions lists active sessions. No ambiguity or overlap exists.
All tool names are single, imperative verbs (reset, run, sessions) following a consistent, simple pattern.
Three tools is well-scoped for a bash sandbox server, covering essential operations without superfluous or missing tools.
The set covers core actions (run, reset, list sessions) but lacks explicit session creation; likely sessions are created implicitly, leaving a minor gap.
Maintenance
Related MCP Connectors
Linux microVM sandboxes for AI agents: run commands, files, processes, pause and wake.
Operate Linux, macOS and Windows from your LLM. Every action runs through an auditable allowlist.
Scoped agent execution. Server-side credentials, policy, budgets and verifiable receipts.
- mcp-serverOAuthai.cdbx
Build Apps and run code in 30 languages — sandboxed, with persistent sessions for agent loops.
Related MCP Servers
- AlicenseNot gradedqualityBmaintenanceA throwaway Docker sandbox for agents to run code and shell commands safely.156 npmMIT
- AlicenseAqualityBmaintenanceStateful, structured, safe shell sessions for AI agents, on real infrastructure.78Apache 2.0
- AlicenseAqualityAmaintenanceExposes persistent, stateful remote Bash sessions to AI agents via SSH, enabling command execution with preserved working directory and environment.746 npm1MIT
- AlicenseAqualityDmaintenanceSandboxed bash execution MCP server for AI agents, using an in-memory virtual filesystem overlay to prevent real filesystem damage, with configurable network access, timeouts, and optional Python/JS runtimes.924 npmApache 2.0