agent-shuttle-mcp
# Agent Shuttle
[](https://github.com/Plartex/agent-shuttle/actions/workflows/tests.yml)
[](LICENSE)
[](https://a2a-protocol.org/latest/)
[](https://modelcontextprotocol.io/)
[](https://www.python.org/)
[Russian version / Русская версия](README.ru.md)
Agent Shuttle gives Python applications one way to work with **Codex**, **Antigravity**, **OpenCode**, and **Claude Code**. It runs local agent tasks, preserves multi-turn sessions, and exposes the agents through [A2A 1.0 JSON-RPC](https://a2a-protocol.org/latest/) and [MCP](https://modelcontextprotocol.io/).
It allows agents and external applications to delegate tasks to peer agents, reuse multi-turn conversations, query live model catalogs and account quotas, and enforce tool permission boundaries—all on local loopback (`127.0.0.1`) without sharing cloud API keys.
---
## Supported Agent Harnesses
| Harness | Primary Integration Mechanism | Auth & Model Access | Key Features |
|---|---|---|---|
| **Codex** | Official `openai-codex` Python SDK | Local Codex App Server sign-in | Sandboxes (`workspace_write`, `read_only`, `full_access`), model & reasoning effort catalog, quota reporting via `account/rateLimits/read`. |
| **Antigravity** | Official `agy` CLI in headless mode (`-p` / `stream-json`) | Signed-in Antigravity account | Real-time models, efforts, and `/usage` quotas. Default settings or full access (`--dangerously-skip-permissions`). Optional legacy SDK backend. |
| **OpenCode** | Managed local HTTP server (`--pure serve`) | Local Ollama or cloud providers | Configured via JSON profile, fine-grained tool policies (`no_tools`, `read_only`, `workspace_write`, `full_access`), model variants. |
| **Claude Code** | Managed CLI in print mode (`claude -p`) | Local Ollama endpoint or Anthropic | Configured via JSON profile, isolated temporary session configs, resume support, safe mode vs full access. |
---
## Installation
Agent Shuttle requires **Python 3.11+** and runs on Windows, Linux, and macOS.
### Install from GitHub
```powershell
python -m venv .venv
.\.venv\Scripts\python.exe -m pip install "git+https://github.com/Plartex/agent-shuttle.git"
```
The package is not published on PyPI yet. For development, clone the repository and install it in editable mode.
### Install from Local Checkout
You can install Agent Shuttle directly from its repository checkout into your project's virtual environment:
```powershell
# Create and activate your virtual environment
python -m venv .venv
.\.venv\Scripts\Activate.ps1
# Install in standard or editable mode
pip install C:\path\to\agent-shuttle
# or editable mode during development:
# pip install -e C:\path\to\agent-shuttle
```
Once installed, the CLI tools (`agent-shuttle`, `agent-shuttle-mcp`) and Python API (`agent_shuttle`) are fully accessible inside that virtual environment. The original checkout directory does not need to stay in place for runtime imports. Existing `agent_bridge` imports and `agent-bridge` commands remain supported as compatibility aliases; A2A metadata keys under `agent_bridge.*` are unchanged.
---
## Quickstart (Windows PowerShell)
Use standard Python packaging, including when developing from a checkout:
```powershell
python -m venv .venv
& .\.venv\Scripts\python.exe -m pip install -e .
& .\.venv\Scripts\agent-shuttle.exe discover
```
Point your MCP client's `command` at the installed `.venv\Scripts\agent-shuttle-mcp.exe` (use its absolute path). MCP starts a local A2A peer when a request needs one; there is no separate server startup step. See the [getting started guide](docs/getting-started.md) for MCP configuration.
For a persistent server, run `agent-shuttle serve codex --workspace . --port 8765` (or `serve antigravity` on port 8766) in a terminal and stop it with Ctrl+C.
---
## Minimal Examples
### 1. Python API
```python
import asyncio
from pathlib import Path
from agent_shuttle import ShuttleClient, HarnessLaunch, connect_harness
async def main():
client = ShuttleClient()
# Query live model catalog and quota information
info = await client.info("http://127.0.0.1:8766")
print("Selected model:", info["capabilities"]["selected_model"])
# One-shot task
result = await client.ask(
"http://127.0.0.1:8766",
"Explain the project structure and list entry points.",
model="gemini-3.8-flash-medium",
reasoning_effort="medium",
)
print(f"[{result.state}] Task {result.task_id}:\n{result.text}")
# Persistent multi-turn session (preserves conversation context)
async with client.session(
"http://127.0.0.1:8765",
model="gpt-5.6-terra",
reasoning_effort="high",
) as session:
step1 = await session.ask("What database migrations are pending?")
step2 = await session.ask("Generate SQL to apply the first migration.")
print("Step 2 response:", step2.text)
print("Step 2 token usage:", step2.usage)
# Managed peer lifecycle: reuse existing server or start a temporary one
launch = HarnessLaunch(
name="antigravity",
url="http://127.0.0.1:8766",
workspace=Path.cwd(),
)
async with connect_harness(launch) as conn:
print("Connected to:", conn.url, "(spawned temporary:", conn.started, ")")
asyncio.run(main())
```
### 2. Command-Line Interface (CLI)
```powershell
# Discover locally installed harnesses without starting models
agent-shuttle discover
# Start an A2A server for Codex
agent-shuttle serve codex --workspace . --port 8765
# Start an A2A server for Antigravity (CLI mode)
agent-shuttle serve antigravity --workspace . --port 8766
# Start an A2A server from a profile (OpenCode or Claude Code)
agent-shuttle serve profile --profile .\examples\opencode-ollama.json --workspace . --port 8767
# Query server capabilities, models, and quota limits
agent-shuttle info http://127.0.0.1:8765
# Send a task from the command line
agent-shuttle ask http://127.0.0.1:8765 "Summarize recent changes" --model gpt-5.6-terra
```
### 3. Model Context Protocol (MCP)
Start the stdio MCP server:
```powershell
agent-shuttle-mcp
# or: python -m agent_shuttle.mcp_server
```
Available MCP tools:
- `ask_agent(agent_id, prompt, model?, reasoning_effort?, tool_policy?, workspace?)`: Reuses a matching local A2A server or starts a temporary one. Built-in IDs are `codex`, `antigravity`, `opencode`, and `claude_code`; the latter two need a profile or an Ollama model.
- `get_agent_info(agent_id, workspace?)`: Fetches live models and quotas, starting a temporary server if needed.
- `ask_antigravity(prompt, model?, reasoning_effort?, workspace?, tool_policy?, turn_timeout_seconds=300)`: Starts an Antigravity server if one is not running.
- `ask_codex(prompt, model?, reasoning_effort?, workspace?)`: Starts a Codex server if one is not running.
- `get_antigravity_info(workspace?)` & `get_codex_info()`: Read live capabilities and quota without burning model turns.
- `submit_task(agent_id, prompt, model?, reasoning_effort?, tool_policy?, workspace?, request_id?)`: Start a long task and return its ID immediately.
- `check_task(task_id)`, `wait_task(task_id, timeout_seconds?)`, `cancel_task(task_id)`: Inspect, wait for, or stop a task.
- `get_result(task_id, cursor?, limit?)`, `get_transcript(task_id, cursor?, limit?)`: Read bounded pages of output and history.
---
## Key Concepts
### Harness Discovery
Run `agent-shuttle discover` (or `discover_harnesses()` in Python) to inspect local executables without launching processes or loading weights. It checks `PATH` and platform-specific standard installation directories (`%LOCALAPPDATA%\agy\bin`, npm global directories, etc.). Custom paths can be specified via environment variables (`BRIDGE_AGY_COMMAND`) or CLI flags (`--agy-command`, `--opencode-command`, `--claude-command`).
### Model & Reasoning Selection
Model parameters are passed as A2A metadata keys (`agent_bridge.model`, `agent_bridge.reasoning_effort`):
- **Antigravity:** Reasoning effort is embedded in model IDs (e.g. `gemini-3.8-flash-medium`). If both `--model` and `--effort` are passed, they must match.
- **Codex:** Model and reasoning effort are configured independently according to the catalog returned by `get_codex_info`.
- **OpenCode & Claude Code:** Profiles define `allowed_models` and optional `reasoning_efforts` (such as model variants for Ollama or CLI flags).
### Workspaces & Session Isolation
- Every server binds to a strictly validated, canonical workspace directory.
- `connect_harness()` verifies that an existing server's workspace matches the caller's target workspace before reusing it.
- **Sessions:** `BridgeSession` maintains a stateful conversation across multiple `ask()` calls. Conversation settings (model, effort, tool policy) are pinned at session creation and cannot be changed mid-session. Idle sessions are cleaned up automatically after 30 minutes.
- **Library task lifecycle:** `TaskManager` runs agents directly from Python, with no A2A or MCP server. It owns task IDs, sessions, a SQLite event journal, cancellation, result paging, preferences, and interruption recovery. A2A projects the same task ID and result through its protocol; MCP continues to reach those tasks through managed A2A peers. See the [Python API](docs/api.md#taskmanager-library-api).
- **Remote task lifecycle:** `ShuttleClient.submit()` returns a remote `TaskHandle` immediately. Use `status()`, bounded `wait(timeout)`, `events()`, `result_page()`, `transcript()`, `result()`, or `cancel()`; reopen a task by ID with `client.task(url, task_id)`. A wait timeout does not stop the agent. An optional UUID `request_id` deduplicates retried submissions. Standalone servers can persist tasks with `--task-db`; MCP task tools do this automatically in the workspace's `.agent-shuttle` directory. Completed results survive restart; interrupted work is marked failed without replay. See the [API reference](docs/api.md#taskhandle-and-bridgeevent).
### Safety & Tool Policies
Agent Shuttle defines four standardized tool policies:
- `no_tools`: Disables tool invocations entirely.
- `read_only`: Permits non-mutating search and file reading.
- `workspace_write`: Allows editing files within the designated workspace.
- `full_access`: Explicitly unclamps all tool restrictions and approval prompts.
> [!WARNING]
> `full_access` grants the worker unrestricted tool access for that task. `--agy-dangerously-skip-permissions` enables that capability for the Antigravity server. Keep servers on loopback. Antigravity's explicit scoped `read_only` policy is enforced by its verified `PreToolUse` hook; an implicit/default CLI policy does not provide the same guarantee.
---
## Testing
Agent Shuttle provides a comprehensive offline test suite using fake backends that execute without network access, credentials, or model quota consumption:
```powershell
python -m unittest discover -s tests -v
```
Opt-in integration tests against real models can be executed by specifying target environments (e.g. `BRIDGE_LIVE_OLLAMA_MODEL=qwen3.5:9b` or `BRIDGE_LIVE_AGY_FULL_ACCESS=1`). See [CONTRIBUTING.md](CONTRIBUTING.md) for full instructions.
---
## Documentation Index
- [Getting Started Guide](docs/getting-started.md)
- [API Reference](docs/api.md)
- [Permissions & Safety Guide](docs/permissions.md)
- [Architecture & Protocol Design](docs/architecture.md)
- [Troubleshooting & Diagnostics](docs/troubleshooting.md)
- [Contributing Guidelines](CONTRIBUTING.md)
- [Security Policy](SECURITY.md)
TDQS
Scored across 6 tools
The three ask_* tools target different agents (generic configured profiles, Antigravity, Codex), and the three get_*_info tools mirror that structure. However, the generic ask_agent/get_agent_info could be confused with the specific ask_antigravity/get_antigravity_info or ask_codex/get_codex_info if those agents are also configured, requiring careful reading of descriptions.
All tool names follow a clear ask_<target> or get_<target>_info pattern in consistent snake_case. The generic 'agent' target versus specific agent names is a meaningful distinction rather than a naming inconsistency.
Six tools is well-scoped for a delegation server, covering ask and info operations for both generic and specific agents without redundancy. Each tool earns its place.
The core ask and info operations are present, but there is no list_agents tool to discover which profiles are available in BRIDGE_AGENTS_JSON. This is a notable gap for a server centered on configured agents, as users must already know the agent names.