Skip to main content
Glama
README.md
# Agent Shuttle

[![Tests](https://github.com/Plartex/agent-shuttle/actions/workflows/tests.yml/badge.svg)](https://github.com/Plartex/agent-shuttle/actions/workflows/tests.yml)
[![MIT license](https://img.shields.io/badge/license-MIT-blue.svg)](LICENSE)
[![A2A Protocol 1.0](https://img.shields.io/badge/A2A-1.0_JSON--RPC-blue)](https://a2a-protocol.org/latest/)
[![MCP](https://img.shields.io/badge/MCP-tools-green)](https://modelcontextprotocol.io/)
[![Python 3.11+](https://img.shields.io/badge/python-3.11+-blue.svg)](https://www.python.org/)

[Russian version / Русская версия](README.ru.md)

Agent Shuttle gives Python applications one way to work with **Codex**, **Antigravity**, **OpenCode**, and **Claude Code**. It runs local agent tasks, preserves multi-turn sessions, and exposes the agents through [A2A 1.0 JSON-RPC](https://a2a-protocol.org/latest/) and [MCP](https://modelcontextprotocol.io/).

It allows agents and external applications to delegate tasks to peer agents, reuse multi-turn conversations, query live model catalogs and account quotas, and enforce tool permission boundaries—all on local loopback (`127.0.0.1`) without sharing cloud API keys.

---

## Supported Agent Harnesses

| Harness | Primary Integration Mechanism | Auth & Model Access | Key Features |
|---|---|---|---|
| **Codex** | Official `openai-codex` Python SDK | Local Codex App Server sign-in | Sandboxes (`workspace_write`, `read_only`, `full_access`), model & reasoning effort catalog, quota reporting via `account/rateLimits/read`. |
| **Antigravity** | Official `agy` CLI in headless mode (`-p` / `stream-json`) | Signed-in Antigravity account | Real-time models, efforts, and `/usage` quotas. Default settings or full access (`--dangerously-skip-permissions`). Optional legacy SDK backend. |
| **OpenCode** | Managed local HTTP server (`--pure serve`) | Local Ollama or cloud providers | Configured via JSON profile, fine-grained tool policies (`no_tools`, `read_only`, `workspace_write`, `full_access`), model variants. |
| **Claude Code** | Managed CLI in print mode (`claude -p`) | Local Ollama endpoint or Anthropic | Configured via JSON profile, isolated temporary session configs, resume support, safe mode vs full access. |

---

## Installation

Agent Shuttle requires **Python 3.11+** and runs on Windows, Linux, and macOS.

### Install from GitHub

```powershell
python -m venv .venv
.\.venv\Scripts\python.exe -m pip install "git+https://github.com/Plartex/agent-shuttle.git"
```

The package is not published on PyPI yet. For development, clone the repository and install it in editable mode.

### Install from Local Checkout

You can install Agent Shuttle directly from its repository checkout into your project's virtual environment:

```powershell
# Create and activate your virtual environment
python -m venv .venv
.\.venv\Scripts\Activate.ps1

# Install in standard or editable mode
pip install C:\path\to\agent-shuttle
# or editable mode during development:
# pip install -e C:\path\to\agent-shuttle
```

Once installed, the CLI tools (`agent-shuttle`, `agent-shuttle-mcp`) and Python API (`agent_shuttle`) are fully accessible inside that virtual environment. The original checkout directory does not need to stay in place for runtime imports. Existing `agent_bridge` imports and `agent-bridge` commands remain supported as compatibility aliases; A2A metadata keys under `agent_bridge.*` are unchanged.

---

## Quickstart (Windows PowerShell)

Use standard Python packaging, including when developing from a checkout:

```powershell
python -m venv .venv
& .\.venv\Scripts\python.exe -m pip install -e .
& .\.venv\Scripts\agent-shuttle.exe discover
```

Point your MCP client's `command` at the installed `.venv\Scripts\agent-shuttle-mcp.exe` (use its absolute path). MCP starts a local A2A peer when a request needs one; there is no separate server startup step. See the [getting started guide](docs/getting-started.md) for MCP configuration.

For a persistent server, run `agent-shuttle serve codex --workspace . --port 8765` (or `serve antigravity` on port 8766) in a terminal and stop it with Ctrl+C.

---

## Minimal Examples

### 1. Python API

```python
import asyncio
from pathlib import Path
from agent_shuttle import ShuttleClient, HarnessLaunch, connect_harness

async def main():
    client = ShuttleClient()

    # Query live model catalog and quota information
    info = await client.info("http://127.0.0.1:8766")
    print("Selected model:", info["capabilities"]["selected_model"])

    # One-shot task
    result = await client.ask(
        "http://127.0.0.1:8766",
        "Explain the project structure and list entry points.",
        model="gemini-3.8-flash-medium",
        reasoning_effort="medium",
    )
    print(f"[{result.state}] Task {result.task_id}:\n{result.text}")

    # Persistent multi-turn session (preserves conversation context)
    async with client.session(
        "http://127.0.0.1:8765",
        model="gpt-5.6-terra",
        reasoning_effort="high",
    ) as session:
        step1 = await session.ask("What database migrations are pending?")
        step2 = await session.ask("Generate SQL to apply the first migration.")
        print("Step 2 response:", step2.text)
        print("Step 2 token usage:", step2.usage)

    # Managed peer lifecycle: reuse existing server or start a temporary one
    launch = HarnessLaunch(
        name="antigravity",
        url="http://127.0.0.1:8766",
        workspace=Path.cwd(),
    )
    async with connect_harness(launch) as conn:
        print("Connected to:", conn.url, "(spawned temporary:", conn.started, ")")

asyncio.run(main())
```

### 2. Command-Line Interface (CLI)

```powershell
# Discover locally installed harnesses without starting models
agent-shuttle discover

# Start an A2A server for Codex
agent-shuttle serve codex --workspace . --port 8765

# Start an A2A server for Antigravity (CLI mode)
agent-shuttle serve antigravity --workspace . --port 8766

# Start an A2A server from a profile (OpenCode or Claude Code)
agent-shuttle serve profile --profile .\examples\opencode-ollama.json --workspace . --port 8767

# Query server capabilities, models, and quota limits
agent-shuttle info http://127.0.0.1:8765

# Send a task from the command line
agent-shuttle ask http://127.0.0.1:8765 "Summarize recent changes" --model gpt-5.6-terra
```

### 3. Model Context Protocol (MCP)

Start the stdio MCP server:
```powershell
agent-shuttle-mcp
# or: python -m agent_shuttle.mcp_server
```

Available MCP tools:
- `ask_agent(agent_id, prompt, model?, reasoning_effort?, tool_policy?, workspace?)`: Reuses a matching local A2A server or starts a temporary one. Built-in IDs are `codex`, `antigravity`, `opencode`, and `claude_code`; the latter two need a profile or an Ollama model.
- `get_agent_info(agent_id, workspace?)`: Fetches live models and quotas, starting a temporary server if needed.
- `ask_antigravity(prompt, model?, reasoning_effort?, workspace?, tool_policy?, turn_timeout_seconds=300)`: Starts an Antigravity server if one is not running.
- `ask_codex(prompt, model?, reasoning_effort?, workspace?)`: Starts a Codex server if one is not running.
- `get_antigravity_info(workspace?)` & `get_codex_info()`: Read live capabilities and quota without burning model turns.
- `submit_task(agent_id, prompt, model?, reasoning_effort?, tool_policy?, workspace?, request_id?)`: Start a long task and return its ID immediately.
- `check_task(task_id)`, `wait_task(task_id, timeout_seconds?)`, `cancel_task(task_id)`: Inspect, wait for, or stop a task.
- `get_result(task_id, cursor?, limit?)`, `get_transcript(task_id, cursor?, limit?)`: Read bounded pages of output and history.

---

## Key Concepts

### Harness Discovery
Run `agent-shuttle discover` (or `discover_harnesses()` in Python) to inspect local executables without launching processes or loading weights. It checks `PATH` and platform-specific standard installation directories (`%LOCALAPPDATA%\agy\bin`, npm global directories, etc.). Custom paths can be specified via environment variables (`BRIDGE_AGY_COMMAND`) or CLI flags (`--agy-command`, `--opencode-command`, `--claude-command`).

### Model & Reasoning Selection
Model parameters are passed as A2A metadata keys (`agent_bridge.model`, `agent_bridge.reasoning_effort`):
- **Antigravity:** Reasoning effort is embedded in model IDs (e.g. `gemini-3.8-flash-medium`). If both `--model` and `--effort` are passed, they must match.
- **Codex:** Model and reasoning effort are configured independently according to the catalog returned by `get_codex_info`.
- **OpenCode & Claude Code:** Profiles define `allowed_models` and optional `reasoning_efforts` (such as model variants for Ollama or CLI flags).

### Workspaces & Session Isolation
- Every server binds to a strictly validated, canonical workspace directory.
- `connect_harness()` verifies that an existing server's workspace matches the caller's target workspace before reusing it.
- **Sessions:** `BridgeSession` maintains a stateful conversation across multiple `ask()` calls. Conversation settings (model, effort, tool policy) are pinned at session creation and cannot be changed mid-session. Idle sessions are cleaned up automatically after 30 minutes.
- **Library task lifecycle:** `TaskManager` runs agents directly from Python, with no A2A or MCP server. It owns task IDs, sessions, a SQLite event journal, cancellation, result paging, preferences, and interruption recovery. A2A projects the same task ID and result through its protocol; MCP continues to reach those tasks through managed A2A peers. See the [Python API](docs/api.md#taskmanager-library-api).
- **Remote task lifecycle:** `ShuttleClient.submit()` returns a remote `TaskHandle` immediately. Use `status()`, bounded `wait(timeout)`, `events()`, `result_page()`, `transcript()`, `result()`, or `cancel()`; reopen a task by ID with `client.task(url, task_id)`. A wait timeout does not stop the agent. An optional UUID `request_id` deduplicates retried submissions. Standalone servers can persist tasks with `--task-db`; MCP task tools do this automatically in the workspace's `.agent-shuttle` directory. Completed results survive restart; interrupted work is marked failed without replay. See the [API reference](docs/api.md#taskhandle-and-bridgeevent).

### Safety & Tool Policies
Agent Shuttle defines four standardized tool policies:
- `no_tools`: Disables tool invocations entirely.
- `read_only`: Permits non-mutating search and file reading.
- `workspace_write`: Allows editing files within the designated workspace.
- `full_access`: Explicitly unclamps all tool restrictions and approval prompts.

> [!WARNING]
> `full_access` grants the worker unrestricted tool access for that task. `--agy-dangerously-skip-permissions` enables that capability for the Antigravity server. Keep servers on loopback. Antigravity's explicit scoped `read_only` policy is enforced by its verified `PreToolUse` hook; an implicit/default CLI policy does not provide the same guarantee.

---

## Testing

Agent Shuttle provides a comprehensive offline test suite using fake backends that execute without network access, credentials, or model quota consumption:

```powershell
python -m unittest discover -s tests -v
```

Opt-in integration tests against real models can be executed by specifying target environments (e.g. `BRIDGE_LIVE_OLLAMA_MODEL=qwen3.5:9b` or `BRIDGE_LIVE_AGY_FULL_ACCESS=1`). See [CONTRIBUTING.md](CONTRIBUTING.md) for full instructions.

---

## Documentation Index

- [Getting Started Guide](docs/getting-started.md)
- [API Reference](docs/api.md)
- [Permissions & Safety Guide](docs/permissions.md)
- [Architecture & Protocol Design](docs/architecture.md)
- [Troubleshooting & Diagnostics](docs/troubleshooting.md)
- [Contributing Guidelines](CONTRIBUTING.md)
- [Security Policy](SECURITY.md)

TDQS

B3.1/5.0

Scored across 6 tools

Disambiguation4/5

The three ask_* tools target different agents (generic configured profiles, Antigravity, Codex), and the three get_*_info tools mirror that structure. However, the generic ask_agent/get_agent_info could be confused with the specific ask_antigravity/get_antigravity_info or ask_codex/get_codex_info if those agents are also configured, requiring careful reading of descriptions.

Naming Consistency5/5

All tool names follow a clear ask_<target> or get_<target>_info pattern in consistent snake_case. The generic 'agent' target versus specific agent names is a meaningful distinction rather than a naming inconsistency.

Tool Count5/5

Six tools is well-scoped for a delegation server, covering ask and info operations for both generic and specific agents without redundancy. Each tool earns its place.

Completeness3/5

The core ask and info operations are present, but there is no list_agents tool to discover which profiles are available in BRIDGE_AGENTS_JSON. This is a notable gap for a server centered on configured agents, as users must already know the agent names.

Maintenance

ActivityMaintained
ResponsivenessNo issues