Skip to main content
Glama
README.md
# AgentNave

<p align="center">
  <img src="docs/assets/agentnave-banner.png" alt="AgentNave — a thin local bridge to CLI subagents" width="100%">
</p>

**A thin local bridge from Agent Managers to CLI subagents.**

AgentNave lets an Agent Manager launch Antigravity CLI, Claude Code, CodeBuddy Code, Codex CLI, or
Grok CLI through three lifecycle tools and one read-only provider-description tool. Each provider keeps its native authentication,
configuration, permissions, and session model; AgentNave supplies the adapter and process
supervision around it.

The boundary is intentional. AgentNave does not plan tasks, assign roles, build DAGs, choose
parallelism, review results, synthesize answers, retry work, or manage worktrees. Those decisions
belong to the calling Agent Manager, where the full task context already exists.

AgentNave has no human-facing CLI. The `agentnave-mcp` command is only the STDIO entry point used by
a compatible MCP host.

## What AgentNave owns

| AgentNave | Calling Agent Manager |
| --- | --- |
| Provider command adapters | Planning and task decomposition |
| In-memory invocation lifecycle | Provider and model selection |
| POSIX process groups / Windows Job Objects | Parallelism, review, and synthesis |
| Normalized terminal results | Retries, permissions, and worktrees |
| Private conversation history and read-only Dashboard | Publication decisions and acceptance |

## Requirements

- Windows, macOS, or Linux
- `uv` and Git (for installation from a release tag)
- At least one authenticated provider CLI

`uv` installs AgentNave in an isolated Python 3.12 environment. A separately managed system
Python is not required.

AgentNave uses dedicated POSIX process-group supervisors on macOS/Linux and kill-on-close Windows
Job Objects on Windows. Windows providers are created suspended, assigned to the Job Object, and
only then resumed so descendants cannot start outside the owned process tree.

## Install

Install the MCP runtime and companion [agentnave-manager Skill](skills/agentnave-manager/SKILL.md)
as a pair. The runtime exposes the tools; the Skill teaches the calling Agent how to use other
CLIs. Planning, scheduling, review, retry decisions, and synthesis remain with the calling Agent.

Follow the [installation guide](docs/installation.md) to:

1. Check the existing runtime's source, version and launcher; reuse it when suitable, or install
   the missing runtime with `uv tool`.
2. Inspect the target host's effective MCP registration; reuse matching settings and add only
   a missing entry, with the host's provider exclusions.
3. Reuse the matching Skill or install its complete directory, including `references/`, from
   the same release. Existing shared sources need only a discovery entry for the new host.
4. Verify both MCP tools and Skill discovery; restart connections or sessions after changes.

Hosts under the same OS user and `uv` tool directories share installation files, but each MCP
connection runs its own STDIO service process. Adding a host does not require reinstalling the
runtime or authorize upgrading it for other hosts. The guide covers differences and shared updates.

The guide covers host registration, provider paths, Skill installation, paired upgrades, rollback,
and removal. Provider CLIs must be installed and authenticated separately. `uv` manages only the
runtime; it does not install the Skill or modify provider permissions and configuration.

**The paired release is `v0.13.0`.** Install the runtime and complete Skill directory from that
tag. `uv tool` installs only the runtime; Skill discovery is a separate step. Hosts without
Skill support can still use MCP alone. The older `v0.5.0` tag contains only the runtime.
`v0.10.0` is the first paired release with native Windows support.

AgentNave saves private conversation history under `~/.agentnave/` by default, including ordinary
task results. Provider authentication and configuration remain owned by their respective CLIs.

## Optional discussion rooms

A calling Agent can act as director: choose a seat, send unread public dialogue, inspect the
candidate reply, then publish or discard it. A shared public board collects published dialogue for every participant. A separate director desk
mirrors that board and additionally shows private instructions, drafts, and turn state. The same mechanism works for technical discussions and an AI show.

Participant profiles support all five CLI adapters, including the director host’s own CLI. Each seat
resumes its own native session, receiving only unread public messages and the current instruction.
Provider-specific discussion profiles restrict tools without changing ordinary invocation defaults.
The director can speak publicly as the host. A tavern-style chat displays original avatars, speaker filters,
copyable messages and a separate private director desk. Each show uses its own directory, public history
and seat sessions; the director can switch among opened shows. Participants receive only their own show.
These profiles combine public-context projection and native tool restrictions; they are not an OS sandbox.
Ordinary CLI calls retain their existing defaults. No scheduler, automatic retries, game engine,
new web framework, account system or cloud service is added.

See [discussion-room tools and limitations](docs/discussion-rooms.md). Workbench and discussion
tools are available from v0.11.0.

For independent answers, use `open_discussion(mode="blind")`. Replies are sealed automatically and
revealed together through `decide_discussion_round`, so later speakers cannot copy earlier answers.

### Unified local workbench

`open_workbench` opens a private list of conversations: **一起聊** (discussion), **各自答** (independent answers), and **做任务** (ordinary CLI work). Ordinary `start_agent` calls automatically collect output into task conversations; use `title` for a new conversation and `conversation_id` to group related invocations. Grouping does not share context or alter native tool permissions. Task pages have no public counterpart. The browser remains read-only.

History defaults to `~/.agentnave/`, or the absolute directory in `AGENTNAVE_DATA_DIR`. Records survive restart; active processes and invocation handles do not. `list_conversations`, `read_conversation` and `update_conversation` inspect, rename, end or reopen saved conversations. Several MCP hosts can share one data directory: each indexes saved records read-only and locks a conversation only on its first write, so another service's conversation stays readable (`writable=false` in `list_conversations`) but is never taken over.

Since v0.12.0, `delete_conversation(room_id)` lets the calling Agent, after user authorization,
permanently remove an archived, inactive conversation's AgentNave history, index, dashboard views
and finished invocation handles. Native CLI files and nonempty directories are retained and
reported. Conversation roots retain their stable `.lock` file for single-writer protection;
external root directories are preserved. Archiving alone deletes nothing. There is no
automatic retention policy.

Independent rounds can explicitly mark unanswered seats with `mark_discussion_absent`, after resolving their active invocation. Revealing a partial round labels it incomplete. After reveal, `continue_discussion` continues together in the same native sessions. See [the conversation guide](docs/discussion-rooms.md) and the [manager Skill](skills/agentnave-manager/SKILL.md).

## The MCP surface

This section describes `v0.13.0`, which retains native Windows process-tree supervision and
the fixed five-minute wait contract introduced in `v0.9.0`.
Start, wait and cancel use flat lifecycle responses; callers upgrading from v0.6.0 must also
update their response handling and paired Skill.
Restart the MCP connection after updating the runtime to refresh its schemas.

AgentNave exposes four core tools, plus thirteen conversation and workbench tools. The initial tool metadata contains a compact provider directory;
provider-specific options are returned only when requested. Model defaults live in the Skill's
per-CLI reference files, loaded only for the selected CLI.

### `describe_provider`

Call with the selected `provider` before its first use in the current context. Returns that
provider's permitted status and supported options.
Reuse the result for later calls with the same provider. This tool neither launches a CLI nor
checks installation or authentication; it does not consume provider quota or change the tool list.

For example: `describe_provider({"provider": "grok"})` → `start_agent(...)` → `wait_agent(...)`.

### `start_agent`

Starts one provider invocation and immediately returns an in-memory `invocation_id`. It requires
`provider`, `prompt`, and an absolute existing `cwd`; `session_id` and explicit
`provider_options` are optional.

Supported providers are `antigravity`, `claude`, `codebuddy`, `codex`, `grok`, and `pi`. The Skill provides model and effort defaults for the Manager to pass explicitly through
allowlisted options. User choices override that guidance; omitted options still inherit native
settings. Exclusions are configured per host process, independently of the model it uses. For Codex calls outside a
Git repository, the Manager must pass
`{"skip_git_repo_check": true}` in `provider_options`.

Pi Coding Agent uses `pi --print --mode json`, with prompts on stdin and native session IDs
for continuation. Install it with `npm install -g --ignore-scripts @earendil-works/pi-coding-agent`,
then run `pi` and `/login`. Omitted model and effort options use Pi's native settings; explicit
`effort` maps to `--thinking`. Pi runs with the permissions of its launching process.

### Choosing a model and reasoning effort

For Codex, explicitly pass `{"dangerously_bypass_approvals_and_sandbox": true}` in
`provider_options` to enable native YOLO (`--dangerously-bypass-approvals-and-sandbox`).
This skips approval prompts and disables sandboxing for that invocation, including when
resuming a session. Only enable it when the user authorizes both effects. Passing `false`
or omitting the option adds no flag and inherits native settings; it does not force a sandbox
or approvals back on. The CLI's saved configuration is unchanged.

For Antigravity, explicitly pass `{"dangerously_skip_permissions": true}` in
`provider_options` to enable native YOLO (`--dangerously-skip-permissions`) for one invocation.
This auto-approves all native tool permission requests for that invocation without changing
the CLI's saved configuration. Omit the option to inherit native settings, or pass `false`
to explicitly disable the flag. Only enable it when the user authorizes permission bypass;
`mode` and `sandbox` are separate options and do not imply YOLO.

On macOS, AgentNave prefers the executable CLI bundled with the ChatGPT desktop app,
then the Codex desktop app, checking `/Applications` before `~/Applications` for each.
This follows desktop app updates instead of selecting an independently installed CLI by PATH.
If no executable bundle is found, or on other platforms, AgentNave uses `codex` from PATH.
AgentNave does not install or upgrade Codex, compare version numbers, or change its model settings.

To override the defaults for one task, tell your calling Agent the provider, model ID, and
reasoning effort. For example: “Use Codex CLI with model `gpt-6-astra` and effort `medium`.”
The Agent passes `{"model": "gpt-6-astra", "effort": "medium"}` in `provider_options`.
Only specified fields override the Skill guidance. To keep your choices across tasks, put the
same preference in your calling Agent's personal instructions. To use the CLI's native settings,
explicitly ask the Agent to omit the corresponding options.

When a new model becomes available, use its exact ID from that provider's model list; updating
AgentNave is not required to pass a new model ID. If you maintain a source installation and want
to change the bundled defaults, edit the selected CLI file in
`skills/agentnave-manager/references/`, then reload the Skill in a fresh context. An unavailable
model should be reported rather than silently replaced.

### `wait_agent`

Each request has a fixed five-minute (300-second) wait window; `wait_agent` accepts only
`invocation_id`. Completion or a recognized execution blocker returns early. Expiry leaves the
invocation running; call again with the same ID to continue. If a host yields a background
call handle, resume that call using the host's wait mechanism before issuing another request.

Start, wait and cancel use one flat response: `invocation_id`, `status`, `reason`, `elapsed_ms`,
plus applicable `activity`, `error`, `output`, `output_age_ms` and `session_id` fields. Reasons are
`started`, `wait_elapsed`, `execution_blocked` and `finished`. Running replies include the latest
public reply tail (at most 1,000 Unicode characters), retained across tool events, and its age.
No new public reply means the same tail can recur; unavailable fields are omitted. There is no
cursor, pagination or separate output-reading tool. Final replies are not subject to the tail limit.
Provider usage/cost, native event names, tool call IDs, tool payloads and thinking are not returned.
Public replies may still contain task data; this is not a redaction service.

`execution_blocked` leaves `status=running`: the Manager decides whether to keep waiting or cancel.
Each recognized blocker category wakes once per invocation, avoiding repeated immediate returns
from the same retry loop. Ordinary tool failures, transient retries and silence do not imply a
blocker; unrecognized errors may only become visible in output or the final result. There is no
unsolicited completion/error push without a pending wait request.

Wait expiry never terminates the invocation. AgentNave has no total runtime deadline and
`start_agent` has no runtime-limit parameter. Keep waiting for the result or use `cancel_agent`
to stop work explicitly. Provider-native limits still apply.

### `cancel_agent`

Stops an invocation and returns its terminal result. Use it only when the Manager intends to end
active provider work; `wait_agent` observes without cancelling.

All tools publish input and output JSON Schemas. Agent-correctable request errors are MCP Tool
errors with retry guidance; provider launch and execution outcomes remain structured Invocation
Results.

## Lifecycle and security

Invocation handles live only in the current MCP server process. When the server stops, AgentNave
makes a best-effort attempt to terminate processes that remain in the provider process tree. A
restart cannot recover old handles, but a retained provider `session_id` can be supplied to a new
`start_agent` call.

Running responses help inspect health and direction using a bounded public reply and native
activity. They do not show every active operation or guarantee progress. Raw streams remain
bounded in process memory; no output log/database is added. A 1,000-character reply tail is kept
per invocation for waiting responses. Final results remain available for this server process.

Use the companion Skill's working-directory and target-path guidance to select `cwd` and describe
the task in the prompt. Project rules load through the CLI's native mechanism.
Antigravity can choose a different terminal `Cwd`:
state the absolute project directory in the handoff, require that terminal directory explicitly,
and verify it with `pwd` before project operations. This is a behavioral instruction, not enforced
directory isolation. `provider_options.project` selects a native project ID or name; it is not a
working-directory override. The companion Skill guides task handoffs and keeps
intermediate files in OS temporary storage without imposing Markdown or a result-file format.
AgentNave itself uses stdin/in-memory output except for Grok's temporary prompt file, which is
removed after use. Provider-owned history and caches remain under provider control.

AgentNave is not a sandbox. On POSIX, a same-user provider with command permission can deliberately
daemonize, kill its supervisor, or otherwise escape ordinary process-group cleanup. Windows Job
Objects provide tree ownership but do not isolate the provider from the user account or the rest of
the machine. Provider-native permissions remain the security boundary; use OS-level isolation when
adversarial containment is required.

## Verify

These checks are for a development checkout, not the `uv tool` installation above. See
[CONTRIBUTING.md](CONTRIBUTING.md) for the complete contributor workflow.

```bash
uv sync --locked --all-groups
uv run ruff format --check .
uv run ruff check .
uv run pyright
uv run pytest
```

## Contributing and security

Contributions are welcome through GitHub Issues and pull requests. See
[CONTRIBUTING.md](CONTRIBUTING.md) for the development workflow and validation requirements.

Do not report security vulnerabilities in a public Issue. Follow [SECURITY.md](SECURITY.md) to use
the repository's private vulnerability reporting channel.

## License

AgentNave is licensed under the [MIT License](LICENSE).

TDQS

A4.5/5.0

Scored across 4 tools

Disambiguation5/5

Each tool has a distinct lifecycle role: start_agent launches, wait_agent observes, cancel_agent stops, and describe_provider inspects provider configuration. There is no meaningful overlap between any two tools.

Naming Consistency4/5

Three tools follow the verb_agent pattern (start_agent, wait_agent, cancel_agent), while describe_provider uses a different object. The imperative verb_noun style is consistent and readable, with only a minor deviation.

Tool Count5/5

Four tools is well-scoped for an agent lifecycle management server. Each tool covers a necessary operation without redundancy or unnecessary surface area.

Completeness4/5

The core lifecycle is covered: start, wait, cancel, and provider capability discovery. A dedicated listing/status tool is absent, but wait_agent and cancel_agent can retrieve state by ID, so agents can work around that minor gap.

Maintenance

ActivityMaintained
ResponsivenessNo issues