Skip to main content
Glama

MCP Agent Mail

Agent Mail Showcase

"It's like gmail for your coding agents!"

A mail-like coordination layer for coding agents, exposed as an HTTP-only FastMCP server. It gives agents memorable identities, an inbox/outbox, searchable message history, and voluntary file reservation "leases" to avoid stepping on each other.

Think of it as asynchronous email + directory + change-intent signaling for your agents, backed by Git (for human-auditable artifacts) and SQLite (for indexing and queries).

Status: Under active development. The design is captured in detail in docs/planning/project_idea_and_guide.md (start with the original prompt at the top of that file).

Why this exists

Modern projects often run multiple coding agents at once (backend, frontend, scripts, infra). Without a shared coordination fabric, agents:

  • Overwrite each other's edits or panic on unexpected diffs

  • Miss critical context from parallel workstreams

  • Require humans to "liaison" messages across tools and teams

This project provides a lightweight, interoperable layer so agents can:

  • Register a temporary-but-persistent identity (e.g., GreenCastle)

  • Send/receive GitHub-Flavored Markdown messages with images

  • Search, summarize, and thread conversations

  • Declare advisory file reservations (leases) on files/globs to signal intent

  • Inspect a directory of active agents, programs/models, and activity

It's designed for: FastMCP clients and CLI tools (Claude Code, Codex, Gemini CLI, Factory Droid, etc.) coordinating across one or more codebases.

Related MCP server: Shift MCP Server

From Idea Spark to Shipping Swarm

If a blank repo feels daunting, follow the field-tested workflow we documented in docs/planning/project_idea_and_guide.md (“Appendix: From Blank Repo to Coordinated Swarm”):

  • Ideate fast: Write a scrappy email-style blurb about the problem, desired UX, and any must-have stack picks (≈15 minutes).

  • Promote it to a plan: Feed that blurb to GPT-5 Pro (and optionally Grok4 Heavy / Opus 4.1) until you get a granular Markdown plan, then iterate on the plan file while it’s still cheap to change. The Markdown Web Browser sample plan shows the level of detail to aim for.

  • Codify the rules: Clone a tuned AGENTS.md, add any tech-specific best-practice guides, and let Codex scaffold the repo plus Beads tasks straight from the plan.

  • Spin up the swarm: Launch multiple Codex panes (or any agent mix), register each identity with Agent Mail, and have them acknowledge AGENTS.md, the plan document, and the Beads backlog before touching code.

  • Keep everyone fed: Reuse the canned instruction cadence from the tweet thread or, better yet, let the commercial Companion app’s Message Stacks broadcast those prompts automatically so you never hand-feed panes again.

Watch the full 23-minute walkthrough (https://youtu.be/68VVcqMEDrs?si=pCm6AiJAndtZ6u7q) to see the loop in action.

Productivity Math & Automation Loop

One disciplined hour of GPT-5 Codex—when it isn’t waiting on human prompts—often produces 10–20 “human hours” of work because the agents reason and type at machine speed. Agent Mail multiplies that advantage in two layers:

  1. Base OSS server: Git-backed mailboxes, advisory file reservations, Typer CLI helpers, and searchable archives keep independent agents aligned without babysitting. Every instruction, lease, and attachment is auditable.

  2. Companion stack (commercial): The iOS app + host automation can provision, pair, and steer heterogeneous fleets (Claude Code, Codex, Gemini CLI, Factory Droid, etc.) from your phone using customizable Message Stacks, Human Overseer broadcasts, Beads awareness, and plan editing tools—no manual tmux choreography required. The automation closes the loop by scheduling prompts, honoring Limited Mode, and enforcing Double-Arm confirmations for destructive work.

Result: you invest 1–2 hours of human supervision, but dozens of agent-hours execute in parallel with clear audit trails and conflict-avoidance baked in.

TLDR Quickstart

One-line installer

curl -fsSL "https://raw.githubusercontent.com/Dicklesworthstone/mcp_agent_mail/main/scripts/install.sh?$(date +%s)" | bash -s -- --yes

What this does:

  • Installs uv if missing and updates your PATH for this session

  • Installs jq if missing (needed for safe config merging; auto-detects your package manager)

  • Creates a Python 3.14 virtual environment and installs dependencies with uv

  • Runs the auto-detect integration to wire up supported agent tools

  • Starts the MCP HTTP server on port 8765 and prints a masked bearer token

  • Creates helper scripts under scripts/ (including run_server_with_token.sh)

  • Adds an am shell alias to your .zshrc or .bashrc for quick server startup (just type am in a new terminal!)

  • Installs Beads Rust (br), a Rust reimplementation of the Beads task tracker, and creates a bd shell alias pointing to br for backwards compatibility. This replaces any existing bd (Go) installation. Pass --skip-beads to opt out. See beads_rust for details on CLI differences.

  • Installs/updates the Beads Viewer bv TUI for interactive task browsing and AI-friendly robot commands (pass --skip-bv to opt out)

  • Prints a short on-exit summary of each setup step so you immediately know what changed

Prefer a specific location or options? Add flags like --dir <path>, --project-dir <path>, --no-start, --start-only, --port <number>, or --token <hex>.

Already have Beads or Beads Viewer installed? Append --skip-beads and/or --skip-bv to bypass automatic installation.

Important: Beads Rust (br) Replaces Beads Go (bd)

The installer automatically replaces bd (the original Go-based Beads CLI) with br (Beads Rust):

  1. br is the actively maintained version. Beads Rust is a complete reimplementation with ongoing development, while the original Go version is no longer actively maintained.

  2. Existing bd users get automatic aliasing. The installer creates a shell alias so that bd commands continue to work by redirecting to br. Your existing workflows and muscle memory are preserved.

  3. A migration skill is installed for agents. AI coding agents receive a bd-br-migration skill that helps them adapt to any CLI differences between the two implementations.

  4. Same data format, compatible workflows. Both implementations use the same .beads/issues.jsonl format, so your existing Beads data remains fully compatible.

To opt out of this replacement, pass --skip-beads to the installer. See the beads_rust repository for detailed documentation on CLI differences.

Starting the server in the future

After installation, you can start the MCP Agent Mail server from anywhere by simply typing:

am

That's it! The am alias (added to your .zshrc or .bashrc during installation) automatically:

  1. Changes to the MCP Agent Mail directory

  2. Runs the server startup script (which uses uv run to handle the virtual environment)

  3. Loads your saved bearer token from .env and starts the HTTP server

Note: If you just ran the installer, open a new terminal or run source ~/.zshrc (or source ~/.bashrc) to load the alias.

Port conflicts? Use --port to specify a different port (default: 8765):

Install with custom port:

curl -fsSL "https://raw.githubusercontent.com/Dicklesworthstone/mcp_agent_mail/main/scripts/install.sh?$(date +%s)" | bash -s -- --port 9000 --yes

Or use the CLI command after installation: uv run python -m mcp_agent_mail.cli config set-port 9000


### If you want to do it yourself

Clone the repo, set up and install with uv in a python 3.14 venv (install uv if you don't have it already), and then run `scripts/automatically_detect_all_installed_coding_agents_and_install_mcp_agent_mail_in_all.sh`. This will automatically set things up for your various installed coding agent tools and start the MCP server on port 8765. If you want to run the MCP server again in the future, simply run `scripts/run_server_with_token.sh`:

Install uv (if you don't have it already):

```bash
curl -LsSf https://astral.sh/uv/install.sh | sh
export PATH="$HOME/.local/bin:$PATH"

# Clone the repo
git clone https://github.com/Dicklesworthstone/mcp_agent_mail
cd mcp_agent_mail

# Create a Python 3.14 virtual environment and install dependencies
# Note: If you have an older uv version, run `uv self update` first
uv python install 3.14
uv venv -p 3.14
source .venv/bin/activate
uv sync

# Detect installed coding agents, integrate, and start the MCP server on port 8765
scripts/automatically_detect_all_installed_coding_agents_and_install_mcp_agent_mail_in_all.sh

# Later, to run the MCP server again with the same token
scripts/run_server_with_token.sh

# Now, simply launch Codex-CLI or Claude Code or other agent tools in other consoles; they should have the mail tool available. See below for a ready-made chunk of text you can add to the end of your existing AGENTS.md or CLAUDE.md files to help your agents better utilize the new tools.

# Change port after installation
uv run python -m mcp_agent_mail.cli config set-port 9000

Ready-Made Blurb to Add to Your AGENTS.md or CLAUDE.md Files:

## MCP Agent Mail: coordination for multi-agent workflows

What it is
- A mail-like layer that lets coding agents coordinate asynchronously via MCP tools and resources.
- Provides identities, inbox/outbox, searchable threads, and advisory file reservations, with human-auditable artifacts in Git.

Why it's useful
- Prevents agents from stepping on each other with explicit file reservations (leases) for files/globs.
- Keeps communication out of your token budget by storing messages in a per-project archive.
- Offers quick reads (`resource://inbox/...`, `resource://thread/...`) and macros that bundle common flows.

How to use effectively
1) Same repository
   - Register an identity: call `ensure_project`, then `register_agent` using this repo's absolute path as `project_key`.
   - Reserve files before you edit: `file_reservation_paths(project_key, agent_name, ["src/**"], ttl_seconds=3600, exclusive=true)` to signal intent and avoid conflict.
   - Communicate with threads: use `send_message(..., thread_id="FEAT-123")`; check inbox with `fetch_inbox` and acknowledge with `acknowledge_message`.
   - Read fast: `resource://inbox/{Agent}?project=<abs-path>&limit=20&agent_token=<registration_token>` or `resource://thread/{id}?project=<abs-path>&agent=<Agent>&agent_token=<registration_token>&include_bodies=true` unless the current MCP session already authenticated as that agent.
   - Tip: set `AGENT_NAME` in your environment so the pre-commit guard can block commits that conflict with others' active exclusive file reservations.

2) Across different repos in one project (e.g., Next.js frontend + FastAPI backend)
   - Option A (single project bus): register both sides under the same `project_key` (shared key/path). Keep reservation patterns specific (e.g., `frontend/**` vs `backend/**`).
   - Option B (separate projects): each repo has its own `project_key`; use `macro_contact_handshake` or `request_contact`/`respond_contact` to link agents, then message directly. Keep a shared `thread_id` (e.g., ticket key) across repos for clean summaries/audits.

Macros vs granular tools
- Prefer macros when you want speed or are on a smaller model: `macro_start_session`, `macro_prepare_thread`, `macro_file_reservation_cycle`, `macro_contact_handshake`.
- Use granular tools when you need control: `register_agent`, `file_reservation_paths`, `send_message`, `fetch_inbox`, `acknowledge_message`.

Common pitfalls
- "from_agent not registered": always `register_agent` in the correct `project_key` first.
- "FILE_RESERVATION_CONFLICT": adjust patterns, wait for expiry, or use a non-exclusive reservation when appropriate.
- Auth errors: if JWT+JWKS is enabled, include a bearer token with a `kid` that matches server JWKS; static bearer is used only when JWT is disabled.

Integrating with Beads (dependency-aware task planning)

Beads is a lightweight task planner that complements Agent Mail by keeping status and dependencies in one place while Mail handles messaging, file reservations, and audit trails.

Note on implementations: The MCP Agent Mail installer installs Beads Rust (br), a Rust reimplementation, and creates a bd alias for backwards compatibility. The original Go implementation is at steveyegge/beads. Both share the same data format (.beads/issues.jsonl) but have some CLI differences. Use --skip-beads during installation if you prefer to manage this yourself.

Highlights:

  • Beads owns task prioritization; Agent Mail carries the conversations and artifacts.

  • Shared identifiers (e.g., bd-123) keep Beads issues, Mail threads, and commits aligned.

  • The br CLI (aliased as bd) provides similar functionality to the original with some enhancements.

Copy/paste blurb for agent-facing docs (leave as-is for reuse):


## Integrating with Beads (dependency-aware task planning)

Beads provides a lightweight, dependency-aware issue database and a CLI (`bd`) for selecting "ready work," setting priorities, and tracking status. It complements MCP Agent Mail's messaging, audit trail, and file-reservation signals. Project: [steveyegge/beads](https://github.com/steveyegge/beads)

Recommended conventions
- **Single source of truth**: Use **Beads** for task status/priority/dependencies; use **Agent Mail** for conversation, decisions, and attachments (audit).
- **Shared identifiers**: Use the Beads issue id (e.g., `bd-123`) as the Mail `thread_id` and prefix message subjects with `[bd-123]`.
- **Reservations**: When starting a `bd-###` task, call `file_reservation_paths(...)` for the affected paths; include the issue id in the `reason` and release on completion.

Typical flow (agents)
1) **Pick ready work** (Beads)
   - `bd ready --json` → choose one item (highest priority, no blockers)
2) **Reserve edit surface** (Mail)
   - `file_reservation_paths(project_key, agent_name, ["src/**"], ttl_seconds=3600, exclusive=true, reason="bd-123")`
3) **Announce start** (Mail)
   - `send_message(..., thread_id="bd-123", subject="[bd-123] Start: <short title>", ack_required=true)`
4) **Work and update**
   - Reply in-thread with progress and attach artifacts/images; keep the discussion in one thread per issue id
5) **Complete and release**
   - `bd close bd-123 --reason "Completed"` (Beads is status authority)
   - `release_file_reservations(project_key, agent_name, paths=["src/**"])`
   - Final Mail reply: `[bd-123] Completed` with summary and links

Mapping cheat-sheet
- **Mail `thread_id`** ↔ `bd-###`
- **Mail subject**: `[bd-###] …`
- **File reservation `reason`**: `bd-###`
- **Commit messages (optional)**: include `bd-###` for traceability

Event mirroring (optional automation)
- On `bd update --status blocked`, send a high-importance Mail message in thread `bd-###` describing the blocker.
- On Mail "ACK overdue" for a critical decision, add a Beads label (e.g., `needs-ack`) or bump priority to surface it in `bd ready`.

Pitfalls to avoid
- Don't create or manage tasks in Mail; treat Beads as the single task queue.
- Always include `bd-###` in message `thread_id` to avoid ID drift across tools.

Prefer automation? Run uv run python -m mcp_agent_mail.cli docs insert-blurbs to scan your code directories for AGENTS.md/CLAUDE.md files and append the latest Agent Mail + Beads snippets with per-project confirmation. The installer also offers to launch this helper right after setup so you can take care of onboarding docs immediately.

Beads Viewer (bv) — AI-Friendly Task Analysis

The Beads Viewer (bv) is a fast terminal UI for Beads projects that also provides robot flags designed specifically for AI agent integration. Project: Dicklesworthstone/beads_viewer

Why bv for Agents?

While bd (Beads CLI) handles task CRUD operations, bv provides precomputed graph analytics that help agents make intelligent prioritization decisions:

  • PageRank scores: Identify high-impact tasks that unblock the most downstream work

  • Critical path analysis: Find the longest dependency chain to completion

  • Cycle detection: Spot circular dependencies before they cause deadlocks

  • Parallel track planning: Determine which tasks can run concurrently

Instead of agents parsing .beads/issues.jsonl directly or attempting to compute graph metrics (risking hallucinated results), they can call bv's deterministic robot flags and get JSON output they can trust. Legacy .beads/beads.jsonl is deprecated in this repo.

Robot Flags for AI Integration

Flag

Output

Agent Use Case

bv --robot-help

All AI-facing commands

Discovery / capability check

bv --robot-insights

PageRank, betweenness, HITS, critical path, cycles

Quick triage: "What's most impactful?"

bv --robot-plan

Parallel tracks, items per track, unblocks lists

Execution planning: "What can run in parallel?"

bv --robot-priority

Priority recommendations with reasoning + confidence

Task selection: "What should I work on next?"

bv --robot-recipes

Available filter presets (actionable, blocked, etc.)

Workflow setup: "Show me ready work"

bv --robot-diff --diff-since <ref>

Changes since commit/date, new/closed items, cycles

Progress tracking: "What changed?"

Example: Agent Task Selection Workflow

# 1. Get priority recommendations with reasoning
bv --robot-priority
# Returns JSON with ranked tasks, impact scores, and confidence levels

# 2. Check what completing a task would unblock
bv --robot-plan
# Returns parallel tracks showing dependency chains

# 3. After completing work, check what changed
bv --robot-diff --diff-since "1 hour ago"
# Returns new items, closed items, cycle changes

When to Use bv vs bd

Tool

Best For

bd

Creating, updating, closing tasks; bd ready for simple "what's next"

bv

Graph analysis, impact assessment, parallel planning, change tracking

Rule of thumb: Use bd for task operations, use bv for task intelligence.

Integration with Agent Mail

Combine bv insights with Agent Mail coordination:

  1. Agent A runs bv --robot-priority → identifies bd-42 as highest-impact

  2. Agent A reserves files: file_reservation_paths(..., reason="bd-42")

  3. Agent A announces: send_message(..., thread_id="bd-42", subject="[bd-42] Starting high-impact refactor")

  4. Other agents see the reservation and Mail announcement, pick different tasks

  5. Agent A completes, runs bv --robot-diff to report downstream unblocks

This creates a feedback loop where graph intelligence drives coordination.

Core ideas (at a glance)

  • FastMCP server over Streamable HTTP (primary transport; no SSE). A STDIO transport (serve-stdio) is also available for direct CLI integration.

  • Dual persistence model:

    • Human-readable markdown in a per-project Git repo for every canonical message and per-recipient inbox/outbox copy

    • SQLite with FTS5 for fast search, directory queries, and file reservations/leases

  • "Directory/LDAP" style queries for agents; memorable adjective+noun names

  • Advisory file reservations for editing surfaces; optional pre-commit guard

  • Resource layer for convenient reads (e.g., resource://inbox/{agent})

Typical use cases

  • Multiple agents splitting a large refactor across services while staying in sync

  • Frontend and backend teams of agents coordinating thread-by-thread

  • Protecting critical migrations with exclusive file reservations and a pre-commit guard

  • Searching and summarizing long technical discussions as threads evolve

  • Discovering and linking related projects (e.g., frontend/backend) through AI-powered suggestions

Workflow FAQ

Do I still need the tmux broadcast script to “feed” every Codex pane?

No. The historical zsh loop from the tweet thread is still handy if you are running the OSS stack by itself, but the AgentMail Companion system now automates that cadence with Message Stacks. Once the companion host services are installed, you queue presets (builder loop, reviewer sweep, test focus, etc.) from the iOS app or CLI and the automation fans those instructions out to every enrolled agent—without touching tmux.

Architecture

graph LR
  A[Agents]
  S[Server]
  G[Git repo]
  Q[SQLite FTS5]

  A -->|HTTP tools/resources| S
  S -->|writes/reads| G
  S -->|indexes/queries| Q

  subgraph GitTree["Git tree"]
    GI1[agents/profile.json]
    GI2[agents/mailboxes/...]
    GI3[messages/YYYY/\nMM/\nid.md]
    GI4[file_reservations/\nsha1.json]
    GA[attachments/xx/\nsha1.webp]
  end

  G --- GI1
  G --- GI2
  G --- GI3
  G --- GI4
  G --- GA

Web UI (human-facing mail viewer)

The server ships a lightweight, server-rendered Web UI for humans. It lets you browse projects, agents, inboxes, single messages, attachments, file reservations, and perform full-text search with FTS5 when available (with an automatic LIKE fallback).

  • Where it lives: built into the HTTP server in mcp_agent_mail.http under the /mail path.

  • Who it's for: humans reviewing activity; agents should continue to use the MCP tools/resources API.

Launching the Web UI

Recommended (simple):

scripts/run_server_with_token.sh
# then open http://127.0.0.1:8765/mail

Advanced (manual commands):

uv run python -m mcp_agent_mail.http --host 127.0.0.1 --port 8765
# or:
uv run uvicorn mcp_agent_mail.http:create_app --factory --host 127.0.0.1 --port 8765

Auth notes:

  • GET pages in the UI are not gated by the RBAC middleware (it classifies POSTed MCP calls only), but if you set a bearer token the separate BearerAuth middleware protects all routes by default.

  • For local dev, set HTTP_ALLOW_LOCALHOST_UNAUTHENTICATED=true (and optionally HTTP_BEARER_TOKEN), so localhost can load the UI without headers.

  • Health endpoints are always open at /health/*.

Routes and what you can do

  • /mail (Unified inbox + Projects + Related Projects Discovery)

    • Shows a unified, reverse-chronological inbox of recent messages across all projects with excerpts, relative timestamps, sender/recipients, and project badges.

    • Below the inbox, lists all projects (slug, human name, created time) with sibling suggestions.

    • Suggests likely sibling projects when two slugs appear to be parts of the same product (e.g., backend vs. frontend). Suggestions are ranked with heuristics and, when LLM_ENABLED=true, an LLM pass across key docs (README.md, AGENTS.md, etc.).

    • Humans can Confirm Link or Dismiss suggestions from the dashboard. Confirmed siblings become highlighted badges but do not automatically authorize cross-project messaging; agents must still establish AgentLink approvals via request_contact/respond_contact.

  • /mail/projects (Projects index)

    • Dedicated projects list view; click a project to drill in.

  • /mail/{project} (Project overview + search + agents)

    • Rich search form with filters:

      • Scope: subject/body/both, Order: relevance or time, optional "boost subject".

      • Query tokens: supports subject:foo, body:"multi word", quoted phrases, and bare terms.

      • Uses FTS5 bm25 scoring when available; otherwise falls back to SQL LIKE on subject/body with your chosen scope.

    • Results show subject, sender, created time, thread id, and a highlighted snippet when using FTS.

    • Agents panel shows registered agents for the project with a link to each inbox.

    • Quick links to File Reservations and Attachments for the project header.

  • /mail/{project}/inbox/{agent} (Inbox for one agent)

    • Reverse-chronological list with subject, sender, created time, importance badge, thread id.

    • Pagination (?page=N&limit=M).

  • /mail/{project}/message/{id} (Message detail)

    • Shows subject, sender, created time, importance, recipients (To/Cc/Bcc), thread messages.

    • Body rendering:

      • If the server pre-converted markdown to HTML, it's sanitized with Bleach (limited tags/attributes, safe CSS via CSSSanitizer) and then displayed.

      • Otherwise markdown is rendered client-side with Marked + Prism for code highlighting.

    • Attachments are referenced from the message frontmatter (WebP files or inline data URIs).

  • /mail/{project}/search?q=... (Dedicated search page)

    • Same query syntax as the project overview search, with a token "pill" UI for assembling/removing filters.

  • /mail/{project}/file_reservations (File Reservations list)

    • Displays active and historical file reservations (exclusive/shared, path pattern, timestamps, released/expired state).

  • /mail/{project}/attachments (Messages with attachments)

    • Lists messages that contain any attachments, with subject and created time.

  • /mail/unified-inbox (Cross-project activity)

    • Shows recent messages across all projects with thread counts and sender/recipients.

Human Overseer: Sending Messages to Agents

Sometimes a human operator needs to guide or redirect agents directly, whether to handle an urgent issue, provide clarification, or adjust priorities. The Human Overseer feature provides a web-based message composer that lets humans send high-priority messages to any combination of agents in a project.

Access: Click the prominent "Send Message" button (with the Overseer badge) in the header of any project view (/mail/{project}), or navigate directly to /mail/{project}/overseer/compose.

What Makes Overseer Messages Special

  1. Automatic Preamble: Every message includes a formatted preamble that clearly identifies it as coming from a human operator and instructs agents to:

    • Pause current work temporarily

    • Prioritize the human's request over existing tasks

    • Resume original plans afterward (unless modified by the instructions)

  2. High Priority: All overseer messages are automatically marked as high importance, ensuring they stand out in agent inboxes.

  3. Policy Bypass: Overseer messages bypass normal contact policies, so humans can always reach any agent regardless of their contact settings.

  4. Special Sender Identity: Messages come from a special agent named "HumanOverseer" (automatically created per project) with:

    • Program: WebUI

    • Model: Human

    • Contact Policy: open

The Message Preamble

Every overseer message begins with this preamble (automatically prepended):

---

🚨 MESSAGE FROM HUMAN OVERSEER 🚨

This message is from a human operator overseeing this project. Please prioritize
the instructions below over your current tasks.

You should:
1. Temporarily pause your current work
2. Complete the request described below
3. Resume your original plans afterward (unless modified by these instructions)

The human's guidance supersedes all other priorities.

---

[Your message body follows here]

Using the Composer

The composer interface provides:

  • Recipient Selection: Checkbox grid of all registered agents (with "Select All" / "Clear" shortcuts)

  • Subject Line: Required, shown in agent inboxes

  • Message Body: GitHub-flavored Markdown editor with preview

  • Thread ID (optional): Continue an existing conversation or start a new one

  • Preamble Preview: See exactly how your message will appear to agents

Example Use Cases

Urgent Issue:

Subject: Urgent: Stop migration and revert changes

The database migration in PR #453 is causing data corruption in staging.

Please:
1. Immediately stop any migration-related work
2. Revert commits from the last 2 hours
3. Wait for my review before resuming

I'm investigating the root cause now.

Priority Adjustment:

Subject: New Priority: Security Vulnerability

A critical security vulnerability was just disclosed in our auth library.

Drop your current tasks and:
1. Update `auth-lib` to version 2.4.1 immediately
2. Review all usages in src/auth/
3. Run the full security test suite
4. Report status in thread #892

This takes precedence over the refactoring work.

Clarification:

Subject: Clarification on API design approach

I see you're debating REST vs. GraphQL in thread #234.

Go with REST for now because:
- Our frontend team has more REST experience
- GraphQL adds complexity we don't need yet
- We can always add GraphQL later if needed

Resume the API implementation with REST.

How Agents See Overseer Messages

When agents check their inbox (via fetch_inbox or resource://inbox/{name}), overseer messages appear like any other message but with:

  • Sender: HumanOverseer

  • Importance: high (displayed prominently)

  • Body: Starts with the overseer preamble, followed by the human's message

  • Visual cues: In the Web UI, these messages may have special highlighting (future enhancement)

Agents can reply to overseer messages just like any other message, continuing the conversation thread.

Technical Details

  • Storage: Overseer messages are stored identically to agent-to-agent messages (Git + SQLite)

  • Git History: Fully auditable; message appears in messages/YYYY/MM/{id}.md with commit history

  • Thread Continuity: Can be part of existing threads or start new ones

  • No Authentication Bypass: The overseer compose form still requires proper HTTP server authentication (if enabled)

Design Philosophy

The Human Overseer feature is designed to be:

  • Explicit: Agents clearly know when guidance comes from a human vs. another agent

  • Respectful: Instructions acknowledge agents have existing work and shouldn't just "drop everything" permanently

  • Temporary: Agents are told to resume original plans once the human's request is complete

  • Flexible: Humans can override this guidance directly in their message body

This creates a clear hierarchy (human → agents) while maintaining the collaborative, respectful tone of the agent communication system.

The Projects index (/mail) features an AI-powered discovery system that intelligently suggests which projects should be linked together, such as frontend + backend or related microservices.

How Discovery Works

1. Smart Analysis The system uses multiple signals to identify relationships:

  • Pattern matching: Compares project names and paths (e.g., "my-app-frontend" ↔ "my-app-backend")

  • AI understanding (when LLM_ENABLED=true): Reads README.md, AGENTS.md, and other docs to understand each project's purpose and detect natural relationships

  • Confidence scoring: Ranks suggestions from 0-100% with clear rationales

2. Beautiful Suggestions Related projects appear as polished cards on your dashboard with:

  • 🎯 Visual confidence indicators showing match strength

  • 💬 AI-generated rationales explaining the relationship

  • ✅ Confirm Link - accept the suggestion

  • ✖️ Dismiss - hide irrelevant matches

3. Quick Navigation Once confirmed, both projects display interactive badges for instant navigation between related codebases.

Why Suggestions, Not Auto-Linking?

TL;DR: We keep you in control. Discovery helps you find relationships; explicit approvals control who can actually communicate.

Agent Mail uses agent-centric messaging: every message follows explicit permission chains:

Send Message → Find Recipient → Check AgentLink Approval → Deliver

This design ensures:

  • Security: No accidental cross-project message delivery

  • Transparency: You always know who can talk to whom

  • Audit trails: All communication paths are explicitly approved

Why not auto-link with AI? If we let an LLM automatically authorize messaging between projects, we'd be:

  • ❌ Bypassing contact policies without human oversight

  • ❌ Risking message misdelivery to unintended recipients

  • ❌ Creating invisible routing paths that are hard to audit

  • ❌ Potentially linking ambiguously-named projects incorrectly

Instead, we give you discovery + control:

  • ✅ AI suggests likely relationships (safe, read-only analysis)

  • ✅ You confirm what makes sense (one click)

  • ✅ Agents still use request_contact / respond_contact for actual messaging permissions

  • ✅ Clear separation: discovery ≠ authorization

The Complete Workflow

1. System suggests: "These projects look related" (AI analysis)
           ↓
2. You confirm: "Yes, link them" (updates UI badges)
           ↓
3. Agents request: request_contact(from_agent, to_agent, to_project)
           ↓
4. You approve: respond_contact(accept=true)
           ↓
5. Messages flow: Agents can now communicate across projects

Think of it like LinkedIn: The system suggests connections, but only you decide who gets to send messages.

Search syntax (UI)

The UI shares the same parsing as the API's _parse_fts_query:

  • Field filters: subject:login, body:"api key"

  • Phrase search: "build plan"

  • Combine terms: login AND security (FTS)

  • Fallback LIKE: scope determines whether subject, body, or both are searched

Prerequisites to see data

The UI reads from the same SQLite + Git artifacts as the MCP tools. To populate content:

  1. Ensure a project exists (via tool call or CLI):

    • Ensure/create project: ensure_project(human_key)

  2. Register one or more agents: register_agent(project_key, program, model, name?)

  3. Send messages: send_message(...) (attachments and inline images are supported; images may be converted to WebP).

Once messages exist, visit /mail, click your project, then open an agent inbox or search.

Implementation and dependencies

  • Templates live in src/mcp_agent_mail/templates/ and are rendered by Jinja2.

  • Markdown is converted with markdown2 on the server where possible; HTML is sanitized with Bleach (with CSS sanitizer when available).

  • Tailwind CSS, Lucide icons, Alpine.js, Marked, and Prism are loaded via CDN in base.html for a modern look without a frontend build step.

  • All rendering is server-side; there's no SPA router. Pages degrade cleanly without JavaScript.

Security considerations

  • HTML sanitization: Only a conservative set of tags/attributes are allowed; CSS is filtered. Links are limited to http/https/mailto/data.

  • Auth: Use bearer token or JWT when exposing beyond localhost. For local dev, enable localhost bypass as noted above.

  • Rate limiting (optional): Token-bucket limiter can be enabled; UI GET requests are light and unaffected by POST limits.

Troubleshooting the UI

  • Blank page or 401 on localhost: Either unset HTTP_BEARER_TOKEN or set HTTP_ALLOW_LOCALHOST_UNAUTHENTICATED=true.

  • No projects listed: Create one with ensure_project.

  • Empty inbox: Verify recipient names match exactly and messages were sent to that agent.

  • Search returns nothing: Try simpler terms or the LIKE fallback (toggle scope/body).

Static Mailbox Export (Share & Distribute Archives)

The share command group exports a project’s mailbox into a portable, read‑only bundle that anyone can review in a browser. It’s designed for auditors, stakeholders, or teammates who need to browse threads, search history, or prove delivery timelines without spinning up the full MCP Agent Mail stack.

Why export to static bundles?

Compliance and audit trails: Deliver immutable snapshots of project communication to auditors or compliance officers. The static bundle includes cryptographic signatures for tamper-evident distribution.

Stakeholder review: Share conversation history with product managers, executives, or external consultants who don't need write access. They can browse messages, search threads, and view attachments in their browser without authentication.

Offline access: Create portable archives for air-gapped environments, disaster recovery backups, or situations where internet connectivity is unreliable.

Long-term archival: Preserve project communication in a format that will remain readable decades from now. Static HTML requires no database server, no runtime dependencies, and survives software obsolescence better than proprietary formats.

Secure distribution: Encrypt bundles with age for confidential projects. Only recipients with the private key can decrypt and view the contents.

What's included in an export

Each bundle contains:

  • Self-contained: Everything ships in a single directory (HTML, CSS/JS, SQLite snapshot, attachments). Drop it on a static host or open it locally.

  • Rich reader UI: Gmail-style inbox with project filters, search, and full-thread rendering—each message is shown with its metadata and Markdown body, just like in the live web UI.

  • Fast search & filters: FTS-backed search and precomputed per-message summaries keep scrolling and filtering responsive even with large archives.

  • Verifiable integrity: SHA-256 hashes for every asset plus optional Ed25519 signing make authenticity and tampering checks straightforward.

  • Chunk-friendly archives: Large databases can be chunked for httpvfs streaming; a companion chunks.sha256 file lists digests for each chunk so clients can trust streamed blobs without recomputing hashes.

  • One-click hosting: The interactive wizard can publish straight to GitHub Pages or Cloudflare Pages, or you can serve the bundle locally with the CLI preview command.

Disaster Recovery Archives (archive commands)

Use the archive subcommands when you need a restorable snapshot (not just a read-only share bundle). Each ZIP under ./archived_mailbox_states/ includes:

  • A SQLite snapshot processed by the same cleanup pipeline as share, but using the archive scrub preset so ack/read state, recipients, attachments, and message bodies remain untouched.

  • A byte-for-byte copy of the storage Git repo (STORAGE_ROOT), preserving markdown artifacts, attachments, and hook scripts.

Quick ref

# Save current state (defaults to the lossless preset)
uv run python -m mcp_agent_mail.cli archive save --label nightly

# List available restore points (JSON is handy for scripts)
uv run python -m mcp_agent_mail.cli archive list --json

# Restore after a disaster (backs up any existing DB/storage before overwriting)
uv run python -m mcp_agent_mail.cli archive restore archived_mailbox_states/<file>.zip --force

During restore the CLI:

  1. Extracts the ZIP into a temp directory.

  2. Moves any existing storage.sqlite3, WAL/SHM siblings, and STORAGE_ROOT into timestamped .backup-<ts> folders so nothing is lost.

  3. Copies the snapshot back to the configured database path and rebuilds the storage repo from the archive contents.

Every archive writes a metadata.json manifest describing the projects captured, scrub preset, and a friendly reminder of the exact archive restore … command to run later.

Reset safety net

clear-and-reset-everything now offers to create one of these archives before deleting anything. By default it prompts interactively; pass --archive/--no-archive to force a choice, and pair with --force --no-archive for non-interactive automation. When an archive is created successfully, the CLI prints both the path and the restore command so you can undo the reset later.

Mailbox Health: am doctor

The doctor command group provides comprehensive diagnostics and repair capabilities for maintaining mailbox health. It emphasizes data safety—creating backups before any destructive operation and using semi-automatic repair to prevent accidental data loss.

Why am doctor?

Over time, mailbox state can drift:

  • Stale locks: Process crashes leave behind .archive.lock or .commit.lock files that block operations

  • Orphaned records: Agents get deleted but their message recipients remain in the database

  • FTS index desync: Full-text search index falls out of sync with actual messages

  • Expired file reservations: Reservations expire but aren't cleaned up

  • WAL files: SQLite WAL/SHM files accumulate (normal during operation, but worth monitoring)

The doctor commands detect these issues and offer safe, automatic repair.

Diagnostic checks

Run comprehensive diagnostics on your mailbox:

# Check all projects
uv run python -m mcp_agent_mail.cli doctor check

# Check with verbose output
uv run python -m mcp_agent_mail.cli doctor check --verbose

# JSON output for automation
uv run python -m mcp_agent_mail.cli doctor check --json

What it checks:

Check

Status

Description

Locks

OK/WARN

Detects stale archive and commit locks from crashed processes

Database

OK/ERROR

Runs PRAGMA integrity_check for SQLite corruption

Orphaned Records

OK/WARN

Finds message recipients without corresponding agents

FTS Index

OK/WARN

Compares message count vs FTS index entries

File Reservations

OK/INFO

Counts expired reservations pending cleanup

WAL Files

OK/INFO

Reports presence of SQLite WAL/SHM files

Example output:

MCP Agent Mail Doctor - Diagnostic Report
==================================================

Check         Status    Details
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
Locks         OK        No stale locks found
Database      OK        Database integrity check passed
Orphaned      OK        No orphaned records found
FTS Index     OK        FTS index synchronized (1,234 messages)
File Res.     INFO      4 expired reservation(s) pending cleanup
WAL Files     OK        No orphan WAL/SHM files

All checks passed!

Semi-automatic repair

The repair command uses a semi-automatic approach:

  • Safe repairs (locks, expired reservations) are applied automatically

  • Data-affecting repairs (orphan cleanup) require confirmation

  • A backup is created before any changes

# Preview what would be repaired (dry-run)
uv run python -m mcp_agent_mail.cli doctor repair --dry-run

# Run repairs with prompts for data changes
uv run python -m mcp_agent_mail.cli doctor repair

# Auto-confirm all repairs (for automation)
uv run python -m mcp_agent_mail.cli doctor repair --yes

# Specify custom backup location
uv run python -m mcp_agent_mail.cli doctor repair --backup-dir /path/to/backups

Repair workflow:

  1. Create backup — Git bundle + SQLite copy created before any changes

  2. Safe repairs (auto-applied):

    • Heal stale locks (removes orphaned .archive.lock, .commit.lock files)

    • Release expired file reservations (marks released_ts in database)

  3. Data repairs (require confirmation):

    • Delete orphaned message recipients

    • Rebuild FTS index (if needed)

Backup management

Doctor creates timestamped backups before repairs. You can also manage backups directly:

# List all available backups
uv run python -m mcp_agent_mail.cli doctor backups

# JSON output for scripting
uv run python -m mcp_agent_mail.cli doctor backups --json

Backup contents:

Each backup includes:

  • database.sqlite3 — Complete SQLite database copy

  • database.sqlite3-wal, database.sqlite3-shm — WAL files if present

  • archive.bundle or {project}.bundle — Git bundle of the archive repository

  • manifest.json — Metadata: when created, why, what's included, restore instructions

Directory structure:

{storage_root}/backups/
  2026-01-06T12-30-45_doctor-repair/
    manifest.json
    database.sqlite3
    archive.bundle

Restore from backup

If something goes wrong, restore from any backup:

# Preview what would be restored
uv run python -m mcp_agent_mail.cli doctor restore /path/to/backup --dry-run

# Restore (prompts for confirmation)
uv run python -m mcp_agent_mail.cli doctor restore /path/to/backup

# Skip confirmation prompt
uv run python -m mcp_agent_mail.cli doctor restore /path/to/backup --yes

Restore process:

  1. Validates backup manifest exists and is readable

  2. Shows backup metadata (creation time, reason, contents)

  3. Creates a pre-restore backup of current state (safety net)

  4. Restores SQLite database from backup

  5. Restores Git archive from bundle

  6. Reports any errors encountered

Safety features:

  • Current database saved as *.sqlite3.pre-restore before overwrite

  • Current archive saved as *.pre-restore directory before overwrite

  • Errors during restore are captured and reported

Best practices

  1. Run diagnostics regularly: am doctor check is fast and non-destructive

  2. Review before repair: Use --dry-run first to see what would change

  3. Keep backups: Don't delete old backups until you've verified the system is healthy

  4. Automate checks: Include am doctor check --json in your CI/monitoring for early warning

Quick Start: Interactive Deployment Wizard

The easiest way to export and deploy is the interactive wizard, which supports both GitHub Pages and Cloudflare Pages:

# Via CLI (recommended)
uv run python -m mcp_agent_mail.cli share wizard

# Or run the script directly
./scripts/share_to_github_pages.py

What the wizard does

The wizard provides a fully automated end-to-end deployment experience:

  1. Session resumption: Detects interrupted sessions and offers to resume exactly where you left off, avoiding re-export

  2. Configuration management: Remembers your last settings and offers to reuse them, saving time on repeated exports

  3. Deployment target selection: Choose between GitHub Pages, Cloudflare Pages, or local export

  4. Automatic CLI installation: Detects and installs missing tools (gh for GitHub, wrangler for Cloudflare)

  5. Guided authentication: Step-by-step browser login flows for GitHub and Cloudflare

  6. Smart project selection:

    • Shows all available projects in a formatted table

    • Supports multiple selection modes: all, single number (1), lists (1,3,5), or ranges (1-3, 2-5,8)

    • Remembers your previous selection for quick re-export

  7. Redaction configuration: Choose between standard (scrub secrets like API keys/tokens, keep agent names) or strict (redact all message bodies)

  8. Cryptographic signing: Optional Ed25519 signing with automatic key generation or reuse of existing keys. Generated keys are written to ~/.mcp-agent-mail/signing-keys/ (mode 0700, never to the cwd) so they cannot be accidentally git add -f'd into a repository

  9. Pre-flight validation: Checks that GitHub repo names are available before starting the export

  10. Deployment summary: Shows what will be deployed (project count, bundle size, target, signing status) and asks for confirmation

  11. Export and preview: Exports the bundle and launches an interactive preview server with automatic port detection (tries 9000-9100)

  12. Interactive preview controls:

    • Press 'r' to force browser refresh (manual cache bust)

    • Press 'd' to skip preview and deploy immediately

    • Press 'q' to quit preview server

  13. Automatic viewer asset refresh: Always ensures latest HTML/JS/CSS from source tree are used, even when reusing bundles

  14. Real-time deployment: Streams git and deployment output in real-time so you can follow the progress

  15. Automatic deployment: Creates repos, enables Pages, pushes code, and gives you the live URL

Session resumption (new in latest version)

If you interrupt the wizard (close terminal, Ctrl+C during preview, etc.), it saves your progress to ~/.mcp-agent-mail/wizard-session/. When you run the wizard again:

Incomplete session detected
  Projects: 3 selected
  Stage: preview
  Workspace: ~/.mcp-agent-mail/wizard-session/bundle

Resume where you left off? (Y/n):

What gets saved:

  • Selected projects and scrub preset

  • Deployment configuration (target, repo name, etc.)

  • Signing key preferences and paths

  • Exported bundle (in session workspace)

  • Current stage (preview, deploy)

Resume scenarios:

  • Closed terminal during preview: Resume → Skip re-export → Launch preview immediately

  • Changed your mind after export: Resume → "Reuse bundle?" → Preview or re-export

  • Want to deploy later: Resume → Press 'd' in preview → Deploy without re-exporting

  • Made viewer code changes: Resume → Assets auto-refresh from source tree

After successful deployment, the session state is automatically cleared. Sessions also clear if they become invalid (workspace deleted, projects removed, etc.).

Configuration persistence

The wizard saves your configuration to ~/.mcp-agent-mail/wizard-config.json after each successful deployment. On subsequent runs, it will show:

Previous Configuration Found
  Projects: 3 selected
  Redaction: standard
  Target: github-new

Use these settings again? (Y/n):

This allows rapid re-deployment with the same settings. The saved configuration includes:

  • Selected project indices (validates against current project list)

  • Redaction preset

  • Deployment target and parameters (repo name, privacy, project name)

  • Signing preferences (whether to sign, whether to generate new key)

  • Last used signing key path (offered as default when not generating new key)

Configuration is project-agnostic: if you add or remove projects, the wizard validates saved indices and prompts for re-selection if needed.

Difference between session and config:

  • Session state (wizard-session/): Temporary, for resuming interrupted runs, includes exported bundle

  • Config file (wizard-config.json): Persistent, for "use last settings" across fresh runs, no bundle

Multi-project selection

The project selector supports flexible selection syntax:

Available Projects:
#  Slug                Path
1  backend-abc123      /abs/path/backend
2  frontend-xyz789     /abs/path/frontend
3  infra-def456        /abs/path/infra
4  scripts-ghi789      /abs/path/scripts

Select projects to export (e.g., 'all', '1,3,5', or '1-3'):

Selection modes:

  • all: Export all projects (default)

  • 1: Export project #1 only

  • 1,3,5: Export projects #1, #3, and #5

  • 1-3: Export projects #1, #2, and #3 (inclusive range)

  • 2-4,7: Export projects #2, #3, #4, and #7 (combined range and list)

Invalid selections (out of range, malformed) are rejected with helpful error messages and the wizard prompts again.

Dynamic port allocation

The preview server automatically detects an available port in the range 9000-9100 instead of failing if port 9000 is in use. The actual port is displayed:

Launching preview server...
Using port 9001 (Ctrl+C to stop server)
Waiting for server to start...
✓ Server ready, opening browser at http://127.0.0.1:9001

This prevents port conflicts when multiple previews are running or when port 9000 is used by other services.

Deployment summary panel

Before starting the export, the wizard shows a comprehensive summary:

═══ Deployment Summary ═══

Projects: 3 selected
Bundle size: ~32 MB
Redaction: standard
Target: GitHub Pages
  Repository: mailbox-viewer-2024
  Visibility: Private
Signing: Enabled (Ed25519)

Proceed with export and deployment? (Y/n):

This gives you a final chance to review all settings and cancel if needed. The bundle size is estimated based on ~10 MB per project plus ~2 MB for static assets.

Real-time deployment streaming

Git operations and Cloudflare deployments stream output in real-time so you can see exactly what's happening:

Initializing git repository and pushing...
Initializing repository...
  Initialized empty Git repository in /tmp/mailbox-preview-abc123/.git/
✓ Initializing repository complete
Adding files...
✓ Adding files complete
Creating commit...
  [main (root-commit) 1a2b3c4] Initial mailbox export
   425 files changed, 123456 insertions(+)
✓ Creating commit complete
Pushing to GitHub...
  Enumerating objects: 430, done.
  Counting objects: 100% (430/430), done.
  Delta compression using up to 8 threads
  Compressing objects: 100% (425/425), done.
  Writing objects: 100% (430/430), 12.34 MiB | 5.67 MiB/s, done.
✓ Pushing to GitHub complete

✓ Successfully pushed to owner/mailbox-viewer-2024

This provides transparency and helps diagnose issues if deployment fails.

Platform-specific details

For GitHub Pages:

  • Wizard detects your package manager (brew/apt/dnf) and offers automated installation of gh CLI

  • For apt/dnf, shows complete manual installation instructions (including repo setup) since automation requires sudo

  • Runs gh auth login interactively to authenticate via browser

  • Creates new repository with your specified name and visibility (public/private)

  • Initializes git, commits, and pushes with streaming output

  • Enables GitHub Pages automatically via the GitHub API

  • Provides the GitHub Pages URL (may take 1-2 minutes to become live)

For Cloudflare Pages:

  • Detects npm and offers automated installation of wrangler CLI

  • Runs wrangler login interactively to authenticate via browser

  • Deploys directly to Cloudflare's global CDN (no git repository needed)

  • Streams wrangler output in real-time

  • Provides the .pages.dev URL immediately (site is live instantly)

  • Benefits: instant deployment, 275+ global locations, automatic HTTPS, unlimited requests on free tier

For local export:

  • Saves bundle to specified directory

  • No CLI installation or authentication required

  • Suitable for manual deployment to custom hosting or inspection

Error handling and recovery

The wizard includes comprehensive error handling:

  • Pre-flight validation: Checks GitHub repo availability before starting export to avoid conflicts

  • Port conflict resolution: Automatically finds an available port for preview server

  • Invalid selection handling: Validates project selections and prompts for correction

  • CLI installation failures: Shows manual installation instructions if automatic installation fails

  • Git operation failures: Each git step is validated; stops on first failure with clear error message

  • Deployment failures: Distinguishes between repo creation, push, and Pages enablement failures

If deployment fails after export, the bundle remains in the temp directory and can be deployed manually using the git commands shown in the manual deployment section below.

The wizard handles all operations automatically. For manual control or advanced options, see the detailed workflows below.

Basic export workflow (manual)

1. Export a bundle

# Export all projects to a directory
uv run python -m mcp_agent_mail.cli share export --output ./my-bundle

# Export specific projects only
uv run python -m mcp_agent_mail.cli share export \
  --output ./my-bundle \
  --project backend-abc123 \
  --project frontend-xyz789

# Export with Ed25519 signing for tamper-evident distribution
uv run python -m mcp_agent_mail.cli share export \
  --output ./my-bundle \
  --signing-key ./keys/signing.key \
  --signing-public-out ./keys/signing.pub

# Export and encrypt with age for secure distribution
uv run python -m mcp_agent_mail.cli share export \
  --output ./my-bundle \
  --age-recipient age1abc...xyz \
  --age-recipient age1def...uvw

The export process:

  1. Creates a snapshot of the SQLite database (read-only, no WAL/SHM files)

  2. Copies message bodies, attachments, and metadata into the bundle structure

  3. Applies redaction rules based on the scrub preset (default: standard)

  4. Generates manifest.json with SHA-256 hashes for all assets

  5. Optionally signs the manifest with Ed25519 (produces manifest.sig.json)

  6. Packages everything into a ZIP archive (optional, enabled by default)

  7. If chunking is enabled, writes the segmented database plus a chunks.sha256 manifest so streamed pages can be verified cheaply

  8. Optionally encrypts the ZIP with age (produces bundle.zip.age)

Refresh an existing bundle

Once you have published a bundle you can refresh it in place without re-running the full wizard. Every export records the settings that were used (projects, scrub preset, attachment thresholds, chunking config) inside manifest.json. The new share update command reads those defaults, regenerates the SQLite snapshot and viewer assets in a temporary directory, and then replaces the bundle atomically—removing obsolete chunked files or attachments along the way.

# Refresh bundle using the originally recorded settings
uv run python -m mcp_agent_mail.cli share update ./my-bundle

# Override one or more export options while updating
uv run python -m mcp_agent_mail.cli share update ./my-bundle \
  --project backend-abc123 \
  --inline-threshold 16384 \
  --chunk-threshold 104857600

# Re-sign and package the refreshed bundle
uv run python -m mcp_agent_mail.cli share update ./my-bundle \
  --zip \
  --signing-key ./keys/signing.key

When chunking was enabled previously but the refreshed snapshot no longer needs it, share update cleans up the chunks/ directory, chunks.sha256, and mailbox.sqlite3.config.json automatically, ensuring the bundle tree matches the new manifest. You can still tweak any setting at update time; overrides are written back into the export_config section of manifest.json for the next refresh.

2. Preview locally

# Serve the bundle on localhost:9000
uv run python -m mcp_agent_mail.cli share preview ./my-bundle

# Custom port and auto-open browser
uv run python -m mcp_agent_mail.cli share preview ./my-bundle \
  --port 8080 \
  --open-browser

This launches a lightweight HTTP server that serves the static files. Open http://127.0.0.1:9000/viewer/ in your browser to explore the archive.

Interactive preview controls:

  • 'r': Force browser reload (bumps manual cache-bust token, triggers viewer refresh)

  • 'd': Request deployment (exits with code 42; wizard detects and proceeds to deploy)

  • 'q': Quit preview server

  • Ctrl+C: Stop preview server

The preview server automatically refreshes viewer assets from the source tree if available, ensuring you always see the latest HTML/JS/CSS during development.

3. Verify integrity

# Verify SRI hashes and signature
uv run python -m mcp_agent_mail.cli share verify ./my-bundle

# Verify with explicit public key (overrides manifest.sig.json)
uv run python -m mcp_agent_mail.cli share verify ./my-bundle \
  --public-key AAAA...base64...

Verification checks:

  • SHA-256 hashes for all vendor libraries (Marked.js, DOMPurify, SQL.js)

  • SHA-256 hashes for the SQLite database and attachments

  • Ed25519 signature over the canonical manifest (if present)

4. Decrypt (if encrypted)

# Decrypt with age identity file (private key)
uv run python -m mcp_agent_mail.cli share decrypt bundle.zip.age \
  --identity ~/.age/key.txt

# Decrypt with passphrase (interactive prompt)
uv run python -m mcp_agent_mail.cli share decrypt bundle.zip.age \
  --passphrase

# Specify custom output path
uv run python -m mcp_agent_mail.cli share decrypt bundle.zip.age \
  --output ./decrypted-bundle.zip \
  --identity ~/.age/key.txt

After decryption, unzip the archive and use share preview to view it.

Export options reference

Option

Type

Default

Description

--output, -o

Path

(required)

Directory where the static bundle will be written

--project, -p

List

All projects

Limit export to specific project slugs or human keys (repeatable)

--inline-threshold

Bytes

65536 (64KB)

Inline attachments smaller than this as base64 data URIs

--detach-threshold

Bytes

26214400 (25MB)

Mark attachments larger than this as external (not bundled)

--scrub-preset

String

standard

Redaction preset: standard or strict (see Redaction presets section)

--chunk-threshold

Bytes

20971520 (20MB)

Split SQLite database into chunks if it exceeds this size

--chunk-size

Bytes

4194304 (4MB)

Chunk size when splitting large databases

--dry-run

Flag

false

Generate security summary and preview without writing files

--zip / --no-zip

Flag

true

Package the bundle into a ZIP archive

--signing-key

Path

None

Path to Ed25519 signing key (32-byte raw seed)

--signing-public-out

Path

None

Write the Ed25519 public key to this file after signing

--age-recipient

String

None

age public key for encryption (repeatable for multiple recipients)

--interactive, -i

Flag

false

Launch interactive wizard (prints guidance; full wizard TBD)

Security features

XSS protection (DOMPurify + CSP)

Message bodies are rendered using a defense-in-depth pipeline:

  1. Marked.js parses GitHub-Flavored Markdown into HTML

  2. DOMPurify sanitizes the HTML, removing dangerous tags and attributes

  3. Content Security Policy restricts script sources, blocks inline event handlers, and limits network access

This prevents malicious content in message bodies from executing JavaScript or exfiltrating data.

CSP configuration notes:

  • script-src: Allows self, CDNs (Tailwind, Alpine.js), and 'unsafe-eval' (required for SQL.js WebAssembly)

  • connect-src: Allows * (all origins) to support preview mode polling and flexible deployment scenarios

  • style-src: Allows self, inline styles (for Tailwind), and font CDNs

  • Trusted Types removed for browser compatibility (Firefox, Safari don't support it yet)

Cryptographic signing (Ed25519)

When you provide a signing key, the export process:

  1. Generates a canonical JSON representation of manifest.json

  2. Signs it with Ed25519 (fast, 64-byte signatures, 128-bit security)

  3. Writes the signature and public key to manifest.sig.json

Recipients can verify the signature using share verify to ensure:

  • The bundle hasn't been modified since signing

  • The bundle was created by someone with the private key

  • All assets match their declared SHA-256 hashes

Requirements and fallback:

  • Requires PyNaCl >= 1.6.0 (installed automatically with this package)

  • If PyNaCl is unavailable or signing fails, export gracefully falls back to unsigned mode

  • Wizard reuses existing signing keys by default (no re-generation unless requested)

  • Private keys are automatically excluded from git via .gitignore (signing-*.key pattern)

Encryption (age)

The age encryption tool (https://age-encryption.org/) provides modern, secure file encryption. When you provide recipient public keys, the export process encrypts the final ZIP archive. Only holders of the corresponding private keys can decrypt it.

Generate keys with:

# Install age (example for macOS)
brew install age

# Generate a new key pair
age-keygen -o key.txt

# Public key is printed to stdout; share it with exporters
# Private key is saved to key.txt; keep it secret

Redaction presets

The export pipeline supports configurable scrubbing to remove sensitive data:

  • standard: Clears acknowledgment/read state, removes file reservations and agent links, scrubs secrets (GitHub tokens, Slack tokens, OpenAI keys, bearer tokens, JWTs) from message bodies and attachment metadata. Retains agent names (which are already meaningless pseudonyms like "BlueMountain"), full message bodies, and attachments.

  • strict: All standard redactions plus replaces entire message bodies with "[Message body redacted]" placeholder and removes all attachments from the bundle.

All presets apply redaction to message subjects, bodies, and attachment metadata before the bundle is written.

Static viewer features

The bundled HTML viewer provides:

Dashboard layout:

  • Gmail-style three-pane interface: Projects sidebar, message list (center), and detail pane (right)

  • Bundle metadata header: Shows bundle creation time, export settings, and scrubbing preset

  • Summary panels: Side-by-side panels displaying projects included, attachment statistics, and redaction summary

  • Message list: Virtual-scrolled message list with sender, subject, snippet, and importance badges

  • Raw manifest viewer: Collapsible JSON display of the complete manifest for verification

Advanced boolean search (new): Powered by SQLite FTS5 with LIKE fallback, supports complex queries:

  • Boolean operators: (auth OR login) AND NOT admin

  • Quoted phrases: "build plan" (exact match)

  • Parentheses: Control precedence like (A OR B) AND (C OR D)

  • Operator precedence: NOT > AND > OR (e.g., A OR B AND C = A OR (B AND C))

  • Automatic debouncing: 140ms delay avoids hammering the database on every keystroke

  • Performance: FTS5 search is 10-100x faster than LIKE on large datasets

Lazy message loading (performance optimization):

  • Initial load fetches only 280-character snippets for all messages (3-6x faster)

  • Full message body loaded on-demand when you click a message

  • Dramatically reduces memory usage and initial load time

Virtual scrolling (new): Clusterize.js-powered virtual list rendering:

  • Smoothly handles 100,000+ messages without slowdown

  • Only ~30 DOM nodes exist at any time (visible rows + buffers)

  • Maintains native scrollbar feel with keyboard navigation

Markdown rendering: Message bodies are rendered with GitHub-Flavored Markdown, supporting code blocks, tables, task lists, and inline images.

Opportunistic OPFS caching: The SQLite database is cached in Origin Private File System (OPFS) in the background:

  • First load: Downloads from network, caches to OPFS during idle time

  • Subsequent loads: Instant from OPFS (even faster than IndexedDB)

  • Automatic cache key validation prevents stale data

Dark mode: Toggle between light and dark themes with localStorage persistence. Dark mode state is managed by the main viewer controller for consistency.

Attachment preview: Inline images render directly in message bodies. External attachments show file size and download links.

Message detail view: Click any message in the list to load its full body (lazy), view metadata (sender, recipients, timestamp, importance), and browse attachments.

No server required: After the initial HTTP serving (which can be a static file host like S3, GitHub Pages, or Netlify), all functionality runs client-side. No backend, no API calls, no authentication.

Browser compatibility: Works in all modern browsers (Chrome, Firefox, Safari, Edge) with graceful fallbacks for missing features (OPFS, FTS5).

Deployment options

Option 1: GitHub Pages (automated via wizard)

# Use the wizard for fully automated deployment
uv run python -m mcp_agent_mail.cli share wizard
# Select: GitHub Pages → provide repo name → wizard handles everything

Or manually:

# Export and unzip
uv run python -m mcp_agent_mail.cli share export --output ./bundle --no-zip
cd bundle

# Initialize git and push to GitHub
git init
git add .
git commit -m "Initial export"
git remote add origin git@github.com:yourorg/project-archive.git
git push -u origin main

# Enable GitHub Pages in repo settings (source: main branch, root directory)

Option 2: Cloudflare Pages (automated via wizard)

# Use the wizard for instant global CDN deployment
uv run python -m mcp_agent_mail.cli share wizard
# Select: Cloudflare Pages → provide project name → wizard deploys directly

Or manually with wrangler CLI:

# Export and deploy
uv run python -m mcp_agent_mail.cli share export --output ./bundle --no-zip
npx wrangler pages deploy ./bundle --project-name=project-archive

# Your site is live at: https://project-archive.pages.dev

Benefits of Cloudflare Pages:

  • Instant deployment (no git repo required)

  • Global CDN with 275+ locations

  • Automatic HTTPS and DDoS protection

  • Zero-downtime updates

  • Generous free tier (500 builds/month, unlimited requests)

Option 3: S3 + CloudFront

# Export and unzip
uv run python -m mcp_agent_mail.cli share export --output ./bundle --no-zip

# Upload to S3
aws s3 sync ./bundle s3://your-bucket/archives/project-2024/ --acl public-read

# Access via CloudFront
# https://d123abc.cloudfront.net/archives/project-2024/

Option 4: Nginx static site

server {
  listen 443 ssl;
  server_name archives.example.com;

  ssl_certificate /etc/letsencrypt/live/archives.example.com/fullchain.pem;
  ssl_certificate_key /etc/letsencrypt/live/archives.example.com/privkey.pem;

  root /var/www/archives/project-2024;
  index index.html;

  # Enable gzip for efficient transfer
  gzip on;
  gzip_types text/html text/css application/javascript application/json application/wasm;

  # Cache static assets
  location ~* \.(js|css|wasm|png|jpg|webp)$ {
    expires 1y;
    add_header Cache-Control "public, immutable";
  }

  # CSP headers are already in index.html meta tag
  # Add HTTPS-only and frame protection
  add_header Strict-Transport-Security "max-age=31536000; includeSubDomains" always;
  add_header X-Frame-Options "DENY" always;
  add_header X-Content-Type-Options "nosniff" always;
}

Option 5: Encrypted distribution via file sharing

For confidential archives:

# Export with age encryption
uv run python -m mcp_agent_mail.cli share export \
  --output ./bundle \
  --signing-key ./signing.key \
  --age-recipient age1auditor... \
  --age-recipient age1manager...

# This produces bundle.zip.age
# Upload to Dropbox, Google Drive, or send via secure file transfer

# Recipients decrypt locally
uv run python -m mcp_agent_mail.cli share decrypt bundle.zip.age \
  --identity ~/.age/key.txt

# Verify integrity
unzip bundle.zip
uv run python -m mcp_agent_mail.cli share verify ./bundle

# Preview locally
uv run python -m mcp_agent_mail.cli share preview ./bundle

Example workflows

Quarterly audit package

# Export Q4 2024 communications for audit
uv run python -m mcp_agent_mail.cli share export \
  --output ./audit-q4-2024 \
  --scrub-preset strict \
  --signing-key ./audit-signing.key \
  --signing-public-out ./audit-signing.pub \
  --age-recipient age1auditor@firm.example

# Produces: audit-q4-2024.zip.age
# Send to auditor with audit-signing.pub

# Auditor verifies:
age -d -i auditor-key.txt audit-q4-2024.zip.age > audit-q4-2024.zip
unzip audit-q4-2024.zip
python -m mcp_agent_mail.cli share verify ./audit-q4-2024 \
  --public-key $(cat audit-signing.pub)

Executive summary for stakeholders

# Export high-importance threads only
# (filter in UI after export, or use SQL to create filtered snapshot)
uv run python -m mcp_agent_mail.cli share export \
  --output ./exec-summary \
  --project backend-prod \
  --scrub-preset standard

# Host on internal web server
cp -r ./exec-summary /var/www/exec-archives/2024-12/
# Share link: https://internal.example.com/exec-archives/2024-12/

Disaster recovery backup

# Monthly encrypted backup
uv run python -m mcp_agent_mail.cli share export \
  --output ./backup-$(date +%Y-%m) \
  --scrub-preset none \
  --signing-key ./dr-signing.key \
  --age-recipient age1dr@company.example

# Store in off-site backup system
aws s3 cp backup-2024-12.zip.age s3://dr-backups/mcp-mail/ \
  --storage-class GLACIER_IR

# Restore procedure documented in runbook

Troubleshooting exports

Export fails with "Database locked"

The export takes a snapshot using SQLite's Online Backup API. If the server is actively writing, wait a few seconds and retry. For large databases, consider temporarily stopping the server during export.

Bundle size is too large

Use --detach-threshold to mark large attachments as external references. These won't be included in the bundle but will show file metadata in the viewer.

# Bundle files under 1MB, mark larger files as external
uv run python -m mcp_agent_mail.cli share export \
  --output ./bundle \
  --detach-threshold 1048576

Alternatively, filter to specific projects with --project.

Encrypted bundle won't decrypt

Verify you're using the correct identity file:

# Get your public key from your identity file
age-keygen -y identity.txt

# Ensure this public key was included in the --age-recipient values during export
# If you have multiple identity files, try each one
age -d -i identity.txt bundle.zip.age > bundle.zip

Signature verification fails

Signature verification requires:

  1. The original manifest.json (unmodified)

  2. The manifest.sig.json file (contains signature and public key)

  3. All assets referenced in the manifest with matching SHA-256 hashes

If verification fails, the bundle may have been tampered with or corrupted during transfer. Re-export and re-transfer.

Viewer shows blank page or errors

Check browser console for errors. Common issues:

  • OPFS not supported: Older browsers may not support Origin Private File System. The viewer will fall back to in-memory mode (slower).

  • Database too large: Browsers limit in-memory database size to ~1-2GB. Use chunking (--chunk-threshold) for very large archives.

  • CSP violations: If hosting the bundle, ensure the web server doesn't add conflicting CSP headers. The viewer's CSP is defined in index.html and should not be overridden.

On-disk layout (per project)

<store>/projects/<slug>/
  agents/<AgentName>/profile.json
  agents/<AgentName>/inbox/YYYY/MM/<msg-id>.md
  agents/<AgentName>/outbox/YYYY/MM/<msg-id>.md
  messages/YYYY/MM/<msg-id>.md
  messages/threads/<thread-id>.md  # optional human digest maintained by the server
  file_reservations/<sha1-of-path>.json
  attachments/<xx>/<sha1>.webp

Message file format

Messages are GitHub-Flavored Markdown with JSON frontmatter (fenced by ---json). Attachments are either WebP files referenced by relative path or inline base64 WebP data URIs.

---json
{
  "id": 1234,
  "thread_id": "TKT-123",
  "project": "/abs/path/backend",
  "project_slug": "backend-abc123",
  "from": "GreenCastle",
  "to": ["BlueLake"],
  "cc": [],
  "created": "2025-10-23T15:22:14Z",
  "importance": "high",
  "ack_required": true,
  "attachments": [
    {"type": "file", "media_type": "image/webp", "path": "projects/backend-abc123/attachments/2a/2a6f.../diagram.webp"}
  ]
}
---

# Build plan for /api/users routes

... body markdown ...

Data model (SQLite)

  • projects(id, human_key, slug, created_at)

  • agents(id, project_id, name, program, model, task_description, inception_ts, last_active_ts, attachments_policy, contact_policy)

  • messages(id, project_id, sender_id, thread_id, subject, body_md, created_ts, importance, ack_required, attachments)

  • message_recipients(message_id, agent_id, kind, read_ts, ack_ts)

  • file_reservations(id, project_id, agent_id, path_pattern, exclusive, reason, created_ts, expires_ts, released_ts)

  • agent_links(id, a_project_id, a_agent_id, b_project_id, b_agent_id, status, reason, created_ts, updated_ts, expires_ts)

  • project_sibling_suggestions(id, project_a_id, project_b_id, score, status, rationale, created_ts, evaluated_ts, confirmed_ts, dismissed_ts)

  • fts_messages(message_id UNINDEXED, subject, body) + triggers for incremental updates

Concurrency and lifecycle

  • One request/task = one isolated operation

  • Archive writes are guarded by a per-project .archive.lock under projects/<slug>/

  • Git index/commit operations are serialized across the shared archive repo by a repo-level .commit.lock

  • DB operations are short-lived and scoped to each tool call; FTS triggers keep the search index current

  • Artifacts are written first, then committed as a cohesive unit with a descriptive message

  • Attachments are content-addressed (sha1) to avoid duplication

How it works (key flows)

  1. Create an identity

  • register_agent(project_key, program, model, name?, task_description?) → creates/updates a named identity, persists profile to Git, and commits.

  1. Send a message

  • send_message(project_key, sender_name, to[], subject, body_md, cc?, bcc?, attachment_paths?, convert_images?, importance?, ack_required?, thread_id?, auto_contact_if_blocked?)

  • Writes a canonical message under messages/YYYY/MM/, an outbox copy for the sender, and inbox copies for each recipient; commits all artifacts.

  • Optionally converts images (local paths or data URIs) to WebP and embeds small ones inline.

sequenceDiagram
  participant Agent
  participant Server
  participant DB
  participant Git

  Agent->>Server: call send_message
  Server->>DB: insert message and recipients
  DB-->>Server: ok
  Server->>Git: write canonical markdown
  Server->>Git: write outbox copy
  Server->>Git: write inbox copies
  Server->>Git: commit
  Server-->>Agent: result
  1. Check inbox

  • fetch_inbox(project_key, agent_name, since_ts?, urgent_only?, unread_only?, include_bodies?, limit?) returns recent messages, preserving thread_id where available. Pass unread_only=true to skip messages this recipient has already explicitly marked read (cuts token-burn for polling agents).

  • acknowledge_message(project_key, agent_name, message_id) marks acknowledgements.

  1. Avoid conflicts with file reservations (leases)

  • file_reservation_paths(project_key, agent_name, paths[], ttl_seconds, exclusive, reason) records an advisory lease in DB and writes JSON reservation artifacts in Git; conflicts are reported if overlapping active exclusives exist (reservations are still granted; conflicts are returned alongside grants).

  • release_file_reservations(project_key, agent_name, paths? | file_reservation_ids?) releases active leases (all if none specified). JSON artifacts remain for audit history.

  • Optional: install a pre-commit hook in your code repo that blocks commits conflicting with other agents' active exclusive file reservations.

sequenceDiagram
  participant Agent
  participant Server
  participant DB
  participant Git

  Agent->>Server: call file_reservation_paths
  Server->>DB: expire old leases and check overlaps
  DB-->>Server: conflicts or grants
  Server->>DB: insert file reservation rows
  Server->>Git: write file reservation JSON files
  Server->>Git: commit
  Server-->>Agent: granted paths and any conflicts
  1. Search & summarize

  • search_messages(project_key, query, limit?) uses FTS5 over subject and body.

  • summarize_thread(project_key, thread_id, include_examples?) extracts key points, actions, and participants from the thread.

  • reply_message(project_key, message_id, sender_name, body_md, ..., sender_token?) creates a subject-prefixed reply, preserving or creating a thread.

Semantics & invariants

  • Identity

    • Names are memorable adjective+noun and unique per project; name_hint is sanitized (alnum) and used if available

    • whois returns the stored profile; list_agents can filter by recent activity

    • last_active_ts is bumped on relevant interactions (messages, inbox checks, etc.)

  • Threads

    • Replies inherit thread_id from the original; if missing, the reply sets thread_id to the original message id

    • Subject lines are prefixed (e.g., Re:) for readability in mailboxes

  • Attachments

    • Image references (file path or data URI) are converted to WebP; small images embed inline when policy allows

    • Non-absolute attachment_paths (and markdown image paths) resolve relative to the project archive root under STORAGE_ROOT/projects/<slug>/, not the code repo root

    • Absolute attachment paths are disabled by default. Enabling ALLOW_ABSOLUTE_ATTACHMENT_PATHS=true on a networked deployment turns message sending into a filesystem read primitive for whatever paths the server process can access.

    • Stored under attachments/<xx>/<sha1>.webp and referenced by relative path in frontmatter

  • File Reservations

    • TTL-based; exclusive means "please don't modify overlapping surfaces" for others until expiry or release

    • Conflict detection is per exact path pattern; shared reservations can coexist, exclusive conflicts are surfaced

    • JSON artifacts remain in Git for audit even after release (DB tracks release_ts)

  • Search

    • External-content FTS virtual table and triggers index subject/body on insert/update/delete

    • Queries are constrained to the project id and ordered by created_ts DESC

Goal: make coordination "just work" without spam across unrelated agents. The server enforces per-project isolation by default and adds an optional consent layer within a project so agents only contact relevant peers.

Isolation by project

  • All tools require a project_key. Agents only see messages addressed to them within that project.

  • An agent working in Project A is invisible to agents in Project B unless explicit cross-project contact is established (see below). This avoids distraction between unrelated repositories.

Policies (per agent)

  • open: accept any targeted messages in the project.

  • auto (default): allow messages when there is obvious shared context (e.g., same thread participants; recent overlapping active file reservations; recent prior direct contact within a TTL); otherwise requires a contact request.

  • contacts_only: require an approved contact link first.

  • block_all: reject all new contacts (errors with CONTACT_BLOCKED).

Use set_contact_policy(project_key, agent_name, policy) to update.

Request/approve contact

  • request_contact(project_key, from_agent, to_agent, reason?, ttl_seconds?, registration_token?) creates or refreshes a pending link and sends a small ack_required "intro" message to the recipient.

  • respond_contact(project_key, to_agent, from_agent, accept, ttl_seconds?, registration_token?) approves or denies; approval grants messaging until expiry.

  • list_contacts(project_key, agent_name) surfaces outbound links with target projects and audit flags for expiry/messageability.

Auth note: these tools require either an agent already authenticated in the current MCP session, or the relevant registration_token. send_message(..., auto_contact_if_blocked=true) only auto-approves when both agents are already authenticated in the same MCP session; otherwise it creates a pending request_contact and fails loud with CONTACT_REQUIRED (the underlying message body is not queued — once the recipient approves, the sender must re-call send_message to actually deliver the payload). macro_contact_handshake(..., auto_accept=true) follows the same rule for new approvals: it only auto-approves when the target agent is already authenticated in the current MCP session or when target_registration_token is supplied explicitly. If the link is already approved, the macro reuses that approval. Otherwise the request remains pending and the macro reports a response_error.

Auto-allow heuristics (no explicit request required)

  • Same thread: replies or messages to thread participants are allowed.

  • Recent overlapping file reservations: if sender and recipient hold active file reservations in the project, messaging is allowed.

  • Recent prior contact: a sliding TTL allows follow-ups between the same pair.

These heuristics minimize friction while preventing cold spam.

Cross-project coordination (frontend vs backend repos)

When two repos represent the same underlying project (e.g., frontend and backend), you have two options:

  1. Use the same project_key across both workspaces. Agents in both repos operate under one project namespace and benefit from full inbox/outbox coordination automatically.

  2. Keep separate project_keys and establish explicit contact:

    • In backend, agent GreenCastle calls:

      • request_contact(project_key="/abs/path/backend", from_agent="GreenCastle", to_agent="BlueLake", reason="API contract changes", registration_token="<GreenCastle token>")

    • In frontend, BlueLake calls:

      • respond_contact(project_key="/abs/path/backend", to_agent="BlueLake", from_agent="GreenCastle", accept=true, registration_token="<BlueLake token>")

    • After approval, messages can be exchanged; in default auto policy the server allows follow-up threads/reservation-based coordination without re-requesting.

Important: You can also create reciprocal links or set open policy for trusted pairs. The consent layer is on by default (CONTACT_ENFORCEMENT_ENABLED=true) but is designed to be non-blocking in obvious collaboration contexts.

Resource layer (read-only URIs)

Expose common reads as resources that clients can fetch. See API Quick Reference → Resources for the full list and parameters.

Example (conceptual) resource read:

{
  "method": "resources/read",
  "params": {"uri": "resource://inbox/BlueLake?project=/abs/path/backend&limit=20&agent_token=<registration_token>"}
}
sequenceDiagram
  participant Client
  participant Server
  participant DB

  Client->>Server: read inbox resource
  Server->>DB: select messages for agent
  DB-->>Server: rows
  Server-->>Client: inbox data

File Reservations and the optional pre-commit guard

  • Guard status and pre-push

    • Print guard status:

      • mcp-agent-mail guard status /path/to/repo

    • Install both guards (pre-commit + pre-push):

      • mcp-agent-mail guard install <project_key> <repo_path> --prepush

    • Pre-commit honors WORKTREES_ENABLED and AGENT_MAIL_GUARD_MODE (warn advisory).

    • Pre-push enumerates to-be-pushed commits (rev-list) and uses diff-tree with --no-ext-diff.

    • Composition-safe install (chain-runner):

      • A Python chain-runner is written to .git/hooks/pre-commit and .git/hooks/pre-push.

      • It executes hooks.d/<hook>/* in lexical order, then <hook>.orig if present (existing hooks are preserved, not overwritten).

      • Agent Mail installs its guard as hooks.d/pre-commit/50-agent-mail.py and hooks.d/pre-push/50-agent-mail.py.

      • Windows shims (pre-commit.cmd/.ps1, pre-push.cmd/.ps1) are written to invoke the Python chain-runner.

    • Matching and safety details:

      • Renames/moves are handled: both the old and new names are checked (git diff --cached --name-status -M -z).

      • NUL-safe end-to-end: paths are collected and forwarded as NUL-delimited to avoid ambiguity.

      • Git-native matching: reservations are checked using Git wildmatch pathspec semantics against repo-root relative paths; core.ignorecase is honored.

      • Emergency bypass (use sparingly): set AGENT_MAIL_BYPASS=1, or use native Git --no-verify. In warn mode the guard never blocks.

Git-based project identity (opt-in)

  • Gate: WORKTREES_ENABLED=1 or GIT_IDENTITY_ENABLED=1 enables git-based identity features. Default off.

  • Identity modes (default dir): dir, git-remote, git-toplevel, git-common-dir.

  • Inspect identity for a path:

    • Resource (MCP): resource://identity/%2Fabs%2Fpath (absolute paths must be URL-encoded inside the resource segment)

    • CLI (diagnostics): mcp-agent-mail mail status /abs/path

  • Precedence (when gate is on):

    1. Committed marker .agent-mail-project-id (recommended)

    2. Discovery YAML .agent-mail.yaml with project_uid:

    3. Private marker under Git common dir .git/agent-mail/project-id

    4. Remote fingerprint: normalized origin URL + default branch

    5. git-common-dir hash; else dir hash

  • Migration helpers:

    • Write committed marker: mcp-agent-mail projects mark-identity . --commit

    • Scaffold discovery file: mcp-agent-mail projects discovery-init . --product <product_uid>

Example identity payload (resource):

{
  "project_uid": "c5b2c86b-7c36-4de6-9a0a-2c4e1c3a1c4a",
  "slug": "repo-a1b2c3d4e5",
  "identity_mode_used": "git-remote",
  "canonical_path": "github.com/owner/repo",
  "human_key": "/abs/worktree/path",
  "repo_root": "/abs/repo",
  "git_common_dir": "/abs/repo/.git",
  "branch": "feature/x",
  "worktree_name": "repo-wt-x",
  "core_ignorecase": true,
  "normalized_remote": "github.com/owner/repo"
}

Adopt/Merge legacy projects (optional)

Consolidate legacy per-worktree projects into a canonical one (safe, explicit, and auditable).

  • Plan the merge (no changes):

    • mcp-agent-mail projects adopt <from> <to> --dry-run

  • Apply the merge (moves artifacts and re-keys DB rows):

    • mcp-agent-mail projects adopt <from> <to> --apply

  • Safeguards and behavior:

    • Requires both projects be in the same repository (validated via git-common-dir).

    • Moves archived Git artifacts from projects/<old-slug>/… to projects/<new-slug>/… while preserving history.

    • Re-keys database rows (agents, messages, file_reservations) from source to target project.

    • Records aliases.json under the target with "former_slugs": [...] for discoverability.

    • Aborts if agent-name conflicts would break uniqueness in the target (fix names, then retry).

    • Idempotent where possible; dry-run always prints a clear plan before apply.

Build slots and helpers (opt-in)

  • amctl env prints helpful environment keys:

    • SLUG, PROJECT_UID, BRANCH, AGENT, CACHE_KEY, ARTIFACT_DIR

    • Example: mcp-agent-mail amctl env --path . --agent AliceDev

  • am-run wraps a command with those keys set:

    • Example: mcp-agent-mail am-run frontend-build -- npm run dev

    • Auth: when talking to the HTTP server, am-run now auto-loads the local agent registration_token from the database when possible. You can also pass --registration-token or set AGENT_MAIL_REGISTRATION_TOKEN.

  • Build slots (advisory, per-project coarse locking):

    • Flags:

      • --ttl-seconds: lease duration (default 3600)

      • --shared/--exclusive: non-exclusive or exclusive lease (default exclusive)

      • --block-on-conflicts: exit non-zero if exclusive conflicts are detected before starting

    • Acquire:

      • Tool: acquire_build_slot(project_key, agent_name, slot, ttl_seconds=3600, exclusive=true)

    • Renew:

      • Tool: renew_build_slot(project_key, agent_name, slot, extend_seconds=1800)

    • Release (non-destructive; marks released):

      • Tool: release_build_slot(project_key, agent_name, slot)

    • Notes:

      • Slots are recorded under the project archive build_slots/<slot>/<agent>__<branch>.json

      • exclusive=true reports conflicts if another active exclusive holder exists

      • Intended for long-running tasks (dev servers, watchers); pair with am-run and amctl env

Product Bus

Group multiple repositories (e.g., frontend, backend, infra) under a single product for product‑wide inbox/search and shared threads.

  • Ensure a product:

    • mcp-agent-mail products ensure MyProduct --name "My Product"

  • Link a project (slug or path) into the product:

    • mcp-agent-mail products link MyProduct .

  • Inspect product and linked projects:

    • mcp-agent-mail products status MyProduct

  • Product‑wide message search (FTS):

    • mcp-agent-mail products search MyProduct "urgent AND deploy" --agent Alice --registration-token "$AGENT_MAIL_REGISTRATION_TOKEN" --limit 50

  • Product‑wide inbox:

    • mcp-agent-mail products inbox MyProduct Alice --registration-token "$AGENT_MAIL_REGISTRATION_TOKEN" --limit 50 --urgent-only --include-bodies --since-ts "2025-11-01T00:00:00Z"

  • Product‑wide thread summarization:

    • mcp-agent-mail products summarize-thread MyProduct "bd-123" --agent Alice --registration-token "$AGENT_MAIL_REGISTRATION_TOKEN" --per-thread-limit 100 --no-llm

  • Auth note:

    • Product-wide search, inbox, and thread summarization now require an authenticated agent identity. These commands accept --registration-token / AGENT_MAIL_REGISTRATION_TOKEN and will auto-use a single unambiguous locally stored token when that is possible.

Containers

  • Build and run locally:

    docker build -t mcp-agent-mail .
    docker run --rm -p 8765:8765 \
      -e HTTP_HOST=0.0.0.0 \
      -e STORAGE_ROOT=/data/mailbox \
      -v agent_mail_data:/data \
      mcp-agent-mail
  • Or with Compose:

    docker compose up --build
  • Notes:

    • Runs as an unprivileged user (appuser, uid 10001).

    • Includes a HEALTHCHECK against /health/liveness.

    • The server reads config from .env via python-decouple. You can mount it read-only into the container at /app/.env.

    • Default bind host is 0.0.0.0 in the container; port 8765 is exposed.

    • Persistent archive lives under /data/mailbox (mapped to the agent_mail_data volume by default).

    • Bind-mount uid mismatches: if you bind-mount a host directory (instead of a named volume) into /data/mailbox, the files on the host are typically owned by uid 1000 rather than the container's appuser (uid 10001). Git would normally refuse to operate on such a "dubious ownership" directory, which surfaces as the non-obvious error Unknown parameter: --cached on diff/status. The image now whitelists /data/mailbox (and, as a catch-all inside the container only, *) via git config --global --add safe.directory. If you rebuild from a fork that omits that step, either re-add the safe.directory lines or chown -R 10001:10001 the host path (see issue #143).

Notes

  • A unique product_uid is stored for each product; you can reference a product by uid or name.

  • Server tools also exist for orchestration: ensure_product, products_link, search_messages_product, and resource://product/{key}.

Exclusive file reservations are advisory but visible and auditable:

  • A reservation JSON is written to file_reservations/<sha1(path)>.json capturing holder, pattern, exclusivity, created/expires

  • The pre-commit guard scans active exclusive reservations and blocks commits that touch conflicting paths held by another agent

  • Agents must set AGENT_NAME so the guard knows who "owns" the commit

  • The server continuously evaluates reservations for staleness (agent inactivity + mail/filesystem/git silence) and releases abandoned locks automatically; the force_release_file_reservation tool uses the same heuristics and notifies the previous holder when another agent clears a stale lease

Install the guard into a code repo (conceptual tool call):

{
  "method": "tools/call",
  "params": {
    "name": "install_precommit_guard",
    "arguments": {
      "project_key": "/abs/path/backend",
      "code_repo_path": "/abs/path/backend"
    }
  }
}

Configuration

Configuration is loaded from an existing .env via python-decouple. Do not use os.getenv or auto-dotenv loaders.

Changing the HTTP Port

If port 8765 is already in use (e.g., by Cursor's Python extension), you can change it:

Option 1: During installation One-liner with custom port:

curl -fsSL "https://raw.githubusercontent.com/Dicklesworthstone/mcp_agent_mail/main/scripts/install.sh?$(date +%s)" | bash -s -- --port 9000 --yes

Or with local script: ./scripts/install.sh --port 9000 --yes


**Option 2: After installation (CLI)**
```bash
# Change port using CLI command
uv run python -m mcp_agent_mail.cli config set-port 9000

# View current port configuration
uv run python -m mcp_agent_mail.cli config show-port

# Restart server for changes to take effect
scripts/run_server_with_token.sh

Option 3: Manual .env edit

# Edit .env file manually with your text editor (recommended)
nano .env  # or vim, code, etc.

# Or append (⚠️ warning: will create duplicate if HTTP_PORT already exists)
echo "HTTP_PORT=9000" >> .env

Option 4: CLI server override

# Override port at server startup (doesn't modify .env)
uv run python -m mcp_agent_mail.cli serve-http --port 9000
from decouple import Config as DecoupleConfig, RepositoryEnv

decouple_config = DecoupleConfig(RepositoryEnv(".env"))

STORAGE_ROOT = decouple_config("STORAGE_ROOT", default="~/.mcp_agent_mail_git_mailbox_repo")
HTTP_HOST = decouple_config("HTTP_HOST", default="127.0.0.1")
HTTP_PORT = int(decouple_config("HTTP_PORT", default=8765))
HTTP_PATH = decouple_config("HTTP_PATH", default="/mcp/")

Common variables you may set:

Configuration reference

Name

Default

Description

STORAGE_ROOT

~/.mcp_agent_mail_git_mailbox_repo

Root for per-project repos and SQLite DB

HTTP_HOST

127.0.0.1

Bind host for HTTP transport

HTTP_PORT

8765

Bind port for HTTP transport

HTTP_PATH

/mcp/

Preferred MCP endpoint mount path (/api and /mcp aliases are also mounted)

HTTP_JWT_ENABLED

false

Enable JWT validation middleware

HTTP_JWT_SECRET

HMAC secret for HS* algorithms (dev)

HTTP_JWT_JWKS_URL

JWKS URL for public key verification

HTTP_JWT_ALGORITHMS

HS256

CSV of allowed algs

HTTP_JWT_AUDIENCE

Expected aud (optional)

HTTP_JWT_ISSUER

Expected iss (optional)

HTTP_JWT_ROLE_CLAIM

role

JWT claim name containing role(s)

HTTP_RBAC_ENABLED

true

Enforce read-only vs tools RBAC

HTTP_RBAC_READER_ROLES

reader,read,ro

CSV of reader roles

HTTP_RBAC_WRITER_ROLES

writer,write,tools,rw

CSV of writer roles

HTTP_RBAC_DEFAULT_ROLE

reader

Role used when none present

HTTP_RBAC_READONLY_TOOLS

health_check,fetch_inbox,whois,search_messages,summarize_thread

CSV of read-only tool names

HTTP_RATE_LIMIT_ENABLED

false

Enable token-bucket limiter

HTTP_RATE_LIMIT_BACKEND

memory

memory or redis

HTTP_RATE_LIMIT_PER_MINUTE

60

Legacy per-IP limit (fallback)

HTTP_RATE_LIMIT_TOOLS_PER_MINUTE

60

Per-minute for tools/call

HTTP_RATE_LIMIT_TOOLS_BURST

0

Optional burst for tools (0=auto=rpm)

HTTP_RATE_LIMIT_RESOURCES_PER_MINUTE

120

Per-minute for resources/read

HTTP_RATE_LIMIT_RESOURCES_BURST

0

Optional burst for resources (0=auto=rpm)

HTTP_RATE_LIMIT_REDIS_URL

Redis URL for multi-worker limits

HTTP_REQUEST_LOG_ENABLED

false

Print request logs (Rich + JSON)

LOG_JSON_ENABLED

false

Output structlog JSON logs

MCP_AGENT_MAIL_OUTPUT_FORMAT

Default output format for tools/resources (json or toon)

TOON_DEFAULT_FORMAT

Global default output format fallback (json or toon)

TOON_STATS

false

Emit TOON token stats (uses tru --stats)

TOON_TRU_BIN

Explicit path/command for the tru encoder (overrides TOON_BIN)

TOON_BIN

tru

Path/command for toon_rust encoder (Node toon is rejected)

INLINE_IMAGE_MAX_BYTES

65536

Threshold (bytes) for inlining WebP images during send_message

CONVERT_IMAGES

true

Convert images to WebP (and optionally inline small ones)

KEEP_ORIGINAL_IMAGES

false

Also store original image bytes alongside WebP (attachments/originals/)

ALLOW_ABSOLUTE_ATTACHMENT_PATHS

false

Allow absolute filesystem paths in attachments and markdown image references. Keep this disabled on networked deployments unless you explicitly want server-side file reads.

LOG_LEVEL

INFO

Server log level

HTTP_CORS_ENABLED

false

Enable CORS middleware when true

HTTP_CORS_ORIGINS

CSV of allowed origins (e.g., https://app.example.com,https://ops.example.com)

HTTP_CORS_ALLOW_CREDENTIALS

false

Allow credentials on CORS

HTTP_CORS_ALLOW_METHODS

*

CSV of allowed methods or *

HTTP_CORS_ALLOW_HEADERS

*

CSV of allowed headers or *

HTTP_BEARER_TOKEN

Static bearer token (only when JWT disabled)

HTTP_ALLOW_LOCALHOST_UNAUTHENTICATED

true

Allow localhost requests without auth (dev convenience)

HTTP_OTEL_ENABLED

false

Enable OpenTelemetry instrumentation

OTEL_SERVICE_NAME

mcp-agent-mail

Service name for telemetry

OTEL_EXPORTER_OTLP_ENDPOINT

OTLP exporter endpoint URL

APP_ENVIRONMENT

development

Environment name (development/production)

DATABASE_URL

sqlite+aiosqlite:///./storage.sqlite3

SQLAlchemy async database URL

DATABASE_ECHO

false

Echo SQL statements for debugging

DATABASE_POOL_SIZE

50 (sqlite) / 25 (other)

Base SQLAlchemy pool size (optional override)

DATABASE_MAX_OVERFLOW

4 (sqlite) / 25 (other)

Extra connections allowed beyond pool_size

DATABASE_POOL_TIMEOUT

45 (sqlite) / 30 (other)

Seconds to wait for a pool connection before failing

GIT_AUTHOR_NAME

mcp-agent

Git commit author name

GIT_AUTHOR_EMAIL

mcp-agent@example.com

Git commit author email

LLM_ENABLED

true

Enable LiteLLM for thread summaries and discovery

LLM_DEFAULT_MODEL

gpt-5-mini

Default LiteLLM model identifier

LLM_TEMPERATURE

0.2

LLM temperature for text generation

LLM_MAX_TOKENS

512

Max tokens for LLM completions

LLM_CACHE_ENABLED

true

Enable LLM response caching

LLM_CACHE_BACKEND

memory

LLM cache backend (memory or redis)

LLM_CACHE_REDIS_URL

Redis URL for LLM cache (if backend=redis)

LLM_COST_LOGGING_ENABLED

true

Log LLM API costs and token usage

FILE_RESERVATIONS_CLEANUP_ENABLED

false

Enable background cleanup of expired file reservations

FILE_RESERVATIONS_CLEANUP_INTERVAL_SECONDS

60

Interval for file reservations cleanup task

FILE_RESERVATION_INACTIVITY_SECONDS

1800

Inactivity threshold (seconds) before a reservation is considered stale

FILE_RESERVATION_ACTIVITY_GRACE_SECONDS

900

Grace window for recent mail/filesystem/git activity to keep a reservation active

FILE_RESERVATIONS_ENFORCEMENT_ENABLED

true

Block message writes on conflicting file reservations targeting mail archive paths (agents/, messages/, attachments/)

ACK_TTL_ENABLED

false

Enable overdue ACK scanning (logs/panels; see views/resources)

ACK_TTL_SECONDS

1800

Age threshold (seconds) for overdue ACKs

ACK_TTL_SCAN_INTERVAL_SECONDS

60

Scan interval for overdue ACKs

ACK_ESCALATION_ENABLED

false

Enable escalation for overdue ACKs

ACK_ESCALATION_MODE

log

log or file_reservation escalation mode

ACK_ESCALATION_CLAIM_TTL_SECONDS

3600

TTL for escalation file reservations

ACK_ESCALATION_CLAIM_EXCLUSIVE

false

Make escalation file reservation exclusive

ACK_ESCALATION_CLAIM_HOLDER_NAME

Ops agent name to own escalation file reservations

CONTACT_ENFORCEMENT_ENABLED

true

Enforce contact policy before messaging. When false, explicitly addressed cross-project recipients (project:X#name / name@project) also no longer require an approved contact link; block_all is still honored

CONTACT_AUTO_TTL_SECONDS

86400

TTL for in-session auto-approved contact links and the "recent contact" recency window (1 day)

CONTACT_PENDING_TTL_SECONDS

604800

TTL for the pending contact-request fallback created by send_message(auto_contact_if_blocked=True) when in-session auto-approval is not possible — i.e. how long an async human approver has to respond (7 days)

CONTACT_AUTO_RETRY_ENABLED

true

Auto-retry contact requests on policy violations

MESSAGING_AUTO_REGISTER_RECIPIENTS

true

Automatically create missing local recipients during send_message and retry routing

MESSAGING_AUTO_HANDSHAKE_ON_BLOCK

true

When contact policy blocks delivery, attempt a contact handshake (auto-accept) and retry

TOOLS_LOG_ENABLED

true

Log tool invocations for debugging

LOG_RICH_ENABLED

true

Enable Rich console logging

LOG_INCLUDE_TRACE

false

Include trace-level logs

TOOL_METRICS_EMIT_ENABLED

false

Emit periodic tool usage metrics

TOOL_METRICS_EMIT_INTERVAL_SECONDS

60

Interval for metrics emission

RETENTION_REPORT_ENABLED

false

Enable retention/quota reporting

RETENTION_REPORT_INTERVAL_SECONDS

3600

Interval for retention reports (1 hour)

RETENTION_MAX_AGE_DAYS

180

Max age for retention policy reporting

QUOTA_ENABLED

false

Enable quota enforcement

QUOTA_ATTACHMENTS_LIMIT_BYTES

0

Max attachment storage per project (0=unlimited)

QUOTA_INBOX_LIMIT_COUNT

0

Max inbox messages per agent (0=unlimited)

RETENTION_IGNORE_PROJECT_PATTERNS

demo,test*,testproj*,testproject,backendproj*,frontendproj*

CSV of project patterns to ignore in retention/quota reports

AGENT_NAME_ENFORCEMENT_MODE

coerce

Agent naming policy: strict (reject invalid adjective+noun names), coerce (auto-generate if invalid), always_auto (always auto-generate)

Development quick start

Prerequisite: complete the setup above (Python 3.14 + uv venv + uv sync).

Dev helpers:

# Quick endpoint smoke test (server must be running locally)
bash scripts/test_endpoints.sh

# Pre-commit guard smoke test (no pytest)
bash scripts/test_guard.sh

Database schema (automatic):

# Tables are created from SQLModel definitions on first run.
# If models change, delete the SQLite DB (and WAL/SHM) and run migrate again.
uv run python -m mcp_agent_mail.cli migrate

Run the server (HTTP-only). Use the Typer CLI or module entry:

uv run python -m mcp_agent_mail.cli serve-http
uv run python -m mcp_agent_mail.http --host 127.0.0.1 --port 8765

Connect with your MCP client using the HTTP (Streamable HTTP) transport on the configured host/port. The endpoint tolerates /api, /api/, /mcp, and /mcp/.

Search syntax tips (SQLite FTS5)

  • Basic terms: plan users

  • Phrase search: "build plan"

  • Prefix search: mig*

  • Boolean operators: plan AND users NOT legacy

  • Field boosting is not enabled by default; subject and body are indexed. Keep queries concise. When FTS is unavailable, the UI/API automatically falls back to SQL LIKE on subject/body.

Design choices and rationale

  • Streamable HTTP first: the modern remote transport is the primary deployment mode; serve-stdio is also provided for clients that prefer a local STDIO server

  • Git + Markdown: Human-auditable, diffable artifacts that fit developer workflows (inbox/outbox mental model)

  • SQLite + FTS5: Efficient indexing/search with minimal ops footprint

  • Advisory file reservations: Make intent explicit and reviewable; optional guard enforces reservations at commit time

  • WebP attachments: Compact images by default; inline embedding keeps small diagrams in context

    • Optional: keep original binaries and dedup manifest under attachments/ for audit and reuse

Examples (conceptual tool calls)

This section has been removed to keep the README focused. See API Quick Reference below for canonical method signatures.

Operational notes

  • One async session per request/task; don't share across concurrent coroutines

  • Use explicit loads in async code; avoid implicit lazy loads

  • Use async-friendly file operations when needed; Git operations are serialized with a file lock

  • Clean shutdown should dispose any async engines/resources (if introduced later)

Security and ops

  • Transport

    • HTTP-only (Streamable HTTP). Place behind a reverse proxy (e.g., NGINX) with TLS termination for production

  • Auth

    • Optional JWT (HS*/JWKS) via HTTP middleware; enable with HTTP_JWT_ENABLED=true

    • Static bearer token (HTTP_BEARER_TOKEN) is independent of JWT; when set, BearerAuth protects all routes (including UI). You may use it alone or together with JWT.

    • When JWKS is configured (HTTP_JWT_JWKS_URL), incoming JWTs must include a matching kid header; tokens without kid or with unknown kid are rejected

    • Starter RBAC (reader vs writer) using role configuration; see HTTP_RBAC_* settings

    • Bearer-only RBAC note: when JWT is disabled, requests use HTTP_RBAC_DEFAULT_ROLE (default reader). That means non-localhost tool calls are read-only unless you set HTTP_RBAC_DEFAULT_ROLE=writer, disable RBAC (HTTP_RBAC_ENABLED=false), or switch to JWT roles. Localhost requests with HTTP_ALLOW_LOCALHOST_UNAUTHENTICATED=true are auto-elevated to writer.

  • Reverse proxy + TLS (minimal example)

    • NGINX location block:

      upstream mcp_mail { server 127.0.0.1:8765; }
      server {
        listen 443 ssl;
        server_name mcp.example.com;
        ssl_certificate /etc/letsencrypt/live/mcp.example.com/fullchain.pem;
        ssl_certificate_key /etc/letsencrypt/live/mcp.example.com/privkey.pem;
        location /mcp/ { proxy_pass http://mcp_mail; proxy_set_header Host $host; proxy_set_header X-Forwarded-Proto https; }
      }
  • Backups and retention

    • The Git repos and SQLite DB live under STORAGE_ROOT; back them up together for consistency

  • Observability

    • Add logging and metrics at the ASGI layer returned by mcp.http_app() (Prometheus, OpenTelemetry)

  • Concurrency

    • Archive writes: per-project .archive.lock prevents cross-project head-of-line blocking

    • Commits: repo-level .commit.lock serializes Git index/commit to avoid races across projects

Python client example (HTTP JSON-RPC)

This section has been removed to keep the README focused. Client code samples belong in examples/.

Troubleshooting

  • "sender_name not registered"

    • Create the agent first with register_agent or create_agent_identity, or check the project_key you're using matches the sender's project

  • Pre-commit hook blocks commits

    • Set AGENT_NAME to your agent identity; release or wait for conflicting exclusive file reservations; inspect .git/hooks/pre-commit

  • Inline images didn't embed

    • Ensure convert_images=true; images are automatically inlined if the compressed WebP size is below the server's INLINE_IMAGE_MAX_BYTES threshold (default 64KB). Larger images are stored as attachments instead.

  • Message not found

    • Confirm the project disambiguation when using resource://message/{id}; ids are unique per project

  • Inbox empty but messages exist

    • Check since_ts, urgent_only, and limit; verify recipient names match exactly (case-sensitive)

FAQ

  • Why Git and SQLite together?

    • Git provides human-auditable artifacts and history; SQLite provides fast queries and FTS search. Each is great at what the other isn't.

  • Are file reservations enforced?

    • Yes, optionally. The server can block message writes when a conflicting active exclusive reservation exists for mail archive paths (agents/, messages/, attachments/) when FILE_RESERVATIONS_ENFORCEMENT_ENABLED=true (default). Reservations themselves are advisory and always return both granted and conflicts. The optional pre-commit hook adds local enforcement at commit time in your code repo for project file paths.

  • Why HTTP-only?

    • Streamable HTTP is the modern remote transport for MCP; avoiding extra transports reduces complexity and encourages a uniform integration path.

  • Why JSON-RPC instead of REST or gRPC?

    • MCP defines a tool/resource method call model that maps naturally to JSON-RPC over a single endpoint. It keeps clients simple (one URL, method name + params), plays well with proxies, and avoids SDK lock-in while remaining language-agnostic.

  • Why separate "resources" (reads) from "tools" (mutations)?

    • Clear semantics enable aggressive caching and safe prefetch for resources, while tools remain explicit, auditable mutations. This split also powers RBAC (read-only vs writer) without guesswork.

  • Why canonical message storage in Git, not only in the database?

    • Git gives durable, diffable, human-reviewable artifacts you can clone, branch, and audit. SQLite provides fast indexing and FTS. The combo preserves governance and operability without a heavyweight message bus.

  • Why advisory file reservations instead of global locks?

    • Agents coordinate asynchronously; hard locks create head-of-line blocking and brittle failures. Advisory reservations surface intent and conflicts while the optional pre-commit guard enforces locally where it matters.

  • Why are agent names adjective+noun?

    • Memorable identities reduce confusion in inboxes, commit logs, and UI. The scheme yields low collision risk while staying human-friendly (vs GUIDs) and predictable for directory listings.

  • Why is project_key an absolute path?

    • Using the workspace's absolute path creates a stable, collision-resistant project identity across shells and agents. Slugs are derived deterministically from it, avoiding accidental forks of the same project.

  • Why WebP attachments and optional inlining?

    • WebP provides compact, high-quality images. Small images can be inlined for readability; larger ones are stored as attachments. You can keep originals when needed (KEEP_ORIGINAL_IMAGES=true).

  • Why both static bearer and JWT/JWKS support?

    • Local development should be zero-friction (single bearer). Production benefits from verifiable JWTs with role claims, rotating keys via JWKS, and layered RBAC.

  • Why SQLite FTS5 instead of an external search service?

    • FTS5 delivers fast, relevant search with minimal ops. It’s embedded, portable, and easy to back up with the Git archive. If FTS isn’t available, we degrade to SQL LIKE automatically.

  • Why is LLM usage optional?

    • Summaries and discovery should enhance, not gate, core functionality. Keeping LLM usage optional controls cost and latency while allowing richer UX when enabled.

API Quick Reference

Tools

Tip: to see tools grouped by workflow with recommended playbooks, fetch resource://tooling/directory?format=json.

Output format (all tools/resources):

  • Tools accept optional format = json | toon; resources accept ?format=toon.

  • TOON returns {format:"toon", data:"<TOON>", meta:{...}} (fallback: {format:"json", ... , meta:{toon_error:"..."}}).

  • Defaults to JSON unless MCP_AGENT_MAIL_OUTPUT_FORMAT or TOON_DEFAULT_FORMAT is set.

Name

Signature

Returns

Notes

health_check

health_check()

{status, environment, http_host, http_port, database_url}

Lightweight readiness probe

ensure_project

ensure_project(human_key: str)

{id, slug, human_key, created_at}

Idempotently creates/ensures project

register_agent

register_agent(project_key: str, program: str, model: str, name?: str, task_description?: str, attachments_policy?: str)

Agent profile dict

Creates/updates agent; writes profile to Git

whois

whois(project_key: str, agent_name: str, include_recent_commits?: bool, commit_limit?: int)

Agent profile dict

Enriched profile for one agent (optionally includes recent commits)

create_agent_identity

create_agent_identity(project_key: str, program: str, model: str, name_hint?: str, task_description?: str, attachments_policy?: str)

Agent profile dict

Always creates a new unique agent

sweep_stale_agents

sweep_stale_agents(project_key: str, agent_name: str, threshold_seconds?: int, require_no_active_reservations?: bool, registration_token?: str)

{project_key, requested_by, threshold_seconds, retired[], retired_agents[], count}

Authenticated project-scoped retirement; caller is excluded and active reservations block retirement by default

send_message

send_message(project_key: str, sender_name: str, to: list[str], subject: str, body_md: str, cc?: list[str], bcc?: list[str], attachment_paths?: list[str], convert_images?: bool, importance?: str, ack_required?: bool, thread_id?: str, auto_contact_if_blocked?: bool, sender_token?: str)

{deliveries: list, count: int, attachments?}

Writes canonical + inbox/outbox, converts images. Non-absolute attachment_paths resolve relative to the project archive root.

reply_message

reply_message(project_key: str, message_id: int, sender_name: str, body_md: str, to?: list[str], cc?: list[str], bcc?: list[str], subject_prefix?: str, sender_token?: str)

{thread_id, reply_to, deliveries: list, count: int, attachments?}

Preserves/creates thread, inherits flags

request_contact

request_contact(project_key: str, from_agent: str, to_agent: str, to_project?: str, reason?: str, ttl_seconds?: int, registration_token?: str)

Contact link dict

Request permission to message another agent

respond_contact

respond_contact(project_key: str, to_agent: str, from_agent: str, accept: bool, from_project?: str, ttl_seconds?: int, registration_token?: str)

Contact link dict

Approve or deny a contact request

list_contacts

list_contacts(project_key: str, agent_name: str, registration_token?: str)

list[dict]

List outbound contact links with target-project and expiry audit metadata

set_contact_policy

set_contact_policy(project_key: str, agent_name: str, policy: str, registration_token?: str)

Agent dict

Set policy: open, auto, contacts_only, block_all

fetch_inbox

fetch_inbox(project_key: str, agent_name: str, limit?: int, urgent_only?: bool, include_bodies?: bool, since_ts?: str, topic?: str, unread_only?: bool, registration_token?: str)

list[dict]

Non-mutating inbox read. unread_only=true filters to messages where this recipient's read_ts is NULL.

mark_message_read

mark_message_read(project_key: str, agent_name: str, message_id: int, registration_token?: str)

{message_id, read, read_at}

Per-recipient read receipt

acknowledge_message

acknowledge_message(project_key: str, agent_name: str, message_id: int, registration_token?: str)

{message_id, acknowledged, acknowledged_at, read_at}

Sets ack and read

macro_start_session

macro_start_session(human_key: str, program: str, model: str, task_description?: str, agent_name?: str, registration_token?: str, file_reservation_paths?: list[str], file_reservation_reason?: str, file_reservation_ttl_seconds?: int, inbox_limit?: int)

{project, agent, file_reservations, inbox}

Orchestrates ensure→register→optional file reservation→inbox fetch

macro_prepare_thread

macro_prepare_thread(project_key: str, thread_id: str, program: str, model: str, agent_name?: str, registration_token?: str, task_description?: str, register_if_missing?: bool, include_examples?: bool, inbox_limit?: int, include_inbox_bodies?: bool, llm_mode?: bool, llm_model?: str)

{project, agent, thread, inbox}

Bundles registration, thread summary, and inbox context

macro_file_reservation_cycle

macro_file_reservation_cycle(project_key: str, agent_name: str, paths: list[str], ttl_seconds?: int, exclusive?: bool, reason?: str, auto_release?: bool, registration_token?: str)

{file_reservations, released}

File Reservation + optionally release surfaces around a focused edit block

macro_contact_handshake

`macro_contact_handshake(project_key: str, requester

agent_name: str, target

to_agent: str, to_project?: str, reason?: str, ttl_seconds?: int, auto_accept?: bool, welcome_subject?: str, welcome_body?: str, requester_registration_token?: str, target_registration_token?: str)`

search_messages

search_messages(project_key: str, query: str, limit?: int, agent_name?: str, registration_token?: str)

list[dict]

FTS5 search (bm25) scoped to the authenticated agent's visible messages

summarize_thread

summarize_thread(project_key: str, thread_id: str, include_examples?: bool, llm_mode?: bool, llm_model?: str, per_thread_limit?: int, agent_name?: str, registration_token?: str)

Single: {thread_id, summary, examples} Multi (comma-sep): {threads[], aggregate}

Extracts participants, key points, actions. Use comma-separated thread_id for multi-thread digest.

install_precommit_guard

install_precommit_guard(project_key: str, code_repo_path: str)

{hook}

Install a Git pre-commit guard in a target repo

uninstall_precommit_guard

uninstall_precommit_guard(code_repo_path: str)

{removed}

Remove the guard from a repo

file_reservation_paths

file_reservation_paths(project_key: str, agent_name: str, paths: list[str], ttl_seconds?: int, exclusive?: bool, reason?: str, registration_token?: str)

{granted: list, conflicts: list}

Advisory leases; Git artifact per path

release_file_reservations

release_file_reservations(project_key: str, agent_name: str, paths?: list[str], file_reservation_ids?: list[int], registration_token?: str)

{released, released_at}

Releases agent's active file reservations

force_release_file_reservation

force_release_file_reservation(project_key: str, agent_name: str, file_reservation_id: int, notify_previous?: bool, note?: str, registration_token?: str)

{released, released_at, reservation}

Clears stale reservations using inactivity/mail/fs/git heuristics and notifies the previous holder

renew_file_reservations

renew_file_reservations(project_key: str, agent_name: str, extend_seconds?: int, paths?: list[str], file_reservation_ids?: list[int], registration_token?: str)

{renewed, file reservations[]}

Extend TTL of existing file reservations

Resources

Output format (resources):

  • Append ?format=toon to any resource URI to receive {format:"toon", data:"<TOON>", meta:{...}}.

  • All resources declare format as an optional query parameter (FastMCP templates accept it).

  • For resources without variable path params (e.g., resource://tooling/projects), include ?format=json or ?format=toon.

  • Defaults to JSON unless MCP_AGENT_MAIL_OUTPUT_FORMAT or TOON_DEFAULT_FORMAT is set.

URI

Params

Returns

Notes

resource://config/environment{?format}

—

{environment, database_url, http}

Inspect server settings

resource://tooling/directory{?format}

—

{generated_at, metrics_uri, clusters[], playbooks[]}

Grouped tool directory + workflow playbooks

resource://tooling/schemas{?format}

—

{tools: {<name>: {required[], optional[], aliases{}}}}

Argument hints for tools

resource://tooling/metrics{?format}

—

{generated_at, tools[]}

Aggregated call/error counts per tool

resource://tooling/locks{?format}

—

{locks[], summary}

Active locks and owners (debug only). Categories: archive (per-project .archive.lock) and custom (e.g., repo .commit.lock).

resource://tooling/capabilities/{agent}{?project}

listed

{generated_at, agent, project, capabilities[]}

Capabilities assigned to the agent (see deploy/capabilities/agent_capabilities.json)

resource://tooling/recent/{window_seconds}{?agent,project}

listed

{generated_at, window_seconds, count, entries[]}

Recent tool usage filtered by agent/project

resource://tooling/projects{?format}

—

list[project]

All projects

resource://project/{slug}

slug

{project..., agents[]}

Project detail + agents

resource://file_reservations/{slug}{?active_only}

slug, active_only?

list[file reservation]

File reservations plus staleness metadata (heuristics, last activity timestamps)

resource://message/{id}{?project,agent,agent_token}

id, project, agent?, agent_token?

message

Single message with body; provide agent auth unless this MCP session already authenticated for the project

resource://thread/{thread_id}{?project,agent,agent_token,include_bodies}

thread_id, project, agent?, agent_token?, include_bodies?

{project, thread_id, messages[]}

Thread listing scoped to the authenticated viewer

resource://inbox/{agent}{?project,since_ts,urgent_only,include_bodies,limit,agent_token}

listed

{project, agent, count, messages[]}

Inbox listing

resource://mailbox/{agent}{?project,limit,agent_token}

project, limit, agent_token?

{project, agent, count, messages[]}

Mailbox listing

resource://mailbox-with-commits/{agent}{?project,limit,agent_token}

project, limit, agent_token?

{project, agent, count, messages[]}

Mailbox listing enriched with commit metadata

resource://outbox/{agent}{?project,limit,include_bodies,since_ts,agent_token}

listed

{project, agent, count, messages[]}

Messages sent by the agent

resource://views/acks-stale/{agent}{?project,ttl_seconds,limit,agent_token}

listed

{project, agent, ttl_seconds, count, messages[]}

Ack-required older than TTL without ack

resource://views/urgent-unread/{agent}{?project,limit,agent_token}

listed

{project, agent, count, messages[]}

High/urgent importance messages not yet read

resource://views/ack-required/{agent}{?project,limit,agent_token}

listed

{project, agent, count, messages[]}

Pending acknowledgements for an agent

resource://views/ack-overdue/{agent}{?project,ttl_minutes,limit,agent_token}

listed

{project, agent, ttl_minutes, count, messages[]}

Ack-required older than TTL without ack

Client Integration Guide

  1. Fetch onboarding metadata first. Issue resources/read for resource://tooling/directory?format=json (and optionally resource://tooling/metrics?format=json) before exposing tools to an agent. Use the returned clusters and playbooks to render a narrow tool palette for the current workflow rather than dumping every verb into the UI.

  2. Scope tools per workflow. When the agent enters a new phase (e.g., "Messaging Lifecycle"), remount only the cluster's tools in your MCP client. This mirrors the workflow macros already provided and prevents "tool overload."

  3. Monitor real usage. Periodically pull or subscribe to log streams containing the tool_metrics_snapshot events emitted by the server (or query resource://tooling/metrics?format=json) so you can detect high-error-rate tools and decide whether to expose macros or extra guidance.

  4. Fallback to macros for smaller models. If you're routing work to a lightweight model, prefer the macro helpers (macro_start_session, macro_prepare_thread, macro_file_reservation_cycle, macro_contact_handshake) and hide the granular verbs until the agent explicitly asks for them.

  5. Show recent actions. Read resource://tooling/recent/60?agent=<name>&project=<slug> (adjust window as needed) to display the last few successful tool invocations relevant to the agent/project.

See examples/client_bootstrap.py for a runnable reference implementation that applies the guidance above.

{
  "steps": [
    "resources/read -> resource://tooling/directory?format=json",
    "select active cluster (e.g. messaging)",
    "mount tools listed in cluster.tools plus macros if model size <= S",
    "optional: resources/read -> resource://tooling/metrics?format=json for dashboard display",
    "optional: resources/read -> resource://tooling/recent/60?agent=<name>&project=<slug> for UI hints"
  ]
}

Monitoring & Alerts

  1. Enable metric emission. Set TOOL_METRICS_EMIT_ENABLED=true and choose an interval (TOOL_METRICS_EMIT_INTERVAL_SECONDS=120 is a good starting point). The server will periodically emit a structured log entry such as:

{
  "event": "tool_metrics_snapshot",
  "tools": [
    {"name": "send_message", "cluster": "messaging", "calls": 42, "errors": 1},
    {"name": "file_reservation_paths", "cluster": "file reservations", "calls": 11, "errors": 0}
  ]
}
  1. Ship the logs. Forward the structured stream (stderr/stdout or JSON log files) into your observability stack (e.g., Loki, Datadog, Elastic) and parse the tools[] array.

  2. Alert on anomalies. Create a rule that raises when errors / calls exceeds a threshold for any tool (for example 5% over a 5-minute window) so you can decide whether to expose a macro or improve documentation.

  3. Dashboard the clusters. Group by cluster to see where agents are spending time and which workflows might warrant additional macros or guard-rails.

See docs/observability.md for a step-by-step cookbook (Loki/Prometheus example pipelines included), and docs/GUIDE_TO_OPTIMAL_MCP_SERVER_DESIGN.md for a comprehensive design guide covering tool curation, capability gating, security, and observability best practices.

Operations teams can follow docs/operations_alignment_checklist.md, which links to the capability templates in deploy/capabilities/ and the sample Prometheus alert rules in deploy/observability/.


Deployment quick notes

  • Direct uvicorn: uvicorn mcp_agent_mail.http:create_app --factory --host 0.0.0.0 --port 8765

  • Python module: python -m mcp_agent_mail.http --host 0.0.0.0 --port 8765

  • Gunicorn: gunicorn -c deploy/gunicorn.conf.py mcp_agent_mail.http:create_app --factory

  • Docker: docker compose up --build

CI/CD

  • Lint and Typecheck CI: GitHub Actions workflow runs Ruff and Ty on pushes/PRs to main/develop.

  • Release: Pushing a tag like v0.1.0 builds and pushes a multi-arch Docker image to GHCR under ghcr.io/<owner>/<repo> with latest and version tags.

  • Nightly: A scheduled workflow runs migrations and lists projects daily for lightweight maintenance visibility.

Log rotation (optional)

If not using journald, a sample logrotate config is provided at deploy/logrotate/mcp-agent-mail to rotate /var/log/mcp-agent-mail/*.log weekly, keeping 7 rotations.

Logging (journald vs file)

  • Default systemd unit (deploy/systemd/mcp-agent-mail.service) is configured to send logs to journald (StandardOutput/StandardError=journal).

  • For file logging, configure your process manager to write to files under /var/log/mcp-agent-mail/*.log and install the provided logrotate config.

  • Environment file path for systemd is /etc/mcp-agent-mail.env (see deploy/systemd/mcp-agent-mail.service).

Container build and multi-arch push

Use Docker Buildx for multi-arch images. Example flow:

# Create and select a builder (once)
docker buildx create --use --name mcp-builder || docker buildx use mcp-builder

# Build and test locally (linux/amd64)
docker buildx build --load -t your-registry/mcp-agent-mail:dev .

# Multi-arch build and push (amd64, arm64)
docker buildx build \
  --platform linux/amd64,linux/arm64 \
  -t your-registry/mcp-agent-mail:latest \
  -t your-registry/mcp-agent-mail:v0.1.0 \
  --push .

Recommended tags: a moving latest and immutable version tags per release. Ensure your registry login is configured (docker login).

Systemd manual deployment steps

  1. Copy project files to /opt/mcp-agent-mail and ensure permissions (owner appuser).

  2. Place environment file at /etc/mcp-agent-mail.env based on deploy/env/production.env.

  3. Install service file deploy/systemd/mcp-agent-mail.service to /etc/systemd/system/.

  4. Reload systemd and start:

sudo systemctl daemon-reload
sudo systemctl enable mcp-agent-mail
sudo systemctl start mcp-agent-mail
sudo systemctl status mcp-agent-mail

Optional (non-journald log rotation): install deploy/logrotate/mcp-agent-mail into /etc/logrotate.d/ and write logs to /var/log/mcp-agent-mail/*.log via your process manager or app config.

See deploy/gunicorn.conf.py for a starter configuration. For project direction and planned areas, read docs/planning/project_idea_and_guide.md.

CLI Commands

The project exposes a developer CLI for common operations:

  • serve-http: run the HTTP transport (Streamable HTTP only)

  • migrate: ensure schema and FTS structures exist

  • lint / typecheck: developer helpers

  • list-projects [--include-agents]: enumerate projects

  • guard install <project_key> <code_repo_path>: install the pre-commit guard into a repo

  • guard uninstall <code_repo_path>: remove the guard from a repo

  • share wizard: launch interactive deployment wizard (auto-installs CLIs, authenticates, exports, deploys to GitHub Pages or Cloudflare Pages)

  • share export --output <path> [options]: export mailbox to a static HTML bundle (see Static Mailbox Export section for full options)

  • share update <bundle_path> [options]: refresh an existing bundle using recorded (or overridden) export settings

  • share preview <bundle_path> [--port N] [--open-browser]: serve a static bundle locally for inspection

  • share verify <bundle_path> [--public-key <key>]: verify bundle integrity (SRI hashes and Ed25519 signature)

  • share decrypt <encrypted_path> [--identity <file> | --passphrase]: decrypt an age-encrypted bundle

  • config set-port <port>: change the HTTP server port (updates .env)

  • config show-port: display the current configured HTTP port

  • clear-and-reset-everything [--force] [--archive/--no-archive]: DELETE the SQLite database (incl. WAL/SHM) and WIPE all contents under STORAGE_ROOT after optionally saving a restore point. Without flags it prompts to create an archive first; --force --no-archive skips all prompts for automation.

  • list-acks --project <key> --agent <name> [--limit N]: list messages requiring acknowledgement for an agent where ack is missing

  • acks pending <project> <agent> [--limit N]: show pending acknowledgements for an agent

  • acks remind <project> <agent> [--min-age-minutes N] [--limit N]: highlight pending ACKs older than a threshold

  • acks overdue <project> <agent> [--ttl-minutes N] [--limit N]: list overdue ACKs beyond TTL

  • file_reservations list <project> [--active-only/--no-active-only]: list file reservations

  • file_reservations active <project> [--limit N]: list active file reservations

  • file_reservations soon <project> [--minutes N]: show file reservations expiring soon

  • doctor check [PROJECT] [--verbose] [--json]: run comprehensive diagnostics on mailbox health

  • doctor repair [PROJECT] [--dry-run] [--yes] [--backup-dir PATH]: semi-automatic repair with backup before changes

  • doctor backups [--json]: list available diagnostic backups

  • doctor restore <backup_path> [--dry-run] [--yes]: restore from a diagnostic backup

Examples:

# Interactive wizard: export + deploy to GitHub Pages (easiest)
./scripts/share_to_github_pages.py

# Export a static bundle with signing and encryption
uv run python -m mcp_agent_mail.cli share export \
  --output ./bundle \
  --signing-key ./keys/signing.key \
  --age-recipient age1abc...xyz

# Preview a bundle locally
uv run python -m mcp_agent_mail.cli share preview ./bundle --port 9000 --open-browser

# Verify bundle integrity
uv run python -m mcp_agent_mail.cli share verify ./bundle

# Refresh an existing bundle in place with recorded settings
uv run python -m mcp_agent_mail.cli share update ./bundle

# Change server port
uv run python -m mcp_agent_mail.cli config set-port 9000

# Install guard into a repo
uv run python -m mcp_agent_mail.cli guard install /abs/path/backend /abs/path/backend

# List pending acknowledgements for an agent
uv run python -m mcp_agent_mail.cli acks pending /abs/path/backend BlueLake --limit 10

# Run mailbox health diagnostics
uv run python -m mcp_agent_mail.cli doctor check

# Preview repairs without making changes
uv run python -m mcp_agent_mail.cli doctor repair --dry-run

# Run repairs (creates backup first, prompts for data changes)
uv run python -m mcp_agent_mail.cli doctor repair

# List available backups
uv run python -m mcp_agent_mail.cli doctor backups

# Restore from a backup
uv run python -m mcp_agent_mail.cli doctor restore /path/to/backup --dry-run

# WARNING: Destructive reset (clean slate)
uv run python -m mcp_agent_mail.cli clear-and-reset-everything --force

Client integrations

Use the automated installer to wire up supported tools automatically (e.g., Claude Code, Cline, Windsurf, OpenCode). Run scripts/automatically_detect_all_installed_coding_agents_and_install_mcp_agent_mail_in_all.sh or the one-liner in the Quickstart above.

Tool-specific integration scripts

For manual integration or customization, dedicated scripts are available:

Tool

Script

What it configures

Claude Code

scripts/integrate_claude_code.sh

.claude/settings.json, hooks, MCP server

Codex CLI

scripts/integrate_codex_cli.sh

~/.codex/config.toml, MCP server, .codex/hooks.json PostToolUse inbox hook

Gemini CLI

scripts/integrate_gemini_cli.sh

~/.gemini/settings.json, MCP server, hooks

Factory Droid

scripts/integrate_factory_droid.sh

~/.factory/settings.json, MCP server, hooks

Each script:

  • Detects the MCP server endpoint from your settings

  • Generates or reuses a bearer token for authentication

  • Configures the MCP server connection

  • Installs hooks/notify handlers for inbox reminders

  • Bootstraps your project and agent identity on the server

Automatic inbox reminders

Agents often get absorbed in their work and forget to check their mail. The integration scripts install lightweight hooks that periodically remind agents when they have unread messages.

How it works:

  • A rate-limited hook script (scripts/hooks/check_inbox.sh) runs after certain tool invocations

  • It checks the inbox via a fast curl call (avoids Python import overhead) with unread_only=true, so mail the agent has already read or acknowledged is never re-announced

  • If there are unread messages, it outputs a brief reminder

  • Rate limited to at most once per 2 minutes to avoid noise

Claude Code / Gemini CLI:

The hook is configured as a PostToolUse hook that fires after Bash or shell tool invocations:

{
  "hooks": {
    "PostToolUse": [
      {
        "matcher": "Bash",
        "hooks": [{ "type": "command", "command": "...check_inbox.sh" }]
      }
    ]
  }
}

Codex CLI:

Uses a project-local Codex hook (<project>/.codex/hooks.json) with the same PostToolUse envelope as Claude Code. The installer copies check_inbox.sh to .codex/hooks/, wraps it in a mode-0700 inbox_wrapper.sh that holds the tokens (so no secret lands in hooks.json), and merges this entry into any existing hooks:

{
  "hooks": {
    "PostToolUse": [
      {
        "matcher": "Bash",
        "hooks": [{ "type": "command", "command": "'/path/to/.codex/hooks/inbox_wrapper.sh'", "timeout": 10 }]
      }
    ]
  }
}

Codex asks you to review and trust the hook (pinned by hash) before it runs, and loads project-local hooks only once the project's .codex/ layer is trusted. Because the hook is project-local, each project keeps its own agent identity and rate-limit window.

The older top-level notify = [...] mechanism is no longer installed: current Codex spawns the notify program with its output discarded and never shows it to the model, so a reminder printed there cannot reach the agent. If an earlier install left a notify = [".../notify_wrapper.sh"] line in ~/.codex/config.toml, it is harmless and can be deleted.

Additional hooks (Claude Code only):

Event

What it does

SessionStart

Shows active file reservations and pending acknowledgments

PreToolUse (Edit)

Warns about file reservations expiring within 10 minutes

PostToolUse (send_message)

Lists recent acknowledgment requests

PostToolUse (file_reservation_paths)

Shows current file reservations

Environment variables for hooks

The inbox check hooks accept these environment variables (set automatically by the integration scripts):

Variable

Description

Default

AGENT_MAIL_PROJECT

Project key (absolute path)

required

AGENT_MAIL_AGENT

Agent name

required

AGENT_MAIL_URL

Server URL

http://127.0.0.1:8765/mcp/

AGENT_MAIL_TOKEN

Bearer token

none

AGENT_MAIL_INTERVAL

Seconds between checks

120


About Contributions: Please don't take this the wrong way, but I do not accept outside contributions for any of my projects. I simply don't have the mental bandwidth to review anything, and it's my name on the thing, so I'm responsible for any problems it causes; thus, the risk-reward is highly asymmetric from my perspective. I'd also have to worry about other "stakeholders," which seems unwise for tools I mostly make for myself for free. Feel free to submit issues, and even PRs if you want to illustrate a proposed fix, but know I won't merge them directly. Instead, I'll have Claude or Codex review submissions via gh and independently decide whether and how to address them. Bug reports in particular are welcome. Sorry if this offends, but I want to avoid wasted time and hurt feelings. I understand this isn't in sync with the prevailing open-source ethos that seeks community contributions, but it's the only way I can move at this velocity and keep my sanity.

Available Tools

41 tools
acknowledge_messageA

Acknowledge a message addressed to an agent (and mark as read).

Behavior

  • Sets both read_ts and ack_ts for the (agent, message) pairing

  • Safe to call multiple times; subsequent calls will return the prior timestamps

Idempotency

  • If acknowledgement already exists, the previous timestamps are preserved and returned.

When to use

  • Respond to messages with ack_required=true to signal explicit receipt.

  • Agents can treat an acknowledgement as a lightweight, non-textual reply.

Returns

dict { message_id, acknowledged: bool, acknowledged_at: iso8601 | null, read_at: iso8601 | null }

Example

{"jsonrpc":"2.0","id":"9","method":"tools/call","params":{"name":"acknowledge_message","arguments":{
  "project_key":"/abs/path/backend","agent_name":"BlueLake","message_id":1234
}}}
ParametersJSON Schema
NameRequiredDescriptionDefault
formatNo
agent_nameYes
message_idYes
project_keyYes
registration_tokenNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and largely meets it: it discloses that it sets both read_ts and ack_ts, that it is idempotent, and that prior timestamps are preserved and returned. It does not mention auth or whether registration_token is required, which is a real gap for a 5-param mutation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Well-structured with clear headers, front-loaded purpose, and a compact example. Slightly padded: the Returns section restates what the output schema already provides, and the example is long, but nothing is egregiously wasteful.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no annotations and 0% schema coverage, the description covers behavior, idempotency and return shape well. However, two of five parameters (format, registration_token) are completely unaddressed, leaving an agent unable to decide whether registration_token is needed to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% and there are 5 parameters, so the description must compensate. It offers no parameter semantics anywhere: format and registration_token are entirely undocumented in both schema and description, and project_key, agent_name and message_id appear only as literal values in the example without explanation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (acknowledge) and resource (a message addressed to an agent), and the parenthetical 'and mark as read' clarifies scope. The behavior section distinguishes it from the sibling mark_message_read by noting it sets both read_ts and ack_ts. No explicit sibling naming, so 4 rather than 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

A dedicated 'When to use' section says to respond to messages with ack_required=true and that it acts as a lightweight non-textual reply. This is clear context, but it never names alternatives such as reply_message or mark_message_read, so the routing is inferred rather than explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

archive_projectA

Soft-delete a project: mark it as archived so it is hidden from active project lists. All messages are preserved and the project can be restored with unarchive_project.

ParametersJSON Schema
NameRequiredDescriptionDefault
project_keyYes
registration_tokenNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the behavioral burden well: it discloses that this is a reversible soft-delete, that messages are preserved, and that the effect is hiding from active lists. It omits permissions/auth requirements and anything about the registration_token's role, so it is strong but not exhaustive.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, zero filler, with the core semantics (soft-delete, reversible, data preserved) front-loaded. Every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be described. The description covers reversibility and data retention, which are the critical facts for a mutation tool, but leaves the registration_token parameter undocumented.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and neither parameter is mentioned in the description. project_key is inferable from context, but registration_token is a non-obvious optional parameter left entirely unexplained in both schema and description, so the description fails to compensate for the coverage gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('soft-delete a project: mark it as archived') and immediately distinguishes itself from the destructive sibling by specifying that data is preserved. The mention of unarchive_project as the reversal path makes it unmistakable against hard_delete_project and unarchive_project.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear context for when to use it (to hide a project from active lists while keeping messages) and names the complementary tool for undoing it. It does not explicitly state when to prefer this over hard_delete_project, so the exclusion is implied rather than stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_agent_identityA

Create a new, unique agent identity and persist its profile to Git.

How this differs from register_agent

  • Always creates a new identity with a fresh unique name (never updates an existing one).

  • name_hint, if provided, MUST be a valid adjective+noun combination and must be available, otherwise an error is raised. Without a hint, a random adjective+noun name is generated.

CRITICAL: Agent Naming Rules

  • Agent names MUST be randomly generated adjective+noun combinations

  • Examples: "GreenCastle", "BlueLake", "RedStone", "PurpleBear"

  • Names should be unique, easy to remember, and NOT descriptive

  • INVALID examples: "BackendHarmonizer", "DatabaseMigrator", "UIRefactorer"

  • Best practice: Omit name_hint to auto-generate a valid name (RECOMMENDED)

When to use

  • Spawning a brand new worker agent that should not overwrite an existing profile.

  • Temporary task-specific identities (e.g., short-lived refactor assistants).

Parameters

return_registration_token : bool, default True When True (default, current behaviour), the response includes the freshly-minted registration_token. When False, the token is omitted from the tool result so transcript-visible MCP sessions can satisfy a "do not echo secrets into scrollback" contract; the agent is still bound to the current MCP session via _bind_session_agent, so follow-up calls in the same session can authenticate without ever surfacing the token. The token still exists on the server and can be retrieved or rotated through the normal admin paths. See issue #154.

Returns

dict { id, name, program, model, task_description, inception_ts, last_active_ts, project_id, registration_token? }

Examples

Auto-generate name (RECOMMENDED):

{"jsonrpc":"2.0","id":"c2","method":"tools/call","params":{"name":"create_agent_identity","arguments":{
  "project_key":"/data/projects/backend","program":"claude-code","model":"opus-4.1"
}}}

With valid name hint:

{"jsonrpc":"2.0","id":"c1","method":"tools/call","params":{"name":"create_agent_identity","arguments":{
  "project_key":"/data/projects/backend","program":"codex-cli","model":"gpt5-codex","name_hint":"GreenCastle",
  "task_description":"DB migration spike"
}}}

Transcript-safe creation (issue #154) — omit the token from the visible tool result and rely on session binding for follow-ups:

{"jsonrpc":"2.0","id":"c3","method":"tools/call","params":{"name":"create_agent_identity","arguments":{
  "project_key":"/data/projects/backend","program":"codex-cli","model":"gpt5",
  "return_registration_token":false
}}}
ParametersJSON Schema
NameRequiredDescriptionDefault
modelYes
formatNo
programYes
name_hintNo
project_keyYes
task_descriptionNo
attachments_policyNoauto
return_registration_tokenNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden and does so: it states this never updates an existing profile, that an invalid/unavailable name_hint raises an error, that random adjective+noun generation occurs otherwise, that profiles are persisted to Git, and it explains the registration-token/session-binding contract (issue #154). This is well beyond what any structured field supplies.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose and sibling differentiation are correctly front-loaded, but the definition is bloated: three full JSON-RPC examples, and the third example re-demonstrates the return_registration_token behavior already described at length in the Parameters section. The naming-rules block with invalid examples is long for what it conveys.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a creation tool with an output schema (so return values need not be re-explained) and 8 params, the description covers naming rules, error conditions, sibling differences, and the token contract. It falls short only on the undocumented format and attachments_policy parameters.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, and it does for name_hint (validation rules, error behavior, recommended omission) and return_registration_token (default, transcript-safety, session binding). However, format and attachments_policy are never mentioned, and project_key/program/model/task_description are only demonstrated inside examples rather than explained.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Create a new, unique agent identity and persist its profile to Git') and immediately distinguishes it from the closest sibling, register_agent. An agent can tell it apart from register_agent/deregister_agent without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Has an explicit 'How this differs from register_agent' section plus a 'When to use' list (brand-new worker, temporary task-specific identity) and a best-practice recommendation to omit name_hint. The alternative and the conditions that select each path are named outright.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

deregister_agentB

Remove an agent from a project. Marks the agent as inactive and removes it from the active roster. Messages from/to the agent are preserved for audit but the agent can no longer send or receive new messages.

ParametersJSON Schema
NameRequiredDescriptionDefault
agent_nameYes
project_keyYes
registration_tokenNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden, and it does disclose useful facts: messages are preserved for audit, the agent is marked inactive, and it can no longer send or receive messages. It omits reversibility (does unretire_agent restore a deregistered agent?), any auth requirements for the undocumented registration_token, and side effects on existing reservations or contacts.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, front-loaded with the action and immediately followed by the concrete effects; every sentence earns its place. Slight deduction because the effect sentences could have been compressed into one without losing information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need no explanation, but the definition still leaves gaps: 0% parameter documentation, no annotations, and no differentiation from the retire/hard_delete siblings. For a destructive-adjacent lifecycle tool this is only minimally adequate.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across three parameters, so the description must compensate and largely does not. It implicitly maps to agent_name and project_key, but registration_token (optional, nullable, defaulting to null) is never explained, and no format or identifier conventions are given for any parameter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource ('Remove an agent from a project') and clarifies the effect (marked inactive, dropped from the active roster). However, it never distinguishes itself from the very similar siblings retire_agent, unretire_agent, and hard_delete_agent, leaving the agent to guess which removal semantics apply.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance and no mention of the obvious alternatives in the sibling set (retire_agent, hard_delete_agent). An agent cannot tell from this text whether deregister is soft, reversible, or terminal relative to those tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ensure_projectA

Idempotently create or ensure a project exists for the given human key.

When to use

  • First call in a workflow targeting a new repo/path identifier.

  • As a guard before registering agents or sending messages.

How it works

  • Validates that human_key is an absolute path-like project key (typically the agent's working directory). It need not exist on the local filesystem: it is an opaque project KEY, and collaborating agents may not share a filesystem.

  • Computes a stable slug from human_key (lowercased, safe characters) so multiple agents can refer to the same project consistently.

  • Ensures DB row exists and that the on-disk archive is initialized (e.g., messages/, agents/, file_reservations/ directories).

CRITICAL: Project Identity Rules

  • The human_key MUST be an absolute path-like project key (typically the agent's working directory path)

  • Two agents working in the SAME directory path are working on the SAME project

  • Example: Both agents in /data/projects/smartedgar_mcp → SAME project

  • Sibling projects are DIFFERENT directories (e.g., /data/projects/smartedgar_mcp vs /data/projects/smartedgar_mcp_frontend)

Parameters

human_key : str An absolute path-like project key (e.g., "/data/projects/backend"), typically the agent's working directory. This MUST be an absolute path, not a relative path or arbitrary slug, but it does NOT need to exist on the local filesystem - it is an opaque project KEY (collaborating agents may not share a filesystem). This is the canonical identifier for the project - all agents using the same key share the same project identity. identity_mode : str, optional Per-call override of the server's PROJECT_IDENTITY_MODE setting; one of "dir", "git-remote", "git-common-dir", "git-toplevel". Only takes effect when worktree-friendly identity is enabled (WORKTREES_ENABLED=1).

Returns

dict Minimal project descriptor: { id, slug, human_key, created_at }.

Examples

JSON-RPC:

{
  "jsonrpc": "2.0",
  "id": "2",
  "method": "tools/call",
  "params": {"name": "ensure_project", "arguments": {"human_key": "/data/projects/backend"}}
}

Common mistakes

  • Passing a relative path (e.g., "./backend") instead of an absolute path

  • Using arbitrary slugs instead of the actual working directory path

  • Creating separate projects for the same directory with different slugs

Idempotency

  • Safe to call multiple times. If the project already exists, the existing record is returned and the archive is ensured on disk (no destructive changes).

ParametersJSON Schema
NameRequiredDescriptionDefault
formatNo
human_keyYes
identity_modeNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden and does so thoroughly: idempotent, non-destructive, validates the key, computes a stable slug, and initializes on-disk directories (messages/, agents/, file_reservations/). It further discloses the return shape and the identity-mode precondition (WORKTREES_ENABLED=1).

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with a strong opening sentence and well-sectioned headers. But the absolute-path/canonical-identifier rule is restated three times (identity rules, Parameters, Common mistakes), and the raw JSON-RPC example adds little. Redundancy dilutes an otherwise structured description.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex identity-establishing tool with an output schema present, the description covers purpose, workflow placement, validation behavior, side effects on disk, idempotency, and example invocations. Nothing an agent needs to call it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 0% schema description coverage, the description must compensate, and it does for two of three params — human_key (absolute, path-like, opaque, need not exist locally) and identity_mode (enum values plus the WORKTREES_ENABLED=1 gate). However, the third schema parameter 'format' is never mentioned anywhere.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource: 'Idempotently create or ensure a project exists for the given human key.' It is clearly distinguishable from siblings like archive_project, hard_delete_project, and register_agent, which do different things. The 'ensure/project identity' framing is precise.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'When to use' explicitly says 'First call in a workflow targeting a new repo/path identifier' and 'As a guard before registering agents or sending messages,' giving clear workflow placement. It does not, however, name a specific alternative or state when-not-to-use relative to siblings like archive_project.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

expire_windowC

Mark a window identity as expired.

Parameters

project_key : str Project identifier. window_uuid : str The UUID of the window identity to expire.

Returns

dict { window_uuid, expired: bool, expired_at }

ParametersJSON Schema
NameRequiredDescriptionDefault
formatNo
project_keyYes
window_uuidYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden, yet it discloses nothing beyond the verb: it does not say whether expiry is reversible, what happens to the window's messages or state, whether permissions are required, or what triggers it. The Returns section only restates structure already available in the output schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the purpose, then a compact parameter list that is genuinely useful given the 0% schema coverage. The 'Returns' block duplicates the existing output schema and is minor waste, but overall it is tight and well structured.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The existence of an output schema means return values need not be explained, and the description redundantly does so anyway. For a mutation tool with no annotations it still leaves key gaps: the undocumented 'format' parameter and the absence of any behavioral or usage context.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With schema description coverage at 0%, the description compensates for two of three parameters ('Project identifier', 'the UUID of the window identity to expire'), adding real meaning the schema lacks. It omits the 'format' parameter entirely, leaving that third parameter undocumented anywhere.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Mark ... as expired') and resource ('window identity'), which is clear enough to distinguish from the sibling rename_window and list_window_identities. However, it does no explicit sibling differentiation and never clarifies what 'expired' means to the identity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives no when-to-use guidance, no prerequisites, and no reference to any alternative tool or flow. Usage must be entirely inferred from the tool name and the sibling list.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

fetch_inboxA

Retrieve recent messages for an agent without mutating read/ack state.

Filters

  • urgent_only: only messages with importance in {high, urgent}

  • since_ts: ISO-8601 timestamp string; messages strictly newer than this are returned

  • limit: max number of messages (default 20)

  • include_bodies: include full Markdown bodies in the payloads

  • topic: filter to messages with this topic tag

  • unread_only: when True, restrict to messages this recipient has not yet explicitly marked read via mark_message_read or acknowledge_message. Per-recipient: a message read by Agent A is still unread for Agent B. A bare fetch_inbox call does NOT mark messages read; this filter inspects existing read state without mutating it.

Usage patterns

  • Poll after each editing step in an agent loop to pick up coordination messages.

  • Use since_ts with the timestamp from your last poll for efficient incremental fetches.

  • Use unread_only=True from polling agents (Claude Code, Codex, etc.) to skip messages the agent has already acknowledged — cuts token-burn at scale by avoiding re-running prompt context against already-handled mail.

  • Combine with acknowledge_message if ack_required is true.

Returns

list[dict] Each message includes: { id, subject, from, created_ts, importance, ack_required, kind, read_at, [body_md] } read_at is this recipient's read timestamp (null while unread), so the default view — which includes already-read mail — stays distinguishable.

Example

{"jsonrpc":"2.0","id":"7","method":"tools/call","params":{"name":"fetch_inbox","arguments":{
  "project_key":"/abs/path/backend","agent_name":"BlueLake","since_ts":"2025-10-23T00:00:00+00:00"
}}}
ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
topicNo
formatNo
since_tsNo
agent_nameYes
project_keyYes
unread_onlyNo
urgent_onlyNo
include_bodiesNo
registration_tokenNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does well on the core behavioral trait: it repeatedly states that fetching does NOT mutate read/ack state, and explains per-recipient read semantics. It omits auth/permission requirements and rate-limit behavior, and the registration_token parameter is never mentioned, which is the remaining gap for a no-annotation tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose is front-loaded, then clearly sectioned into Filters / Usage patterns / Returns / Example. It is long, but the length is justified by ten parameters and no schema descriptions; the Returns block is somewhat redundant given an output schema exists.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Covers purpose, filtering semantics, non-mutation guarantees, and a working call example, which is strong for a no-annotation tool. The main omissions are the format and registration_token parameters and any authentication context, which an agent may need to invoke correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It richly documents six of ten parameters (urgent_only's importance enum, since_ts strictness, limit default, include_bodies, topic, unread_only's per-recipient semantics), but leaves format, registration_token, and the meaning of project_key/agent_name (beyond their appearance in the example) undocumented, so the compensation is partial.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Retrieve recent messages for an agent') plus a distinguishing scope qualifier ('without mutating read/ack state'), which cleanly separates it from siblings like search_messages, fetch_topic, and fetch_summary.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The 'Usage patterns' section gives explicit when-to-use guidance: poll after each editing step, use since_ts for incremental fetches, use unread_only to cut token burn, and combine with acknowledge_message when ack_required. Alternatives and conditions are named rather than left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

fetch_summaryA

Retrieve stored project-wide summaries.

Parameters

project_key : str Project identifier. since_hours : float Return summaries whose end_ts is within this window (default 24h). limit : int Maximum summaries to return (default 5). format : str, optional Output format.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
formatNo
project_keyYes
since_hoursNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It does disclose useful behavior: the since_hours window is applied against end_ts, and defaults of 24h and 5 summaries are stated. It omits permissions, ordering of results, and any note on whether summaries are per-project only or also per-agent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose is front-loaded in a single sentence, followed by a scannable parameter block. The parameter block largely restates the schema, which is redundant when the schema is visible, but it is compact and well ordered.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained. What is missing is the relationship to the sibling summarizers: an agent cannot tell from this text whether fetch_summary retrieves summaries produced by summarize_recent/summarize_thread or some independently stored artifact, which matters for calling it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, and it does document all four parameters: project_key as identifier, since_hours as an end_ts window with a 24h default, limit as a max count with a 5 default, and format as the output format. The only weak spot is 'format', whose accepted values are never enumerated.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Retrieve stored project-wide summaries'), which separates it from the generative siblings summarize_thread and summarize_recent by the word 'stored'. It does not explicitly name those siblings, so the distinction must be inferred rather than read.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is only implied: 'stored' hints that this reads previously generated summaries rather than producing new ones, but the description never states when to prefer this over summarize_recent, summarize_thread, or fetch_inbox. No prerequisites or exclusions are given.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

fetch_topicA

Fetch all messages in a project with a given topic tag, regardless of recipient.

Parameters

project_key : str Project identifier. topic_name : str The topic tag to filter by (case-insensitive). limit : int Max number of messages to return (default 50). include_bodies : bool Include full Markdown bodies in the payloads (default true). since_ts : Optional[str] ISO-8601 timestamp; only messages newer than this are returned. unread_only : bool When True, restrict to messages where the viewer has a recipient row that has not been explicitly marked read. This narrows beyond the default sender-or-recipient visibility — messages the viewer sent (but is not a recipient of) and broadcast/thread-visible messages where the viewer has no MessageRecipient row are excluded under this flag, because "unread" is only well-defined for a recipient row. A bare fetch_topic call does NOT mark messages read.

Returns

list[dict] Each message includes: { id, subject, from, created_ts, importance, topic, [body_md] }

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
formatNo
since_tsNo
agent_nameNo
topic_nameYes
project_keyYes
unread_onlyNo
include_bodiesNo
registration_tokenNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does disclose non-obvious behavior: a bare call does not mark messages read, and unread_only narrows visibility in a way that excludes sent and broadcast messages lacking a MessageRecipient row. It omits permissions/auth context that the unexplained registration_token and agent_name parameters imply, and says nothing about truncation when limit is hit.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Numpydoc-style layout is front-loaded with the one-line purpose, then parameters and returns, so an agent can stop reading early. The longer unread_only paragraph earns its length by resolving a genuinely ambiguous flag.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, yet the description still summarizes the returned fields, and it covers the two required parameters plus the main filters. The gap is the auth/identity parameters, which are neither in the schema nor explained, leaving an agent unsure whether credentials are needed for this call.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, and it documents six of nine parameters with real semantics, including the delicate unread_only narrowing rule and the include_bodies payload effect. Three parameters (format, agent_name, registration_token) remain undocumented in both description and schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (fetch), resource (messages), and a scoping rule (with a given topic tag, regardless of recipient) that immediately separates it from recipient-scoped siblings like fetch_inbox. An agent can distinguish it from fetch_inbox and search_messages without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'regardless of recipient' implicitly contrasts with inbox-style retrieval, and the note that a bare call does not mark messages read hints at mark_message_read as the follow-up. However, no sibling is named explicitly and there is no stated when-not condition, leaving routing to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

file_reservation_pathsA

Request advisory file reservations (leases) on project-relative paths/globs.

Semantics

  • Conflicts are reported if an overlapping active exclusive reservation exists held by another agent

  • Glob matching is symmetric (fnmatchcase(a,b) or fnmatchcase(b,a)), including exact matches

  • When granted, a JSON artifact is written under file_reservations/<sha1(path)>.json and the DB is updated

  • TTL must be >= 60 seconds (enforced by the server settings/policy)

  • Server-side enforcement (if enabled) only checks reservations that target mail archive paths such as agents/, messages/, or attachments/; code repo enforcement is via the pre-commit guard

Do / Don't

Do:

  • Reserve files before starting edits to signal intent to other agents.

  • Use specific, minimal patterns (e.g., app/api/*.py) instead of broad globs.

  • Set a realistic TTL and renew with renew_file_reservations if you need more time.

Don't:

  • Reserve the entire repository or very broad patterns (e.g., **/*) unless absolutely necessary.

  • Hold long-lived exclusive reservations when you are not actively editing.

  • Ignore conflicts; resolve them by coordinating with holders or waiting for expiry.

Parameters

project_key : str agent_name : str paths : list[str] File paths or glob patterns relative to the project workspace (e.g., "app/api/*.py"). ttl_seconds : int Time to live for the file_reservation; expired file_reservations are auto-released. exclusive : bool If true, exclusive intent; otherwise shared/observe-only. reason : str Optional explanation (helps humans reviewing Git artifacts).

Returns

dict { granted: [{id, path_pattern, exclusive, reason, expires_ts}], conflicts: [{path, holders: [...]}] }

Example

{"jsonrpc":"2.0","id":"12","method":"tools/call","params":{"name":"file_reservation_paths","arguments":{
  "project_key":"/abs/path/backend","agent_name":"GreenCastle","paths":["app/api/*.py"],
  "ttl_seconds":7200,"exclusive":true,"reason":"migrations"
}}}
ParametersJSON Schema
NameRequiredDescriptionDefault
pathsYes
formatNo
reasonNo
exclusiveNo
agent_nameYes
project_keyYes
ttl_secondsNo
registration_tokenNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations the description carries the full burden and delivers: conflict-detection rules, symmetric glob matching semantics, side effects (JSON artifact written to file_reservations/<sha1>.json plus DB update), the enforced TTL floor of 60s, and the precise scope of server-side enforcement. This is unusually rich disclosure for a mutation/lease tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Sectioned (Semantics, Do/Don't, Parameters, Returns, Example) and front-loaded with the core purpose, so an agent can scan it quickly. It runs long and the JSON-RPC example is somewhat verbose, but nearly every line carries actionable information for an 8-parameter tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, yet the description still helpfully spells out the {granted, conflicts} return shape. Combined with the semantics and conflict behavior, an agent has enough to call it correctly, with the only real gap being the undocumented 'format' and 'registration_token' parameters.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate, and it documents 6 of 8 params (project_key, agent_name, paths, ttl_seconds, exclusive, reason) with real meaning beyond the bare schema. However, 'format' and especially 'registration_token' are never explained, leaving an apparent auth/format parameter undocumented.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Opens with a specific verb+resource+scope: 'Request advisory file reservations (leases) on project-relative paths/globs.' This clearly distinguishes it from the sibling reservation tools (release_file_reservations, renew_file_reservations, force_release_file_reservation) without needing to open any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The Do/Don't section gives explicit when-to-use (reserve before edits), when-not (broad globs, long-lived holds, ignoring conflicts), and names the alternative 'renew with renew_file_reservations' for extending time. This is exactly the when/when-not/alternatives structure that earns a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

force_release_file_reservationB

Force-release a stale file reservation held by another agent after inactivity heuristics.

The tool validates that the reservation appears abandoned (agent inactive beyond threshold and no recent mail/filesystem/git activity). When released, an optional notification is sent to the previous holder summarizing the heuristics.

ParametersJSON Schema
NameRequiredDescriptionDefault
noteNo
formatNo
agent_nameYes
project_keyYes
notify_previousNo
registration_tokenNo
file_reservation_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden, and it does add real value: it discloses the abandonment-heuristic validation and that an optional notification is sent to the previous holder. It omits auth/permission requirements (implied by registration_token), whether the release is reversible, and any failure behavior when the heuristics do not validate.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences that front-load the action before describing validation and notification behavior. Every clause is relevant, with only mild redundancy between 'inactivity heuristics' and the enumerated activity checks.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return values need not be explained, and the description covers the core mutation and its trigger well. But for a destructive action on another agent's resource, the absence of auth/permission and irreversibility details leaves it short of complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% across 7 parameters, so the description must compensate and largely does not. It loosely touches notify_previous ('optional notification') but says nothing about note, format, registration_token, file_reservation_id, or project_key, leaving the majority of parameters undocumented in both schema and description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource ('Force-release a stale file reservation') and crucially scopes it to a reservation 'held by another agent' with abandonment heuristics, which separates it from the plain release siblings. It never names release_file_reservations explicitly, so sibling differentiation is by implication rather than direct contrast.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied through the validation conditions ('agent inactive beyond threshold and no recent mail/filesystem/git activity'), telling the agent this is for abandoned reservations. However, it never says when to prefer this over release_file_reservations or renew_file_reservations, nor any exclusion such as 'do not use on your own reservations'.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

hard_delete_agentA

IRREVERSIBLY delete an agent and ALL associated data: messages sent by the agent, message recipient records, file reservations, agent links, window identities, and on-disk archive files (inbox, outbox, attachments). This is NOT soft-delete — data is permanently destroyed and cannot be recovered. Requires the confirmation parameter to be exactly 'I UNDERSTAND'.

ParametersJSON Schema
NameRequiredDescriptionDefault
agent_nameYes
project_keyYes
confirmationYes
registration_tokenNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.1/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations supplied, the description carries the full behavioral burden and does so well: irreversibility, permanence ('cannot be recovered'), an exhaustive list of destroyed data classes, and the on-disk artifacts (inbox, outbox, attachments) affected. Only the authorization requirements around registration_token are left unstated.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three dense sentences, front-loaded with the destructive action and its scope, followed by the soft-delete disqualifier and the confirmation requirement. No filler or repetition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be described, and the destruction scope plus confirmation gate cover the highest-risk aspects of this operation. The remaining gap is the undocumented registration_token and any auth/permission prerequisites implicit in its presence.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across four parameters, so the description must compensate, yet it only meaningfully explains 'confirmation' (exact literal 'I UNDERSTAND'). Neither project_key, agent_name, nor the optional registration_token is described, and registration_token in particular is not self-explanatory.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('IRREVERSIBLY delete an agent') and then enumerates exactly what is removed: messages, recipient records, file reservations, agent links, window identities, and on-disk archives. The explicit 'This is NOT soft-delete' line distinguishes it from the soft variants (retire_agent, deregister_agent) an agent would otherwise consider.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It conveys the precondition for invocation (confirmation must be exactly 'I UNDERSTAND') and the permanent nature of the operation, which strongly implies when it is appropriate. However, it never names the sibling alternatives (retire_agent, deregister_agent, unretire_agent) or tells the agent to prefer them when reversibility is acceptable, leaving the tool-selection decision to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

hard_delete_projectA

IRREVERSIBLY delete a project and ALL associated data: every agent, message, message recipient, file reservation, agent link, window identity, message summary, sibling suggestion, product link, and the entire on-disk project archive directory. This is NOT soft-delete — data is permanently destroyed and cannot be recovered. Requires the confirmation parameter to be exactly 'I UNDERSTAND'.

ParametersJSON Schema
NameRequiredDescriptionDefault
project_keyYes
confirmationYes
registration_tokenNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does well: it discloses that destruction is permanent and unrecoverable, enumerates every category of data destroyed including the on-disk archive directory, and specifies the required confirmation string. It stops short of stating authorization prerequisites or any rate/audit behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the destructive scope in the first clause and never buries the key constraint. The long enumeration of destroyed entities is justified because each item scopes the blast radius, though the sentence runs long and could be trimmed slightly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a high-risk mutation with no annotations, the description supplies what an agent most needs: irreversibility, blast radius, and the confirmation contract. An output schema exists so return values need not be described, though the unstated role of registration_token is a minor remaining gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It fully explains the 'confirmation' parameter (must be exactly 'I UNDERSTAND'), but leaves 'project_key' and the optional 'registration_token' (default null) unexplained — the latter is entirely opaque and its relationship to the operation is never stated.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Opening words give a specific destructive verb ('IRREVERSIBLY delete') plus the exact resource ('a project and ALL associated data'), and the enumerated cascade list makes clear it is broader than archive_project or hard_delete_agent. An agent can distinguish it from every sibling without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description draws a clear contrast with soft-delete ('This is NOT soft-delete'), which implicitly routes the agent to archive_project for reversible removal. It does not, however, name that sibling or state the exact conditions under which permanent deletion is appropriate, so the guidance is clear but not fully explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

health_checkC

Return basic readiness information for the Agent Mail server.

ParametersJSON Schema
NameRequiredDescriptionDefault
formatNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It implies a non-destructive read, but says nothing about authentication requirements, whether it hits external dependencies, rate limits, or what 'basic' excludes from the result. For a tool with zero annotation coverage this is a meaningful gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler or redundancy. Every word earns its place, though the brevity comes at the cost of the information other dimensions penalize.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained. However, with no annotations, an undocumented parameter, and no usage context, the description is only minimally complete for a tool whose whole value is diagnostic clarity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the single 'format' parameter is typed only as string-or-null with no enumerated values. The description never mentions the parameter or its accepted formats, so it adds no meaning beyond the bare schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Return) and resource (basic readiness information for the Agent Mail server), which is clearly distinguishable from every sibling tool that deals with agents, messages, or reservations. It does not, however, define what 'readiness' actually covers (process liveness vs. dependency checks), so the purpose is clear in category but not in detail.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no statement of when to call this versus alternatives, no prerequisite or auth context, and no indication of whether it should be used as a preflight step. The agent can only infer usage from the name 'health_check'.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

install_precommit_guardA

Install the Agent Mail pre-commit guard into a git repository.

The guard blocks a commit that stages files another agent holds an exclusive file reservation on (set AGENT_MAIL_GUARD_MODE=warn to only warn). Commits must run with AGENT_NAME set to the committing agent. Any existing pre-commit hook is kept and still runs. Does nothing when worktree features are disabled (WORKTREES_ENABLED=0).

Parameters

project_key : str Project whose file reservations the guard enforces. code_repo_path : str Path to the git repository to install the hook into. format : str, optional Output format.

Returns

dict {"hook": "<path to the installed pre-commit hook>"}; empty string when skipped.

ParametersJSON Schema
NameRequiredDescriptionDefault
formatNo
project_keyYes
code_repo_pathYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden and does so: it discloses what the guard blocks, the AGENT_MAIL_GUARD_MODE=warn escape hatch, the AGENT_NAME requirement for commits, that an existing pre-commit hook is preserved, and the disabled-worktree no-op condition. These are exactly the non-obvious side effects an agent needs before mutating a repo.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the action in sentence one, then behavior, then parameters and returns in scannable sections. The params/returns blocks are duplicated effort against the schema, but with 0% schema coverage they are load-bearing rather than padding.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

A mutation tool with no annotations is fully covered: scope, side effects, env-var preconditions, no-op condition, and the return shape (hook path, empty string when skipped). Nothing an agent needs to invoke it safely is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, and it documents all three parameters via the numpydoc block: project_key (whose reservations are enforced), code_repo_path (target repo), and format. The format entry ('Output format') adds little beyond the name, keeping this short of a 5.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Install the Agent Mail pre-commit guard into a git repository') and implicitly distinguishes itself from the sibling uninstall_precommit_guard. An agent knows exactly what artifact is being created and where.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear operating context: the guard applies to commits staging files held under exclusive reservation, and it is a no-op when WORKTREES_ENABLED=0. It does not explicitly name the removal counterpart or spell out prerequisites for installation, so it stops short of full when/when-not routing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_contactsC

List contact links for an agent in a project.

ParametersJSON Schema
NameRequiredDescriptionDefault
formatNo
agent_nameYes
project_keyYes
registration_tokenNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. 'List' implies a read, but nothing is said about permissions, what a 'contact link' represents, or what registration_token does. For a tool with zero annotation coverage this is a significant gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single efficient sentence with the object and scope front-loaded and no filler. It is appropriately sized, though brevity here borders on under-specification.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be described, but the definition still omits when to use it, what a contact link is, and the meaning of two undocumented optional parameters. For a 4-parameter tool with no annotations, more is required.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across 4 parameters, and the description mentions none of them explicitly. 'project' and 'agent' loosely map to project_key and agent_name, but the nullable format and registration_token parameters get no explanation anywhere, leaving them opaque.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('List') and resource ('contact links') with scope ('for an agent in a project'), so the agent knows exactly what comes back. It does not, however, distinguish itself from contact-related siblings such as request_contact, respond_contact, or whois.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance and no mention of alternatives. With siblings like request_contact, respond_contact, and set_contact_policy in the same family, the agent has no signal about which contact tool to pick. Usage is only implied by the name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_window_identitiesC

List active window identities for a project.

Returns all non-expired window identities with their display names, last activity timestamps, and age.

Parameters

project_key : str Project identifier.

Returns

dict { identities: [{ id, window_uuid, display_name, created_ts, last_active_ts, expires_ts }] }

ParametersJSON Schema
NameRequiredDescriptionDefault
formatNo
project_keyYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations present, the description carries the full burden. It does disclose the filtering behavior ('all non-expired window identities') and the returned fields, which is genuine behavioral content for a read-only listing, but it says nothing about ordering, pagination, result-size limits, or permissions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The prose is front-loaded and short, but the tool string re-documents the 'Parameters' section that the input schema already covers and a 'Returns' block even though an output schema exists, so a meaningful fraction of the text is duplicated structured data.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple read-only list tool with an output schema, the definition covers the essentials of scope and return fields. It still omits ordering, pagination, and whether 'format' affects the output shape, which are the remaining gaps an agent would face before calling it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% for two parameters. The description only glosses project_key as 'Project identifier', which adds essentially nothing beyond the name, and the 'format' parameter is never mentioned anywhere, leaving it entirely unexplained across schema and prose.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The first sentence gives a specific verb and resource ('List active window identities for a project') and the second narrows scope to non-expired entries with named fields. It's immediately understandable, but it never distinguishes itself from lookalike siblings such as whois or create_agent_identity, so an agent must infer which identity-related tool applies.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance: no statement of which situations call for this listing versus whois or the agent-registry tools, and no prerequisites or exclusions. The only usage hint is implicit in the word 'active' (non-expired).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

macro_contact_handshakeC

Request contact permissions and optionally auto-approve plus send a welcome message.

ParametersJSON Schema
NameRequiredDescriptionDefault
modelNo
formatNo
reasonNo
targetNo
programNo
to_agentNo
requesterNo
thread_idNo
agent_nameNo
to_projectNo
auto_acceptNo
project_keyYes
ttl_secondsNo
welcome_bodyNo
welcome_subjectNo
task_descriptionNo
register_if_missingNo
target_registration_tokenNo
requester_registration_tokenNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description must carry the full behavioral burden. It discloses the main actions (request, auto-approve, welcome message) but omits critical side effects: what auto-approval entails, how registration works, token requirements, TTL behavior, and whether the operation is reversible.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single concise sentence, but for a 19-parameter macro it is under-specified rather than appropriately concise. It fails to provide the structural detail needed to navigate the tool's complexity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the high complexity (19 parameters, no annotations, no parameter descriptions), the description is grossly incomplete. It lacks usage guidance, parameter meanings, and behavioral details, leaving the agent unable to reliably invoke the tool despite the presence of an output schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and there are 19 parameters. The description only loosely hints at auto_accept and welcome_subject/welcome_body, leaving the vast majority of parameters (e.g., ttl_seconds, register_if_missing, target_registration_token) completely undocumented for the agent.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific action: request contact permissions, optionally auto-approve, and send a welcome message. This clearly identifies the macro's combined purpose. However, it does not differentiate this tool from siblings like request_contact or respond_contact, leaving the agent to infer why the macro is preferable.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit guidance on when to use this tool versus alternatives such as request_contact or respond_contact. The description implies a handshake flow but does not state conditions, prerequisites, or exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

macro_file_reservation_cycleC

Reserve a set of file paths and optionally release them at the end of the workflow.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathsYes
formatNo
reasonNomacro-file_reservation
exclusiveNo
agent_nameYes
project_keyYes
ttl_secondsNo
auto_releaseNo
registration_tokenNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden and largely fails. It does not disclose locking exclusivity, TTL/expiry behavior, auto-release semantics, conflict handling, or the need for a registration token — all of which are critical for a mutation tool that reserves shared resources.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The single sentence is front-loaded and waste-free, but for a 9-parameter macro it is under-specified rather than genuinely concise, leaving the agent without the operational detail it needs.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained. But with no annotations, 0% parameter coverage, and a mutation that reserves shared resources, the description is far too thin to let an agent call the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across 9 parameters, so the description must compensate but only gestures at 'paths' and 'release'. Critical parameters like exclusive (default true), ttl_seconds, auto_release, and registration_token receive no explanation in either the schema or the description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb (reserve) and resource (file paths) and adds the cycle concept ('optionally release them at the end of the workflow'), so an agent can grasp the purpose. However, it never distinguishes this macro from the closely related siblings file_reservation_paths and release_file_reservations, which is the differentiator a 5 would require.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit when-to-use guidance and no comparison against alternatives. The phrase 'end of the workflow' faintly implies this is a combined reserve-then-release macro, but the agent is left to infer whether to call this or the individual reservation/release tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

macro_prepare_threadC

Macro helper that aligns an agent with an existing thread by ensuring registration, summarising the thread, and fetching recent inbox context.

ParametersJSON Schema
NameRequiredDescriptionDefault
modelYes
formatNo
programYes
llm_modeNo
llm_modelNo
thread_idYes
agent_nameNo
inbox_limitNo
project_keyYes
include_examplesNo
task_descriptionNo
registration_tokenNo
register_if_missingNo
include_inbox_bodiesNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries full burden, and it does disclose the composite side effects (may register the agent, summarises the thread, fetches inbox context), which is genuinely useful for a macro whose internal steps are otherwise opaque. However it omits permission/token requirements, what registration does if it fires, and any mutation/reversibility context an agent would want before invoking.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One front-loaded sentence that enumerates the three internal steps without filler, well sized for a macro. It could be marginally tightened but nothing is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 14-parameter macro with an output schema, no annotations, and 0% schema coverage, the description is too thin: most parameters are undocumented anywhere, there is no permission/auth context, and no indication of which optional params matter. The output schema covers return values, but the input side is materially under-specified.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across 14 parameters, so the description must compensate and largely does not. It only loosely gestures at a few params (registration -> register_if_missing/registration_token, summarising -> thread_id/format, inbox -> inbox_limit/include_inbox_bodies) while leaving model, program, llm_mode, llm_model, task_description, agent_name, and project_key entirely unexplained.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific composite verb+resource: it 'aligns an agent with an existing thread' by ensuring registration, summarising the thread, and fetching inbox context. This tells an agent exactly what the macro orchestrates, though it does not distinguish itself from siblings like summarize_thread, fetch_inbox, or macro_start_session that perform overlapping steps.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'existing thread' implies the precondition (a thread must already exist, as opposed to macro_start_session which presumably creates one), but there is no explicit when-to-use/when-not statement, no guidance on choosing this over summarize_thread + fetch_inbox + register_agent individually, and no prerequisite/ordering notes for a 14-parameter macro.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

macro_start_sessionB

Macro helper that boots a project session: ensure project, register agent, optionally file_reservation paths, and fetch the latest inbox snapshot.

ParametersJSON Schema
NameRequiredDescriptionDefault
modelYes
formatNo
programYes
human_keyYes
agent_nameNo
inbox_limitNo
task_descriptionNo
registration_tokenNo
file_reservation_pathsNo
file_reservation_reasonNomacro-session
file_reservation_ttl_secondsNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden. It does disclose the composite side effects (project creation, agent registration) and marks file reservations as 'optional', which is useful. It says nothing about failure/partial-rollback semantics, idempotency, or whether repeated calls re-register the agent, which matters for a multi-step mutating macro.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single efficient sentence with the macro framing front-loaded and the sub-steps following in execution order. No padding, though it is dense enough that the enumerated steps blur together slightly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need no explanation. But for an 11-parameter, multi-effect macro with zero annotation coverage and zero schema descriptions, the one-sentence description leaves meaningful gaps around required credentials and parameter meaning.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across 11 parameters, so the description must compensate. It only conceptually covers the reservation paths and the inbox snapshot; the required keys (human_key, program, model), plus registration_token, format, and task_description are left entirely unexplained. An agent must guess the meaning and format of most inputs.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('boots a project session') and enumerates the composite steps it performs: ensure project, register agent, optional file reservations, inbox snapshot. This makes it distinguishable from the single-purpose siblings (ensure_project, register_agent, fetch_inbox), though it never explicitly names them as alternatives.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'Boots a project session' implies the bootstrap use case, and the macro framing suggests it replaces several calls at once. However, there is no explicit when-to-use vs. when-not, no mention of prerequisites (e.g. registration_token requirements), and no guidance on when to prefer the atomic siblings instead.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

mark_message_readA

Mark a specific message as read for the given agent.

Notes

  • Read receipts are per-recipient; this only affects the specified agent.

  • This does not send an acknowledgement; use acknowledge_message for that.

  • Safe to call multiple times; later calls return the original timestamp.

Idempotency

  • If mark_message_read has already been called earlier for the same (agent, message), the original timestamp is returned and no error is raised.

Returns

dict { message_id, read: bool, read_at: iso8601 | null }

Example

{"jsonrpc":"2.0","id":"8","method":"tools/call","params":{"name":"mark_message_read","arguments":{
  "project_key":"/abs/path/backend","agent_name":"BlueLake","message_id":1234
}}}
ParametersJSON Schema
NameRequiredDescriptionDefault
formatNo
agent_nameYes
message_idYes
project_keyYes
registration_tokenNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.3/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden and does so: it discloses per-recipient semantics, that no acknowledgement is emitted, idempotent re-call behavior returning the original timestamp, and the exact return shape. These are exactly the traits an agent needs before invoking a mutation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the purpose, which is good, but the Idempotency section restates the Notes bullet ('Safe to call multiple times; later calls return the original timestamp') almost verbatim. The Returns section is also partly redundant given a declared output schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a message-state mutation the description covers behavior, the alternative tool, idempotency and return fields, and an output schema exists so return values need no further elaboration. The only real gap is the unexplained 'format' and 'registration_token' parameters.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across 5 parameters, so the description must compensate and largely does not. The worked example incidentally shows project_key, agent_name and message_id, but 'format' and 'registration_token' are never explained in either place, leaving required vs optional inputs ambiguous.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Mark ... as read') plus resource and scope ('a specific message ... for the given agent'), and explicitly distinguishes itself from the sibling acknowledge_message. An agent can tell exactly what this does without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly names the alternative tool and the condition that selects it: 'This does not send an acknowledgement; use acknowledge_message for that.' It also clarifies the per-recipient scoping so the agent knows which agent the read state applies to.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

purge_old_messagesC

Delete messages older than the configured retention period. Defaults to retention_max_age_days from config (180 days). Returns count of messages purged.

ParametersJSON Schema
NameRequiredDescriptionDefault
dry_runNo
project_keyYes
max_age_daysNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries full behavioral burden. It discloses the default retention source (retention_max_age_days, 180 days) and that it returns a purged count, but omits critical safety facts for a destructive mutation: that dry_run defaults to true, whether the delete is reversible, and what authorization it requires.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences with the core action front-loaded and no filler. Every clause carries information, though the return-count sentence is somewhat redundant given an output schema exists.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values needn't be restated. However, for a destructive purge tool with zero annotations and 0% param coverage, the description is under-complete: it avoids the dry_run safety default and irreversibility, which an agent needs before calling.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate for all three params. It partially explains max_age_days via the retention default, but says nothing about project_key (the required param) or the pivotal dry_run flag, leaving most parameter semantics undocumented.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb + resource ('Delete messages older than the configured retention period'), which is clearly distinguishable from sibling read tools like search_messages or fetch_inbox. It does not explicitly name a sibling it competes with, but the destructive retention-scoped scope is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit when-to-use guidance, no prerequisites (e.g., permissions), and no mention of alternatives such as archiving vs purging. The description only implies the retention-cleanup context without stating conditions for invoking it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

register_agentA

Create or update an agent identity within a project and persist its profile to Git.

When to use

  • At the start of a coding session by any automated agent.

  • To update an existing agent's program/model/task metadata and bump last_active.

Semantics

  • If name is omitted, a random adjective+noun name is auto-generated.

  • Reusing the same name updates the profile (program/model/task) and refreshes last_active_ts.

  • A profile.json file is written under agents/<Name>/ in the project archive.

Agent Identity

Two naming modes are supported:

  1. Explicit identity — pass a stable ID like alpha-one, cc-0, or worker_42. Must match [A-Za-z0-9][A-Za-z0-9._-]{0,127}. Useful for swarm workflows where agents are relaunched onto the same identity.

  2. Auto-generated — omit name to get a random adjective+noun identity like GreenLake or BlueDog (RECOMMENDED for ad-hoc use).

Invalid examples: "BackendHarmonizer", "DatabaseMigrator" (descriptive role names are rejected in strict mode).

Parameters

project_key : str The same human key you passed to ensure_project (or equivalent identifier). program : str The agent program (e.g., "codex-cli", "claude-code"). model : str The underlying model (e.g., "gpt5-codex", "opus-4.1"). name : Optional[str] A valid explicit identity (e.g., "alpha-one", "cc-0") or adjective+noun combination (e.g., "BlueLake"). If omitted, a random valid name is auto-generated (RECOMMENDED). Names are unique per project; passing the same name updates the profile. task_description : str Short description of current focus (shows up in directory listings).

Returns

dict { id, name, program, model, task_description, inception_ts, last_active_ts, project_id }

Examples

Register with auto-generated name (RECOMMENDED):

{"jsonrpc":"2.0","id":"3","method":"tools/call","params":{"name":"register_agent","arguments":{
  "project_key":"/data/projects/backend","program":"codex-cli","model":"gpt5-codex","task_description":"Auth refactor"
}}}

Register with explicit valid name:

{"jsonrpc":"2.0","id":"4","method":"tools/call","params":{"name":"register_agent","arguments":{
  "project_key":"/data/projects/backend","program":"claude-code","model":"opus-4.1","name":"BlueLake","task_description":"Navbar redesign"
}}}

Pitfalls

  • Names MUST match the adjective+noun format or an error will be raised

  • Names are case-insensitive unique. If you see "already in use", pick another or omit name.

  • Use the same project_key consistently across cooperating agents.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameNo
modelYes
formatNo
programYes
project_keyYes
task_descriptionNo
attachments_policyNoauto
registration_tokenNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.2/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden and does so well: it discloses the on-disk artifact (agents/<Name>/profile.json), the case-insensitive uniqueness rule, last_active_ts refresh on reuse, the name regex, and strict-mode rejection rules.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Well-organized with headers, but heavily padded: two near-duplicate JSON-RPC examples, repeated 'RECOMMENDED' notes, and a Pitfalls bullet that contradicts the earlier identity section by claiming names MUST be adjective+noun when explicit IDs were just declared valid.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Rich enough for a mutation tool with no annotations, covering naming, update semantics, and side effects. The unaddressed params and redundant Returns block (an output schema exists) keep it short of complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and there are 8 parameters, but the description documents only 5 (project_key, program, model, name, task_description). format, attachments_policy, and registration_token are left undocumented in both prose and schema, so the agent must guess their meaning.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Create or update an agent identity within a project') plus the persistence side effect ('persist its profile to Git'). An agent can distinguish this from deregister_agent/retire_agent/hard_delete_agent, which are the destructive siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Has an explicit 'When to use' section (session start, metadata update) and a Semantics section for the create-vs-update branch. It does not, however, address overlap with the sibling create_agent_identity, leaving ambiguity about which identity-creation path to pick.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

release_file_reservationsA

Release active file reservations held by an agent.

Behavior

  • If both paths and file_reservation_ids are omitted, all active reservations for the agent are released

  • Otherwise, restricts release to matching ids and/or path patterns

  • JSON artifacts stay in Git for audit; DB records get released_ts

Returns

dict { released: int, released_at: iso8601 }

Idempotency

  • Safe to call repeatedly. Releasing an already-released (or non-existent) reservation is a no-op.

Examples

Release all active reservations for agent:

{"jsonrpc":"2.0","id":"13","method":"tools/call","params":{"name":"release_file_reservations","arguments":{
  "project_key":"/abs/path/backend","agent_name":"GreenCastle"
}}}

Release by ids:

{"jsonrpc":"2.0","id":"14","method":"tools/call","params":{"name":"release_file_reservations","arguments":{
  "project_key":"/abs/path/backend","agent_name":"GreenCastle","file_reservation_ids":[101,102]
}}}
ParametersJSON Schema
NameRequiredDescriptionDefault
pathsNo
formatNo
agent_nameYes
project_keyYes
registration_tokenNo
file_reservation_idsNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden and does so well: it discloses idempotency ('safe to call repeatedly', already-released reservations are a no-op), the side effect on Git artifacts vs DB records ('JSON artifacts stay in Git for audit; DB records get released_ts'), and the return shape. This is exactly the behavioral context annotations would otherwise supply.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with a one-line purpose followed by clearly headed Behavior/Returns/Idempotency/Examples sections, so an agent can scan to the relevant part. The two JSON-RPC examples are verbose but genuinely illustrative of the omitted-vs-ids invocation patterns.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so the Returns section is largely redundant but harmless. Behavior, idempotency and usage examples are all present; the only real omissions are the undocumented 'format' and 'registration_token' parameters and any note on permissions/auth requirements for release.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across 6 parameters, so the description must compensate. It meaningfully explains paths, file_reservation_ids, project_key and agent_name (via examples and the scoping rules), but leaves 'format' and 'registration_token' completely unexplained, which is a notable gap given auth-related parameters are involved.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Release active file reservations held by an agent'), which is clearly distinguishable from reservation-creating siblings like file_reservation_paths and renew_file_reservations. It does not, however, differentiate itself from the very close sibling force_release_file_reservation, leaving a real ambiguity unresolved.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The Behavior section explains the parameter-driven scoping rules (omit both = release all; supply either = restrict), which implies usage. But it never states when to prefer this tool over force_release_file_reservation or what the difference is, so the agent must infer the choice between near-identical siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

rename_windowB

Update the display name of a window identity.

Parameters

project_key : str Project identifier. window_uuid : str The UUID of the window identity to rename. new_display_name : str New display name (must be a valid adjective+noun agent name).

Returns

dict Updated window identity record.

ParametersJSON Schema
NameRequiredDescriptionDefault
formatNo
project_keyYes
window_uuidYes
new_display_nameYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are supplied, so the description must carry the behavioral burden. It does disclose one important constraint — the new display name must be a valid adjective+noun agent name — and confirms it returns the updated record, but says nothing about permissions, idempotency, reversibility, or what triggers the naming validation to fail.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The summary sentence is front-loaded and clear, but the numpy-style Parameters and Returns blocks largely restate the input schema, which is redundant given the structured fields already present. Reasonably sized but not economical.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so the Returns section is redundant, and the mutation's side effects, authorization needs, and the undocumented 'format' parameter remain unexplained. Adequate for a simple rename, but incomplete for a state-changing tool with zero annotations.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate, and it documents all three required parameters with meaning plus a validation rule on new_display_name that the schema does not express. It misses the optional 'format' parameter entirely, but does well on the parameters that matter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Update the display name of a window identity'), which is precise enough to distinguish it from siblings like create_agent_identity or list_window_identities. It stops short of explicitly naming an alternative, but the purpose itself is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to rename versus creating a new identity, expiring a window, or any prerequisite workflow. With a dense sibling set including create_agent_identity and expire_window, the absence of routing guidance is a real gap.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

renew_file_reservationsB

Extend expiry for active file reservations held by an agent without reissuing them.

Parameters

project_key : str Project slug or human key. agent_name : str Agent identity who owns the reservations. extend_seconds : int Seconds to extend from the later of now or current expiry (min 60s). paths : Optional[list[str]] Restrict renewals to matching path patterns. file_reservation_ids : Optional[list[int]] Restrict renewals to matching reservation ids.

Returns

dict { renewed: int, file_reservations: [{id, path_pattern, old_expires_ts, new_expires_ts}] }

ParametersJSON Schema
NameRequiredDescriptionDefault
pathsNo
formatNo
agent_nameYes
project_keyYes
extend_secondsNo
registration_tokenNo
file_reservation_idsNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden, and it does disclose useful traits: extension is measured from the later of now or current expiry, and there is a 60s minimum. It omits what happens with expired or non-held reservations, permission/auth requirements, and the fact that a registration_token option exists at all.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The one-line purpose statement is front-loaded and the structured Parameters/Returns sections are justified given zero schema-level docs. The Returns block is slightly redundant since an output schema exists, but overall there is little waste.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not have been restated, and the description still frames the renewal semantics adequately for a stateful mutation. The undocumented 'registration_token' and 'format' parameters and undefined behavior for already-expired reservations are the only real omissions.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, and it documents 5 of 7 parameters with real semantics (project slug, owning agent, extension origin with 60s minimum, path and id filters). It leaves 'format' and 'registration_token' completely unexplained, which is the remaining gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The opening sentence states a specific verb (extend) and resource (expiry of active file reservations) and adds a qualifier ('without reissuing them') that separates it from file_reservation_paths. However, it never names a sibling tool, so the differentiation from release_file_reservations and force_release_file_reservation is left implicit.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description never says when to renew versus release and re-acquire, nor does it note prerequisites or exclusions. 'Extend expiry ... without reissuing' implies a use case but gives no explicit routing guidance against the several sibling reservation tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

reply_messageA

Reply to an existing message, preserving or establishing a thread.

Behavior

  • Inherits original importance and ack_required flags

  • thread_id is taken from the original message if present; otherwise, the original id is used

  • Subject is prefixed with subject_prefix if not already present

  • Defaults to to the original sender if not explicitly provided

Parameters

project_key : str Project identifier. message_id : int The id of the message you are replying to. sender_name : str Your agent name (must be registered in the project). body_md : str Reply body in Markdown. to, cc, bcc : Optional[list[str]] Recipients by agent name. If omitted, to defaults to original sender. subject_prefix : str Prefix to apply (default "Re:"). Case-insensitive idempotent.

Do / Don't

Do:

  • Keep the subject focused; avoid topic drift within a thread.

  • Reply to the original sender unless new stakeholders are strictly required.

  • Preserve importance/ack flags from the original unless there is a clear reason to change.

  • Use CC for FYI only; BCC sparingly and with intention.

Don't:

  • Change thread_id when continuing the same discussion.

  • Escalate to many recipients; prefer targeted replies and start a new thread for new topics.

  • Attach large binaries in replies unless essential; reference prior attachments where possible.

Returns

dict Message payload including thread_id and reply_to.

Examples

Minimal reply to original sender: If the caller has not already authenticated as sender_name in this MCP session, include sender_token.

{"jsonrpc":"2.0","id":"6","method":"tools/call","params":{"name":"reply_message","arguments":{
  "project_key":"/abs/path/backend","message_id":1234,"sender_name":"BlueLake",
  "body_md":"Questions about the migration plan...","sender_token":"<registration_token>"
}}}

Reply with explicit recipients and CC:

{"jsonrpc":"2.0","id":"6c","method":"tools/call","params":{"name":"reply_message","arguments":{
  "project_key":"/abs/path/backend","message_id":1234,"sender_name":"BlueLake",
  "body_md":"Looping ops.","to":["GreenCastle"],"cc":["RedCat"],"subject_prefix":"RE:",
  "sender_token":"<registration_token>"
}}}
ParametersJSON Schema
NameRequiredDescriptionDefault
ccNo
toNo
bccNo
formatNo
body_mdYes
message_idYes
project_keyYes
sender_nameYes
sender_tokenNo
subject_prefixNoRe:

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and largely meets it: it discloses inherited importance/ack flags, thread_id derivation, subject prefixing, default recipient, and (via the example) the sender_token requirement. It doesn't state whether the original message is mutated, permission requirements, or rate limits, so a small gap remains.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The purpose is front-loaded and sectioned clearly (Behavior, Parameters, Do/Don't, Returns, Examples), so it is easy to scan. It is on the long side and the two full JSON-RPC examples are verbose relative to what they convey.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so the return-value section is surplus, but behavior, auth, and recipient defaults are all covered for a mutation tool with no annotations. The undocumented 'format' parameter is the main completeness gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate, and it documents 8 of 10 parameters (project_key, message_id, sender_name, body_md, to/cc/bcc, subject_prefix) with useful semantics like the case-insensitive idempotent prefix. The 'format' parameter is never explained and sender_token is only shown in examples, leaving two gaps.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The opening sentence gives a specific verb and resource (reply to an existing message) and adds scope ('preserving or establishing a thread') that separates it from a plain send. It never names the sibling send_message, so the distinction from that tool is left implicit rather than explicit.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The Do/Don't block is genuine routing guidance: reply to the original sender, prefer a new thread for new topics, don't change thread_id, don't escalate recipients. It tells the agent both when to reply and when to choose an alternative (start a new thread), which is exactly what a usage section should do.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

request_contactA

Request contact approval to message another agent.

Creates (or refreshes) a pending AgentLink and sends a small ack_required intro message.

Discovery

To discover available agent names, use: resource://agents/{project_key} Agent names are NOT the same as program names or user names.

Parameters

project_key : str Project slug or human key. from_agent : str Your agent name (must be registered in the project). to_agent : str Target agent name (use resource://agents/{project_key} to discover names). to_project : Optional[str] Target project if different from your project (cross-project coordination). reason : str Optional explanation for the contact request. ttl_seconds : int Time to live for the contact approval request (default: 7 days).

ParametersJSON Schema
NameRequiredDescriptionDefault
modelNo
formatNo
reasonNo
programNo
to_agentYes
from_agentYes
to_projectNo
project_keyYes
ttl_secondsNo
task_descriptionNo
registration_tokenNo
register_if_missingNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries full behavioral burden. It does disclose the core mutation (creates/refreshes an AgentLink, sends an ack_required message), which is genuinely useful. But it omits material side effects visible in the schema: register_if_missing defaults to true (implying this call may auto-register an agent) and registration_token exists — neither is mentioned, nor is any auth or rate-limit context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the action and effect, then cleanly sectioned under 'Discovery' and 'Parameters' headings. Slightly padded by restating the discovery resource both in prose and in the to_agent description, but overall well-sized and scannable.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need no explanation, and the core workflow is covered. But for a 12-parameter mutation tool with zero schema coverage, leaving six parameters — notably register_if_missing and registration_token — undocumented means an agent cannot fully predict the call's side effects, which is a meaningful completeness gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, and it documents only 6 of 12 parameters (project_key, from_agent, to_agent, to_project, reason, ttl_seconds) with useful semantics such as ttl default of 7 days and cross-project intent. It leaves model, format, program, task_description, registration_token, and register_if_missing entirely undocumented, so half the surface — including a behavior-changing flag — has no explanation anywhere.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Request contact approval to message another agent') and immediately describes the concrete effect: 'Creates (or refreshes) a pending AgentLink and sends a small ack_required intro message.' This distinguishes it from siblings like send_message, respond_contact, list_contacts, and set_contact_policy without needing to open any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear context for use (cross-project coordination, discovering target agents) and points at resource://agents/{project_key} for name discovery. However, it never names an alternative tool or states when NOT to use it — e.g. that respond_contact handles the approval side or send_message handles the post-approval messaging — so the agent must infer the workflow boundary.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

respond_contactC

Approve or deny a contact request.

ParametersJSON Schema
NameRequiredDescriptionDefault
acceptYes
formatNo
to_agentYes
from_agentYes
project_keyYes
ttl_secondsNo
from_projectNo
registration_tokenNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden yet discloses almost nothing beyond the verb. It never mentions permissions, whether the 'registration_token' parameter implies an auth requirement, what ttl_seconds controls, or how approval/denial affects the requesting agent's state. Only the implication that this is a state-changing decision is conveyed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The single sentence is front-loaded and waste-free, but for a tool with eight parameters it is under-specified rather than genuinely concise. There is no structural help (no bullets, no parameter hints) to aid invocation.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained, but for an 8-parameter mutation tool with zero schema coverage and no annotations, the description is far too thin. An agent cannot determine required inputs, auth needs, or side effects from this text.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 8 parameters and 0% schema description coverage, the description must compensate and instead says nothing about any parameter. It does not clarify that 'accept' is the approve/deny switch, nor explain the roles of ttl_seconds, registration_token, from_project, or from_agent.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a specific verb pair ('Approve or deny') and resource ('contact request'), so the operation is immediately unambiguous. It does not, however, distinguish this from siblings like request_contact or macro_contact_handshake, which is the only thing keeping it from a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no statement of when to use this tool versus alternatives such as request_contact, list_contacts, set_contact_policy, or macro_contact_handshake. The single sentence conveys the action but leaves all routing decisions to the agent.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

retire_agentA

Soft-delete an agent: mark it as retired so it stops accepting new messages while preserving message history. Retired agents are hidden from active agent lists but visible in 'all agents' views.

ParametersJSON Schema
NameRequiredDescriptionDefault
agent_nameYes
project_keyYes
registration_tokenNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are present, so the description carries the burden, and it does well: it discloses that this is a soft (non-destructive) delete, that message history is preserved, that new messages are rejected, and how retired agents appear in listings. It omits auth/permission requirements and does not explicitly state reversibility, which keeps it below 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, no filler, with the key behavioral fact (soft-delete, history preserved) front-loaded. Every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need no explanation, and the behavioral surface is reasonably covered. However, with three undocumented parameters and no annotations, the definition leaves project_key, registration_token, and the authorization context unexplained for a state-mutating tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across three parameters. The description mentions 'an agent' but adds no meaning for project_key or registration_token, and does not explain the optional registration_token's role or when to supply it. It does not compensate for the coverage gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Soft-delete an agent') and immediately clarifies the mechanism ('mark it as retired'), which cleanly distinguishes it from hard_delete_agent and deregister_agent in the sibling set. An agent can tell what this does without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies when the operation is appropriate by describing the effect (stops accepting messages, history preserved), and the existence of unretire_agent hints at reversibility, but it never names an alternative or states when to choose retire vs deregister vs hard_delete. Usage is inferable but not guided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_messagesA

Full-text search over subject and body for a project.

Tips

  • SQLite FTS5 syntax supported: phrases ("build plan"), prefix (mig*), boolean (plan AND users)

  • Results are ordered by bm25 score (best matches first)

  • Limit defaults to 20; raise for broad queries

Query examples

  • Phrase search: "build plan"

  • Prefix: migrat*

  • Boolean: plan AND users

  • Require urgent: urgent AND deployment

Parameters

project_key : str Project identifier. query : str FTS5 query string. limit : int Max results to return.

Returns

list[dict] Each entry: { id, subject, importance, ack_required, created_ts, thread_id, from }

Example

{"jsonrpc":"2.0","id":"10","method":"tools/call","params":{"name":"search_messages","arguments":{
  "project_key":"/abs/path/backend","query":""build plan" AND users", "limit": 50
}}}
ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
queryYes
formatNo
agent_nameNo
project_keyYes
registration_tokenNo

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It usefully discloses bm25 ordering, the default limit of 20, and the exact return fields, but says nothing about authentication/registration needs even though parameters like registration_token and agent_name exist and are unexplained.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose is front-loaded and the structure is scannable, but it is padded: the same query examples appear in both the Tips list and the Query examples section, and the large JSON-RPC example duplicates information already given in the Parameters block.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema the description correctly supplies a Returns section with field names, which is a real strength. However it leaves three parameters undocumented and gives no guidance on auth/registration context, so it is only partially complete for a six-parameter tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It does document project_key, query and limit in one-liners, but three of six parameters (format, agent_name, registration_token) are undocumented in both the schema and the description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb, resource and scope: full-text search over message subject and body within a project. It is clearly distinguishable from list-oriented siblings like fetch_inbox/fetch_topic and summarization tools like summarize_thread.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The Tips and Query examples sections explain how to build FTS5 queries (phrases, prefix, boolean) and advise raising limit for broad queries, but never says when to prefer this tool over fetch_inbox, fetch_topic, or summarize_thread. Usage is implied rather than routed.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

send_messageA

Send a Markdown message to one or more recipients and persist canonical and mailbox copies to Git.

Discovery

To discover available agent names for recipients, use: resource://agents/{project_key} Agent names are NOT the same as program names or user names.

What this does

  • Stores message (and recipients) in the database; updates sender's activity

  • Writes a canonical .md under messages/YYYY/MM/

  • Writes sender outbox and per-recipient inbox copies

  • Optionally converts referenced images to WebP and embeds small images inline

  • Supports explicit attachments via attachment_paths in addition to inline references

Parameters

project_key : str Project identifier (same used with ensure_project/register_agent). sender_name : str Must match an agent registered in the project. to : list[str] Primary recipients (agent names). At least one of to/cc/bcc must be non-empty. subject : str Short subject line that will be visible in inbox/outbox and search results. body_md : str GitHub-Flavored Markdown body. Image references can be file paths or data URIs. cc, bcc : Optional[list[str]] Additional recipients by name. attachment_paths : Optional[list[str]] Extra file paths to include as attachments; will be converted to WebP and stored. convert_images : Optional[bool] Overrides server default for image conversion/inlining. If None, server settings apply. Note: sender attachments_policy "inline"/"file" always forces conversion/inlining. importance : str One of {"low","normal","high","urgent"} (free form tolerated; used by filters). ack_required : bool If true, recipients should call acknowledge_message after reading. thread_id : Optional[str] If provided, message will be associated with an existing thread. broadcast : bool If true and to is empty, expand recipients to all registered agents in the project (excluding the sender). Mutually exclusive with explicit to recipients. Respects contact_policy settings and is best-effort across contact boundaries: expanded recipients that are retired, set block_all, or would need contact approval are skipped rather than blocking the send, and are reported in broadcast_skipped ([{"agent": name, "reason": ...}]). No contact request is created for them — request contact explicitly if you want them included. auto_contact_if_blocked therefore never fires for broadcast-expanded recipients; it still applies to explicitly named to/cc/bcc names. topic : Optional[str] Optional topic tag (max 64 chars). Must start with a letter or digit and may otherwise contain alphanumerics, '.', '_', or '-' — so beads_rust hierarchical IDs like br-abc.1 can be used verbatim. Stored on the message for topic-based filtering via fetch_inbox(topic=...) or fetch_topic(). auto_contact_if_blocked : Optional[bool] When True (and contact policy blocks delivery to one or more recipients), the server will attempt to resolve the block automatically:

- If the recipient is already authenticated in the **same MCP session**, run
  ``macro_contact_handshake(..., auto_accept=True)`` to approve the link in-band.
  The current send proceeds normally and the message is delivered.
- Otherwise, fall back to creating a **pending** ``request_contact`` aimed at the
  recipient. This call then **fails loud** with ``CONTACT_REQUIRED`` carrying
  ``auto_contact_requested`` in ``data``. **The message body is not queued** —
  once the recipient approves the contact (``respond_contact(..., accept=True)``),
  the sender must re-call ``send_message`` to actually deliver the payload.

Defaults to the server-wide ``MESSAGING_AUTO_HANDSHAKE_ON_BLOCK`` setting (true
unless overridden). The pending-request TTL is governed by
``CONTACT_PENDING_TTL_SECONDS`` (default 7 days, separate from the in-session
auto-approval TTL ``CONTACT_AUTO_TTL_SECONDS``).

Returns

dict { "deliveries": [ { "project": str, "payload": { ... message payload ... } } ], "count": int }

Edge cases

  • If no recipients are given, the call fails.

  • Unknown recipient names fail fast; register them first.

  • Non-absolute attachment paths are resolved relative to the project archive root.

Do / Don't

Do:

  • Keep subjects concise and specific (aim for ≤ 80 characters).

  • Use thread_id (or reply_message) to keep related discussion in a single thread.

  • Address only relevant recipients; use CC/BCC sparingly and intentionally.

  • Prefer Markdown links; attach images only when they materially aid understanding. The server auto-converts images to WebP and may inline small images depending on policy.

Don't:

  • Send large, repeated binaries—reuse prior attachments via attachment_paths when possible.

  • Change topics mid-thread—start a new thread for a new subject.

  • Broadcast to "all" agents unnecessarily—target just the agents who need to act.

Examples

  1. Simple message:

{"jsonrpc":"2.0","id":"5","method":"tools/call","params":{"name":"send_message","arguments":{
  "project_key":"/abs/path/backend","sender_name":"GreenCastle","to":["BlueLake"],
  "subject":"Plan for /api/users","body_md":"See below."
}}}
  1. Inline image (auto-convert to WebP and inline if small):

{"jsonrpc":"2.0","id":"6a","method":"tools/call","params":{"name":"send_message","arguments":{
  "project_key":"/abs/path/backend","sender_name":"GreenCastle","to":["BlueLake"],
  "subject":"Diagram","body_md":"![diagram](docs/flow.png)","convert_images":true
}}}
  1. Explicit attachments:

{"jsonrpc":"2.0","id":"6b","method":"tools/call","params":{"name":"send_message","arguments":{
  "project_key":"/abs/path/backend","sender_name":"GreenCastle","to":["BlueLake"],
  "subject":"Screenshots","body_md":"Please review.","attachment_paths":["shots/a.png","shots/b.png"]
}}}
ParametersJSON Schema
NameRequiredDescriptionDefault
ccNo
toYes
bccNo
topicNo
formatNo
body_mdYes
subjectYes
broadcastNo
thread_idNo
importanceNonormal
project_keyYes
sender_nameYes
ack_requiredNo
sender_tokenNo
convert_imagesNo
attachment_pathsNo
auto_contact_if_blockedNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden and does so: it discloses DB writes, Git canonical/outbox/inbox copies, image conversion/inlining policy, contact-policy blocking behavior, fail-loud CONTACT_REQUIRED semantics with the explicit warning that the body is not queued, broadcast skipping with reasons, and edge cases for unknown/non-absolute paths.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Well structured with front-loaded purpose and clear section headers, and most content earns its place given the tool's complexity. However, the three full JSON-RPC examples and the Returns block (already covered by the output schema) add length that could be trimmed.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 17-parameter mutation tool with no annotations and 0% schema description coverage, the description supplies everything an agent needs: parameter semantics, failure modes, policy interactions, and discovery path. The Returns section duplicates the output schema but does not create a gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, yet the description documents 15 of 17 parameters with semantics, constraints, and interactions (topic charset rules, importance enum, thread_id association, broadcast vs explicit to, auto_contact_if_blocked TTL behavior). Only format and sender_token go unexplained, which is a minor omission against otherwise thorough compensation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The opening line states a specific verb (send), resource (Markdown message), recipient scope (one or more), and the side effect (persist canonical and mailbox copies to Git). This distinguishes it clearly from siblings like reply_message and fetch_inbox.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit Do/Don't guidance covers when to thread vs start new, when to broadcast vs target, when to attach images, and when to use reply_message instead. It also names the discovery resource (resource://agents/{project_key}) and warns agent names differ from program/user names.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

set_contact_policyC

Set contact policy for an agent: open | auto | contacts_only | block_all.

ParametersJSON Schema
NameRequiredDescriptionDefault
formatNo
policyYes
agent_nameYes
project_keyYes
registration_tokenNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden and falls short. It does not say whether changing a policy is reversible, whether it affects existing contacts or pending requests, whether it requires elevated permissions (registration_token is a parameter), or how the four policies differ behaviorally.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single tight sentence with the key discriminator (the four policy values) front-loaded. It is efficient, though slightly under-specified given the tool's complexity rather than verbose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return values need not be described, but for a five-parameter mutation tool with zero annotation coverage and zero schema description coverage, the description is far too thin. Scope, permissions, and policy semantics are all missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% for all five parameters, so the description must compensate and does not. It lists the policy values (useful, since the schema declares no enum for policy), but project_key, agent_name, format, and registration_token are entirely unexplained.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (set) and resource (contact policy for an agent) and even enumerates the valid policy values, so the core action is unambiguous. It does not, however, distinguish itself from contact-related siblings such as request_contact, respond_contact, or list_contacts, leaving the boundary implicit.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description says what the tool does but never when to use it: no prerequisites, no conditions, and no mention of alternatives like request_contact or macro_contact_handshake. The agent must infer usage entirely from the name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

summarize_recentA

Summarize all recent project messages within a time window.

Fetches messages from the last since_hours hours, groups them by thread, and produces a combined project-wide summary. Results are stored in the message_summaries table for fast retrieval via fetch_summary.

Idempotent: if a summary already exists for the same time window (within 5-minute tolerance) it is returned from cache.

Parameters

project_key : str Project identifier (slug or human key). since_hours : float How far back to look (default 1 hour). llm_mode : bool Use LLM to refine the summary (default True). llm_model : str, optional Override LLM model name. max_messages : int Maximum messages to include (default 500, capped at 500). format : str, optional Output format (json or toon).

ParametersJSON Schema
NameRequiredDescriptionDefault
formatNo
llm_modeNo
llm_modelNo
project_keyYes
since_hoursNo
max_messagesNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does well: it discloses a write side effect (results stored in the message_summaries table), idempotency with a 5-minute cache tolerance, LLM usage, and thread grouping. It omits permissions/auth requirements and LLM cost implications, but the behavioral disclosure is well above average.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the core purpose in the first sentence, then layers storage/idempotency and parameters in a scannable numpy-style layout. Some default values are repeated from the schema, but given 0% schema coverage the parameter block earns its space.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need no explanation, and the description covers storage location, idempotency, and every parameter. For a 6-param, no-annotation tool it is nearly complete, missing only permission/authorization context for what is effectively a mutating, LLM-costing operation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate and largely does: it documents all six parameters with types, defaults, and meaning (llm_mode 'refine the summary', max_messages 'capped at 500', format 'json or toon'). Gaps remain — llm_model lists no valid values and project_key's format is only loosely described — so it stops short of fully compensating.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (summarize) plus resource (recent project messages) and scope (time window), and implicitly distinguishes itself from summarize_thread by producing a 'project-wide summary' rather than a thread-level one. An agent can identify the operation without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Clear context: fetch messages over a window, group by thread, and store for retrieval via fetch_summary — so it routes the agent to the sibling for reading results. However, it never explicitly states when to prefer this over summarize_thread or search_messages, leaving that comparison to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

summarize_threadA

Extract participants, key points, and action items for one or more threads.

Single-thread mode (thread_id is a single ID):

  • Returns detailed summary with optional example messages

  • Response: { thread_id, summary: {participants[], key_points[], action_items[]}, examples[] }

Multi-thread mode (thread_id is comma-separated IDs like "TKT-1,TKT-2,TKT-3"):

  • Returns aggregate digest across all threads

  • Response: { threads: [{thread_id, summary}], aggregate: {top_mentions[], key_points[], action_items[]} }

Parameters

project_key : str Project identifier. thread_id : str Single thread ID for detailed summary, OR comma-separated IDs for aggregate digest. include_examples : bool If true (single-thread mode only), include up to 3 sample messages. llm_mode : bool If true and LLM is enabled, refine the summary with AI. llm_model : Optional[str] Override model name for the LLM call. per_thread_limit : int Max messages to consider per thread (multi-thread mode).

Examples

Single thread:

{"thread_id": "TKT-123", "include_examples": true}

Multiple threads:

{"thread_id": "TKT-1,TKT-2,TKT-3"}
ParametersJSON Schema
NameRequiredDescriptionDefault
formatNo
llm_modeNo
llm_modelNo
thread_idYes
agent_nameNo
project_keyYes
include_examplesNo
per_thread_limitNo
registration_tokenNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full load and does well: it discloses the exact response shapes per mode, that include_examples is single-thread only and caps at 3 samples, that llm_mode refines via AI only when enabled, and that per_thread_limit bounds multi-thread work. It omits any statement of read-only nature, auth/permission needs, or side effects.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the one-line purpose, then organized into modes, response shapes, parameters, and examples. It is longer than average, and the response-shape blocks and examples partially overlap, but the structure is scannable and each section adds value for a 9-param tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex, unannotated tool with 9 params, the description covers modes, most parameters, and request examples. An output schema exists so return values need not be restated, but the three undocumented parameters represent a genuine completeness gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, and it documents 6 of 9 parameters (project_key, thread_id, include_examples, llm_mode, llm_model, per_thread_limit) with real semantic meaning plus usage examples. It leaves format, agent_name, and registration_token entirely undocumented.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (extract) and resource set (participants, key points, action items) over one or more threads, and clearly separates single-thread from multi-thread behavior. The purpose is unmistakable, though it never names or differentiates itself from siblings like summarize_recent or fetch_summary.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It clearly explains the two operating modes and how thread_id format selects between them, which implies usage. However, it gives no guidance on when to choose this tool over alternatives such as summarize_recent or fetch_summary, and states no exclusions or prerequisites.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

sweep_stale_agentsB

Retire abandoned agents in the caller's project using the server's conservative inactivity heuristic. The caller is never retired, the threshold has a 60-second floor, and active file reservations block retirement by default.

ParametersJSON Schema
NameRequiredDescriptionDefault
agent_nameYes
project_keyYes
threshold_secondsNo
registration_tokenNo
require_no_active_reservationsNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With zero annotations the description carries the full behavioral burden and does so well: it discloses that the caller is never retired, that the threshold has a 60-second floor, and that active reservations block retirement by default. It omits whether retirement is reversible (unretire_agent exists as a sibling) and what authorization registration_token implies.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single dense sentence, front-loaded with the action and packed with behavioral constraints and no filler. It earns nearly every clause, though the run-on structure slightly reduces scannability.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return shape need not be described, and the behavioral constraints are covered. However, for a destructive batch operation with no annotations and 0% schema coverage, leaving three parameters (notably registration_token) and the reversibility question unanswered is a meaningful gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% and the description only partially compensates: it clarifies threshold_seconds (60-second floor) and require_no_active_reservations (default blocking), but agent_name, project_key, and especially registration_token are left unexplained. Three of five parameters carry no semantics anywhere.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb (retire), resource (abandoned agents), scope (caller's project), and mechanism (conservative inactivity heuristic), which clearly sets it apart from single-target siblings like retire_agent or deregister_agent. It never explicitly names an alternative, but the batch/heuristic framing is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Conditions are given (60-second threshold floor, reservations blocking retirement, caller exempt), which implies when the tool applies, but there is no explicit when-to-use vs an alternative such as retire_agent for a single agent. The agent must infer that this is the bulk-sweep option.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

unarchive_projectC

Restore an archived project back to active status.

ParametersJSON Schema
NameRequiredDescriptionDefault
project_keyYes
registration_tokenNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It states the state transition, but says nothing about required permissions, whether the call is idempotent, or what happens if the project is not archived.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single efficient sentence with the key outcome (active status) front-loaded and no wasted words. Its brevity is appropriate, though it comes at the cost of the missing behavioral and parameter detail noted above.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return values need not be described, but for a mutation tool with zero annotation coverage and zero parameter documentation, the definition omits too much for an agent to call it confidently.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the description never mentions either parameter. In particular registration_token hints at an auth/registration requirement that is completely unexplained, and project_key's format is left undefined in both places.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb (restore/unarchive) and resource (project) and states the resulting state change to active status. It is clearly the inverse of the sibling archive_project, though it does not name that sibling explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this versus alternatives, no prerequisites, and no mention of what condition must hold (e.g. project must currently be archived). The inverse relationship with archive_project is left for the agent to infer.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

uninstall_precommit_guardA

Remove the Agent Mail commit guards from a git repository.

Removes Agent Mail's own pre-commit and pre-push guard plugins (and a legacy single-file Agent Mail hook); other hooks in the repository are left in place.

Parameters

code_repo_path : str Path to the git repository to remove the guard from. format : str, optional Output format.

Returns

dict {"removed": true} when a guard was removed, false when none was installed.

ParametersJSON Schema
NameRequiredDescriptionDefault
formatNo
code_repo_pathYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden and does reasonably well: it discloses exactly what is destroyed and what is left intact, and explains the return contract ({"removed": true} when a guard was removed, false when none was installed), which implies idempotent behavior. It does not address permissions, concurrency, or whether git config is touched, keeping it short of a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The purpose is front-loaded in the first sentence and the numpy-style Parameters/Returns sections are well organized. The structure is slightly heavy for a two-parameter tool and the Returns block partly duplicates the output schema, but little is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no annotations and 0% schema coverage, the description supplies the missing behavioral context (destruction scope, side-effect boundary, return semantics). An output schema already exists, so the explicit Returns section is redundant rather than necessary, and the weak 'format' documentation leaves a small gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It does document both parameters, giving code_repo_path real meaning ('Path to the git repository to remove the guard from'), but 'format : str, optional — Output format' is uninformative, offering no allowed values or effect on output. Partial compensation yields a mid score.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb and resource ('Remove the Agent Mail commit guards from a git repository') and precisely scopes what is removed ('pre-commit and pre-push guard plugins (and a legacy single-file hook)'). It is easily distinguished from its inverse sibling, install_precommit_guard, and even clarifies what is NOT touched (other hooks remain).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied by the purpose: call it to strip Agent Mail guards from a repo. However, the description gives no explicit when-to-use guidance, no prerequisites (e.g., whether the guard must be present), and never names install_precommit_guard as the inverse operation. The scope note about other hooks is a boundary clarification rather than usage routing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

unretire_agentB

Restore a retired agent back to active status. The agent will resume accepting new messages.

ParametersJSON Schema
NameRequiredDescriptionDefault
agent_nameYes
project_keyYes
registration_tokenNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It does disclose the post-condition ('resume accepting new messages'), which is genuinely useful, but it omits authorization requirements, whether the operation is reversible, and how the registration_token affects the call.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences with the state transition front-loaded and the effect second. No filler or redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be described, and the state transition is stated clearly. However, with no annotations and 0% parameter documentation, an agent lacks the auth/permission and token semantics needed to invoke it confidently.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across all three parameters, and the description adds no parameter meaning. The purpose and requirement status of registration_token (auth? optional override?) are left entirely unexplained, and the description does not compensate for the coverage gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (restore) and resource (agent) with a clear target state: 'Restore a retired agent back to active status.' The inverse relationship to the sibling retire_agent is implied by 'retired' and 'active status,' though no sibling is named explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance, no alternatives, and no prerequisites (e.g., authentication or who may unretire). The agent must infer the trigger condition from the state wording alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

whoisA

Return enriched profile details for an agent, optionally including recent archive commits.

Discovery

To discover available agent names, use: resource://agents/{project_key} Agent names are NOT the same as program names or user names.

Parameters

project_key : str Project slug or human key. agent_name : str Agent name to look up (use resource://agents/{project_key} to discover names). include_recent_commits : bool If true, include latest commits touching the project archive authored by the configured git author. commit_limit : int Maximum number of recent commits to include.

Returns

dict Agent profile augmented with { recent_commits: [{hexsha, summary, authored_ts}] } when requested.

ParametersJSON Schema
NameRequiredDescriptionDefault
formatNo
agent_nameYes
project_keyYes
commit_limitNo
registration_tokenNo
include_recent_commitsNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does disclose the return shape (a dict augmented with recent_commits when requested) and the conditional behavior of include_recent_commits. It omits anything about read-only/idempotency, permissions, or auth implications (e.g., what registration_token does), leaving meaningful behavioral gaps.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The definition is front-loaded with the core purpose and organized into Discovery/Parameters/Returns sections that each add value. It is somewhat long, and the Returns section partially duplicates the output schema, but nothing is egregiously wasteful.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is a read-style profile lookup with an output schema present, so return values need not be spelled out (and are anyway). Discovery guidance and commit-option behavior make it largely self-sufficient; the residual gap is the undocumented format and registration_token parameters.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, and it does add real meaning for the core parameters: project_key, agent_name (with a discovery pointer), include_recent_commits, and commit_limit semantics are all explained. It still leaves two parameters (format and registration_token) undocumented, so it is not fully complete.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The opening sentence states a specific verb+resource ('Return enriched profile details for an agent, optionally including recent archive commits'), and the Discovery section explicitly warns that agent names are NOT program or user names, disambiguating it from siblings. An agent can tell this apart from fetch_summary or list_window_identities without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives clear operational context for discovery via resource://agents/{project_key} and clarifies the identity-space constraint. However, it never states when to prefer this over adjacent siblings like fetch_summary or list_window_identities, so it lacks explicit exclusions/alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 41 tool updatesv0.3.4
    • First observedacknowledge_message
    • First observedarchive_project
    • First observedcreate_agent_identity
    • First observedderegister_agent
    • First observedensure_project
    • First observedexpire_window
    • First observedfetch_inbox
    • First observedfetch_summary
    • First observedfetch_topic
    • First observedfile_reservation_paths
    • First observedforce_release_file_reservation
    • First observedhard_delete_agent
    • First observedhard_delete_project
    • First observedhealth_check
    • First observedinstall_precommit_guard
    • First observedlist_contacts
    • First observedlist_window_identities
    • First observedmacro_contact_handshake
    • First observedmacro_file_reservation_cycle
    • First observedmacro_prepare_thread
    • First observedmacro_start_session
    • First observedmark_message_read
    • First observedpurge_old_messages
    • First observedregister_agent
    • First observedrelease_file_reservations
    • First observedrename_window
    • First observedrenew_file_reservations
    • First observedreply_message
    • First observedrequest_contact
    • First observedrespond_contact
    • First observedretire_agent
    • First observedsearch_messages
    • First observedsend_message
    • First observedset_contact_policy
    • First observedsummarize_recent
    • First observedsummarize_thread
    • First observedsweep_stale_agents
    • First observedunarchive_project
    • First observeduninstall_precommit_guard
    • First observedunretire_agent
    • First observedwhois

TDQS

B3.2/5.0

Scored across 41 tools

Disambiguation3/5

Most tools target distinct resources and actions, but several clusters overlap: agent lifecycle (register/create/deregister/retire/unretire), message state (mark read vs acknowledge), summary views (fetch/summarize thread/recent), and macro helpers duplicate lower-level operations. Detailed descriptions disambiguate some overlaps, but the set still requires careful selection.

Naming Consistency4/5

Nearly all names use snake_case and follow verb_noun or resource_action patterns, such as ensure_project, send_message, and release_file_reservations. Minor deviations like whois and a few long macro_* names slightly reduce predictability, but overall the naming is consistent.

Tool Count2/5

41 tools is far above a well-scoped set for one MCP server; it mixes core messaging, admin, contact, reservation, summary, guard, and macro operations. Although each tool may earn its place, the volume increases selection burden and feels heavy for the domain.

Completeness4/5

The surface covers project and agent lifecycle, messaging, contacts, file reservations, summaries, and pre-commit guards comprehensively. Minor gaps include no direct list_projects/list_agents tool and no edit/delete individual message, but resource-based discovery and purge/archive paths mitigate these.

Maintenance

ActivityActive
ResponsivenessResponsive

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    D
    maintenance
    A coordination layer for coding agents that provides identities, message threading, and searchable history. It features file reservation leases to prevent agents from overwriting each other's work in multi-agent environments.
    1
    MIT
  • A
    license
    Not graded
    quality
    A
    maintenance
    A lightweight coordination layer for multiple AI agents working on the same codebase, providing check-in and check-out tools via STDIO or Streamable HTTP.
    70 npm
    1
    Apache 2.0
  • A
    license
    Not graded
    quality
    F
    maintenance
    A mail-like coordination layer for coding agents, providing identities, inbox/outbox, searchable threads, and advisory file reservations to prevent conflicts in multi-agent workflows.
    34 PyPI
    1
    MIT