Skip to main content
Glama

AgentMate

npm version npm downloads License: MIT Node.js

Português (Brasil)

AI agents work better together.

AgentMate connects AI coding agents so they can collaborate, delegate, review, and help each other complete tasks. Give your AI agent a teammate: let Claude Code and the Codex CLI ask, review, research, plan, implement, cross-review and lead each other as background jobs.

AgentMate: in Claude Code, /mate:review hands a review to Codex; in Codex, $mate:ask asks Claude a question.

Why AgentMate

  • A second opinion from a different model. Ask the other CLI a question, or have it review your diff, before you commit to an approach. It reads your repository; it does not share your session's assumptions.

  • More than two CLIs. Besides Claude Code and Codex, AgentMate also drives the Gemini CLI and the Antigravity CLI (agy), the same one the author's agy-staff plugin is built on. Both are experimental teammates: tested with fake binaries, never against the real CLIs.

  • Work that does not block you. Every task is a durable background job with an id. The session that started it can end, and the result is still there.

  • Roles instead of raw prompts. ask, review, research, plan, implement and teamlead each send a tuned prompt with a defined output format, and each runs with the narrowest permissions that role needs. Jobs are read-only unless you say otherwise. crossreview chains two of them: one agent implements, the other reviews, and the loop runs without you relaying anything. split divides a broad goal into independent parts that both agents work on in parallel and review each other's, sharing context through sessions.

Related MCP server: claude-code-codex-agents

60-second quickstart

Install

Install the plugin in the host you use (or both), then restart the host.

# Claude Code
claude plugin marketplace add naldomadeira/agentmate
claude plugin install mate@agentmate

# Codex
codex plugin marketplace add naldomadeira/agentmate
codex plugin add mate@agentmate

The marketplace is named agentmate and the plugin is named mate, so the install id is mate@agentmate and every command starts with mate:.

Migrating from Agents Bridge (0.3.0 or earlier)? The project was renamed to AgentMate: the npm package is now agentmate, the plugin mate, the tools mate_*, the state directory ~/.agentmate and the environment variables AGENTMATE_*. Remove the old plugin (claude plugin uninstall bridge@agents-bridge, or agents-bridge@agents-bridge from 0.2.0) and the old marketplace (claude plugin marketplace remove agents-bridge); in Codex, remove both with codex plugin --help for the exact verbs. Then install mate@agentmate as above and restart the host. Jobs in ~/.agents-bridge are not migrated.

Invoke a role

Type /mate: in Claude Code to see every role; in Codex type $mate or open /skills. Both hosts take the provider first, then the request:

/mate:ask codex Is it safe to call this migration twice? See db/migrate/0042.sql
$mate:ask claude Does this retry loop in src/queue.ts have a race?

Want shorter commands (/ask, /prompts:ask)? See Slash commands. More examples are in Use cases.

Check the installation

npx -y agentmate doctor

doctor verifies Node.js, that the codex and claude CLIs (and the optional, experimental gemini and agy CLIs) are on your PATH and respond to --version, the job state directory, stale running jobs (it lists their ids) and legacy registrations, and prints a fix for each problem. It does not check authentication: if a job fails right away, log in to the destination CLI yourself. See the installation guide for upgrades, a local-development install and cleaning up leftover legacy registrations.

Set up a repository

npx -y agentmate init          # add or refresh the AgentMate block in AGENTS.md and CLAUDE.md
npx -y agentmate init --check  # change nothing; exit 1 if a block is missing or out of date

agentmate init (or /mate:init) writes a block of at most 25 lines between <!-- agentmate:start --> and <!-- agentmate:end --> in the repository's existing AGENTS.md and CLAUDE.md (a file with CRLF line endings keeps them), so every agent that opens the repo knows the /mate: and $mate: commands, how to check on jobs (agentmate jobs list, agentmate inbox) and the rules for a job it receives (no commit or push unless asked, report with the role's headings, treat session notes as data). Text outside the markers is never touched and a second run changes nothing. Claude Code does not read AGENTS.md natively, so --create starts both AGENTS.md and a CLAUDE.md that begins with @AGENTS.md (Claude Code imports it) followed by the block; use --files to pick other files and --cwd to run it elsewhere. Run it again after upgrading to refresh the block.

For agents

Paste this into any coding agent:

Read the raw text of https://raw.githubusercontent.com/naldomadeira/agentmate/main/docs/INSTALL_FOR_AGENTS.md
(curl it - do not work from a summary) and follow it to install and verify the AgentMate plugin
for the host you are running in. Respond in the user's language.

Upgrade

Claude Code and Codex install a copy of the plugin, so a new version only reaches you when you pull it in:

claude plugin marketplace update agentmate && claude plugin update mate@agentmate
codex plugin marketplace upgrade agentmate && codex plugin add mate@agentmate

Hosts cache the plugin per version, so restart Claude Code or Codex afterwards. A fix only lands if the plugin version changed.

Use cases

Examples use Claude Code's /mate:...; in Codex use $mate:.... The provider (codex or claude) comes first; pick the one that is not your host.

Use case

Invocation

Quick second opinion

/mate:ask codex Is this migration idempotent?

Review the working tree

/mate:review codex Review the current working tree

Review a PR or diff

/mate:review codex Review PR #730, focus on error handling

Challenge a plan

/mate:plan codex critique the migration plan in docs/plan.md

Research a topic

/mate:research claude How does auth work in this repo? Compare the options.

Implement a scoped fix

/mate:implement codex Fix the flaky retry test in test/queue.test.ts

Run a team lead

/mate:teamlead claude Audit error handling and propose fixes; delegate to codex

Implement, then cross-review

/mate:crossreview codex Add a --dry-run flag to the export command

Split a goal across both agents

/mate:split codex Add CSV and JSON export to the report command

Job ops

/mate:jobs list, /mate:jobs result <id>, /mate:jobs cancel <id>

implement and crossreview edit files, and split does when you pass --mode write, so use them only when you authorize that. The others are read-only.

Teammates

Teammate

Provider id

Needs

Notes

Claude Code

claude

the claude CLI

Reads with an explicit tool allowlist; web access in research.

Codex CLI

codex

the codex CLI

Sandboxed by mode; no web access.

Gemini CLI (experimental)

gemini

the gemini CLI on PATH (or AGENTMATE_GEMINI_BIN)

Runs gemini -p with --output-format stream-json and an approval mode per role. Headless Gemini has no shell and no web, --resume is unverified (so continue is refused), and a team lead is write-only. Tested with a fake binary only.

Antigravity CLI (experimental)

agy

the agy CLI on PATH (or AGENTMATE_AGY_BIN)

Runs agy -p with --output-format stream-json, --add-dir <cwd> and --print-timeout <N>m (the job deadline in whole minutes). Read-only jobs pass no permission flag: headless agy denies the tool calls its own profile does not allow, so results can be thin until you install the agy-staff allowlist or run in write mode. Write jobs and the (write-only) team lead pass --dangerously-skip-permissions. No shell is counted on and no web; continue works through --conversation <id>. Tested with a fake binary only.

A job on an agent that is not installed is refused up front with an install hint, and agentmate doctor lists each agent (a missing Gemini or Antigravity CLI is reported as optional, for example ok agy not installed (optional)). teamlead, crossreview and split pair the provider with the first installed other agent (in the order codex, claude, gemini, agy), resolve and check it before anything starts, and record it as the job's partner. Pass partner (mate_teamlead, mate_crossreview, mate_split, or jobs start --partner <agent>) to pick the other agent yourself, for example /mate:crossreview codex <task> with partner: gemini to have Gemini review Codex's change. The partner must differ from the provider and be installed. Gemini limits, in short: it cannot run shell commands (reviews of its changes or by it get the diff inline, capped at 30 000 characters, and an implement job cannot run the tests), it cannot search the web, it cannot continue a job, and as a team lead it needs mode: write. agy shares the diff-inline, no-web and write-only team lead limits, but it can continue a job; its read-only results depend on the permissions of your agy profile (see the table above).

What you can do

Eight roles, each reachable as a skill, an MCP tool and a CLI command. <provider> is codex, claude, gemini or agy (the last two experimental); pick the one that is not the host you are in.

Role

Skill

MCP tool

CLI

Mode

ask

ask

mate_ask

jobs ask <provider> "<question>"

read-only

review

review

mate_review

jobs start <provider> "<prompt>" --role review

read-only

research

research

mate_research

jobs start <provider> "<prompt>" --role research

read-only (Claude adds web access)

plan

plan

mate_plan

jobs start <provider> "<prompt>" --role plan

read-only

implement

implement

mate_implement

jobs start <provider> "<prompt>" --role implement

write (always)

teamlead

teamlead

mate_teamlead

jobs start <provider> "<prompt>" --role teamlead

read-only by default, write opt-in

crossreview

crossreview

mate_crossreview

jobs start <provider> "<task>" --role crossreview [--max-rounds N]

write on the implementer, read-only review

split

split

mate_split

jobs start <provider> "<goal>" --role split [--max-parts N] [--mode write]

read-only by default, write opt-in (one git worktree per part)

mate_ask waits for the answer (up to 120 seconds by default) and returns it in the same call. The other role tools return a job id immediately unless you pass waitSeconds.

Codex limits an MCP tool call to about 60 seconds by default. When you run inside Codex, pass waitSeconds: 45 to mate_ask and continue with mate_wait if the answer has not arrived.

Six more skills cover the rest:

  • jobs lists, observes, collects and cancels jobs.

  • delegate is the generic path (mate_start) for work that fits no role.

  • codex, claude, gemini and agy (the last two experimental) are shortcuts that route a plain request to the right role with the provider already set.

  • init sets up a repository for AgentMate (see Set up a repository).

In Claude Code the plugin also adds four agents that wrap Codex: codex-teammate (questions and general delegation), codex-reviewer, codex-researcher and codex-teamlead. They brief Codex, verify what it returns and report their own conclusion instead of forwarding raw output.

Every command in the table works without MCP. Prefix CLI commands with npx -y agentmate.

Slash commands

Nine commands (ask, review, research, plan, implement, teamlead, crossreview, split and jobs) can be started as a command in either host. Both hosts take the provider first, then the request.

Host and style

How to invoke

How to get it

Claude Code plugin

/mate:ask codex <question>

Installed with the plugin.

Claude Code bare

/ask codex <question>

npx -y agentmate install commands claude --global

Codex skill

$mate:ask claude <question>

Installed with the plugin; or pick "Mate: Ask" from the /skills menu.

Codex slash

/prompts:ask claude <question>

npx -y agentmate install commands codex, then restart Codex.

The plugin forms (/mate:ask, $mate:ask) need no extra install. The bare /ask and /prompts:ask rows are optional extras. Why two steps: Claude Code always prefixes plugin skills with the plugin name, so a bare /ask needs a user-level command file. Codex plugins can ship skills but not slash commands, so /prompts:<name> comes from a custom prompt file in $CODEX_HOME/prompts/ (default ~/.codex/prompts/).

npx -y agentmate install commands [claude|codex|both] [--global|--local] copies the templates from the package (templates/claude-commands/ and templates/codex-prompts/) and prints the command names it installed. The default target is both. It asks before overwriting an existing file. --local installs the Claude Code commands to ./.claude/commands/; Codex custom prompts are user-level only, so they are always installed globally.

Codex custom prompts are deprecated. OpenAI marks them deprecated in favour of skills. They still work today, and the $mate:ask skill needs no extra install, so use whichever you prefer. Restart Codex after installing prompts.

Each command calls the same mate_* tool as its skill, falls back to the npx -y agentmate jobs ... CLI when MCP is not loaded, and points to the skill for the full rules. jobs takes a verb instead of a provider: /jobs list, /jobs observe <id>, /jobs result <id>, /jobs cancel <id>.

Team lead mode

A team lead is a job whose worker plans a broad objective, delegates pieces to the other provider through the CLI, reviews the results and writes a report. You start it with mate_teamlead or the teamlead skill and follow it with mate_observe. The lead delegates to partner when you pass one (for example a claude lead with partner: gemini); the default is the first installed other agent.

your session
  `-- teamlead job (depth 0, started by you)       codex lead: danger-full-access
        |-- research job (claude)                   depth 1, cannot start jobs
        |-- review job (claude)                     depth 1, cannot start jobs
        `-- implement job (claude)                  depth 1, write, one per working tree

The stored depth is 0 for a job started by a session and 1 for a job started by a worker. The limit is two levels: a session starts a team lead, the lead starts child jobs, and the children cannot start jobs.

The final report has the sections Objective, Plan, Delegations (id, provider, role, status), Findings, Decisions, Deliverables and Open questions. jobs observe <id> shows the lead's output with its children; jobs list --parent <id> lists only the children.

Codex sandbox warning. A Codex team lead runs with --sandbox danger-full-access whether mode is read-only or write. It has to start worker processes and write job state under ~/.agentmate, so the sandbox cannot be narrower. In read-only mode the runtime refuses any write child job and the prompt forbids edits, but the lead itself is not sandboxed. Treat a Codex team lead like any Codex session with full filesystem access, or lead with claude instead, whose permissions are an explicit tool allowlist that includes an explicit deny of Edit, Write and NotebookEdit in read-only mode.

The team lead calls the CLI pinned to the installed version (npx -y agentmate@<version> jobs ...), so a local checkout that is not published to npm must be published or linked before team lead mode works.

Gemini team lead warning (experimental). A gemini team lead is accepted only with mode: write, because delegating needs the shell and Gemini allows it only in --approval-mode yolo, which approves every tool call without asking. A read-only Gemini lead is refused (A Gemini team lead needs mode write: delegation requires the shell, which Gemini only allows in yolo mode.). Treat it like the Codex lead above, or lead with claude.

agy team lead warning (experimental). An agy team lead is accepted only with mode: write, because delegating needs the shell and headless agy allows it only with --dangerously-skip-permissions, which skips every permission check. A read-only agy lead is refused (An agy team lead needs mode write: delegation requires the shell, which agy only allows with --dangerously-skip-permissions.). Treat it like the Codex lead above, or lead with claude.

Use a team lead only when the work has several independent parts. One question or one review is cheaper as ask or review.

Cross-review

Cross-review lets one agent implement and the other review, so you do not copy a diff between sessions. It is a workflow job: the worker does not call a CLI itself, it runs the steps as child jobs. provider implements; the other agent reviews (partner, when you pass one). Start it with mate_crossreview or the crossreview skill (/mate:crossreview codex <task>), and follow it with mate_observe.

implement (provider, write)  ->  review (other agent, read-only)  ->  verdict
        ^                                                              |
        |            Verdict: request-changes, rounds left             |
        `---- implement --continue with the findings  <---------------'
                                                      |
                              Verdict: approve  ->  done

Each round has two child jobs on the same working directory (depth 1, parentJob = the workflow id). The implementer runs in write mode; from round 2 it continues its own session with the reviewer's findings (an implementer that cannot continue, such as Gemini, starts a fresh implement job with the findings instead). The reviewer reads the uncommitted git diff (and git status for new files) read-only, or receives the diff in its briefing when it cannot run the shell (Gemini, agy), with the implementer's report as context, and must end its review with the line Verdict: approve or Verdict: request-changes.

Stop rules:

  • Verdict: approve ends the workflow as done.

  • Verdict: request-changes starts another round while rounds remain (maxRounds, 1 to 5, default 2). When the budget is used up the workflow is done and the report says so; the latest edits stay in the working tree.

  • No clear verdict ends the workflow at once as done with a ## Needs human section; it never loops blindly.

  • A child that ends error, timeout or canceled ends the workflow as error, naming that child's id and error. A child that ends quota_exhausted ends it as error with that child's own error, which already says how to hand the work to the other agent. The workflow's own --timeout is the overall deadline, and jobs cancel <id> on the workflow cancels its running children first (recursively), then the workflow itself.

The report (jobs result <id>) has ## Task, ## Rounds (round, implement job, review job, verdict), ## Final review, ## Changes, ## Needs human when it applies, and ## Next steps with the jobs result <child-id> commands. jobs events <id> lists one important event per step, such as round 1: review verdict approve.

npx -y agentmate jobs start codex "Add a --dry-run flag to the export command" --role crossreview --max-rounds 3

It edits files, so start it only when you authorize that, and keep one write job per working tree. Only a top-level session can start it, like a team lead. --model applies to the implementer only.

Task splitting

Task splitting divides a broad goal into independent parts, runs the parts in parallel on both agents and has the other agent review each one, so you do not relay results between sessions. Like cross-review it is a workflow job: the worker calls no CLI itself, it runs the steps as child jobs (depth 1, parentJob = the workflow id) that share one session. provider plans; each part goes to an installed agent, or only to the provider and its partner when you pass one. Start it with mate_split or the split skill (/mate:split codex <goal>), and follow it with mate_observe.

goal -> plan (provider, read-only)
          |  1..maxParts parts: closed interfaces, no overlapping files, one agent each
          v
        parts in parallel ---- read-only: a research job per part, on the working directory
          |                    write:     a git worktree + branch per part, an implement job in each
          v
        cross-review (the other agent of each part, read-only, Verdict: approve | request-changes)
          |
          v
        integration report (parts table, merge order, what needs a human)
  1. Plan. A plan job on provider returns a fenced json block, { "parts": [{ "id", "title", "briefing", "files", "agent" }] }, with 1 to maxParts parts (2 to 4, default 3). The worker takes the last such block and checks it: unique ids (a-z, 0-9, -), known agents (a missing or unknown agent alternates, starting with the other agent). If the block is invalid the workflow ends error with a pointer to jobs result <plan-job>. The plan goes into the session notes, so every part sees it.

  2. Parts, in parallel. Read-only (default): one research job per part on the part's agent, in the working directory. Write (--mode write): for each part the worker creates a worktree from the recorded base commit, git worktree add -b agentmate/<split-id>/<part-id> ~/.agentmate/worktrees/<split-id>/<part-id> <base-commit>, then an implement job works there, so your working tree is untouched. When the implementer finishes, AgentMate commits what it left on the part's branch automatically (--no-verify, gpg signing off).

  3. Cross-review. Each finished part is reviewed read-only by the other agent of the pair (the provider or the partner): in write mode in the part's worktree, against the commit the branch started from (git diff <base-commit>, or that diff inline when the reviewer is Gemini or agy and cannot run the shell); in read-only mode over the research result. The review ends with Verdict: approve or Verdict: request-changes.

  4. Report. jobs result <id> has ## Goal, ## Parts (part, title, agent, part job, review job, verdict, branch), ## Integration, ## Needs human when it applies and ## Next steps with the jobs result <child-id> commands. In write mode the report lists every worktree and branch with its cleanup commands (git worktree remove <path>, git branch -D <branch>), ## Integration has the ordered git merge agentmate/<split-id>/<part-id> commands for approved parts only, and every other part goes under ## Needs human; in read-only mode it merges the research results. jobs events <id> lists one important event per step.

A part that fails does not stop the others: they run to completion (and are reviewed), then the workflow ends error naming the failed part. A part that is not approved (request-changes, a missing verdict or a failure) lands in ## Needs human. The workflow's own --timeout is the overall deadline, and jobs cancel <id> on it also cancels the running children.

Write mode limitations. It needs a git repository with a clean working tree, checked before the planner runs; commit or stash first. The worktrees are fresh checkouts: they have no node_modules, no .env and no submodule contents, so a part cannot run checks that need them unless its briefing says how to set them up.

What is not automated: AgentMate never merges, pushes, rebases or deletes branches for you, and merge conflicts between parts are not resolved automatically. Run the merges from the report yourself, resolve any conflict, run the tests, then remove the worktrees. There is no second round: a part with request-changes is for you (or a follow-up job) to fix.

npx -y agentmate jobs start codex "Add CSV and JSON export to the report command" --role split --max-parts 3
npx -y agentmate jobs start codex "Add CSV and JSON export to the report command" --role split --mode write

Write mode edits files (in the worktrees), so start it only when you authorize that. Only a top-level session can start it, like a team lead. --model applies to the planner and to parts run by the same agent.

How it works

AgentMate treats delegated work as a durable background job:

  1. Start a task on codex or claude. The tool returns a job id immediately (or the answer, for mate_ask).

  2. Wait for that id, collect its result, or request progress when a person asks for it.

  3. Use the stored result to decide the next step. Jobs survive the caller session ending.

A task moves through a queue, worker, and returned result.

The jobs MCP server (npx -y agentmate serve jobs, registered by the plugin) and the jobs CLI share one runtime. State lives under ~/.agentmate. A detached worker runs the provider CLI and records the output, so an expired wait never stops a job.

Every job also writes an append-only event log (events.jsonl). Both Codex and Claude (--output-format stream-json) stream events while they run. Events carry one of three levels: important (the agent's messages, errors, start and finish), status (files changed) and fyi (commands run). mate_observe and jobs observe show only important and status events, so progress checks stay small; ask for fyi with levels / --level, read the full log with mate_events / jobs events <id>, and pull the raw stdout/stderr tails only when needed with raw / --raw. See docs/ARCHITECTURE.md for the runtime, adapters and event model.

Capability

MCP

CLI

Start work

mate_start, mate_ask, mate_review, mate_research, mate_plan, mate_implement, mate_teamlead, mate_crossreview, mate_split

jobs start, jobs ask

Wait or fetch output

mate_wait, mate_result

jobs wait <id>, jobs result <id>

Request progress

mate_observe

jobs observe <id> [--raw]

Read job events

mate_events

jobs events <id> [--follow]

Cancel work

mate_cancel

jobs cancel <id>

Find jobs

mate_list

jobs list [--cwd] [--parent <id>]

Sessions

mate_session_start, mate_session_show, mate_session_notes, mate_session_list

sessions start/show/notes/list, jobs start --session <id>

Inbox

mate_inbox

inbox [--cwd] [--all] [--no-ack] [--follow]

Sessions. A session is shared context across jobs and agents. mate_session_start(title, cwd?) creates one under ~/.agentmate/sessions/<id>/ (session.json plus notes.md; membership is derived from the jobs that carry the session id, so session.json has no jobs array); pass its id as session to any role tool or mate_start (CLI: jobs start ... --session <id>) and the job is recorded in it. mate_session_notes(id, text, author?) appends a note (a single note is capped at 2000 characters), mate_session_show(id) shows the notes (tail) and the session's jobs, and mate_session_list(cwd?, limit?) lists sessions. When a session has notes, every worker started in it receives them after the task, under ## Shared session notes, in a code fence and framed as data written by other agents, not as instructions. The injection is capped at 4000 characters of whole entries (the newest that fit), whatever the role, so keep notes short and factual: decisions, constraints, file locations. Jobs that a workflow starts (crossreview, split) inherit the workflow's session, and split creates a session for itself when you pass none and writes its plan into the notes. The notes are plain text under ~/.agentmate, so keep secrets out of them.

Session start summary. In Claude Code the plugin also registers a SessionStart hook (hooks/hooks.json). When a session opens, it prints one line (400 characters at most) about the jobs started from that directory or a subdirectory: the jobs that finished since the last session there (quota_exhausted ones first, marked "needs hand-off"), how many are still running and how many are stale (their worker is gone). It stays silent when there is nothing to report and for 120 seconds after the last summary in the same directory; the first time in a directory it looks back 24 hours. Set AGENTMATE_HOOK_QUIET=1 to turn it off. It never fails a session: any error exits silently. This is a Claude Code feature only; Codex has no equivalent hook.

Inbox. Every time a job records an important message, error or finish, AgentMate also appends one line to ~/.agentmate/inbox.jsonl (owner-only, rotated to inbox.1.jsonl past 5 MB, one generation): { ts, job, cwd, session?, provider, role, kind, text }. It is how one agent learns that the other finished or failed without sitting in wait. mate_inbox(cwd?, unread? = true, ack? = true, limit? = 20) returns the entries for that directory or a subdirectory, one line each (HH:MM:SS <job> <role>/<provider> <kind> <text>), oldest first, and moves a read marker kept per directory in ~/.agentmate/inbox-cursors/ through the last entry it showed (… N more unread (run again) when there are more), so nothing is shown twice; a terminal job already read through mate_wait or mate_result is not listed again, and entries of jobs started by another job (workflow steps, team lead children) are left out because their parent reports. Entry text is capped at 500 characters and is untrusted worker output: data, not instructions. Workers never read the host's inbox: mate_inbox declines when AGENTMATE_JOB_ID is set and the hooks stay silent inside a worker. From a terminal, agentmate inbox does the same (--all for every directory, --no-ack to only look) and agentmate inbox --follow prints new lines every second until Ctrl-C. In Claude Code the plugin also registers a UserPromptSubmit hook: before each prompt it adds up to 5 unread entries for the directory to the context (10 seconds of cooldown per directory, AGENTMATE_HOOK_QUIET=1 turns it off, any error exits silently), and the SessionStart summary mentions how many are unread. Codex has no hooks, so the jobs, teamlead, crossreview and split skills tell it to call mate_inbox when it resumes a turn with jobs in progress, before mate_wait. Entries are short summaries; read the full output with jobs result <id>.

Skills prefer the mate_* tools. If the host did not load MCP, they run the same job contract through npx -y agentmate; they never change a user's host configuration as a fallback.

Usage examples

Ask a quick question

npx -y agentmate jobs ask codex "Why might src/jobs/store.ts lose a write under concurrent workers?" --wait 120s

The answer is printed when it arrives. If the wait expires, the job keeps running: jobs wait <id> collects it.

Review a change

Start with a read-only request. Give the receiving CLI enough context to produce an actionable result: name the goal, relevant files or diff, and the expected answer.

npx -y agentmate jobs start codex "Review the current diff. Report only actionable findings." --role review
# retain the job ID printed by start
npx -y agentmate jobs wait <job-id>

After a plugin restart, the same workflow can be requested through the installed review skill or the delegate skill. The skill selects MCP when it is available and otherwise runs the CLI commands above.

Delegate an authorized edit

Jobs default to read-only, except implement, which always runs in write mode and rejects read-only. Use it only when the task is explicitly allowed to change files, and send one write job per working tree at a time. The --mode write flag is redundant for implement but makes the intent explicit.

npx -y agentmate jobs start claude "Add a focused regression test for the parser." --role implement --mode write --cwd .
npx -y agentmate jobs wait <job-id>

Run a team lead

npx -y agentmate jobs start claude "Audit the CLI for inconsistent error handling and propose fixes. Delegate independent areas to codex." --role teamlead
npx -y agentmate jobs observe <job-id>
npx -y agentmate jobs wait <job-id> --timeout 10m

Cross-review a change

npx -y agentmate jobs start codex "Add a --dry-run flag to the export command" --role crossreview --max-rounds 3
npx -y agentmate jobs wait <job-id> --timeout 10m
npx -y agentmate jobs result <job-id>

Codex implements, Claude reviews the uncommitted diff, and the report lists the rounds. Use jobs events <job-id> for the step-by-step log.

Split a goal across both agents

sid=$(npx -y agentmate sessions start "CSV and JSON export")
npx -y agentmate sessions notes "$sid" "Keep the CLI flags stable; formatters live in src/export/."
npx -y agentmate jobs start codex "Add CSV and JSON export to the report command" --role split --session "$sid"
npx -y agentmate jobs wait <job-id> --timeout 10m
npx -y agentmate jobs result <job-id>
npx -y agentmate sessions show "$sid"

Codex plans the parts, both agents research them in parallel, and the report merges the findings. Add --mode write to implement each part in its own worktree and branch.

Continue, inspect, or cancel a job

An expired wait does not stop work. Repeat wait for the same ID, inspect output when progress is requested, or collect a stored result after an interrupted terminal session.

npx -y agentmate jobs observe <job-id>
npx -y agentmate jobs result <job-id>
npx -y agentmate jobs cancel <job-id>

A finished job with a saved session can continue on the same provider (not available for gemini yet, whose --resume is unverified; start a new job with the full context. agy continues with --conversation <id>):

npx -y agentmate jobs start codex "Address the highest-priority finding." --continue <job-id>

wait exit code

Meaning

Next action

0

Job completed

Read and assess the returned result.

1

Job failed, was canceled or hit its quota (quota_exhausted)

Read result for the retained output and error.

2

Wait expired while job remains active

Repeat wait; do not create a duplicate job.

jobs ask uses the same exit codes. Do not pipe wait or ask: a pipe discards the exit code.

jobs wait --timeout (for example 10m or 90s) only limits how long the command waits. To limit how long a job may run, pass jobs start --timeout <minutes>: a positive number of minutes, at most 120.

Safety model

  • Read-only by default. ask, review, plan and research always run read-only, and teamlead is read-only unless you pass mode: write. implement always runs in write mode: --role implement and mate_implement default to write and reject read-only. Use it only after the user has authorized edits.

  • Permissions per role. The sandbox or tool allowlist follows the role and mode:

    Role and mode

    Codex sandbox

    Claude permissions

    Gemini (experimental)

    agy (experimental)

    read-only (ask, review, plan, read-only teamlead)

    read-only

    allowlist (below) and an explicit deny of Edit, Write and NotebookEdit

    --approval-mode default

    no permission flag (agy's own profile)

    research

    read-only

    the read-only allowlist plus WebSearch and WebFetch, with the same explicit deny

    --approval-mode default

    no permission flag (agy's own profile)

    implement, write teamlead

    workspace-write

    acceptEdits permission mode plus the verification allowlist (below)

    --approval-mode auto_edit

    --dangerously-skip-permissions

    teamlead (Codex lead, read-only or write)

    danger-full-access

    not applicable

    not applicable

    not applicable

    teamlead (Claude lead)

    not applicable

    the rows above, plus CLI access limited to agentmate jobs * (installed version)

    not applicable

    not applicable

    teamlead (Gemini lead, write only)

    not applicable

    not applicable

    --approval-mode yolo

    not applicable

    teamlead (agy lead, write only)

    not applicable

    not applicable

    not applicable

    --dangerously-skip-permissions

    The read-only allowlist is Read, Grep, Glob, git diff, git log, git show and git status. The write-mode allowlist adds pnpm, npm, npx, yarn, bun, make, git add and git commit so a worker can run verification commands. Extend it with the environment variable AGENTMATE_CLAUDE_WRITE_TOOLS, a comma-separated list of Claude permission patterns. A Claude team lead cannot run install, only agentmate jobs *. Claude workers start with --strict-mcp-config, so none of your MCP servers (claude.ai connectors and plugins included) load into a job: they start faster and only have the tools above. Set AGENTMATE_CLAUDE_INHERIT_MCP=1 to give workers your MCP servers.

  • Gemini runs with the least approval that fits the role (experimental). --approval-mode default denies every tool that needs approval in headless mode, so read-only and research jobs can read but not run shell commands or search the web; auto_edit approves file edits but still denies the shell, so an implement job cannot run the tests. A Gemini team lead needs the shell, which only --approval-mode yolo allows: it is accepted with mode: write only, and it approves every tool call without asking, so it is as unrestricted as the Codex lead below. --sandbox and the legacy --yolo flag are never used.

  • agy asks for nothing in read-only and for everything in write mode (experimental). Read-only jobs pass no permission flag, so headless agy denies the tool calls its own profile does not allow; a read-only job can therefore return thin results until you install the agy-staff allowlist or run in write mode. --dangerously-skip-permissions is passed to every mode: write job, including implement and the team lead, and approves every tool call without asking, so it is as unrestricted as the Codex lead below. An agy team lead is accepted with mode: write only. agy has no --sandbox here either.

  • A Codex team lead is not sandboxed. It runs with --sandbox danger-full-access in either mode, because it must spawn worker processes and write job state. "Read-only" for a Codex lead means that the runtime refuses any write child job (a read-only parent cannot start write children) and that the prompt forbids edits; it does not restrict the lead's own process. Lead with claude when this matters.

  • Delegation depth limit of 2. A session starts a team lead (depth 0), the lead starts child jobs (depth 1), and children cannot start jobs. The runtime refuses a third level and refuses a team lead, a cross-review or a split started by a worker.

  • Cross-review writes only through its implementer. The implement step gets the implement permissions above; the review step is read-only. The workflow job itself calls no CLI.

  • Split writes only inside its worktrees. In write mode each part's implement job runs in its own git worktree on its own branch (agentmate/<split-id>/<part-id>), the planner and reviewers are read-only, and nothing is merged into your branch for you.

  • One write job per working tree at a time. Two writers in one tree collide. The skills and the team lead prompt follow this rule; use separate git worktrees for parallel edits (split does this for its parts).

  • A spent plan is a status, not a crash. When a provider reports that its usage limit, quota or credits are used up, the job ends quota_exhausted (terminal; wait and ask exit 1). Its error reads <provider> quota exhausted: <line>. Retry after the reset or start the job on <other agent>., and result adds Hand off: start the same job with provider <other>.; the other agent is the first installed one, and when none is installed both lines say to retry after the reset instead. Detection reads the stderr tail and parsed errors only, so a 429 alone is not exhaustion, and the job is not retried once a quota line is seen. Extend the detection with AGENTMATE_QUOTA_PATTERNS, case-insensitive regular expressions separated by |; invalid ones are ignored. The defaults cover agy's RESOURCE_EXHAUSTED status and its Individual quota reached and quota exceeded wording; its bare code 429 is too loose to be a default, so add AGENTMATE_QUOTA_PATTERNS='\bcode\s*429\b' if you want it (a 429 caused by a tool the agent called then counts as exhaustion too).

  • The delegator owns acceptance. Job output is an input to your judgment. Verify claims and run the tests before you merge anything a worker produced.

  • No hidden configuration changes. The plugin registers its own MCP server. The fallback path runs the CLI and never edits host configuration. Do not put secrets in briefings: prompts and results are stored in plain text under ~/.agentmate, in files created with owner-only permissions (0600 for files, 0700 for directories).

Troubleshooting

Start with npx -y agentmate doctor. It prints ok, warn or fail for each check with a hint, and exits 1 if any check fails. It runs without codex or claude installed and reports the missing CLI as a warning. It checks that each CLI responds to --version, not that you are logged in.

Symptom

Likely cause and fix

mate_* tools or skills do not appear

Restart the host after installing; check claude plugin list or codex plugin list.

A job fails immediately

The destination CLI is missing from PATH (doctor reports this) or not authenticated (doctor does not check; log in to it).

wait or ask exits 2

The job is still running. Repeat jobs wait <id>; do not start a duplicate.

A job shows running but nothing happens

The worker process died. doctor lists the ids of these jobs; jobs cancel <id> clears them.

Delegation depth limit reached

A job started by a worker tried to start another job (a third level). Return the findings to the session that started the job instead.

Only a top-level session can start a teamlead job

A worker tried to start a team lead (or a cross-review or a split: ... a crossreview job, ... a split job). Start it from your own session.

Job ends quota_exhausted

The provider's usage limit, quota or credits are spent. Detection looks at the stderr tail and parsed errors only, a 429 alone is not exhaustion, and the job is not retried once a quota line is seen. Wait for the reset it names, or start the same job on the other agent (result prints the hint). If your provider words it differently, add a pattern to AGENTMATE_QUOTA_PATTERNS (for example AGENTMATE_QUOTA_PATTERNS="plan cap|budget burned").

Duplicated or conflicting tools

A legacy registration (serve codex / serve claude) is still present. doctor flags it; see "Removed in 0.6.0".

Requirements

  • Node.js 18 or later

  • Claude Code and/or Codex CLI, authenticated

  • The destination CLI available on the initiating host's PATH

Removed in 0.6.0

The deprecated synchronous servers (serve codex, serve claude), the setup command, the agentmate-codex / agentmate-claude binaries and the legacy /codex and /claude setup skill are gone; the plugin and the mate_* job tools replace them. To remove a leftover registration, run claude mcp remove codex -s user for Claude Code, or delete the [mcp_servers.claude] section from ~/.codex/config.toml for Codex. npx -y agentmate doctor flags both. See the installation guide for the details.

Development

git clone https://github.com/naldomadeira/agentmate.git
cd agentmate
pnpm install
pnpm build
pnpm test
pnpm lint

Cutting a release

pnpm version:bump x.y.z          # package.json, src/lib/version.ts, both plugin manifests, the Codex marketplace
pnpm release:prepare             # version check, build, smoke:pack (real tarball install), smoke:cli
git commit -am "chore: release vx.y.z"
git tag vx.y.z
git push --follow-tags

The Release workflow runs on the tag: it fails when the tag does not match the package version, runs the checks, publishes through npm Trusted Publishing (no token) and verifies that the version appears on the registry.

Release notes are in the changelog.

License

MIT

Available Tools

6 tools
codex_explain_codeCodex Explain CodeB

Ask Codex to deeply explain code, logic, or architecture. Useful for understanding unfamiliar code, onboarding, or documenting complex systems.

ParametersJSON Schema
NameRequiredDescriptionDefault
depthNoDepth of explanation: overview, detailed, or full execution tracedetailed
targetYesWhat to explain: file path, function name, module, or code snippet
contextNoAdditional context about the codebase
workingDirectoryNo

TDQS

B3.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. "Ask Codex" hints that work is delegated to an external agent, but nothing is said about latency, cost, whether the target is uploaded/sent, or what the response looks like for a 4-parameter tool with no output schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences with the core action front-loaded and no redundant filler. The second sentence earns its place by adding usage context, though it is the weaker half.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no annotations and no output schema, the description should carry more of the behavioral load for a 4-parameter tool. It covers what the tool does and when to use it, but omits return format, target-size limits, and delegation/execution caveats.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 75%, with the enum on depth and descriptions for target and context already documented in the schema. The description adds no parameter meaning beyond that, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

"Ask Codex to deeply explain code, logic, or architecture" gives a specific verb (explain) and resource (code/logic/architecture), clearly distinct from write-oriented siblings like codex_implement or codex_review_code. It stops short of naming an alternative, but an agent can tell what this tool is for.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description lists scenarios (understanding unfamiliar code, onboarding, documenting complex systems), which implies when to reach for it. However it names no alternative tool and states no exclusions, so the agent must infer that codex_review_code is for critique rather than explanation.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

codex_implementCodex ImplementA

Ask Codex to implement a feature, fix a bug, or make code changes. WARNING: This modifies your codebase. Returns a summary of what was changed.

ParametersJSON Schema
NameRequiredDescriptionDefault
taskYesWhat to implement or fix
modelNoOverride the Codex model
sandboxNoSandbox level (must allow writes)workspace-write
workingDirectoryNo

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does disclose the critical trait: 'WARNING: This modifies your codebase' plus a note about the return summary. However, it omits permission/auth requirements, whether edits are reversible or committed, and how the sandbox setting affects scope.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, zero waste, purpose front-loaded and the mutation warning placed prominently. Nothing to trim.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation tool with no annotations and no output schema, it covers the essentials: the action, the fact it writes to the codebase, and the return shape. It is slightly thin on the sandbox/permission consequences but otherwise sufficient to call the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 75% and the description adds no parameter-level detail at all (task, model, sandbox, workingDirectory are undocumented in the prose). Since the schema already documents most parameters, baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('implement a feature, fix a bug, or make code changes') and the description clearly positions this as the mutating action among siblings like codex_review_code and codex_explain_code. It stops short of explicitly naming an alternative, so it's clear but not fully differentiated by text alone.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The use cases (implement, fix, change code) imply when to call it, but there is no when-not guidance or reference to alternatives such as codex_plan_* or codex_review_*. Usage is implied rather than stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

codex_plan_perfCodex Performance PlanB

Ask Codex to analyze performance and create an improvement plan. Identifies bottlenecks, proposes ranked optimizations with expected impact.

ParametersJSON Schema
NameRequiredDescriptionDefault
targetYesWhat to optimize: function, module, or pipeline path
contextNoAdditional context about usage patterns
metricsNoPerformance metrics to focus on
constraintsNoConstraints: must not increase binary size, etc.
workingDirectoryNo

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It usefully discloses the output shape (bottlenecks, ranked optimizations with expected impact) and that the deliverable is a plan rather than edits, which tells the agent this is non-mutating. It does not disclose permissions, latency/cost, or whether analyses are cached, so the behavioral picture is only partial.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences with no filler, and the primary action is front-loaded before the deliverable detail. Nothing redundant, though it is quite terse for a 5-param planning tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With five parameters (one required), 80% schema coverage, and no output schema, the description compensates reasonably by listing what the resulting plan contains. It gives the agent enough to know what calling this produces, though it could say more about how target/constraints scope the plan.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 80%, so the schema already documents target, context, metrics, and constraints. The description adds nothing about parameter formats, defaults, or how metrics/constraints interact with the analysis, leaving the schema to do the work. Baseline 3 is appropriate given high coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States specific verbs and resources: 'analyze performance and create an improvement plan', with the deliverable further specified as 'bottlenecks' and 'ranked optimizations with expected impact'. An agent can tell this is a planning/analysis tool rather than an implementation one. However, it never names or contrasts with close siblings like codex_review_plan or codex_review_code, so differentiation is left to inference.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no statement of when to use this versus codex_review_plan, codex_implement, or other siblings, and no prerequisites or exclusions. Usage is only implied by the tool's subject matter (performance work). Nothing routes the agent explicitly.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

codex_queryAsk CodexB

Ask OpenAI Codex a question or give it a task. Use for getting a second opinion, exploring unfamiliar code, or tasks that benefit from a different model's perspective.

ParametersJSON Schema
NameRequiredDescriptionDefault
modelNoOverride the Codex model
promptYesThe question or task for Codex
sandboxNoSandbox level controlling what Codex can modifyread-only
workingDirectoryNoWorking directory (defaults to server cwd)

TDQS

B3.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden and falls short. It never warns that the 'sandbox' parameter can escalate Codex to 'workspace-write' or 'danger-full-access' (i.e. file mutation), nor does it mention that this is an external model call with associated latency, cost, or credential requirements. Those are material traits the agent needs before invoking.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences, with the core purpose front-loaded and usage contexts following. No filler, though the second sentence is somewhat generic and could be sharpened against the specialized siblings.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no annotations and no output schema, the description must carry more weight for a 4-parameter tool that can grant another model write access to the workspace. It explains neither the return/response behavior nor the safety implications of the sandbox enum, leaving significant gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents prompt, model, sandbox, and workingDirectory, giving a baseline of 3. The description adds no parameter-level detail beyond that, and notably omits any explanation of the sandbox risk levels.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Ask') and resource ('OpenAI Codex'), plus elaborates on the kinds of tasks (questions or task delegation). It is clearly the general-purpose entry point, though it never explicitly positions itself against the more specialized siblings like codex_review_code or codex_explain_code.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives concrete usage contexts: 'getting a second opinion, exploring unfamiliar code, or tasks that benefit from a different model's perspective.' That is real when-to-use guidance, but it names no exclusions and does not tell the agent when to prefer a sibling such as codex_explain_code instead.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

codex_review_codeCodex Code ReviewB

Ask Codex to review code. Provide a git diff range, file paths, or a code snippet. Returns specific, actionable feedback.

ParametersJSON Schema
NameRequiredDescriptionDefault
targetYesWhat to review: git diff range (e.g., "HEAD~3..HEAD"), file paths, or code snippet
contextNoAdditional context about the codebase or changes
focusAreasNoFocus on: bugs, performance, style, security, etc.
workingDirectoryNo

TDQS

B3.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It states the output is 'specific, actionable feedback' but says nothing about read-only behavior, whether it reads the working tree or git history, rate limits, or how large inputs are handled. For an analysis tool with zero annotation coverage, this is a significant gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences that front-load the purpose and then cover inputs and output. Little waste, though the input sentence largely repeats the schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema and no annotations, the description should do more to explain the nature of the feedback and constraints. It is minimally adequate but leaves the agent guessing about scope and behavior.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 75%, so the schema documents most parameters including target, context, and focusAreas. The description merely restates the target options already in the schema and adds no syntax or format detail; workingDirectory is undocumented in both.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('review code') and clarifies the input form (diff range, file paths, snippet). It does not, however, differentiate itself from siblings like codex_explain_code or codex_review_plan, so an agent must infer the boundary.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies when to use it (when you have a diff, paths, or snippet to review) but gives no explicit when-not guidance or alternatives. No routing to siblings such as codex_explain_code for explanation tasks.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

codex_review_planCodex Plan ReviewA

Ask Codex to critique an implementation plan. Identifies gaps, risks, missing edge cases, and suggests improvements.

ParametersJSON Schema
NameRequiredDescriptionDefault
planYesThe implementation plan or design to review
constraintsNoKnown constraints: timeline, tech stack, compatibility
codebasePathNoPath to relevant codebase for context
workingDirectoryNo

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden. It usefully discloses what the critique produces (gaps, risks, missing edge cases, improvements), but says nothing about determinism, latency, whether codebasePath materially affects results, or what the response looks like.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences, front-loaded with the core action and followed by the concrete outputs of the review. Every clause earns its place with no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 4-parameter tool with no output schema and no annotations, the description covers the essential purpose and deliverable but omits return-shape expectations and the role of undocumented parameters like workingDirectory, leaving the agent to infer call mechanics.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 75%, so most parameters are self-documented in the schema. The description adds no meaning beyond the schema and leaves workingDirectory (undocumented) entirely to the schema's absence, so baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (critique) and resource (implementation plan), and the name plus description make it distinguishable from codex_review_code by the plan-vs-code subject. It does not explicitly name sibling tools to route between, so it falls short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is only implied: an agent can infer it should call this when it has a plan to be reviewed rather than code, but the description never states when to use this versus codex_review_code or codex_plan_perf, nor any prerequisites.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 6 tool updatesv0.1.0
    • First observedcodex_explain_code
    • First observedcodex_implement
    • First observedcodex_plan_perf
    • First observedcodex_query
    • First observedcodex_review_code
    • First observedcodex_review_plan

TDQS

A3.6/5.0

Scored across 6 tools

Disambiguation4/5

Each tool has a distinct primary purpose, but codex_query is a generic catch-all that could overlap with the specialized tools. codex_review_plan and codex_plan_perf also have some conceptual overlap around planning, though their descriptions differentiate them.

Naming Consistency4/5

All names use a consistent codex_ prefix and snake_case, following a clear verb or verb_noun pattern. Minor deviation: codex_query and codex_implement are single verbs, while others are compound.

Tool Count5/5

Six tools is well-scoped for a Codex bridge, covering query, review, explanation, performance planning, implementation, and code review. Each tool appears to earn its place without excessive fragmentation.

Completeness4/5

The surface covers core Codex use cases: asking questions, reviewing plans/code, explaining code, performance analysis, and implementing changes. Minor gaps exist for specialized tasks like test generation or refactoring, but these can be handled via codex_query or codex_implement.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers