Skip to main content
Glama

codex-router-mcp

CI License: MIT

An MCP server that puts guardrails around delegating work to OpenAI Codex.

Codex already ships its own MCP server (codex mcp-server), and it exposes two tools: codex and codex-reply. If all you want is "run a Codex session from another agent", use that — it is first-party and costs you nothing to maintain.

This project exists for what happens around the delegation:

  • No task can get stuck "running". Every turn is supervised: a lost completion is recovered from the thread's real state, a turn blocked on an approval nobody can give is interrupted, and an interrupt Codex ignores is forced.

  • Quota is checked before a thread is started, normalized by window duration, and an exhausted account returns a structured handoff instead of a failure.

  • Models are picked by policy — fast, balanced or frontier — with reasoning effort capped per model.

  • Images can be generated, so an agent that cannot draw can still ship a real raster asset, and look at it before using it.

  • Risky work runs in a dedicated git worktree, so a bad turn cannot touch your working tree.

  • Every turn is bracketed by checkpoints, so you can roll one back.

  • Reviews run read-only, in both directions, optionally with a second model.

  • Changes are read from disk, not trusted from Codex. Files Codex writes through the shell are reported; rejected writes are reported as failed.

Claude Code  ──MCP──▶  codex-router-mcp  ──JSON-RPC──▶  codex app-server

One persistent codex app-server child process serves every thread, so the second delegation does not pay the startup cost again. Concurrent delegations each get their own thread and never cross results.

Requirements

  • Node 20+

  • The codex CLI on PATH, logged in (codex login) or configured with an API key. The gpt-6-* models need Codex CLI 0.155 or newer (codex update).

Related MCP server: Hydra

Install

claude mcp add codex-router -s user -- npx -y codex-router-mcp

Or from a clone:

npm install && npm run build
claude mcp add codex-router -s user -- node "/absolute/path/to/dist/index.js"

-s user makes it available in every project. A relative path only resolves from the directory the client was started in, so use an absolute one.

Give the model a policy

Installing the tools gives the model the ability to delegate. It still needs a policy for when. Put something like this in your global agent instructions — without it the model sees twelve tools and no guidance:

You are the tech lead. Codex is an external subagent; you decide what to delegate.

  • Delegate work that is self-contained, mechanical and narrowly scoped — migrations, refactors that follow a pattern, filling in tests, boilerplate. Do it yourself when it needs architectural judgement, context from the conversation, or is small enough that delegating costs more than it saves.

  • Pick the model by difficulty: luna for mechanical work, sol (the default) for everyday work, astra for the hardest problems. Check codex_get_limits before anything large.

  • Use isolation: "worktree" for large, risky or experimental work, or when there is uncommitted work that must not be lost.

  • status: "running" means Codex is still working. Wait with codex_task_status({ taskId, waitSeconds }); never delegate the same task twice.

  • Always review what Codex produced — changed files, diff, then the code — and check it stayed inside scope. Send corrections through codex_continue.

  • When you need a real image, use codex_generate_image, then look at the preview before using it.

  • On quota_exhausted, read remainingWork and finish the task yourself. Never wait for a quota reset and never retry in a loop.

  • If failedFileChanges is present, those files were not written. Do not review them.

CLAUDE.md is the full version this repository runs on (in Polish).

Tools

Tool

Purpose

codex_get_models({ refresh? })

Recommended models by tier, with the efforts the policy allows; read live.

codex_get_limits()

Quota windows normalized by duration, with a delegation verdict.

codex_delegate({ task, workingDirectory, scope?, model?, reasoningEffort?, isolation?, branch?, waitSeconds?, timeoutSeconds? })

Run a task in a fresh Codex thread.

codex_continue({ taskId, instruction, ... })

Follow-up instruction on an existing thread, context intact.

codex_task_status({ taskId?, waitSeconds?, refresh? })

Status and live progress; waitSeconds blocks until the task finishes.

codex_interrupt(taskId)

Stop the in-flight turn — guaranteed to leave running; the thread survives.

codex_review({ workingDirectory?, taskId?, target?, ... })

Read-only review of your work or of a Codex task.

codex_generate_image({ prompt, outputPath?, count?, size?, referenceImages?, ... })

Generate images; files on disk plus a preview.

codex_checkpoints(taskId)

List working-tree snapshots taken around turns.

codex_restore({ taskId, checkpointId, removeUntracked? })

Roll the working tree back.

codex_worktree({ taskId, action, message?, force? })

Commit or remove a task's isolated worktree.

codex_server({ action? })

App-server health and running tasks, or restart a wedged app-server.

Models

The catalogue is always read live from Codex. On top of it sits a small policy that steers work towards three models and caps how hard each may think:

Model

Alias

Tier

Reasoning

Use for

gpt-6-luna

luna, fast

fast

low → xhigh

Mechanical edits, boilerplate, pattern-following tests, image generation. Lowest quota use.

gpt-6-sol

sol, balanced

balanced

low → high

Everyday feature work, bug fixes with a known cause, most review. The default.

gpt-6-astra

astra, frontier, best

frontier

low → high

Hard debugging, algorithm design, cross-cutting changes, second opinions on critical code. Heaviest on quota.

Every model defaults to medium effort. A request above a model's cap is clamped, not refused — an over-eager effort is a cost decision, not an error, and a refusal would waste a round trip — and the clamp is reported in notes. A model outside the policy still runs if named explicitly, with a note steering back. A policy model missing from the live catalogue is reported as unavailable rather than assumed to exist.

Override the policy with AGENT_ROUTER_MODEL_POLICY (a JSON array of { id, aliases, tier, maxEffort, defaultEffort, summary, useFor }; only id is required) and the default with AGENT_ROUTER_DEFAULT_MODEL.

Keeping turns under control

A delegated turn can go wrong in more ways than "it failed": the completion notification can be lost, Codex can block on an approval that a headless router can never give, an interrupt can go unanswered, or the app-server can wedge. None of these may leave a task in running.

  • One outcome, processed once. A turn's result is handled by whoever settles it — Codex's completion, an error, a reconcile, the watchdog or an interrupt — never only by the caller that happened to be waiting. A turn that outlives waitSeconds still finishes and records its result.

  • Waiting instead of polling. codex_task_status({ taskId, waitSeconds }) blocks until the task finishes. Running results carry progress: health (active, quiet, stalled, blocked), running and idle seconds, deadline, current plan step, last message and last command.

  • Calls that fit the client. Many MCP clients cancel a request after 60 s — it is the SDK default. Every blocking call therefore returns within 50 s by default, handing back running and a taskId; the task itself keeps going. While a call blocks, the server sends MCP progress notifications, which clients that honour them use to reset their timeout and show progress. If your client allows longer calls, raise AGENT_ROUTER_DEFAULT_WAIT_SECONDS.

  • Reconciliation. When a thread goes quiet, or a status check finds it silent, the router reads the thread back with thread/read and settles any turn Codex says has ended. A dropped notification cannot wedge a task.

  • Watchdog. Every few seconds it checks running turns: past its time limit (timeoutSeconds, default one hour) a turn is interrupted; blocked on an approval or user input for more than a minute, it is interrupted; silent for three minutes, it is re-read and flagged stalled — but silence alone never kills, because a long command can be quiet and still working.

  • Interrupts that always land. codex_interrupt asks Codex first; if Codex does not confirm within ten seconds it checks the thread's real state, and failing that marks the turn interrupted itself, reported as forced: true. Late completions of a turn that was already settled are ignored, so they can never be mistaken for the next turn on the same thread.

  • Restart. codex_server({ action: "restart" }) replaces a wedged app-server. Turns in flight are reported interrupted; their threads survive on disk and codex_continue resumes them. On Windows the whole process tree is killed, so no orphaned codex.exe is left behind.

Every intervention — reconcile, stall, auto-interrupt, forced stop, restart — is recorded on the task under interventions, with its reason.

Lean results

Results are read by a model, often once per poll, so they are kept small:

  • A running result carries progress and nextStep, not a replay of the request, the command history or a quota snapshot.

  • A finished result carries a compact limits (each window once); codex_get_limits still returns the full detail.

  • The diff is capped at 20,000 characters, the summary at 24,000, each command at 400 and changedFiles at 200 entries, with the truncation stated.

Several sessions at once

Every Claude Code session starts its own router, and they share one state file. Task ids are unique across processes, each router writes back only the tasks it created or changed, and a session can look up a task another session started. A task whose turn is running in a different router shows up as interrupted there: only the router that started a turn can follow it. Finished tasks are kept for 30 days, up to 200 of them.

The router exits when its client closes the connection, taking its app-server with it.

Image generation

codex_generate_image drives Codex's built-in image tool.

codex_generate_image({
  prompt: "A flat app icon: a white paper plane on a deep blue rounded square",
  outputPath: "assets/icon.png",
})
  • The Codex thread runs read-only: generation happens server-side and the router writes the files itself, so images work even where the local sandbox cannot write.

  • Files default to <workingDirectory>/generated-images/<prompt-slug>.png. outputPath may name a file or a directory; a .jpg path is transcoded rather than mislabelled. Existing files are never overwritten unless you pass overwrite: true — a numeric suffix is added and the real path returned.

  • count (1–4), size and transparentBackground are passed to Codex as requirements. size is a hint, not a guarantee.

  • referenceImages attaches local PNG/JPEG files for edits, variations or style.

  • The result lists every image with its path, dimensions, size, the prompt Codex actually used, and how transparent it is. A downscaled JPEG preview is attached so the calling model can look at the result (preview: "full" for the original, "none" for paths only). Transparency is drawn over a checkerboard, so a translucent area reads as transparent rather than as a white shape.

  • Codex sometimes returns an RGBA image whose alpha channel is largely translucent even when no transparency was asked for. That is reported as a warning — the file is left untouched, since there is no single correct way to flatten a broken alpha channel.

  • Defaults to gpt-6-luna at low effort: the turn is a single tool call.

  • Image generation has its own quota. When it runs out the result is quota_exhausted with the reset time — and, unlike a code task, it does not tell the caller to finish the job itself, since it cannot.

Quota

Limits are normalized by window duration, not by slot name:

windowDurationMins

label

60

1h

300

5h

1440

daily

10080

weekly

43200

monthly

other

derived (3h, 2w, 90min, …)

primary is not assumed to be the 5h window — the API is free to put the weekly window there. Per window you get usedPercent, remainingPercent, resetsAt (ISO + epoch), resetsInMinutes and rateLimitReached, plus tightest (the window that actually gates the next turn) and a verdict: ok, low (delegate, with a warning attached) or exhausted.

The handoff

When quota runs out — at preflight or mid-turn — you get this instead of a failure:

{
  "status": "quota_exhausted",
  "taskId": "codex-20260829173437-001",
  "originalTask": "...",
  "threadId": "01a04e96-7474-79a1-a173-8c3cc2919eeb",
  "changedFiles": ["src/a.ts (update)"],
  "summary": "Codex ran out of quota mid-task. ...",
  "remainingWork": "Unfinished plan steps reported by Codex: ...",
  "limits": { "windows": [ ... ] },
  "nextStep": "Do NOT wait for the quota to reset ... finish it yourself."
}

The router never retries and never waits for a reset. Partial work is reported so the calling agent can continue from where Codex stopped. usageLimitExceeded, rateLimitExceeded and sessionBudgetExceeded are all treated as quota.

A non-quota failure returns status: "failed" instead — the two are kept distinct so a compile error is not mistaken for a billing problem.

Isolation: git worktrees

isolation: "worktree" creates a linked worktree on a dedicated branch (agent-router/<taskId> unless you pass branch) and points Codex at it. Your working tree is never touched, whatever the turn does.

Worktrees are created under ~/.agent-router/worktrees/ — outside the repository, so they never appear in git status. If workingDirectory was a subdirectory of the repo, Codex is placed in the matching subdirectory.

For an isolated task, changedFiles and diff are computed against the commit the branch started from, so they show the cumulative result across every turn.

Integration is deliberately manual:

codex_worktree({ taskId, action: "commit" })   # work lands on the task branch
git merge agent-router/<taskId>                # you run this, not the router
codex_worktree({ taskId, action: "remove" })   # clean up

The router never writes to your branch.

Checkpoints

Inside a git repository, the working tree is snapshotted before and after every turn, capturing tracked and untracked files while respecting .gitignore.

The snapshot is built through a throwaway GIT_INDEX_FILE, so it never disturbs what you have staged. git stash create is the obvious primitive but it silently omits untracked files — exactly what a delegated agent tends to produce.

codex_checkpoints(taskId)
codex_restore({ taskId, checkpointId: "cp-1" })

codex_restore rewrites file contents with git restore --worktree --overlay, leaving the index alone and deleting nothing on its own. Files created after the checkpoint are reported as leftoverFiles and only deleted when removeUntracked: true is passed. Every restore first captures the current state and returns it as safetyCheckpoint, so a restore is itself undoable; if that snapshot cannot be taken, the restore is refused.

Snapshots start from a copy of your index, so a file that is tracked although .gitignore matches it is captured like any other.

Checkpoints are dangling commits, not refs, so they never show up in your branches or git log --all. The flip side: git's garbage collection prunes unreferenced commits, by default after two weeks. A restore checks that the commit still exists and refuses, changing nothing, if it does not.

Files inside a git submodule are not captured: the snapshot records only the submodule's commit. Delegate from inside the submodule if Codex should work there with checkpoints.

Review

codex_review uses Codex's native review/start with inline delivery and a read-only sandbox — the reviewer cannot edit what it reviews.

  • Pass workingDirectory to have Codex review your uncommitted work.

  • Pass taskId (optionally with a different model) to have Codex review a previous Codex task. The reviewer is given the original task and its scope, so it also flags work that went out of bounds.

target selects what to review: uncommittedChanges (default), baseBranch, commit, or custom.

Honest change reporting

Codex only tracks edits made through its patch tool. A file it writes with a shell command — which it does often — produces no change event, so Codex's own list can come back empty while the files sit on disk. The router therefore takes changedFiles and diff from the most truthful source available, and says which one in changeSource:

  • worktree — an isolated task, diffed against the commit its branch started from;

  • working-tree — a task in a git repository, diffed between the snapshot taken before its first turn and the one after its last, so shell writes are included;

  • codex-reported — outside git, only what Codex itself tracked. Shell writes may be missing here, and the result says so rather than guessing.

Files Codex reports that git cannot see (ignored paths) are merged in.

The opposite failure matters too: Codex can finish a turn cleanly while every write it attempted was rejected — a misconfigured sandbox does exactly that. The rejected patches are listed under failedFileChanges. In a git repository the router can confirm nothing reached disk and says so, telling the caller not to review files that were never written; a rejected patch whose file was written anyway by a later command is not reported as a failure. Outside git it cannot tell, and the warning asks the caller to check the files instead of claiming either way.

Known issue: the Codex sandbox on Windows

AGENT_ROUTER_SANDBOX defaults to workspace-write. On Windows that sandbox needs a helper binary, codex-windows-sandbox-setup.exe, that some Codex installations do not ship (still the case in 0.155.1). When it is missing every write is silently rejected.

Reproduce it without this server:

codex sandbox cmd /c "echo hi > test.txt"

A healthy install writes the file; a broken one prints orchestrator_helper_launch_failed: ... program not found. Note that windowsSandbox/readiness still reports ready, so it does not catch this.

Repair the Codex installation if you can. AGENT_ROUTER_SANDBOX=danger-full-access works around it but removes the sandbox entirely — pair it with isolation: "worktree" at minimum. Image generation and review are unaffected: both run read-only.

Updating Codex on Windows

If codex update fails with tar (child): Cannot connect to C: resolve failed, it was run from Git Bash, where tar is GNU tar and reads C:\… as a remote host. If it fails with Get-FileHash is not recognized, it was run from PowerShell 7, whose module path breaks the Windows PowerShell 5.1 installer it spawns. Run it from PowerShell with the module path reset:

Remove-Item Env:PSModulePath; codex update

Configuration

All optional, set as environment variables on the MCP server entry. Values are validated at startup: an unknown AGENT_ROUTER_ISOLATION or AGENT_ROUTER_SANDBOX, or a number out of range, stops the server with an error naming the variable, rather than being silently ignored.

Variable

Default

Purpose

AGENT_ROUTER_CODEX_BIN

codex

Executable to spawn.

AGENT_ROUTER_CODEX_ARGS

app-server

Args; a JSON array is accepted for paths with spaces.

AGENT_ROUTER_SANDBOX

workspace-write

Sandbox for delegations (reviews and images are always read-only).

AGENT_ROUTER_APPROVAL_POLICY

never

Codex runs headless; nobody can answer prompts.

AGENT_ROUTER_AUTO_APPROVE

false

Accept an approval request that arrives anyway.

AGENT_ROUTER_DEFAULT_MODEL

gpt-6-sol

Model used when a caller names none.

AGENT_ROUTER_MODEL_POLICY

built in

JSON array replacing the model policy.

AGENT_ROUTER_IMAGE_MODEL

gpt-6-luna

Model for image generation.

AGENT_ROUTER_IMAGE_EFFORT

low

Reasoning effort for image generation.

AGENT_ROUTER_IMAGE_PREVIEW_MAX_EDGE

768

Long edge of the preview, in pixels.

AGENT_ROUTER_ISOLATION

none

Default isolation: none or worktree.

AGENT_ROUTER_WORKTREE_ROOT

~/.agent-router/worktrees

Where linked worktrees are created.

AGENT_ROUTER_CHECKPOINTS

on

Set to off to stop snapshotting around turns.

AGENT_ROUTER_QUOTA_PREFLIGHT

on

Set to off to skip the pre-delegation quota check.

AGENT_ROUTER_QUOTA_LOW_PERCENT

15

Remaining percent that triggers the low warning.

AGENT_ROUTER_QUOTA_BLOCK_PERCENT

2

Remaining percent that blocks delegation.

AGENT_ROUTER_DEFAULT_WAIT_SECONDS

50

Blocking window before returning running; kept under the common 60 s client timeout.

AGENT_ROUTER_MAX_WAIT_SECONDS

1800

Ceiling on waitSeconds.

AGENT_ROUTER_PROGRESS_INTERVAL_SECONDS

10

How often a blocking call sends an MCP progress notification.

AGENT_ROUTER_TURN_TIMEOUT_SECONDS

3600

Hard limit per turn; 0 disables it.

AGENT_ROUTER_STALL_SECONDS

180

Silence after which a turn is re-read and flagged stalled.

AGENT_ROUTER_BLOCKED_TIMEOUT_SECONDS

60

How long a turn may wait on an approval or input before it is interrupted.

AGENT_ROUTER_WATCHDOG_INTERVAL_SECONDS

15

How often the watchdog checks running turns.

AGENT_ROUTER_INTERRUPT_GRACE_SECONDS

10

How long an interrupt waits for Codex to confirm before forcing.

AGENT_ROUTER_STATE_FILE

~/.agent-router/tasks.json

Task metadata file, written atomically and shared safely between sessions.

AGENT_ROUTER_DEBUG

false

Mirror app-server stderr and protocol traffic to stderr.

Tests

npm test

275 assertions. The real MCP server is booted over stdio but pointed at test/fake-app-server.mjs instead of codex app-server, so the whole router — JSON-RPC client, notification wiring, turn supervision, quota and model policy, image handling, task store, git plumbing — is exercised without a Codex account and without spending quota. The fake reproduces the failure modes that matter: a completion that arrives after the caller stopped waiting, a completion that never arrives, an interrupt that is ignored, a turn blocked on an approval, a start reply that arrives after its turn was already settled and a newer one begun, two image generations racing for one file name, a file written through the shell that Codex never reports, and an image with a broken alpha channel. Git cases run against throwaway repositories. This is what CI runs on Linux, macOS and Windows.

npm run smoke

Read-only check against the real codex app-server: prints the recommended models, the Codex version and current limits. Starts no turn, so it spends no quota.

Stability and terms

This server talks to codex app-server, which the Codex CLI marks [experimental]. Its protocol has no public documentation or stability guarantee; the type definitions in src/protocol.ts mirror only the subset used here and were derived from codex app-server generate-ts. A Codex release can change it. Regenerate and re-check if something breaks:

codex app-server generate-ts --out ./generated-ts

On terms: this server does not fork Codex, does not touch authentication, and does not reimplement any OpenAI client. It spawns the official codex binary you installed and logged into yourself. OpenAI has stated that the Codex CLI is Apache-2.0 and that forking is permitted, but has not clarified whether third-party tools driving a ChatGPT-plan session are covered by the Terms of Use, and their docs recommend API keys for automation. If you are automating heavily, or building anything commercial on this, use an API key and take your own legal advice.

License

MIT — see LICENSE.

Available Tools

10 tools
codex_checkpointsList task checkpointsA

List the working-tree snapshots taken around a task's turns. Each checkpoint captures tracked and untracked files without touching the user's index, and can be restored with codex_restore. Requires the working directory to be inside a git repository.

ParametersJSON Schema
NameRequiredDescriptionDefault
taskIdYesTask whose checkpoints should be listed.

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the behavioral disclosure burden. It adds useful transparency by stating that checkpoints capture tracked and untracked files "without touching the user's index" and that restoration happens via codex_restore. It does not detail edge cases like empty git repos or corrupted checkpoints, but for a listing tool this is adequate.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is three sentences with no wasted words. It front-loads the core action, then adds important behavioral details and the prerequisite. Every sentence contributes necessary information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple one-parameter listing tool, the description is complete enough: it states what is listed, what checkpoints capture, that the operation does not touch the index, how to restore, and the git requirement. The absence of an output schema is acceptable because the listing semantics are clear from the description and tool name.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and the single required parameter taskId is already documented in the schema as "Task whose checkpoints should be listed." The description adds contextual framing around "a task's turns" but no additional parameter-level semantics, so baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: "List the working-tree snapshots taken around a task's turns." It clearly distinguishes itself from the sibling codex_restore by noting that checkpoints "can be restored with codex_restore," making the tool's listing-only role unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear context for when to use the tool: to list checkpoints for a task, and it points to codex_restore as the restoration alternative. It also gives a prerequisite (must be inside a git repository). It stops short of explicitly stating when not to use it versus alternatives, so it is not a perfect 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

codex_continueContinue a Codex taskA

Send a follow-up instruction into an existing Codex thread, keeping all of its prior context. Use it to iterate on review feedback instead of re-delegating from scratch.

ParametersJSON Schema
NameRequiredDescriptionDefault
modelNoOptionally switch model for this turn onward.
taskIdYestaskId returned by a previous codex_delegate.
instructionYesThe follow-up instruction for Codex.
waitSecondsNoHow long to block before returning a pollable taskId.
reasoningEffortNoOptionally switch reasoning effort for this turn onward.

TDQS

A4.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are present, so the description carries the full burden. It does disclose that this is a mutating follow-up that preserves prior thread context, but it doesn't disclose the asynchronous execution model, that a pollable taskId is returned, or how invalid/expired threads are handled. Useful but incomplete behavioral coverage.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences with zero waste: purpose and identifying trait are front-loaded, and the usage guidance follows immediately. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Complete enough to invoke correctly: required params, purpose, and the continue-vs-delegate choice are clear, and the waitSeconds schema hint covers the pollable-taskId return flow. However, with no output schema and no annotations, the description could do more to state how results are obtained (e.g., polling via codex_task_status).

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all five parameters (taskId, instruction, model, waitSeconds, reasoningEffort) with meaning. The description only reinforces that instruction is a follow-up and taskId refers to an existing thread — baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific action ('Send a follow-up instruction into an existing Codex thread') and the defining trait (keeping all prior context). It differentiates from the sibling codex_delegate by explicitly rejecting re-delegating from scratch, so an agent can tell them apart without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives an explicit when-to-use ('Use it to iterate on review feedback') and names the alternative behavior to avoid ('instead of re-delegating from scratch'), which maps to codex_delegate. The selection condition is concrete and actionable.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

codex_delegateDelegate a task to CodexA

Hand a self-contained coding task to Codex as a subagent. Starts a fresh Codex thread, runs the task, and returns the result plus the files it changed. Checks quota first: if Codex has no quota left it returns status 'quota_exhausted' with a handoff so you can finish the work yourself instead of waiting for a reset.

ParametersJSON Schema
NameRequiredDescriptionDefault
taskYesThe full task instruction for Codex. Be specific and self-contained.
modelNoCodex model id from codex_get_models. Omit for the account default.
scopeNoScope boundary: what Codex may and may not touch.
branchNoBranch name for the worktree. Default: "agent-router/<taskId>".
isolationNo"worktree" runs Codex in a dedicated git worktree on its own branch, so a bad turn cannot touch the user's working tree. "none" edits in place. Default: none.
waitSecondsNoHow long to block before returning a pollable taskId (default 240s, max 1800s).
reasoningEffortNoReasoning effort supported by the chosen model (see codex_get_models).
workingDirectoryYesAbsolute path Codex should treat as its working directory.

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden, and it delivers: it discloses that the tool checks quota before running, starts a fresh thread, returns changed files, and returns status 'quota_exhausted' with a handoff on failure. These are behavioral traits an agent cannot infer from the schema alone. It stops short of describing error behavior beyond quota or cost/long-running implications, but the disclosed traits are substantive.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences with zero fluff: the core purpose is front-loaded first, followed by execution behavior, then the key edge case. Every sentence earns its place — the quota-exhausted sentence describes a real decision-relevant scenario for the agent rather than filler. Appropriately sized for an 8-parameter tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with no annotations and no output schema, the description explains the essential return behavior (result plus changed files) and the most likely failure mode (quota exhaustion with handoff). The pollable taskId return is only hinted at through the waitSeconds parameter description, and the full return shape beyond 'result plus files' is underspecified. Given the moderate complexity — spawning a subagent that edits files — this is slightly better than adequate but not fully complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all 8 parameters, establishing the baseline of 3. The description adds no parameter-level meaning beyond what the schema provides — it mentions output (files changed) and quota, but neither maps to a specific parameter. This is adequate because the schema carries the parameter documentation burden.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb-resource pair ('Hand a self-contained coding task to Codex as a subagent') and describes a concrete outcome: runs the task and returns the result plus changed files. 'Starts a fresh Codex thread' distinguishes this from codex_continue, and the quota-first behavior distinguishes it from codex_get_limits. An agent can tell what this tool does and roughly how it differs from nearby siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The qualifier 'self-contained coding task' implies when this tool is appropriate, and the quota_exhausted handoff describes a fallback action. However, the description never explicitly names alternatives or exclusion conditions (e.g., 'use codex_continue for ongoing conversations, codex_review for reviewing changes'), which is a real gap given nine closely related siblings. Usage context is present but only implied, not stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

codex_get_limitsRead Codex rate limitsA

Read Codex usage limits, normalized by window duration (300 min -> '5h', 10080 min -> 'weekly'), with usedPercent, remainingPercent, resetsAt and rateLimitReached per window, plus a delegation verdict. Check this before delegating anything large.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden. It discloses normalization behavior with concrete examples (300 min -> '5h', 10080 min -> 'weekly'), lists the returned fields, and mentions a delegation verdict. While it doesn't detail error behavior or auth requirements, it is transparent enough for a read-only limits check.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two dense sentences with no filler. The first sentence front-loads the operation and return details; the second gives actionable guidance. Every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a parameterless read tool with no output schema, the description is complete: it explains normalization, names the returned fields, signals read-only intent, and provides usage context. An agent has enough information to invoke and interpret the result correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters and schema coverage is 100%, so there is no parameter information the description needs to add. The baseline for 0-parameter tools is 4, and the description appropriately focuses on output and usage rather than parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Read') and resource ('Codex usage limits'), and details what the tool returns: normalized windows, percentages, reset times, and a delegation verdict. It also distinguishes itself from siblings like codex_delegate by framing this as a pre-flight check.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly states when to use the tool: 'Check this before delegating anything large.' This gives clear usage context relative to the delegation workflow, though it does not explicitly name alternatives or provide when-not-to-use cases.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

codex_get_modelsList Codex modelsA

List the Codex models available to this account, with the reasoning-effort levels each one supports. Read live from the Codex model catalogue — never hardcoded. Use it before codex_delegate when you want to match model strength to task difficulty.

ParametersJSON Schema
NameRequiredDescriptionDefault
refreshNoBypass the 60s catalogue cache and re-read from Codex.

TDQS

A4.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the burden of behavioral disclosure. It adds useful context by saying it reads live from the Codex model catalogue and is never hardcoded, but it does not mention the 60-second cache behavior, and 'read live' slightly overstates freshness for a tool whose refresh parameter implies cached results by default.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, each earning its place: the main action, the data-source caveat, and the usage routing. It is front-loaded with the core purpose and contains no fluff.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only list utility with one optional parameter already fully documented in the schema, the description covers purpose, account scope, output content (models plus reasoning-effort levels), and when to use it. The output expectation is clear even without an output schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%; the refresh parameter is already documented as bypassing the 60s catalogue cache and re-reading from Codex. The description adds no parameter-level meaning, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('List'), a clear resource ('Codex models available to this account'), and a key included detail ('reasoning-effort levels each one supports'). This clearly distinguishes it from sibling tools like codex_get_limits or codex_delegate without needing to inspect their schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly tells the agent when to call it: before codex_delegate, when matching model strength to task difficulty. This gives concrete, decision-relevant routing to a sibling tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

codex_interruptInterrupt a Codex taskA

Stop the turn Codex is currently running for a task. The thread survives, so codex_continue can pick it back up.

ParametersJSON Schema
NameRequiredDescriptionDefault
taskIdYesTask whose in-flight turn should be stopped.

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It adds important non-obvious context: the thread survives interruption and can be resumed via codex_continue. It does not cover edge cases like interrupting when no turn is running, but the core behavior is transparent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two short sentences with no filler. The primary action is front-loaded, and the key consequence (thread survives, resumable) is delivered in the second sentence efficiently.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter tool with no output schema, the description covers the essential facts: what is interrupted and what happens afterward. It does not explain behavior when there is no in-flight turn, but given the tool's simplicity and the sibling context, it is sufficiently complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and the parameter taskId is already clearly described as 'Task whose in-flight turn should be stopped.' The tool description adds no new parameter detail, which is acceptable since the schema fully documents the only parameter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Stop') and resource ('the turn Codex is currently running for a task'), which clearly identifies the tool's action. It also distinguishes itself from related siblings like codex_continue by noting the thread survives, making the purpose unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implicitly provides usage context by stating that codex_continue can pick the thread back up, which signals this tool is for pausing rather than terminating a task. It does not explicitly list exclusions or compare with siblings like codex_task_status, but the context is clear enough for most agents.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

codex_restoreRestore a checkpointA

Roll the working tree back to a checkpoint — use it when Codex made things worse. This overwrites files on disk, so confirm with the user before calling it unless they already asked for the rollback. The pre-restore state is always captured as a new checkpoint first, so the operation is itself undoable.

ParametersJSON Schema
NameRequiredDescriptionDefault
taskIdYesTask the checkpoint belongs to.
checkpointIdYesCheckpoint id from codex_checkpoints, e.g. "cp-1".
removeUntrackedNoAlso delete files created after the checkpoint. Default false — they are reported as leftovers instead.

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It states that files on disk are overwritten (destructive), requires user confirmation, and reveals that the pre-restore state is captured as a new checkpoint, making the operation undoable. This is thorough and directly useful for safe invocation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences with zero waste: purpose/trigger first, then the destructive warning and consent requirement, then the undo guarantee. Every sentence earns its place and the most critical safety information is front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive mutation tool with no annotations and no output schema, the description covers the essential context: what it does, when to use it, side effects, user-consent requirements, and undoability. Schema handles parameter details. Nothing critical is missing for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and the schema already explains taskId, checkpointId, and removeUntracked. The description reinforces the checkpoint concept but adds no parameter-specific syntax or format detail beyond the schema. Baseline 3 is appropriate when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific action ('roll the working tree back to a checkpoint') and a clear trigger ('use it when Codex made things worse'). This distinguishes it from siblings like codex_checkpoints, which lists checkpoints, and codex_review, which reviews work. The verb and resource are unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives an explicit when-to-use signal ('when Codex made things worse') and a pre-call requirement: confirm with the user unless they already asked for the rollback. It doesn't explicitly name alternatives or state when not to use it, but the context is clear enough for an agent to route correctly.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

codex_reviewHave Codex review codeA

Ask Codex to review changes and report findings. Use it on YOUR OWN work for a second opinion before you ship, or on a Codex task's output with a different model. Codex reviews read-only and changes nothing. Returns a review task you can poll or extend with codex_continue.

ParametersJSON Schema
NameRequiredDescriptionDefault
modelNoCodex model to review with (see codex_get_models).
branchNoBase branch, for target "baseBranch".
commitNoCommit sha, for target "commit".
targetNoWhat to review. Default: uncommittedChanges.
taskIdNoReview the work of this Codex task, in its own directory or worktree. Combine with a different model for a cross-model second opinion.
waitSecondsNoHow long to block before returning a pollable taskId.
instructionsNoWhat to focus on. Required for target "custom"; otherwise added as extra guidance for the reviewer.
reasoningEffortNoReasoning effort for the review.
workingDirectoryNoAbsolute path to review in. Required unless taskId is given.

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of disclosing side effects, and it does so clearly: 'Codex reviews read-only and changes nothing.' It also discloses the return behavior—a review task that can be polled or extended—which is important since there is no output schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is three tight sentences with no filler. The core action is front-loaded, the primary use cases follow immediately, and the safety guarantee and return contract are packed into short, scannable statements.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no annotations and no output schema, the description covers the essential operational context: read-only behavior, return format as a pollable task, and how to chain with codex_continue. It could slightly expand on what 'report findings' means or how results are retrieved, but the schema and sibling names make the overall workflow reasonably complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and all nine parameters have meaningful descriptions in the schema itself. The tool description adds context around parameters like taskId and model ('on a Codex task's output with a different model'), but the schema already does the heavy lifting for individual parameter meaning.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Ask Codex to review changes and report findings.' It clearly positions the tool as a review action, distinct from the sibling tools that manage checkpoints, worktrees, or delegate tasks, and it explicitly notes the returned artifact is a pollable review task.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives concrete when-to-use guidance: use on your own work before shipping, or on a Codex task's output with a different model for a cross-model second opinion. It also tells the agent how to continue after the call by polling or extending with codex_continue, though it does not explicitly say when not to use this tool or name alternative review workflows.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

codex_task_statusCheck a Codex taskA

Read the current state of a delegated task: status, model, reasoning effort, changed files, commands run, plan, diff, worktree, checkpoints, and timestamps. Poll this when codex_delegate returned status 'running'. Omit taskId to list all known tasks.

ParametersJSON Schema
NameRequiredDescriptionDefault
taskIdNoTask to inspect. Omit to list every task this router knows about.

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the burden of behavioral disclosure. It clearly states this is a read operation, lists the fields the caller will see, and clarifies the list-all behavior when taskId is omitted. For a read-only status tool, this is sufficient even without detailed side-effect or rate-limit notes.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences with no filler. The first sentence front-loads the operation and the full set of returned state fields; the second provides the polling trigger and the list-all variant. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with one optional parameter, no output schema, and no annotations, the description provides all essential context: what is inspected, what fields will be returned, when to poll, and how to list all tasks. Nothing an agent needs to correctly call this tool is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and the schema already explains that taskId is the task to inspect and that omitting it lists all known tasks. The description repeats this omission behavior without adding new syntactic or format details, so the baseline score of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses the specific verb 'Read' and clearly identifies the resource: the current state of a delegated Codex task. It enumerates exactly what state is returned (status, model, reasoning effort, changed files, commands run, plan, diff, worktree, checkpoints, timestamps), which distinguishes it from sibling tools like codex_delegate or codex_interrupt.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives an explicit trigger condition: 'Poll this when codex_delegate returned status running.' It also explains the omit-taskId behavior for listing all tasks. It does not explicitly enumerate alternatives or state when not to use this tool, but the usage context is clear and actionable.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

codex_worktreeManage a task's worktreeA

Commit or remove the isolated git worktree of a task delegated with isolation "worktree". "commit" records the work on the task branch and reports the merge command; the router never merges into the user branch itself. "remove" tears the worktree down and refuses to discard uncommitted work unless forced.

ParametersJSON Schema
NameRequiredDescriptionDefault
forceNoFor "remove": discard uncommitted changes in the worktree.
actionYes"commit" the work onto the task branch, or "remove" the worktree.
taskIdYesTask whose worktree to act on.
messageNoCommit message. Defaults to the task description.

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full disclosure burden and does well: it explains that 'commit' records work and reports a merge command without merging into the user branch, and that 'remove' tears down the worktree and refuses to discard uncommitted work unless forced. It leaves out post-commit worktree state and error conditions, but the critical side effects and safety behavior are transparent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences with no filler: the first states the overall purpose, the second explains commit semantics, the third explains remove semantics including the forced flag. Each sentence earns its place and the key scoping constraint is front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description provides the essential operational context for a two-action tool: target resource, per-action behavior, and safety defaults. It does not describe the full return value shape or error handling, and there is no output schema to supplement, but the available information is sufficient for selection and basic invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. The description adds meaning beyond the schema by explaining what 'commit' produces (a merge command) and that 'force' overrides a refusal to discard uncommitted work. This helps the agent reason about when to use the optional force and message params, though the message param is only documented in the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names two specific actions (commit, remove) applied to a specific resource (the task's isolated git worktree), and scopes it to tasks delegated with isolation "worktree". This clearly distinguishes it from sibling tools like codex_restore or codex_checkpoints, which handle different task lifecycle concerns.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description clearly states the tool applies only to tasks delegated with isolation "worktree", giving a concrete condition for use. It does not explicitly name alternatives or list when-not-to-use scenarios, but the isolation qualifier provides enough guidance for an agent to select it correctly.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 10 tool updatesv0.1.0
    • First observedcodex_checkpoints
    • First observedcodex_continue
    • First observedcodex_delegate
    • First observedcodex_get_limits
    • First observedcodex_get_models
    • First observedcodex_interrupt
    • First observedcodex_restore
    • First observedcodex_review
    • First observedcodex_task_status
    • First observedcodex_worktree

TDQS

A4.3/5.0

Scored across 10 tools

Disambiguation5/5

Each tool targets a clearly distinct action or resource: delegation, continuation, status polling, interrupting, reviewing, checkpoint listing/restoring, worktree management, and preflight checks for models/limits. Even the closely related checkpoint and worktree tools are differentiated by their git semantics and descriptions.

Naming Consistency3/5

All tools share the codex_ prefix, but the action pattern is mixed: codex_delegate, codex_restore, codex_continue, codex_interrupt, and codex_review are bare verbs, codex_get_models and codex_get_limits use get_, while codex_checkpoints, codex_worktree, and codex_task_status are bare nouns. The naming is readable but not predictable enough for an agent to guess tool names confidently.

Tool Count5/5

Ten tools is well-scoped for a Codex routing and task-lifecycle server. Each tool covers a necessary phase—preflight checks, delegation, follow-up, status, interruption, review, checkpoint rollback, and worktree handling—without redundant entries.

Completeness5/5

The tool surface covers the full delegation lifecycle: checking models and limits before starting, delegating, continuing, polling status, interrupting, reviewing, and rolling back via checkpoints or worktrees. No obvious dead ends or critical missing operations stand out for the router's stated purpose.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    A
    quality
    A
    maintenance
    Enables Claude Code to delegate coding tasks—investigation, review, long-running background jobs, and optional file edits—to a locally installed OpenAI Codex CLI, with the model and reasoning effort chosen per task and read-only execution by default. Follow-up requests reuse Codex's existing context, and write-enabled delegations can be confined to a managed git worktree so the user's own checkout stays untouched.
    8
    204 npm
    MIT