Skip to main content
Glama
Ryen-LTC

codex-subagent-for-claude

codex-subagent-for-claude

English | 简体中文

codex-subagent-for-claude MCP server ci

Use Codex as a subagent inside Claude Code: Claude delegates a task, Codex works on it in the background, and the result comes back into Claude's conversation on its own.

Two Python files, no third-party dependencies. Currently verified on Windows 11 only.

Features

  • Async dispatch: codex_spawn returns immediately and Claude carries on. Multiple tasks run in parallel inside one long-lived Codex process, no cold start per task

  • Results delivered automatically: when Codex finishes, the result shows up in the conversation on Claude's next action, no polling. Claude collects any outstanding results before ending its turn

  • Mid-flight control: codex_steer sends a follow-up instruction and Codex changes course after its current step; codex_interrupt stops a task and keeps what's done; pass thread_id to continue a conversation

  • Code review: codex_review runs Codex's built-in review mode — read-only, defect-focused, with file and line numbers

  • Per-window isolation: each Claude Code window gets its own Codex process

  • Independent model settings: subagents default to gpt-6-sol / reasoning effort high / standard service tier, unaffected by the Codex desktop chat settings; change the defaults via environment variables or override per task

Related MCP server: Code Worker MCP

How it works

Claude Code ──MCP──> server.py ──JSON-RPC──> codex app-server (long-lived child process)
Claude Code ──hook──> hook.py: injects finished results back into the session that dispatched them

Requirements

  • Windows 11, Python 3.10+ (python on PATH), Claude Code

  • A logged-in Codex. By default the codex.exe bundled with the Codex desktop app is used (under %LOCALAPPDATA%\OpenAI\Codex\bin\, updated with the app); the app itself does not need to be running. The npm package @openai/codex also works, or point CODEX_SUB_BIN at any binary. Both share the login and config in ~/.codex

Install

Clone anywhere; <install-dir> below stands for that path. You can also just hand these steps to Claude or Codex.

  1. Register the MCP server:

claude mcp add codex-sub -s user -- python "<install-dir>\server.py"
  1. Merge the hooks into ~/.claude/settings.json (escape backslashes as \\ in JSON):

{
  "hooks": {
    "UserPromptSubmit": [{"hooks": [{"type": "command", "command": "python \"<install-dir>\\hook.py\" UserPromptSubmit", "timeout": 10}]}],
    "PostToolUse":      [{"hooks": [{"type": "command", "command": "python \"<install-dir>\\hook.py\" PostToolUse", "timeout": 10}]}],
    "Stop":             [{"hooks": [{"type": "command", "command": "python \"<install-dir>\\hook.py\" Stop", "timeout": 10}]}]
  }
}
  1. Optional: add "mcp__codex-sub" to permissions.allow, otherwise every call asks for confirmation.

Restart your Claude Code session. To try it without touching global config: claude --plugin-dir "<install-dir>".

Usage

Just talk to Claude; it knows when to delegate:

Hand the auth module refactor to Codex. Meanwhile fix the API tests yourself, then merge when it's done.

Have Codex review what I just changed.

Run two Codex tasks in parallel: one fixes add, one fixes mul, each in its own worktree.

Tool

Purpose

codex_spawn

Dispatch a task. wait_s waits a while first, thread_id continues a thread, worktree runs in a separate branch, sandbox / model / effort override per task

codex_review

Code review: uncommitted changes / against a base branch / a commit / custom instructions

codex_wait

Wait for results; a timeout only returns the current state, the task keeps running

codex_status

List tasks, or show one task's details and full output

codex_steer

Send a follow-up instruction to a running task

codex_interrupt

Interrupt, keeping what has been produced

  • Tasks are ephemeral by default and don't appear in Codex history; subagents can use the skills and plugins in ~/.codex

  • For fire-and-forget work pass detach=true; Claude won't wait for it when ending its turn

Permissions

Subagents run without a sandbox (sandbox=danger-full-access) and without approvals (approvalPolicy=never) by default: they can edit any file and run any command without stopping to ask, and the occasional question that would need a human answer is skipped so the task never hangs waiting for one. The reason: Codex's Windows sandbox runs commands under a restricted account that cannot see tools installed in the user profile (Python, pnpm, uv, …), so a sandboxed subagent can't even run tests.

If that's not acceptable: set CODEX_SUB_SANDBOX=read-only, or pass sandbox=read-only per task. codex_review is always read-only.

Configuration

Environment variable

Default

CODEX_SUB_BIN

auto-detected

CODEX_SUB_SANDBOX

danger-full-access

CODEX_SUB_MODEL / CODEX_SUB_EFFORT / CODEX_SUB_SERVICE_TIER

gpt-6-sol / high / default

CODEX_SUB_STATE_DIR

%LOCALAPPDATA%\codex-sub (logs and session state; server and hook must use the same value)

Notes

  • Relies on some experimental codex app-server APIs; Codex updates may change the protocol. Verified with Codex 0.157 / 0.158

  • Each Claude session keeps one app-server process alive (~150 MB); it exits with the session

Tests: python tests/handshake.py (no Codex needed, seconds) and python tests/smoke.py (calls Codex for real, about 3 minutes).

Uninstall: claude mcp remove codex-sub -s user, then remove the three hooks from settings.json.

License

MIT

Available Tools

6 tools
codex_interruptA
Idempotent

中断一个正在运行的任务:Codex 收到后立即停止,已完成的部分输出和文件改动保留,任务状态变为 interrupted,结果可用 codex_status 查看。对已结束的任务调用无副作用,直接返回其结果。被中断的 thread 之后仍可用 codex_spawn(thread_id=...) 续问。只想改方向不想停,用 codex_steer。

ParametersJSON Schema
NameRequiredDescriptionDefault
idYes任务的 8 位 id(codex_spawn / codex_review 返回的)

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare the mutation/idempotency profile, but the description discloses the concrete behavioral consequences: Codex halts immediately, partial output and file changes are retained, status transitions to interrupted, and the interrupted thread remains resumable via codex_spawn(thread_id=...). It also reinforces idempotentHint explicitly for already-finished tasks.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Every clause earns its place: the core action is front-loaded, followed by state/retention behavior, the no-op case, resumability, and the sibling alternative. No filler despite the density.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-parameter mutation tool with no output schema, the description supplies everything an agent needs: the state transition, side-effect profile, retention semantics, error-case behavior, and where to read results. Nothing material is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There is a single parameter with 100% schema description coverage, and the schema already explains it as the 8-digit task id returned by codex_spawn/codex_review. The description adds no syntax or format detail beyond the schema, so the baseline 3 for fully-covered schemas applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource (interrupt a running task) and immediately names the relevant siblings — codex_status for viewing results and codex_steer for redirecting without stopping. An agent can distinguish this from all five sibling tools without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly covers when to use it (a running task), what happens when the condition does not hold (calling on a finished task is side-effect-free and returns its result), and the alternative when you do not actually want to stop (codex_steer). This is when/when-not/alternative guidance in full.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

codex_reviewA
Read-only

用 Codex 内置的代码审查模式(等价于 codex review)审查改动,固定只读、以找缺陷为主,返回按严重度排序、带文件行号的问题列表。scope=uncommitted 审工作区未提交改动;base_branch 审当前分支相对基准分支的改动(value=分支名);commit 审某个提交(value=sha);custom 按自定义说明审(value=说明)。异步/等待语义同 codex_spawn,也返回任务 id。想让 Codex 改代码用 codex_spawn,只想让它挑毛病用这个。

ParametersJSON Schema
NameRequiredDescriptionDefault
cwdNogit 仓库目录(绝对路径),默认当前目录
modelNo仅用户明确指定时填
scopeNo审查范围,默认 uncommitted
valueNoscope 对应的值:分支名 / commit sha / 自定义审查说明;uncommitted 不需要
detachNotrue = 不关心结果,Claude 结束回合时不等它
wait_sNo先同步等待的秒数,审查通常 1~3 分钟;0 立即返回

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false, and the description reinforces that with 固定只读 and adds value beyond them: defect-focused output, severity-sorted results with file:line, async/wait semantics identical to codex_spawn, and a task id return. Timing context (审查通常 1~3 分钟) is also supplied.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the core purpose and return shape, then enumerates scopes, then the sibling routing. Dense but every clause carries information; only the repeated scope enumeration overlaps slightly with the enum in the schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description correctly compensates by describing the return (severity-sorted, file:line issue list) and the async task-id behavior. It is complete enough for an agent to call it correctly without opening the schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds real meaning: the scope→value mapping (分支名 / sha / 自定义说明) and the note that uncommitted needs no value. It does not explain detach's relationship to the async task id, leaving one minor gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource (审查改动 via Codex 内置审查模式), the mode's character (固定只读、以找缺陷为主), and the return shape (按严重度排序、带文件行号的问题列表). It explicitly distinguishes itself from the sibling codex_spawn by intent (改代码 vs 挑毛病).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit routing: use codex_spawn to change code, use this only to find faults. It also breaks down when to pick each scope value (uncommitted/base_branch/commit/custom) with the value semantics for each, so the agent can select without inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

codex_spawnA
Destructive

把一个任务派给 Codex 子代理,立即返回 8 位任务 id(wait_s>0 时先等一会儿,等到就直接给结果)。可以连续派多个,它们在同一个常驻 Codex 进程里并行执行。任务完成后结果会自动送回本会话;也可以用 codex_wait 主动等、codex_status 看进度、codex_steer 中途改方向、codex_interrupt 中断。传 thread_id 可在旧任务的上下文上续问;该 thread 上不能有仍在运行的任务。默认无沙箱、不审批:Codex 会直接改文件、跑命令。

ParametersJSON Schema
NameRequiredDescriptionDefault
cwdNo工作目录(绝对路径),默认当前目录
modelNo仅用户明确指定时填,默认 gpt-6-sol
detachNotrue = 不关心结果的后台任务;Claude 结束回合时不会因为它被拦下
effortNo推理强度,默认 high;简单任务可降到 low/medium
promptYes任务说明:目标、改动范围、完成标准(要跑什么验证)、返回格式(结论 + 文件:行号,不贴大段代码)
wait_sNo先同步等待这么多秒;估计几分钟内能完成的任务填 120~300 可直接拿到结果,长任务填 0 立即返回
persistNo把会话持久化到 Codex 历史,以便下次 Claude 会话还能用 thread_id 续问;默认不持久化
sandboxNo默认 danger-full-access(无沙箱);只读分析用 read-only
worktreeNo在独立 git worktree(分支 codex-sub/<id>)里执行,多任务并行改同一仓库时用,完成后自行合并分支
thread_idNo续问:在之前某个任务的 thread 上继续,保留全部上下文,只发增量指令

TDQS

A4.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare destructiveHint=true, openWorldHint=true and readOnlyHint=false, and the description corroborates with the critical behavioral warning that by default there is no sandbox and no approval, so Codex will directly modify files and run commands. It also discloses the detach semantics and the thread_id precondition (no still-running task on that thread), which annotations do not convey. It stops short of covering failure modes or rate/concurrency limits.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single dense paragraph with no filler; the core capability and return value are front-loaded, followed by the sibling routing and then the sandbox warning. It is slightly over-packed with the five sibling name-drops in one breath, but every sentence carries operational information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 10-parameter mutating tool with no output schema, the description still explains the return value (8-digit id), how results arrive back to the session, the persistence option, and the default no-sandbox execution posture. Nothing an agent needs in order to call it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3, but the description adds meaning beyond the schema: the detach flag's effect on turn blocking, the wait_s heuristic on which values yield an immediate result, and the constraint that a reused thread_id cannot have a running task. These are genuine additions rather than restatements of the schema text.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource (dispatch a task to a Codex subagent) and the immediate return value (8-digit task id), plus the wait_s shortcut. It names the sibling tools codex_wait, codex_status, codex_steer and codex_interrupt, so an agent can distinguish this spawn tool from its siblings without opening another schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit when-to-use guidance: multiple concurrent dispatches are allowed and run in parallel in one resident process; task results auto-return to the session; thread_id is for continuing an old task's context but the thread must have no running task. It also surfaces the alternatives for monitoring/steering rather than leaving them to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

codex_statusA
Read-onlyIdempotent

立即返回,不等待。不传 ids:一行一个列出本会话全部任务及状态。传 ids:看详情——运行中的显示最近动作、改动文件、输出预览;已结束的显示完整结果(超过 6000 字截断,full=true 看全文)。被 codex_interrupt 中断的任务,其已产出的部分也从这里取。

ParametersJSON Schema
NameRequiredDescriptionDefault
idsNo任务 id 列表;省略 = 只列清单
fullNotrue = 不截断,返回完整结果文本

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint and idempotentHint, so safety profile is covered. The description adds valuable behavioral context: immediate return (no waiting), truncation at 6000 characters, full=true to see entire text, and that partial output from interrupted tasks is retrievable. This goes beyond annotations and helps an agent predict output size and completeness.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and front-loaded with the immediate-return behavior. Sentences are dense but each conveys necessary behavior (listing vs details, truncation, interrupted tasks). Slightly heavy but no wasted words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only status tool with no output schema, the description covers the key behavioral aspects: return timing, output structure with/without ids, truncation limit, and partial results from interrupted tasks. It does not explain pagination or exact format of the listing, but given no output schema, the coverage is good enough for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so both parameters (ids, full) are fully documented in the schema. The description mirrors the schema semantics (e.g., full=true for untruncated) but adds no syntax or format detail beyond what the schema already provides. Baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states this tool reports status of tasks and details, with explicit behavior for with/without ids. However, its difference from siblings like codex_wait or codex_interrupt is only implicit — an agent might confuse status-checking with waiting. It mentions codex_interrupt, but no direct differentiation from the other siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description says when to omit ids (list all) vs provide ids (detail view). It does not explicitly tell when to use this tool versus codex_wait or codex_interrupt, nor does it exclude cases. Implied usage only.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

codex_steerA

给一个正在运行的任务追加或修正指令,Codex 在当前步骤结束后转向,不用重来,已做的工作保留。只对运行中的任务有效;任务已结束时报错,此时应改用 codex_spawn(thread_id=...) 续问。想让它停下来而不是转向,用 codex_interrupt。

ParametersJSON Schema
NameRequiredDescriptionDefault
idYes运行中任务的 8 位 id(codex_spawn 返回的)
promptYes追加的指令,只写增量,不用重复原任务

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, destructiveHint=false, and idempotentHint=false, so the safety profile is covered. The description adds meaningful behavior beyond that: the run pivots at the end of the current step, prior work is retained, and calling on a finished task errors. It does not address response format or whether repeated steers compose, keeping it short of a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core action and effect, then the precondition and the two sibling alternatives in order of relevance. Every sentence carries distinct information with no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, but the description covers effect, precondition, failure mode, and alternate routes, which is everything needed to invoke this two-parameter steering tool correctly. Annotations carry the safety profile, so nothing material is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%; both parameters are already documented in the schema (id = 8-digit running task id from codex_spawn, prompt = incremental instruction). The description adds only modest framing about writing incremental instructions, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (append/correct instructions) and resource (a running task), then immediately clarifies the effect: Codex pivots after the current step without restarting and preserves completed work. It distinguishes itself from codex_spawn (continuation) and codex_interrupt (stopping) by name.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly bounds usage: only valid for running tasks, errors if the task has ended, and in that case directs the agent to codex_spawn(thread_id=...). It also names the alternative for the opposite intent (codex_interrupt). When/when-not/alternatives are all present.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

codex_waitA
Read-onlyIdempotent

阻塞等待一个或多个任务结束并返回它们的结果。不传 ids 就等本会话全部运行中的任务。超时只是返回当前状态,任务不会被杀,可以再次调用继续等。已结束的任务立即返回。等待期间每 10 秒发一次进度通知。只想看一眼不想等,用 codex_status。

ParametersJSON Schema
NameRequiredDescriptionDefault
idsNo任务 id 列表(codex_spawn / codex_review 返回的 8 位 id);省略 = 全部运行中的任务
modeNoall=全部结束才返回(默认),any=任一结束就返回
timeoutNo最长等待秒数,默认 600,上限 3000

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the readOnly/idempotent annotations, it discloses non-obvious behavior: timeout returns current state without killing the task, the call can be re-issued to keep waiting, finished tasks return immediately, and progress notifications fire every 10 seconds. This is rich context the annotations cannot convey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Dense and front-loaded: core behavior first, then timeout/notification behavior, then the alternative-tool pointer. Every sentence carries distinct information with no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, but the description covers what to expect on return (results, current state on timeout, immediate return for finished tasks) plus the rerun and notification behavior. Nothing an agent needs to invoke it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. The description adds real meaning on top of the schema by explaining the timeout semantics (returns state, does not kill, re-callable) and the no-ids default, going past raw field documentation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: blocking-wait on one or more tasks and returning their results. It also states scope explicitly ('不传 ids 就等本会话全部运行中的任务'), which an agent can use to distinguish it from the other codex_* siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives an explicit alternative and the condition selecting it: '只想看一眼不想等,用 codex_status'. It also clarifies the no-ids default, which effectively tells the agent when this tool is the right choice versus polling status.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 5 tool updatesv0.2.3
    • Changedcodex_interrupt1 field changed
      • addedInput schema / properties / id / description
        Added value: +"任务的 8 位 id(codex_spawn / codex_review 返回的)"
    • Changedcodex_review6 fields changed
      • changedInput schema / properties / cwd / description
        Previous value: -"仓库目录(绝对路径),默认当前目录"New value: +"git 仓库目录(绝对路径),默认当前目录"
      • addedInput schema / properties / detach / description
        Added value: +"true = 不关心结果,Claude 结束回合时不等它"
      • addedInput schema / properties / model / description
        Added value: +"仅用户明确指定时填"
      • changedInput schema / properties / scope / description
        Previous value: -"默认 uncommitted"New value: +"审查范围,默认 uncommitted"
      • changedInput schema / properties / value / description
        Previous value: -"分支名 / commit sha / 自定义审查说明"New value: +"scope 对应的值:分支名 / commit sha / 自定义审查说明;uncommitted 不需要"
      • changedInput schema / properties / wait_s / description
        Previous value: -"先同步等待的秒数,审查通常 1~3 分钟"New value: +"先同步等待的秒数,审查通常 1~3 分钟;0 立即返回"
    • Changedcodex_status2 fields changed
      • addedInput schema / properties / full / description
        Added value: +"true = 不截断,返回完整结果文本"
      • addedInput schema / properties / ids / description
        Added value: +"任务 id 列表;省略 = 只列清单"
    • Changedcodex_steer2 fields changed
      • addedInput schema / properties / id / description
        Added value: +"运行中任务的 8 位 id(codex_spawn 返回的)"
      • addedInput schema / properties / prompt / description
        Added value: +"追加的指令,只写增量,不用重复原任务"
    • Changedcodex_wait3 fields changed
      • addedInput schema / properties / ids / description
        Added value: +"任务 id 列表(codex_spawn / codex_review 返回的 8 位 id);省略 = 全部运行中的任务"
      • changedInput schema / properties / mode / description
        Previous value: -"all=全部完成才返回,any=任一完成就返回"New value: +"all=全部结束才返回(默认),any=任一结束就返回"
      • changedInput schema / properties / timeout / description
        Previous value: -"秒,默认 600"New value: +"最长等待秒数,默认 600,上限 3000"
  2. 6 tool updatesv0.2.1
    • First observedcodex_interrupt
    • First observedcodex_review
    • First observedcodex_spawn
    • First observedcodex_status
    • First observedcodex_steer
    • First observedcodex_wait

TDQS

A4.3/5.0

Scored across 6 tools

Disambiguation4/5

Each tool targets a distinct lifecycle action (spawn, status, wait, steer, interrupt, review), and descriptions actively disambiguate near-neighbors by naming the alternative ('只想看一眼不想等,用 codex_status'). The main residual overlap is codex_status vs codex_wait, since both return task results, though the blocking-immediate distinction is explained.

Naming Consistency5/5

All six tools follow a uniform codex_<action> snake_case pattern (spawn, wait, status, steer, interrupt, review). Verbs are consistent and predictable throughout.

Tool Count5/5

Six tools is well-scoped for a subagent-management server, covering spawn, monitoring, control, and review. Every tool earns its place with no redundant additions.

Completeness4/5

The surface covers the full task lifecycle: dispatch (spawn), monitor (status/wait), control (steer/interrupt), and review, plus thread continuation via thread_id. Minor gap: no explicit tool to enumerate historical/prior threads beyond session-scoped status.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers