Skip to main content
Glama
4714407
by 4714407

Codex app-server MCP 桥接

想直接接入并开始使用?请先看 使用指南。

English documentation

安全问题请阅读 安全策略 并私下报告;参与贡献前请阅读 贡献指南。版本变更记录见 CHANGELOG.md。

本项目让 Claude Code 通过本地 stdio MCP 调用独立的 Codex app-server 会话,默认适用于提问、代码分析和只读审查;用户显式启用 workspace-write 后,也可在任务工作区内修改代码。它不会连接或继承当前 Codex 桌面聊天上下文;没有 threadId 时创建新会话,提供 threadId 时恢复该会话。

Claude Code → MCP stdio → 本桥接程序 → Codex app-server stdio → 独立 Codex thread

兼容性、安装与登录

最低 Node.js 版本为 20。已在 Windows 上核实 Node.js v24.14.1、npm 11.11.0、Codex CLI 0.154.0,并以该 Codex 版本生成协议定义核对字段。实现使用 Node.js 的跨平台 spawn(..., { shell: false });macOS 和 Linux 尚未实际验证,不对此作出已验证声明。

$mcpDir = '<codex-app-server-mcp 所在目录>'
Set-Location $mcpDir
npm.cmd install
codex --version
codex login
node .\server.mjs

桥接优先使用 CODEX_EXECUTABLE;未设置时从 PATH 查找 codex(Windows 也查找 .exe、.cmd 和 .bat)。找不到时工具返回明确配置说明。桥接不会读取、复制、打印或硬编码令牌、API Key 或 Codex 全局配置。

Windows PowerShell:

$env:CODEX_EXECUTABLE = 'C:\Program Files\OpenAI\Codex\bin\codex.exe'
node D:\path\to\codex-app-server-mcp\server.mjs

macOS / Linux:

export CODEX_EXECUTABLE="$(command -v codex)" # 可省略,让程序从 PATH 查找
node /path/to/codex-app-server-mcp/server.mjs

未登录时,请在普通终端执行 codex login 后重启桥接。若本机 Codex 协议不兼容,初始化或对应 JSON-RPC 请求会以方法名和服务端错误明确失败,不会静默降级。

Related MCP server: Claude Code Subagent MCP

MCP 工具

codex_start({ prompt, cwd, threadId?, model? }) 按启动时的沙箱配置启动任务。prompt 必须非空;cwd 必须是存在的绝对目录。它立即返回 jobId、threadId 和已获得时的 turnId。同一桥接进程只允许一个活动任务,第二个调用会收到 busy;恢复既有 thread 时 cwd 必须与该 thread 初始 cwd 相同。

codex_status({ jobId, cursor?, waitMs? }) 返回任务状态 starting、running、waiting_for_input、completed、failed 或 cancelled,以及已收到的助手回复、错误、进度和关联 ID。传入上次结果的 outputCursor 作为 cursor 时,assistantReply 只包含增量;waitMs 支持最多 30 秒的长轮询。启动结果和状态均包含 sandboxMode,工具描述也显示当前模式。最多保留 100 个任务,每项回复最多 256 KiB,截断会标为 outputTruncated: true。

codex_cancel({ jobId }) 使用 app-server turn/interrupt 请求中断。返回 requested 仅表示已发出请求;只有收到匹配 turn/completed 的 interrupted 才确认 cancelled。重复取消和已结束任务有稳定结果。

codex_start({"prompt":"只回复 CODEX_OK,不读取或修改任何文件。","cwd":"D:\\project"})
codex_status({"jobId":"..."})
codex_cancel({"jobId":"..."})

codex_models({ cursor?, limit? }) 从当前 Codex app-server 实时读取可用模型。返回模型 ID、显示名称、描述、默认推理强度、支持的推理强度、可用性提示以及 nextCursor;Claude 可先调用此工具展示列表,再将用户选中的 model 传给下一次 codex_start。

codex_steer({ jobId, prompt }) 可向运行中的普通 turn 补充指令;codex_review({ threadId, target, delivery? }) 启动原生审查;codex_compact({ threadId }) 异步压缩长会话。Codex 需要澄清时,使用 codex_pending_input({ jobId }) 取得问题与 inputId,再由 codex_answer_input({ inputId, answers }) 明确回答;该流程不能批准命令或文件变更。codex_doctor() 显示连接、可执行文件、沙箱、目录白名单和可选任务索引配置。

设置 CODEX_STATE_PATH 后,桥接将以原子写入的本地 JSON 索引保留不含 prompt/回复/凭据的任务元数据;重启时已运行任务标为 unknown,不会自动重试,可使用保存的 threadId 手动恢复。

权限、超时和关闭

未设置 CODEX_SANDBOX_MODE 时使用 read-only。用户可显式设置为 workspace-write,其他值(包括空字符串和拼写错误)会导致启动失败。可写模式必须同时设置 CODEX_ALLOWED_ROOTS(分号分隔的目录列表或 JSON 字符串数组),任务 cwd 必须位于白名单内;未设置白名单时可写任务会拒绝启动。新建与恢复 thread 都明确设定当前沙箱和 approvalPolicy: "never";每个 turn 再指定相应策略:只读为 { type: "readOnly", networkAccess: false };可写为 { type: "workspaceWrite", writableRoots: [cwd], networkAccess: false }。never 代表不发起交互审批,不代表自动批准。服务端提出的命令、文件改动和权限审批一律拒绝;用户输入会进入等待状态,须通过 codex_pending_input 和 codex_answer_input 明确回答,未知服务端请求收到 JSON-RPC -32601,不会悬挂。工作区内写入在 workspace-write 下由沙箱直接允许,无需自动批准请求;超出沙箱的请求仍被拒绝。工具不接受改变沙箱或审批策略的参数,也不会使用危险绕过选项。

初始化超时为 20 秒,普通 RPC 超时为 30 秒。Codex 子进程退出、初始化失败或 JSON-RPC 失败会让活动任务失败,绝不误报成功。stdin 关闭、SIGINT 或 SIGTERM 时只关闭本桥接启动的 app-server,并清理计时器和待处理请求。

注册 Claude Code

已按当前 Claude Code 文档核对:stdio 选项和 --scope 必须放在服务名前,-- 后才是启动命令。请自行执行:

$mcpDir = '<codex-app-server-mcp 所在目录>'
$nodeExe = (Get-Command node.exe).Source
claude mcp add --transport stdio --scope user codex-agent -- "$nodeExe" (Join-Path $mcpDir 'server.mjs')
claude mcp get codex-agent

卸载命令:

claude mcp remove codex-agent --scope user

测试与分发

npm.cmd test
npm.cmd run test:real
npm.cmd pack --dry-run
npm.cmd pack

npm test 运行 JSONL 分块、启动响应前事件、忙状态、取消确认、子进程退出,以及使用官方 MCP 客户端 SDK 从外部进程启动服务的初始化、tools/list、不存在 cwd、未知 job 检查。test:real 额外验证真实 Codex 简短答复与 threadId 上下文恢复;它需要网络、有效登录和可用额度。测试只使用项目内 test-tmp,不会访问业务项目。模拟不会替代真实端到端成功。

npm 的 files 白名单仅包含运行必需的 server.mjs、lib/、README 和 LICENSE,排除测试、测试临时目录、日志、凭据、本地配置和会话数据。npm pack 仅在本地生成 tarball,不会发布到 npm。

两种 Codex 连接模式

MCP 服务应由 Claude Code、Cursor 等 MCP 客户端按其 command、args 和 env 配置自动启动;无需手动运行桥接进程或使用平台专属启动脚本。首次仍需在项目目录执行一次 npm.cmd install(或 npm install)。

两种模式向 Claude Code 暴露相同的 codex_start、codex_status、codex_cancel、codex_models、codex_steer、codex_review、codex_compact、codex_pending_input、codex_answer_input 和 codex_doctor 工具。每次启动只选择一种模式:设置 CODEX_APP_SERVER_URL 即为远程模式;未设置时为本地模式。

本地模式

本地模式由桥接程序以 spawn(codex, ['app-server', '--listen', 'stdio://']) 启动并管理子进程。cwd 是桥接机器上的存在目录,必须是绝对路径。桥接关闭时只停止它自己启动的 app-server。

Windows PowerShell:

Remove-Item Env:CODEX_APP_SERVER_URL -ErrorAction Ignore
$env:CODEX_EXECUTABLE = 'C:\Program Files\OpenAI\Codex\bin\codex.exe' # 或让程序从 PATH 查找
node D:\path\to\codex-app-server-mcp\server.mjs

macOS / Linux:

unset CODEX_APP_SERVER_URL
export CODEX_EXECUTABLE="$(command -v codex)" # 可省略
node /path/to/codex-app-server-mcp/server.mjs

远程模式

远程模式不会启动、重启或终止任何 Codex 进程;它只连接既有 app-server。cwd 是远程 Codex 执行环境中的绝对路径,桥接不会在本机检查它是否存在,远程 app-server 会验证它。

配置 CODEX_APP_SERVER_URL。跨机器连接必须使用 wss://;ws:// 仅接受 localhost、127.0.0.1 或 ::1,用于同机服务或 SSH 隧道。URL 不得内嵌用户名、密码或 Token。

Bearer Token 可任选其一提供:CODEX_REMOTE_BEARER_TOKEN,或 CODEX_REMOTE_TOKEN_FILE 指向只含 Token 的秘密文件。Token 只用于 WebSocket 握手的 Authorization: Bearer … 请求头,不会写入日志、任务状态或 README 示例。优先使用秘密文件。

Windows PowerShell(TLS 反向代理后的远程服务):

$env:CODEX_APP_SERVER_URL = 'wss://codex.example.internal/app-server'
$env:CODEX_REMOTE_TOKEN_FILE = 'C:\secure\codex-app-server.token'
Remove-Item Env:CODEX_EXECUTABLE -ErrorAction Ignore
node D:\path\to\codex-app-server-mcp\server.mjs

macOS / Linux(TLS 反向代理后的远程服务):

export CODEX_APP_SERVER_URL='wss://codex.example.internal/app-server'
export CODEX_REMOTE_TOKEN_FILE="$HOME/.config/codex-app-server.token"
node /path/to/codex-app-server-mcp/server.mjs

使用 SSH 隧道时,远程主机上的 app-server 应仅监听 loopback;在本机建立隧道后使用本地 ws://:

# 远程主机:使用其安全的 Token 文件启动 app-server
codex app-server --listen ws://127.0.0.1:4500 --ws-auth capability-token --ws-token-file /secure/codex-app-server.token
# 本机:保持隧道运行
ssh -N -L 4500:127.0.0.1:4500 user@codex-host
export CODEX_APP_SERVER_URL='ws://127.0.0.1:4500'
export CODEX_REMOTE_TOKEN_FILE="$HOME/.config/codex-app-server.token"
node /path/to/codex-app-server-mcp/server.mjs

远程 WebSocket 在任务运行时断开,任务会显示 status: "unknown" 与断线原因。桥接不会自动重连后重发该任务,避免重复执行;可以在确认远程状态后手动发起新任务。关闭桥接只关闭其 WebSocket 客户端连接,不会向远程 app-server 发送终止信号。

本地模式的真实 Codex 测试仍依赖本机登录和可访问的用户 Home。远程模式已使用本地受控 WebSocket 服务验证握手 Bearer 头、远程 cwd 透传和断线为 unknown;尚未针对外部真实远程 Codex 端点运行集成测试。

沙箱开关:默认只读,用户显式启用写入

在 Claude 项目的 .mcp.json 中,保留现有 codex-agent 的 command、args 和 env,在它的 env 对象中增加:

"CODEX_SANDBOX_MODE": "workspace-write"

保存后重启该项目的 Claude / MCP 连接。恢复只读时改为 read-only;删除该键将回退到进程环境,环境也未设置时才使用只读默认值。

这是一项启动配置,不是 Claude 可以通过 prompt 或 codex_start 参数自行切换的权限。每个桥接实例启动后固定使用该模式,新建、恢复会话及每次任务均重新发送对应权限。支持本地和远程连接;远程的 cwd 与可写目录属于远程执行环境。

写入范围是任务 cwd 工作区及 Codex 标准沙箱允许的临时目录,受服务端和平台的额外限制约束。不会开启网络或全磁盘访问。当前 cwd 仍由调用方提供,因此请明确任务目录。配置为 workspace-write 并不保证外部依赖下载或工作区外的全局缓存写入成功。

若为调试而手动启动,可仅使用环境变量:

$env:CODEX_SANDBOX_MODE = 'workspace-write'
$env:CODEX_ALLOWED_ROOTS = 'D:\projects\my-app'
node .\server.mjs

验证:npm.cmd test 包含默认值、非法配置、新建/恢复会话的双层权限、远程策略、审批拒绝和 MCP 工具描述检查。 真实沙箱文件测试使用独立临时目录(不修改业务仓库):

$env:CODEX_RUN_SANDBOX_REAL = '1'
npm.cmd run test:sandbox-real

真实测试依次检查默认只读不能创建文件、显式开启写入能创建文件,以及恢复可写会话到只读后不能再次创建文件。需有效 Codex 登录和可用网络;未执行真实测试不能视为沙箱执行验证通过。

指定模型

在项目 .mcp.json 的 codex-agent.env 中设置 CODEX_MODEL 为账号可用的模型 ID,保存后重启 Claude / MCP。示例(将占位符替换为实际模型 ID):

"CODEX_MODEL": "YOUR_MODEL_ID",
"CODEX_SANDBOX_MODE": "workspace-write"

未设置 CODEX_MODEL 时不传 model:新会话使用 Codex 默认配置,恢复会话遵循 Codex 对已有会话的默认行为。显式设置后,新建、恢复会话和每个 turn 均传入该模型。空值会报错;不自动替换账号不支持的模型,不修改 Codex 全局配置。该设置适用于本地和远程模式,也可通过启动环境传入。任务状态的 requestedModel 表示请求的模型 ID(未设置时为 null),不冒充服务端确认的实际模型。

通过话术切换模型

codex_start 现在接受可选 model 参数。对 Claude 说“通过 codex-agent,使用模型 <实际模型 ID> 审查当前项目”,Claude 应把模型 ID 传入 model,而不是仅写进 prompt。 优先级为:本次 model 参数 > CODEX_MODEL 启动配置 > Codex 默认。支持在下一次调用中沿用 threadId 并更换 model,兼容性和账号权限由服务端验证;不自动替换不支持的模型。运行中的任务不能中途切换,需等待结束或取消后再调用。未指定 model 且未配置 CODEX_MODEL 时,恢复已有会话遵循服务端已有模型设置。

注意:可通过话术选择模型,但不能通过话术改变沙箱权限;写入权限仍必须由用户配置 CODEX_SANDBOX_MODE。

Available Tools

10 tools
codex_answer_inputA

提交 codex_pending_input 返回的问题答案,并恢复任务;不能用于批准命令或文件变更。

ParametersJSON Schema
NameRequiredDescriptionDefault
answersYes
inputIdYes

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the behavioral burden. It discloses the main side effect ('resumes the task') and a key limitation (cannot be used to approve commands or file changes). It does not cover errors or return behavior, but the essential behavioral boundary is clearly stated.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single focused sentence with no filler. It front-loads the primary action and effect, then adds the critical exclusion. Every word earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description identifies the trigger, the effect, and an important non-goal, which is a good foundation. However, with no annotations, no output schema, and 0% parameter coverage, it leaves inputId provenance and answer object structure unexplained. This is a moderate gap for a simple tool, not a fatal one.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the description adds almost no parameter-level meaning. It does not explain what inputId refers to, where it comes from, or how the answers object should be structured. The phrase 'answers to questions' gives a general hint, but it does not compensate for the lack of schema descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('submit') and resource ('answers to questions returned by codex_pending_input'), and explicitly states it is not for approving commands or file changes. This clearly distinguishes it from sibling tools and leaves no ambiguity about its core purpose.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It clearly states when to use the tool (when codex_pending_input has returned questions needing answers) and gives a firm exclusion (not for approving commands or file changes). It does not name a specific alternative for approvals, but the when-not is explicit enough for an agent to route correctly.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

codex_cancelA

请求中断正在运行的 Codex 任务;终态以后才确认取消。

ParametersJSON Schema
NameRequiredDescriptionDefault
jobIdYes

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the behavioral disclosure burden. It does disclose a meaningful trait: cancellation is confirmed only after the terminal state ('终态以后才确认取消'), suggesting asynchronous confirmation. However, it omits other behavioral details such as what happens to the job, whether cancellation is reversible, error cases, or permission requirements.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single front-loaded sentence that states the action and immediately adds the important terminal-state caveat. Every word is purposeful, with no redundant filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-parameter tool, the description covers the core action and a key timing nuance, but it omits edge cases such as invalid jobId, already-terminated tasks, or exact return/confirmation semantics. The phrase '终态以后才确认取消' is helpful but somewhat ambiguous about whether the call blocks until terminal state or returns immediately.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema has 0% description coverage for the single parameter jobId, and the tool description does not explain it. The parameter's role is inferable from its name and the tool's purpose, but the description adds no explicit meaning beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific action and resource: '请求中断正在运行的 Codex 任务' (request interruption of a running Codex task). This clearly distinguishes it from siblings like codex_start, codex_status, and codex_steer. The purpose is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage when a Codex task is running and needs to be interrupted. However, it does not explicitly state when not to use it, such as for tasks already in a terminal state, nor does it mention any alternative tools or preconditions. Usage context is present but only implicitly.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

codex_compactA

压缩已有 Codex thread 的上下文,适用于长会话;作为一个异步任务返回 jobId。

ParametersJSON Schema
NameRequiredDescriptionDefault
threadIdYes

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Discloses a key behavioral trait: the operation is asynchronous and returns a jobId, which is important for agents. Since no annotations exist, the description carries the burden, but it does not explain what compaction means for thread history or how to track the job beyond the jobId.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The entire purpose, usage context, and async behavior are conveyed in one compact sentence. Every phrase carries information and the core operation is front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-parameter tool, the description is nearly sufficient: it names the input, the purpose, and the return behavior. Missing details such as polling with codex_status or potential side effects are not fatal but would round out the picture for an agent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, but the single parameter threadId is clearly tied to the description's '已有 Codex thread' phrasing. The description adds some meaning by indicating it operates on an existing thread, but it does not explain how to obtain the threadId or its relationship to codex_start output.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description specifies the verb '压缩' (compact) and the resource '已有 Codex thread 的上下文' (existing Codex thread's context), making the tool's purpose unmistakable. It also adds the scope '适用于长会话', helping distinguish it from sibling Codex tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides a clear usage context: '适用于长会话' (suitable for long sessions). It does not explicitly mention when not to use it or which alternative to prefer, so it stops short of strong routing guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

codex_doctorB

检查桥接模式、Codex 可执行文件、初始化状态、沙箱、目录白名单和任务索引配置。

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

B3.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden of behavioral disclosure. It only states that the tool 'checks' the listed items, implying read-only diagnostics, but it does not disclose whether the check has side effects, what the output format is, or whether the tool can repair issues. The listed components add some context, but the safety and result behavior remain unclear.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single compact sentence leads with the verb and enumerates all six check targets with no filler or repeated information. This is appropriately sized for a zero-parameter diagnostic tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the zero-parameter complexity, the description adequately lists what is checked. However, since there is no output schema, a note on what the agent receives after the check (e.g., a report, list of issues, or pass/fail status) would make it fully self-contained. The description is usable but leaves the result behavior to inference.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema defines zero parameters, so there is nothing for the description to explain. The list of checked items effectively serves as the entire parameter surface, making additional parameter semantics unnecessary.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description uses an explicit verb '检查' (check) and lists the exact resources it inspects: bridge mode, Codex executable, initialization status, sandbox, directory whitelist, and task index configuration. This makes the tool's scope clear and distinguishes it from siblings like codex_status, though it does not explicitly contrast itself with that tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to call codex_doctor versus codex_status or other siblings. The listed check areas imply a diagnostic context, but the description never states use cases, prerequisites, or when this tool is not appropriate.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

codex_modelsB

从当前 Codex app-server 实时读取可用模型列表。可使用 nextCursor 继续读取下一页。

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
cursorNo

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the burden. It does disclose that the read is realtime and that pagination is supported, which rules out a purely static list. However, it doesn't describe the return shape, cursor source, or any error/rate-limit behavior, so transparency is only partial.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is short and front-loaded, with the core read action in the first sentence. The pagination note is useful, though the mismatched parameter name costs it a perfect score.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple list tool with no annotations and no output schema, the description should state that it returns a paged list, how cursor is obtained, and how limit controls page size. It covers the paged aspect loosely but omits limit and exact parameter semantics, leaving gaps for a correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description needed to document both parameters. It explains the pagination concept but refers to 'nextCursor' while the actual schema property is 'cursor', and it never mentions the 'limit' parameter or its 1–100 range. This is partial, misleading help.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('read') and a concrete resource ('available model list from the current Codex app-server'). It is clearly distinct from the sibling action tools (start/status/cancel/steer/review/etc.), so an agent can identify it without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is given for when to call this tool versus alternatives. There is no mention of prerequisites such as 'call this before codex_start' or any exclusions, so the agent must infer usage from the name and sibling context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

codex_pending_inputB

列出某个任务正在等待回答的 Codex 问题。

ParametersJSON Schema
NameRequiredDescriptionDefault
jobIdYes

TDQS

B3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

没有提供任何注解,描述需要承担全部行为披露责任。“列出”暗示只读,但未明确说明该操作不会修改任何任务状态、是否会阻塞等待、对无效 jobId 的表现,以及返回内容的格式。

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

描述只有一句,核心功能放在开头,没有冗余内容。结构高效,但过短也导致行为与参数信息不足。

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

对于只有一个参数且是只读列举操作的简单工具,描述给出了输入对象(某个任务)和输出内容(等待回答的 Codex 问题),属于最低可用水平。但没有输出 schema 或注解,也未说明返回问题的具体字段集合,因此仍有明显缺口。

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

schema 参数覆盖率为 0%,而描述没有直接说明 jobId 参数的含义、来源或格式。仅通过“某个任务”间接暗示与 jobId 的对应关系,对唯一参数的语义补偿不足。

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

描述使用具体动词“列出”和明确资源“某个任务正在等待回答的 Codex 问题”,清楚说明功能。虽然没有像最佳示例那样点名与 codex_answer_input、codex_status 等兄弟工具的区分,但语义本身足以避免直接混淆。

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

描述隐含使用场景:当需要查看某个任务尚未回答的 Codex 问题时调用。但没有明确说明何时不应使用,也没有提及应改用 codex_answer_input 或 codex_status 的条件。

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

codex_reviewC

对已有 Codex thread 启动原生代码审查。仅接受明确的审查目标。

ParametersJSON Schema
NameRequiredDescriptionDefault
targetYes
deliveryNo
threadIdYes

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. It states the action (start native code review) and the constraint (only explicit targets), but doesn't disclose what happens to the thread, whether it's a long-running operation, whether it modifies state, or what the delivery modes (inline/detached) mean behaviorally. For a tool that likely mutates or drives a thread, this is a significant gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is very short (one sentence) and front-loads the core action. It's efficient, but the brevity comes at the cost of missing critical parameter semantics and behavioral context. Still, every word earns its place, and the constraint is clearly stated.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has 3 parameters, a nested target object with 4 enum subtypes, no output schema, and no annotations, the description is far from complete. An agent would need to understand what 'delivery' means, what each target type implies, and what the review process entails. The description only covers the 'explicit target' requirement, leaving major gaps for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. The description mentions '明确的审查目标' (explicit review target), which maps to the 'target' parameter, but doesn't explain the meaning of 'threadId', 'delivery', or the target subtypes (uncommittedChanges, baseBranch, commit, custom). The nested object structure with 4 subtypes is completely undocumented in the description, leaving the agent to guess at semantics.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('启动原生代码审查' - start native code review) and resource ('已有 Codex thread' - existing Codex thread), distinguishing it from sibling tools like codex_start and codex_status. It also adds a constraint ('仅接受明确的审查目标' - only accepts explicit review targets), which helps clarify scope. However, it doesn't explicitly name sibling tools for differentiation, so it loses a point.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage context: it operates on an existing thread and requires an explicit target. The constraint '仅接受明确的审查目标' gives some guidance on when to use it (when a clear review target exists). However, it doesn't explicitly state when NOT to use it or mention alternatives like codex_steer or codex_status for other operations on threads.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

codex_startA

在新的或已有 Codex thread 中启动任务,立即返回 jobId。可用 model 参数指定本次任务模型。配置模型:Codex 默认配置 / 已有会话模型。当前沙箱:read-only;只读,禁止修改文件。权限由用户启动配置决定,工具参数不能更改。

ParametersJSON Schema
NameRequiredDescriptionDefault
cwdYes
modelNo
promptYes
threadIdNo

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description discloses that the sandbox is read-only and that permissions are determined by user startup configuration, not tool parameters. However, with no annotations provided, the description carries the burden; it doesn't mention what happens to the job after starting, how to check status, or any side effects beyond returning a jobId.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise and front-loaded with the main action and return value. It includes important behavioral notes about read-only sandbox and permissions. Slightly dense but each sentence adds value.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool that starts a task, the description covers the core action, return value, model selection, and sandbox constraints. However, it lacks details about how to use the returned jobId, how to handle errors, or what happens if the thread doesn't exist. Given the sibling tools for status/cancel, an agent might need more guidance on the lifecycle.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It mentions prompt, cwd, model, and threadId implicitly (new or existing thread), but doesn't explain the format or constraints of each parameter beyond what the schema already shows. It adds some context about model selection but not enough to fully compensate for the 0% coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool starts a task in a new or existing Codex thread and immediately returns a jobId. It also mentions the model parameter for specifying the task model, which distinguishes it from sibling tools like codex_status or codex_cancel.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explains that it can be used in new or existing threads and that the model parameter can specify the model. It doesn't explicitly say when to use alternatives, but the context of starting a task is clear enough to differentiate from status/cancel/steer tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

codex_statusA

查询 Codex 任务状态和已经收到的助手答复。传入 cursor 时只返回新增输出;waitMs 可长轮询最多 30 秒。

ParametersJSON Schema
NameRequiredDescriptionDefault
jobIdYes
cursorNo
waitMsNo

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It goes beyond the schema by explaining that passing cursor returns only new output since the last query, and that waitMs enables long polling up to 30 seconds. This provides meaningful behavioral context about incremental retrieval and polling behavior. It does not explicitly state side effects, but the '查询' wording implies read-only intent, and for a status tool this is reasonably transparent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two short sentences. The first sentence front-loads the primary purpose, and the second succinctly explains the two optional parameters' behavior. There is no redundant or filler content, making it easy to parse and act on.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a polling tool with no output schema, the description provides enough to invoke the tool correctly (jobId, cursor, waitMs) but does not explain the return value structure, possible statuses, or how to detect completion. Given the complexity of a status/reply tool and the absence of output schema, this is a notable gap. The description could be more complete by indicating what the response contains or how to interpret statuses, but it is minimally viable.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It does so for cursor and waitMs by explaining their semantics (incremental output and long polling duration), which are not apparent from the schema alone. jobId is not described explicitly, but its role is self-evident as the task identifier. The description adds significant meaning for two of the three parameters, making the tool more usable.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: '查询 Codex 任务状态和已经收到的助手答复' (query Codex task status and already received assistant replies). This is a specific verb (query) and resource (task status + replies), and it distinguishes itself from sibling tools like codex_start, codex_cancel, and codex_pending_input by focusing on status retrieval. The additional details about cursor and waitMs further clarify its role as a polling mechanism.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage (querying status and replies, especially with incremental updates or long polling), but it does not explicitly state when to use this tool versus alternatives like codex_pending_input or codex_review. There are no exclusions or explicit recommendations, leaving the agent to infer the appropriate context from the tool name and description.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

codex_steerA

向当前运行中的普通 Codex turn 补充指令;不会改变沙箱或审批策略。

ParametersJSON Schema
NameRequiredDescriptionDefault
jobIdYes
promptYes

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full responsibility for behavioral disclosure. It does add one useful guarantee: it will not change sandbox or approval policy. However, it does not disclose whether the prompt is appended or replaces existing instructions, whether multiple calls are allowed, or what effect it has on already-generated content.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence that front-loads the core action and scope, then adds a valuable non-effect. There is no filler, redundancy, or unnecessary detail.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no output schema, no annotations, and zero schema description coverage, the description is too sparse. It explains the general purpose and one non-effect, but it omits explicit parameter mapping, when-to-use guidance, and behavioral side effects. An agent can guess, but it is not fully equipped to call this tool correctly without further documentation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% and the description does not explicitly explain the parameters. It hints that prompt is the instruction and jobId refers to the running turn, but it never names them or describes expected formats or behaviors. The parameter names are self-explanatory, yet the description adds no explicit value beyond the JSON Schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb and resource: '补充指令' (supplement instructions) to the '当前运行中的普通 Codex turn' (currently running normal Codex turn). It also distinguishes itself from siblings by scoping to normal running turns and explicitly noting it does not change sandbox/approval policy. An agent can separate this from codex_start, codex_cancel, codex_review, and the input-related tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies the use case: add instructions to an already running Codex turn. However, it does not explicitly state when not to use it or name alternatives, such as codex_pending_input or codex_answer_input for turns awaiting input. The routing is left to inference, not explicit guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 10 tool updatesv1.0.0
    • First observedcodex_answer_input
    • First observedcodex_cancel
    • First observedcodex_compact
    • First observedcodex_doctor
    • First observedcodex_models
    • First observedcodex_pending_input
    • First observedcodex_review
    • First observedcodex_start
    • First observedcodex_status
    • First observedcodex_steer

TDQS

A3.7/5.0

Scored across 10 tools

Disambiguation5/5

Each tool maps to a distinct Codex operation: start, status, cancel, steer, review, compact, pending input, answer input, model listing, and diagnostics. Even similar actions like codex_steer and codex_answer_input are cleanly separated by their stated purpose.

Naming Consistency4/5

The codex_ prefix and snake_case style are consistent throughout, making the family easy to recognize. Minor deviations exist: codex_models and codex_status are noun-like rather than verb_noun, while most others follow an action-oriented pattern.

Tool Count5/5

Ten tools is a well-scoped set for managing Codex tasks and app-server operations. Each tool covers a meaningful lifecycle or capability without unnecessary redundancy.

Completeness4/5

The surface covers the main task lifecycle: start, monitor, cancel, steer, answer pending input, compact context, and review. Minor gaps exist around broader thread history or explicit thread deletion, but these are not essential for the core workflow.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers