Skip to main content
Glama
benzkittisak

codex-async-mcp

by benzkittisak

codex-async-mcp

codex CLI를 비동기적으로 래핑하는 로컬 MCP 서버입니다. 차단하는 대신 job_id를 즉시 반환하여 Claude가 MCP 프로토콜 타임아웃(-32001)에 걸리지 않도록 합니다.

요구 사항

  • Python 3.11+

  • codex CLI가 설치되어 있고 $PATH에 있어야 함 (v0.125.0+)

  • Claude Code CLI


Related MCP server: codex-mcp-server

설치

cd ~/payroll-mcp   # or wherever this repo lives
pip install -e ".[dev]"

확인:

python -c "from codex_async_mcp.server import mcp; print(mcp.name)"
# → codex-async-mcp

Claude에 등록

전역 (모든 프로젝트)

claude mcp add codex-async -s user -- python -m codex_async_mcp.server

프로젝트 전용

cd ~/payrollservice-thailand   # or any project
claude mcp add codex-async -- python -m codex_async_mcp.server

확인

claude mcp list
# codex-async: python -m codex_async_mcp.server - ✓ Connected

도구 권한 추가 (settings.local.json)

{
  "permissions": {
    "allow": [
      "mcp__codex-async__codex_start",
      "mcp__codex-async__codex_poll",
      "mcp__codex-async__codex_list",
      "mcp__codex-async__codex_cancel"
    ]
  }
}

도구

도구

설명

codex_start(prompt, cwd, approval_policy?)

백그라운드에서 codex 시작 → job_id 즉시 반환

codex_poll(job_id, tail_lines?)

상태 확인 + 출력 끝부분 확인

codex_list(limit?)

최근 작업 나열 (최신순)

codex_cancel(job_id)

실행 중인 작업 종료

approval_policy 값

값

Codex 플래그

동작

suggest

-s read-only

읽기 전용 샌드박스, 쓰기 불가

auto-edit

--full-auto

편집 자동 적용

full-auto

--dangerously-bypass-approvals-and-sandbox

프롬프트 없음, 샌드박스 없음

Claude 자동화의 경우 항상 full-auto를 사용하세요. — suggest 모드는 서브프로세스 내부에서 절대 도착하지 않을 대화형 입력을 기다립니다.

사용 예시

codex_start(
  prompt="In app/services/prorate_calculation_service.rb line 96, change format(...) to number_to_currency(...)",
  cwd="/Users/bbgummybear/payrollservice-thailand",
  approval_policy="full-auto"
)
# → { job_id: "f3a9b2", status: "running", pid: 12345 }

codex_poll(job_id="f3a9b2")
# → { status: "running", output: "Reading file..." }

codex_poll(job_id="f3a9b2")
# → { status: "done", exit_code: 0, output: "Applied changes to prorate_calculation_service.rb" }

작업 상태

작업은 ~/.codex-async/jobs/{job_id}/에 저장됩니다:

~/.codex-async/jobs/f3a9b2/
  meta.json     ← status, pid, timestamps, exit_code
  output.txt    ← stdout + stderr from codex

meta.json 구조:

{
  "job_id": "f3a9b2",
  "status": "running | done | error | cancelled",
  "prompt": "...",
  "cwd": "/path/to/repo",
  "approval_policy": "full-auto",
  "pid": 12345,
  "started_at": "2026-04-29T10:00:00+00:00",
  "finished_at": null,
  "exit_code": null
}

문제 해결

claude mcp list에서 codex-async: ... - ✗ Failed 발생

Python을 찾을 수 없거나 패키지가 올바른 환경에 설치되지 않았습니다.

# Check which python Claude is using
which python

# If using conda, register with the full path
claude mcp add codex-async -s user -- /Users/bbgummybear/miniconda3/bin/python -m codex_async_mcp.server

# Verify the package is installed in that environment
/Users/bbgummybear/miniconda3/bin/python -c "import codex_async_mcp; print('ok')"

codex_start 직후 status: "error" 발생

Codex 시작에 실패했습니다. 원시 출력을 확인하세요:

cat ~/.codex-async/jobs/<job_id>/output.txt

일반적인 원인:

출력 메시지

해결 방법

command not found: codex

codex가 PATH에 없음 — 셸 프로필에 추가하거나 config.py에서 CODEX_BIN 설정

unknown flag: --dangerously-bypass-approvals-and-sandbox

Codex 버전 < 0.125.0 — npm install -g @openai/codex를 실행하여 업그레이드

permission denied

cwd가 존재하지 않거나 Claude에 접근 권한이 없음


status: "running" 상태에서 멈추고 완료되지 않음

서브프로세스가 중단되었습니다(입력을 기다리거나 루프에 빠짐).

# Check if the process is still alive
ps aux | grep codex

# Check live output
tail -f ~/.codex-async/jobs/<job_id>/output.txt

# Cancel the job
codex_cancel(job_id="<job_id>")

가장 흔한 원인: 대화형 승인을 기다리는 approval_policy="suggest"를 사용 중인 경우입니다. 대신 "full-auto"를 사용하세요.


서버 재시작 후 작업이 status: "running"으로 표시됨

MCP 서버가 재시작 시 메모리 내 Popen 레지스트리를 잃어버렸습니다. 다음 codex_poll 호출 시 PID가 죽은 것을 감지하고 상태를 자동으로 업데이트합니다.

codex_poll(job_id="<job_id>")
# → { status: "done", ... }   ← auto-resolved on first poll

오래된 작업이 디스크를 채움

# View all jobs sorted by date
ls -lt ~/.codex-async/jobs/

# Delete jobs older than 7 days
find ~/.codex-async/jobs -maxdepth 1 -type d -mtime +7 -exec rm -rf {} +

프로젝트 구조

codex-async-mcp/
├── README.md
├── pyproject.toml
├── src/
│   └── codex_async_mcp/
│       ├── __init__.py
│       ├── server.py        # MCP entry point, tool definitions
│       ├── job_manager.py   # spawn / poll / cancel / list
│       └── config.py        # JOBS_DIR, CODEX_BIN, defaults
└── tests/
    └── test_job_manager.py

테스트 실행

pytest tests/ -v

Available Tools

4 tools
codex_cancelA

Cancel a running codex job by sending SIGTERM to the subprocess.

ParametersJSON Schema
NameRequiredDescriptionDefault
job_idYesThe job_id returned by codex_start.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It discloses that cancellation is done by sending SIGTERM to the subprocess, which is a key behavioral trait. It does not cover edge cases like job already finished, but the main behavior is well communicated.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single 12-word sentence, front-loaded with the action, and contains no wasted words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's simplicity (one parameter, single action), the description is nearly complete. It could mention what happens if the job is not running, but the output schema likely handles error responses.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema already describes the only parameter (job_id) with 100% coverage. The description adds no new semantic information beyond what the schema provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Cancel') and resource ('running codex job'), and clearly distinguishes from siblings like codex_list, codex_poll, and codex_start.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implicitly states when to use (to cancel a running job with a job_id from codex_start) but does not explicitly state when not to use or provide alternative scenarios. However, the simplicity of the action mitigates the need for extensive guidelines.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

codex_listA

List recent codex jobs with their status and prompt summaries.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMax number of jobs to return (most recent first). Default: 20.

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.5/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations provided, so description carries full burden. It only states it lists jobs with status and summaries, but lacks details on pagination, ordering (though limit param says 'most recent first'), rate limits, or side effects. The description is too minimal for full transparency.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Single sentence, no redundant words, front-loaded with the core action. Every word earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the presence of an output schema (status and prompt summaries mentioned), the description is somewhat complete for a simple list operation. However, it lacks details on error handling, empty results, or additional behavioral context that would fully inform an agent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% for the single parameter 'limit', and the schema itself provides a description including default and ordering. The tool description adds no extra meaning beyond what the schema already conveys.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Clearly states the tool lists recent codex jobs, including status and prompt summaries. The verb 'list' and resource 'recent codex jobs' are specific and distinguish from sibling tools (cancel, poll, start) which are different actions.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No explicit guidance on when to use or avoid this tool. The purpose is implied by the name and description, but no alternatives or exclusions are mentioned. Siblings have distinct purposes, so usage is inferred but not clarified.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

codex_pollA

Poll the status and output of a running (or finished) codex job.

ParametersJSON Schema
NameRequiredDescriptionDefault
job_idYesThe job_id returned by codex_start.
tail_linesNoHow many trailing lines of output to return. Default: 100.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description states the purpose but lacks details on behavioral traits such as whether the tool is idempotent or safe to call repeatedly. Since annotations are absent, the description carries the burden, and it only provides minimal transparency.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence with no extraneous words. It efficiently conveys the tool's purpose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the presence of an output schema, the description does not need to explain return values. It covers the essential purpose and scope, though it could mention that the tool can be called multiple times safely.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, and the description adds no additional meaning beyond what the schema already provides for the two parameters. The description's mention of 'output' hints at tail_lines, but this is redundant with the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb 'Poll' and the resource 'status and output of a running (or finished) codex job', which is specific and distinguishes it from sibling tools like codex_start, codex_cancel, and codex_list.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage after a job is started, but does not explicitly provide when-not-to-use or alternatives. The context is clear enough for an agent to infer appropriate usage.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

codex_startA

Start a codex task asynchronously in the background.

Returns a job_id immediately — does not block or timeout. Use codex_poll(job_id) to check progress.

ParametersJSON Schema
NameRequiredDescriptionDefault
promptYesThe task description to pass to codex.
cwdYesAbsolute path to the working directory for codex.
approval_policyNoOne of 'suggest', 'auto-edit', 'full-auto'. Default: 'suggest'.suggest

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations provided, so description carries full burden. Discloses async, non-blocking, immediate return of job_id. Does not mention side effects or auth, but core behavior is adequately covered.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, no wasted words. Front-loaded with purpose and key behavior. Highly efficient.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Has output schema. Describes async nature and returns job_id. Could mention cancellation via sibling codex_cancel, but sufficient for a simple start tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so baseline 3. Description does not add meaning beyond schema; each parameter is defined in schema. No extra context provided.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Clearly states the verb 'Start', resource 'codex task', and key behavior 'asynchronously in the background'. Distinguishes from siblings by mentioning that codex_poll is used to check progress.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides explicit guidance to use codex_poll for progress checking. Implicitly tells when to use this tool (async tasks) but lacks explicit when-not-to-use or alternatives beyond polling.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 4 tool updatesv0.1.0
    • First observedcodex_cancel
    • First observedcodex_list
    • First observedcodex_poll
    • First observedcodex_start

TDQS

A4.2/5.0

Scored across 4 tools

Disambiguation5/5

Each tool has a distinct purpose: start, list, poll, and cancel. There is no overlap in functionality, and the descriptions clearly differentiate them.

Naming Consistency5/5

All tools follow a consistent 'codex_verb' pattern using snake_case, making it predictable for an agent to infer tool behavior from the name.

Tool Count5/5

Four tools cover the essential operations for managing async jobs (start, list, poll, cancel) without redundancy or missing critical actions.

Completeness5/5

The tool set covers the full lifecycle of an async job: initiating (start), monitoring (poll, list), and termination (cancel). No obvious gaps are present.

Maintenance

ActivityInactive
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    D
    maintenance
    Wraps OpenAI Codex CLI as an MCP server, exposing 8 Codex tools (exec, review, skill list, skill run, status, poll, list jobs, kill) as named tools for use with pi or codex.
    949 npm
    ISC
  • A
    license
    Not graded
    quality
    A
    maintenance
    A local STDIO MCP server that bridges MCP clients to the Codex CLI by sending instructions to a configured workspace, exposing task run, status, and result tools with a read-only sandbox and no remote transport.
    133
    MIT