Skip to main content
Glama

Local Codex Bridge

English · 简体中文

A thin supervisory MCP bridge between external AI supervisors and native Codex.

Current release: V2.3.4

Local Codex Bridge is a lightweight MCP stdio adapter for Windows and macOS:

ChatGPT / external AI supervisor
              ↕
        Local Codex Bridge
              ↕
      native Codex app-server
              ↕
   native Codex threads / turns

It lets an AI that is good at conversation, planning, and sustained supervision oversee real engineering work performed by local native Codex. It does not recreate Codex.

The supervisor owns the goal, resources, boundaries, risks, approvals, and acceptance. Codex retains its native autonomy for coding and execution.

The Bridge stays thin:

  • It creates no second job or task system.

  • It does not copy Codex conversation history.

  • It keeps no parallel thread database.

  • It does not cache the “current model.”

  • It does not rebuild Goal, Queue, or Search.

  • It does not replace Codex session, thread, or turn semantics.

The native Codex thread/session remains the source of truth for execution.

Changelog · Protocol and compatibility assumptions


What V2.3.0 introduced

V2.3.0 expanded the public supervisory surface to 12 MCP tools while retaining the boundary of exposing only the native surface needed for supervision.

The release added or refined:

  • Independent paging of native persistent History.

  • Native Goal: read, set, and clear a persistent objective.

  • Native Queue: manage follow-up input for an active workflow.

  • Native Search: search across threads and locate occurrences within a thread.

  • On-demand capability and lineage metadata.

  • Typed compact observation, drainage, and loss semantics.

  • A single event-driven observe wait of up to 120 seconds.

  • Delivery of every successful tool result through structuredContent.

  • One shared Bridge core for Windows and macOS, with differences confined to native platform boundaries.

Bridge does not reimplement these capabilities. Native Codex still owns History, Goal, Queue, Search, and the thread lifecycle; Bridge provides bounded mappings and supervision.


Related MCP server: Codex Bridge

Responsibilities

External supervisor / ChatGPT

The supervisor is suited to:

  • Understand the user's goal.

  • Break down the work.

  • Set scope, resources, and risk boundaries.

  • Decide when to keep observing, correct, approve, or interrupt.

  • Judge whether the result meets acceptance criteria.

  • Provide oversight when Codex cannot safely decide on its own.

Native Codex

Native Codex continues to own:

  • Thread and turn lifecycles.

  • Workspace files and command execution.

  • Its context and persistent history.

  • Sandbox and approval-policy behavior.

  • The actual model and reasoning effort in use.

  • Persisted execution results.

  • Native Goal, Queue, Search, and lineage.

Local Codex Bridge

Bridge connects the two:

  • MCP stdio ↔ Codex app-server JSONL.

  • Bounded exposure of state needed for supervision.

  • Forwarding explicit control intent.

  • Failing closed at high-risk, ambiguous, or protocol boundaries.

  • No second orchestration runtime.


The 12 MCP tools

Tool

Purpose

Main boundary

codex_threads

List, filter, and read native persistent thread metadata

Metadata only; filters are not ACLs

codex_history

Page native persistent history

No Bridge history store or automatic full read

codex_search

Search native threads and locate within-thread occurrences

Returns locators; no Bridge index or relevance layer

codex_models

Read one native model/list page on demand

No cached catalog or current-model registry

codex_goal

Read, set, or clear a native thread goal

No implicit resume or turn start; distinct from checkpoint

codex_queue

Manage native queued follow-ups

No Bridge scheduler; enqueue does not mean execution or completion

codex_turn

Create or resume a thread and start a turn

Accepted does not mean completed

codex_observe

Read bounded live events, pending requests, terminal state, and cursor

Does not reconstruct live state after runtime loss

codex_steer

Add a semantic correction to the current active turn

Not a timer, poll, or retry

codex_respond

Answer a real pending approval, user-input, or permission request

Must match the original request ID and exact scope

codex_interrupt

Interrupt the exact active thread and turn

Not process control

codex_checkpoint

Keep an optional, concise, bounded supervisor anchor

Not a transcript, job ID, or Codex history

See src/tools.ts for the full schemas and runtime validation.

Content delivery and read policy

History and Search default to content_policy:"protected". Secret-shaped content detection may reject an entire page, including ordinary code that matches the detector. A caller may explicitly choose content_policy:"exact" to deliver unchanged native text to the MCP caller for that call. This may expose sensitive content; it is not saved as a session preference and has no automatic fallback. Protected mode provides default detection and an explicit choice point, not an access boundary that prevents the supervisor from obtaining the original content.

Successful Goal and Queue reads and responses preserve native values without secret-shape filtering of their bodies. Known oversized echoed input is rejected before a write; native-added fields can still produce an acknowledged mutation whose result cannot be delivered. Do not resubmit the mutation for that reason. For an oversized multi-item page, a smaller limit may help. If one item remains oversized, inspect it on the native side; repeatedly making the same Bridge read cannot recover it. Errors and operational diagnostics remain redacted.

These exact responses do not inherit Observe's short-text, internal-array, or field-count budgets. They remain subject to actual byte limits, required-field checks, and defensive serialization checks. Failure never returns a partial page or fabricated cursor. See the protocol details.

In compact view, terminal.final_result_pending: true means final text awaits a later page, and final_result_meta is omitted until then. Follow next_cursor even if the status is terminal. Once the text arrives, inspect final_result_meta.complete: false means source text was not fully received or live retention clipped it. Raw and terminal text also have a 48k cap. If more content is needed, read narrow History for that thread and terminal turn, or inspect native Codex. Ordinary compact-event budgets remain unchanged.


Typical supervision workflow

1. Start or continue a turn

A successful codex_turn response only means native turn/start was accepted. It does not mean the task is complete.

Long-running work should usually continue under codex_observe supervision.

codex_turn
    ↓
codex_observe
    ↓
 ┌───────────────┬────────────────┬─────────────────┐
 │ continue      │ steer          │ respond         │
 │ observing     │ same turn      │ real pending    │
 │               │                │ request         │
 └───────────────┴────────────────┴─────────────────┘
    ↓
terminal state / acceptance

Key rules:

  • A long interval without new command output does not prove Codex is stalled.

  • Use codex_steer for new semantic information or a correction, not a timed nudge.

  • codex_respond can answer only a real pending request.

  • Use codex_interrupt when the current turn actually needs to stop.

  • thread_id is a native Codex thread identity, not a permanent task ID invented by Bridge.

2. Keep live state separate from persistent history

codex_observe serves current live supervision; codex_history reads native persisted turns and items.

After Bridge restarts, loses its runtime ring, or no longer has a live cursor, it does not manufacture a seemingly active runtime from history. Read codex_history to recover persistent content.

codex_search locates content; History reads it. An occurrence's turnCursor can anchor a narrow History read in the same thread, but Search does not replace History.

3. Keep Goal, Queue, and Steer distinct

  • Goal: the persistent objective of a native thread.

  • Queue: follow-up input that native Codex executes after the current active workflow.

  • Steer: an immediate semantic correction to the current active turn.

These are not three spellings of a “next prompt,” and Bridge does not combine them into its own task model.

codex_goal(action:"set") requires an explicit budget intent:

budget_mode

Native tokenBudget

Argument requirement

preserve

Omitted, preserving the existing budget

Do not pass token_budget

unlimited

null, removing the budget ceiling

Do not pass token_budget

fixed

Specified amount

Pass a positive safe-integer token_budget

Bridge sends no extra resume or turn-start request for Goal set, get, or clear. An active Goal may still cause native Codex to continue execution. Goal clear is not an interrupt or a barrier against work already scheduled. Use the exact codex_interrupt when the current turn must stop.

Queue delete likewise does not interrupt a turn that has already begun.


codex_observe

The default view:"compact" delivers bounded, typed supervision facts. Use view:"raw" for a narrow inspection of retained native events.

Optional wait_ms performs one event-driven wait with a fixed deadline, up to 120000 ms. Bridge does not poll in the background or decide on its own that Codex is stalled.

Keep these distinctions in mind:

  • Continue either view with next_cursor.

  • stream_lost reports evicted streaming deltas.

  • facts_lost reports evicted other supervision facts.

  • cursor_lost summarizes either kind of loss; it does not tell you to jump to cursor_floor.

  • With runtime_available:false, live pending, terminal, and cursor fields may only be unavailable placeholders. Do not infer that nothing happened.

When runtime_available:true, compact pending_requests is a complete current snapshot. Compact omits it when empty; absence means there are currently no pending requests, not that this page contains no update.

Compact terminal.final_result is also subject to the delivery window of the final item and terminal cursor. Final text may already have arrived as a message fact, so a compact terminal on one page may contain only status and error, without final_result. An absent field does not mean the raw terminal snapshot or persisted History lacks final text. Do not skip next_cursor or infer status just to force these fields onto one page.


Terminal state: use status

Always classify a terminal by native terminal.status first.

Do not infer terminal status from other fields:

  • final_result does not mean the turn succeeded.

  • terminal.error does not necessarily mean status:"failed".

  • A turn with status:"interrupted" may carry a structured error.

  • Bridge does not need a separate error notification to learn the terminal reason.

Read terminal.status first, then interpret final_result and error as separate facts about that terminal.


Successful results use structuredContent

In V2.3.0, the complete result of every successful tools/call is in:

result.structuredContent

result.content contains only a tiny text marker, not parseable success JSON.

Older clients that parse successful results from text content need to migrate to structuredContent.

Errors still use explicit isError and text-error paths; they are not disguised as successful results.


Do not retry a write just because it timed out

If a native mutating request was written to app-server but its acknowledgement timed out, Bridge reports the outcome as:

UNKNOWN / possibly accepted

This is not a confirmed failure.

Affected requests include:

  • thread/start

  • thread/resume

  • turn/start

  • turn/steer

  • turn/interrupt

  • thread/goal/set / thread/goal/clear

  • thread/queue/add / update / delete / reorder

Observe or read native state first, then decide on any further action.

Do not resubmit a mutating request solely because its acknowledgement timed out.

Bridge does not automatically retry, compensate, or guess the result.


codex_history follows the native history mode:

  • Paginated thread: page the turn index, then read items in a specified turn.

  • Legacy thread: page complete turns.

Bridge creates no item/chunk cursors of its own and does not silently read a full history after a paging failure.

A successful History page must be delivered losslessly. A page that exceeds the transport byte boundary, triggers the selected content policy, or fails defensive JSON serialization is rejected in full. Bridge does not return partial data with a cursor.

codex_search maps native Search only:

  • kind:"threads": locators across threads.

  • kind:"occurrences": occurrence locators inside a specified paginated thread.

Search does not build a Bridge semantic index, score relevance, infer a workspace restriction, or fall back to reading complete History.

See PROTOCOL-ASSUMPTIONS.md for exact paging, cursor, and transport contracts.


Capability and lineage

codex_threads can expose native capability and lineage metadata on demand, including:

  • canAcceptDirectInput

  • sessionId

  • forkedFromId

  • parentThreadId

  • Source information

These are native metadata, not a Bridge writer lease, permission model, or lifecycle database.

parent_thread_id and ancestor_thread_id can filter spawned descendants. Spawn lineage and fork lineage are different concepts. Bridge does not recursively build a relationship graph or automatically resume or take over a thread from lineage.


Model and reasoning effort

Bridge does not take ownership of Codex's “current model” state.

If codex_turn omits model and effort, Bridge:

  • Does not call model/list.

  • Does not infer the current model.

  • Does not send a new model or effort override.

If the supervisor specifies a model, Bridge performs bounded validation against the current native catalog. It does not cache the catalog or create a current-model registry.

Where upstream metadata cannot establish that a model and effort combination is incompatible, Bridge does not guess. Native Codex makes the final decision.


Elicitation is currently unsupported

mcpServer/elicitation/request is not currently part of the response surface supported by codex_respond.

If native Codex emits such a request, Bridge:

  • Retains and exposes it.

  • Does not silently discard it.

  • Does not guess its response schema.

  • Does not fabricate a generic answer.

Support would require an explicit, stable, validated upstream contract.


Quick start

Requirements

  • Windows or macOS

  • Node.js 24+

  • Official Codex executable:

    • Available as codex on PATH, or

    • Specified explicitly with CODEX_EXE.

This project does not bundle or depend on the @openai/codex npm package.

Clone, build, and test

git clone https://github.com/zoeynine/Local-Codex-Bridge.git
cd Local-Codex-Bridge
npm ci
npm run typecheck
npm run build
npm test

Start directly

Windows PowerShell:

$env:CODEX_EXE = 'C:\path\to\codex.exe' # Omit if codex is on PATH
npm start

macOS / POSIX shell:

CODEX_EXE=/path/to/codex npm start

You may omit CODEX_EXE if codex is already on PATH.

Configure an MCP client

A strict MCP stdio client should launch the built Node entry directly:

command: node
args:    <absolute-path-to-repository>/dist/src/index.js
env:     CODEX_EXE=<optional-path-to-codex>

Client configuration formats vary, but the final command should run directly:

node <repository>/dist/src/index.js

Do not put npm start behind Secure MCP Tunnel or another strict JSON-RPC stdio transport. npm lifecycle output can contaminate protocol stdout.

After the Bridge tool set changes, an already connected MCP client usually needs to reconnect or restart to refresh its tool catalog.


Optional: Secure MCP Tunnel

Remote MCP use can place Secure MCP Tunnel in front of Bridge:

remote MCP client
        ↕
Secure MCP Tunnel
        ↕
node <repository>/dist/src/index.js
        ↕
native Codex

Tunnel authentication, profile, port, readiness endpoint, and process lifecycle are external configuration.

This repository:

  • Does not create a Tunnel profile.

  • Does not store production credentials.

  • Does not hard-code a production port.

  • Does not turn the Tunnel control plane into a Bridge HTTP API.


Windows

The optional Tray in windows/ is a lightweight launch and status layer for an installed Tunnel client. It is not required for the Bridge core.

The canonical launcher is LocalCodexBridgeTray.*.

Local settings template:

windows/local-settings.example.json

The actual windows/local-settings.json remains ignored and is not committed.

Configuration precedence:

  1. Explicit command-line arguments.

  2. LOCAL_CODEX_BRIDGE_* environment variables.

  3. Legacy LUMEN_CODEX_V2_* environment variables.

  4. Ignored local settings.

The Tray does not restart Tunnel automatically. It stops a process started by the current Tray instance only after rechecking that process identity, profile, PID, and related information still match.


macOS

Start Mac Codex Bridge.app, launcher/, and bin/start-production-tunnel provide macOS Finder and Tunnel integration.

They are platform layers. The Bridge still runs the same entry:

dist/src/index.js

After changing the launcher or Finder bundle, rebuild and validate on macOS 12+:

launcher/build-launcher.sh
npm run test:macos

Windows and macOS are two platform entry points to one Bridge, not separate implementations.


Security and trust boundaries

Local Codex Bridge does not create a new operating-system sandbox.

Actual file, command, network, and process capabilities still depend on native Codex configuration and each turn's:

  • sandbox

  • approval_policy

For example:

  • danger-full-access broadens the file, command, and process access allowed by the sandbox.

  • approval_policy=never does not expand the OS sandbox by itself, but removes the interactive approval layer.

These are separate risk dimensions.

Also keep in mind:

  • Natural-language instructions sent through codex_turn or codex_steer may lead Codex to use its existing file and command capabilities.

  • The absence of a generic shell MCP tool in Bridge does not mean native Codex cannot execute commands.

  • codex_threads can see persistent threads visible to the same OS user and Codex runtime. Filters do not provide access isolation.

  • Bridge inherits its environment when starting app-server, except that it removes the Tunnel CONTROL_PLANE_API_KEY.

  • Other environment variables remain part of the trusted launch boundary; avoid unnecessary secrets there.

  • Bridge has transport sanitization and bounded projection, but is not a hostile multi-tenant gateway.

  • Keep checkpoints short; do not store full prompts, transcripts, raw events, command output, or final answers.

For remote use, an authenticated and correctly configured Tunnel must provide the connection boundary.


Persistence

Native Codex persists:

  • Threads.

  • Turns.

  • Conversation history.

  • Native execution results.

  • Thread goals.

  • Queued follow-ups.

Bridge's live event ring, active-turn runtime state, and pending requests are primarily bounded in-memory state.

After ring loss, Bridge does not reconstruct fake live state from History. stream_lost, facts_lost, and cursor_lost explicitly report missing live supervision evidence.

Checkpoint

codex_checkpoint is the one intentionally persisted piece of Bridge-side supervisory state, and it remains concise and bounded.

Windows default:

%LOCALAPPDATA%\LocalCodexBridge\checkpoints\<sha256(thread_id)>.json

macOS default:

~/Library/Application Support/LocalCodexBridge/checkpoints/<sha256(thread_id)>.json

Override with:

LOCAL_CODEX_BRIDGE_CHECKPOINT_DIR

The legacy LUMEN_CODEX_V2_CHECKPOINT_DIR is still supported for compatibility. Bridge does not automatically migrate old checkpoints.


Deliberate non-goals

Local Codex Bridge deliberately does not provide:

  • A browser UI.

  • An HTTP control plane or HTTP MCP server.

  • A second task queue or job database.

  • Transcript duplication.

  • A semantic index.

  • A model cache or current-model registry.

  • A queued-message facade.

  • Automatic retries of mutating requests.

  • Automatic app-server restart.

  • A generic shell or command/exec MCP surface.

  • Automatic exposure of every experimental app-server API.

The aim is not to copy all of Codex app-server into MCP, but to expose the smallest validated surface needed for supervision.


Upgrading Codex

Bridge necessarily depends on a small set of native app-server protocol assumptions.

Current dependencies, validation status, corresponding code locations, and checks needed after upstream changes are collected in:

PROTOCOL-ASSUMPTIONS.md

When upgrading the Codex runtime, changing protocol-facing behavior, or investigating a related regression, check that list first. Do not change Bridge solely on the basis of an older implementation assumption.


Development and testing

Common checks:

npm run typecheck
npm run build
npm test

npm test runs shared runtime, compact, History, Goal, Queue, Search, app-server, MCP, checkpoint, platform, shutdown, UX, and related tests, then the tests for the current platform.

The protocol/compact-schema check requires CODEX_EXE to explicitly identify the official Codex executable under test. This script does not use the PATH fallback of the normal Bridge startup path.

Windows PowerShell:

$env:CODEX_EXE = 'C:\path\to\codex.exe'
npm run check:compact-schema

macOS / POSIX shell:

CODEX_EXE=/path/to/codex npm run check:compact-schema

The repository retains the historical npm run smoke:live helper, but the current script has not been adapted to V2.3.0 structuredContent, independent History, or the runtime-loss contract. It is not a valid smoke or cross-platform acceptance command for this release. Do not run it merely to qualify a release. If a live smoke test is needed later, update the script separately and review its persistent-thread side effects first.

Main implementation locations:

  • src/mcp.ts — MCP stdio / JSON-RPC boundary

  • src/app-server.ts — native Codex app-server process / protocol adapter

  • src/tools.ts — 12 tools, schemas, and supervisory semantics

  • src/runtime.ts — bounded live runtime / events / pending requests

  • src/observe-compact.ts — compact typed supervision projection

  • src/history.ts — lossless History delivery checks

  • src/goal.ts — native Goal validation / delivery

  • src/queue.ts — native Queue validation / delivery

  • src/search.ts — native Search validation / delivery

  • src/checkpoint.ts — optional supervisory checkpoint

  • src/platform.ts — Windows / macOS platform boundary

  • src/version.ts — canonical Bridge version

  • windows/ — optional Windows Tray

  • launcher/, bin/, Start Mac Codex Bridge.app — optional macOS integration


License

MIT License — see LICENSE.

Collaborators and acknowledgements

Collaborators: Xiaonian (ChatGPT) and Alden / Xingjian (Codex).

Thank you for helping turn the small idea of letting an external AI genuinely supervise native Codex, step by step, into a Bridge thin enough and clear enough to share and build on. (*╹▽╹*)

And thank you to Yu'an. Without you, I would not have tried to do something at all. ღ( ´・ᴗ・` )

Available Tools

7 tools
codex_checkpointCheckpoint Codex SupervisionA

Optional, bounded supervisor cognition memory keyed to one native Codex thread_id; the key is not a permanent task identity and does not require future work to remain on that thread. Use it to protect the original goal, constraints, and acceptance plus concise supervisor state during long or complex supervision when context dilution or goal drift makes an external anchor worthwhile. Initialization is not tied to crossing a ChatGPT window or round, starting another Codex turn, or switching native threads; initialize early when a task is already expected to be sufficiently long or complex for that protection. Do not use for one-shot work, and do not turn duration into a hard threshold: elapsed time, observe/poll count, token count, or mere silence are not automatic triggers. Later updates remain semantic-event driven and require a material change in understanding or root cause, constraint or scope interpretation, steering decision, user-authorized amendment or effective goal, or acceptance judgment or an explicit decision not to accept yet. Before final acceptance of a checkpointed task, read it once to re-anchor the original goal, constraints, acceptance, and current supervisor frame. This tool is optional and uncoupled from all other tools. Store concise supervisor summaries only; never prompts, transcripts, raw events, command output, final answers, or raw event streams. Updates preserve only immutable original plus bounded previous/current supervisor state.

ParametersJSON Schema
NameRequiredDescriptionDefault
actionYesRead the checkpoint, or initialize/update it at a material supervisor decision point.
next_stepNoSingle next supervision step.
thread_idYesNative Codex thread id; no second task identifier is created.
original_goalNoConcise original user goal. Required only on initialization and immutable thereafter.
effective_goalNoCurrent effective goal after legitimate user amendments; defaults to original_goal on initialization.
current_decisionNoCurrent supervisor decision and why it matters.
acceptance_statusNoConcise acceptance assessment, not a task lifecycle or job status.
current_amendmentNoLatest concise user-authorized requirement amendment, or null to clear it, without changing the immutable original.
original_acceptanceNoConcise original acceptance criteria. Required only on initialization and immutable thereafter.
original_constraintsNoConcise original constraints. Required only on initialization and immutable thereafter.
current_understandingNoCurrent concise root-cause or task understanding.

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only say readOnlyHint=false, but the description adds rich behavioral context: key semantics (not a permanent task identity), update triggers (semantic event, material change), storage restrictions (never prompts/transcripts), and the requirement to read before final acceptance. This goes far beyond the annotations and is consistent with them.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is longer than average (~180 words), but every sentence earns its place by covering purpose, usage, exclusions, storage policy, and update semantics. It is well-structured, with clear statements and prohibitions, though some length could be trimmed without losing content.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (11 parameters, multiple update conditions, acceptance process), the description is complete: it explains when to use, what to store, how updates work, and the read-before-acceptance rule. No output schema is present, but the description focuses on behavior and constraints, which is sufficient for an agent to use it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema already describes all 11 parameters (100% coverage). The description adds semantic meaning beyond the schema by explaining the immutable-vs-mutable distinction (original vs current/effective goal, original constraints, etc.), what should not be stored, and the relationship between parameters like original and effective goals. This adds value without repeating schema details.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: an optional, bounded supervisor cognition memory keyed to a Codex thread_id, used to protect the original goal, constraints, acceptance, and supervisor state. It distinguishes this tool from siblings by emphasizing it is uncoupled and optional, and by being specific about its function as a checkpoint for supervision context.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit when-to-use guidance (long/complex supervision with context dilution or goal drift) and when-not-to-use guidance (one-shot work, no hard duration thresholds). It also notes the tool is optional and uncoupled from other tools, helping an agent decide when to invoke it versus alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

codex_interruptInterrupt Codex TurnA
DestructiveIdempotent

Directly request turn/interrupt for the specified active Codex thread and turn. It does not stop or restart the Bridge or Codex app-server processes.

ParametersJSON Schema
NameRequiredDescriptionDefault
turn_idYesActive Codex turn to interrupt.
thread_idYesActive Codex thread.

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate destructive and idempotent hints, but the description adds valuable context by clarifying that the interrupt does not stop or restart Bridge/app-server processes. This reduces risk of misuse, even though it doesn't specify async behavior or effects on already-completed turns.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with the action, and every part adds value. No filler or redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple two-parameter tool with annotated destructiveness, the description sufficiently covers purpose and boundaries. No output schema means return-value details are not expected. Minor gap: no mention of whether interrupt is asynchronous or what happens if the turn is not active.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% with clear descriptions for thread_id and turn_id. The description merely repeats 'active thread and turn' without adding meaning beyond the schema, so baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb phrase ('Directly request turn/interrupt') and names the exact resource ('specified active Codex thread and turn'). It clearly distinguishes from sibling tools by stating it does not stop or restart processes, matching the interrupt-only scope.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides clear context: it is for sending an interrupt request to an active thread/turn. The exclusion of process-level control is stated, which implicitly differentiates from broader lifecycle tools. However, it does not explicitly name alternatives or situations where a different tool would be preferred.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

codex_observeObserve Codex TurnA
Read-onlyIdempotent

Read bounded incremental sanitized Bridge runtime events, pending requests, and terminal output for a thread. Optional wait_ms performs one bounded event-driven wait only when the live turn is active and the current snapshot has nothing useful; it is not polling or stall detection. After Bridge process loss, falls back to persistent thread/read history and marks live state unreconstructable. A long interval with no new command or output can still mean Codex is actively reasoning; absence of new command activity alone is not evidence of a stall. When actively supervising an in-progress turn, use repeated bounded-wait observe calls until terminal unless the user explicitly pauses or stops; do not end supervision merely because one snapshot is inProgress. After every wake or deadline return, inspect the newly available events/state and decide whether steer, respond, or interruption is needed before starting the next bounded wait.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMaximum runtime events to return.
cursorNoReturn runtime events with a cursor greater than this value.
wait_msNoOptional per-call wait for the next live runtime change when nothing useful is ready; 0 returns immediately. This is event-driven waiting, not stall detection.
thread_idYesCodex thread to observe.

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description goes well beyond the annotations by documenting fallback behavior ('After Bridge process loss, falls back to persistent thread/read history'), the precise semantics of wait_ms ('one bounded event-driven wait only when the live turn is active'), and the caveat that a long interval without new commands is not evidence of a stall. It reveals useful runtime behavior such as 'inProgress' snapshots and live-state unreconstructability without contradicting the read-only/idempotent annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long but front-loaded with the core action, followed by wait semantics, fallback behavior, and a supervision workflow. Every sentence carries actionable information, though there is minor redundancy with the schema's wait_ms explanation; the density is appropriate for the tool's complexity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Despite having no output schema, the description explains the categories of returned data ('Bridge runtime events, pending requests, and terminal output'), the fallback path, and how the agent should react after each wake/deadline. It could be more explicit about the exact shape of the returned 'events/state' snapshot, but it provides strong contextual coverage for a complex observation tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema already fully describes all four parameters (100% coverage), including wait_ms's event-driven nature. The tool description adds meaningful nuance by restricting wait_ms to 'only when the live turn is active' and emphasizing 'one bounded' wait, which clarifies the operational contract beyond the schema text.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Read') and precisely names the resources: 'Bridge runtime events, pending requests, and terminal output for a thread.' It clarifies the observational scope with 'bounded incremental sanitized' and clearly distinguishes this from the sibling turn-management tools by framing it as the supervision/observation primitive.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly states when to use it: 'When actively supervising an in-progress turn, use repeated bounded-wait observe calls until terminal unless the user explicitly pauses or stops.' It also provides exclusions ('it is not polling or stall detection') and warns against treating inactivity as a stall, while naming follow-up actions (steer, respond, interruption).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

codex_respondRespond to Codex RequestA
Destructive

Answer one currently pending app-server request by its original raw JSON-RPC id and exact thread/method scope. Supports only command/file approval methods with concrete response contracts and item/tool/requestUserInput; unsupported methods remain pending and observable.

ParametersJSON Schema
NameRequiredDescriptionDefault
methodYesExact app-server request method.
answersNorequest_user_input question-id to answer-array mapping.
turn_idNoExact turn scope when the pending request has one.
decisionNoCommand or file approval decision.
responseNoExact result object for the known item/tool/requestUserInput method.
thread_idYesExact pending-request thread scope.
request_idYesOriginal app-server JSON-RPC request id, preserving string or integer type.
execpolicy_amendmentNoCommand approval exec-policy amendment; encoded in app-server's native decision shape.

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true and readOnlyHint=false. The description adds useful behavioral context by noting that unsupported methods 'remain pending and observable' and that each call answers exactly one request, which goes beyond the annotation data.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact at two sentences. The first sentence conveys the core action and target; the second adds critical constraints (supported methods, behavior for unsupported ones). No wasted words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (8 parameters, nested objects, no output schema), the description captures the essential purpose, supported method types, and a key behavioral trait. It could emphasize destructive side effects more, but the annotations already indicate destructiveness, so the description is sufficiently complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description adds a little semantic grouping (e.g., 'command/file approval methods' for decision/execpolicy_amendment, 'item/tool/requestUserInput' for answers/response) but does not significantly go beyond the schema's own descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's verb ('Answer') and resource ('one currently pending app-server request') with specific scope markers (original JSON-RPC id, exact thread/method). It also distinguishes itself from siblings by focusing on approval/response methods and explicitly noting that unsupported methods remain pending and observable.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear context on when to use the tool (when there is a pending request) and explicitly excludes unsupported methods ('Supports only...'; unsupported methods remain pending). It does not name alternatives directly, but the scope and exclusions are clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

codex_steerSteer Active Codex TurnA
Destructive

Append text to the same active Codex turn using turn/steer with an expected turn-id precondition. This does not create a new turn. Do not steer merely because reasoning is taking a long time or no new command has appeared; steer only for a semantic redirect or correction based on new evidence or changed user intent.

ParametersJSON Schema
NameRequiredDescriptionDefault
textYesAdditional user text.
thread_idYesActive Codex thread.
expected_turn_idYesExact active turn id required by app-server.

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

While annotations already indicate destructive behavior, the description adds context by specifying the expected_turn_id precondition and clarifying that the operation modifies the existing turn rather than creating one. No contradictions with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise and well-structured: the first sentence states the action, the second clarifies scope, and the third provides targeted usage guidance. Every sentence earns its place with no redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's moderate complexity and lack of output schema, the description covers purpose, usage guidelines, and key behavioral details. It could mention error scenarios or postconditions, but the destructive annotation and clear action make it sufficiently complete for an agent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema already provides 100% coverage with descriptions for all three parameters. The description reinforces the role of expected_turn_id as a precondition and text as appended content, but adds little novel semantic information beyond what the schema includes.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool appends text to the same active Codex turn, using a specific method (turn/steer) and a precondition. It explicitly says it does not create a new turn, distinguishing it from siblings like codex_turn or codex_threads.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides explicit when-to-use and when-not-to-use guidance: steer only for semantic redirect or correction based on new evidence or changed intent, and not merely because reasoning takes long or no new command appeared. This helps the agent choose appropriately among sibling tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

codex_threadsCodex ThreadsA
Read-onlyIdempotent

List or search persistent local Codex threads through thread/list, or read one thread through thread/read. This does not reconstruct live Bridge events.

ParametersJSON Schema
NameRequiredDescriptionDefault
cwdNoOptional exact absolute Windows drive-letter cwd filter for thread/list.
limitNoMaximum threads in the returned page.
cursorNoOpaque cursor returned by a prior thread/list call.
thread_idNoWhen supplied, read this exact Codex thread instead of listing threads.
search_termNoOptional Codex title substring filter for thread/list.
include_turnsNoInclude persisted turns when reading one thread.

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false. The description adds useful behavioral context beyond this by specifying the threads are 'persistent local' and clarifying that the tool does not reconstruct live events. This is substantial but does not cover all edge cases (e.g., pagination errors, data source specifics), so a 4 is appropriate.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with an unambiguous summary, and no wasted words. Every sentence earns its place: the first states the main actions, the second clarifies an important behavioral boundary.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's dual mode (list/read) and the presence of six documented optional parameters plus comprehensive annotations, the description adequately covers the main purpose and an important caveat. It could briefly mention what the read returns when include_turns is false, but the schema handles this. Overall, it is complete enough for an AI agent to select and invoke correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all parameters. The description adds a small amount of context by naming 'thread/list' and 'thread/read' modes, but this is largely redundant with the schema's thread_id description. Baseline 3 applies because the description does not substantially compensate beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses specific action verbs 'List or search' and 'read', names the resource ('persistent local Codex threads'), and explicitly scopes behavior with 'This does not reconstruct live Bridge events.' It clearly distinguishes from siblings by indicating it handles listing/reading rather than per-thread actions.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description states when to use the tool (to list/search/read persistent local threads) and provides an exclusion ('does not reconstruct live Bridge events'). However, it does not explicitly name an alternative sibling or provide a direct contrast with other tools, so it falls short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

codex_turnStart or Continue Codex TurnA
Destructive

Start a persistent Codex thread and turn, or resume an existing thread and start a turn. Prefer continuing the same native thread when its context remains useful, but a fresh thread is allowed; thread_id is not a permanent task identity. Returns as soon as turn/start is accepted; observe separately for events and completion.

ParametersJSON Schema
NameRequiredDescriptionDefault
cwdNoAbsolute Windows drive-letter cwd. Required for a new thread; optional override for resume.
textYesUser text passed directly to Codex as one text input item.
modelNoOptional model identifier passed through to app-server.
effortNoOptional reasoning effort passed through to turn/start.
sandboxNoCodex app-server sandbox mode override.
thread_idNoExisting persistent Codex thread to resume. Omit to create a new thread.
approval_policyNoCodex app-server approval policy override.

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description adds valuable behavioral details beyond annotations: it returns as soon as the turn is accepted, meaning it is asynchronous, and it warns that thread_id is not a permanent task identity. Annotations already mark the tool as destructive/open-world, so this context complements them without contradiction.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is three sentences with the core action front-loaded, followed by important nuances about thread reuse and asynchronous behavior. Every sentence earns its place with no redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no output schema, the description appropriately clarifies the return value ('returns as soon as turn/start is accepted') and where to get actual results ('observe separately'). It doesn't cover error cases or param interactions, but the schema descriptions and annotations adequately cover those aspects.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema already covers all 7 parameters with descriptions (100% coverage), setting a baseline of 3. The description adds a useful caveat about thread_id not being permanent, but does not significantly enrich parameter understanding beyond what the schema provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool starts or resumes a Codex turn on a persistent thread, using a specific verb+resource. It distinguishes itself from siblings by being the entry point to initiate/continue turns, unlike interrupt/observe/steer.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It provides explicit guidance on when to reuse a thread ('Prefer continuing the same native thread when its context remains useful') and when a fresh thread is allowed. It also directs the agent to 'observe separately for events and completion', indicating this tool is not for getting results directly.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 7 tool updatesv2.1.1
    • First observedcodex_checkpoint
    • First observedcodex_interrupt
    • First observedcodex_observe
    • First observedcodex_respond
    • First observedcodex_steer
    • First observedcodex_threads
    • First observedcodex_turn

TDQS

A4.4/5.0

Scored across 7 tools

Disambiguation5/5

Each tool has a clearly distinct role in the Codex Bridge supervision lifecycle: listing/reading threads, starting turns, interrupting, observing events, steering, responding to requests, and checkpointing. Even read-like tools (codex_threads vs codex_observe) are cleanly separated by persistent history vs live runtime events.

Naming Consistency4/5

All tools share the codex_ prefix and use snake_case, but the pattern is not perfectly uniform: most are verb-based (interrupt, observe, steer, respond, turn), while codex_threads is a plural noun and codex_checkpoint is a compound noun. Minor deviation, but predictable and readable.

Tool Count5/5

Seven tools is a well-scoped set for the server's purpose of supervising Codex threads. Each tool addresses a distinct supervision operation without redundancy or bloat, fitting comfortably in the ideal 3-15 range.

Completeness5/5

The tool surface covers the full supervision lifecycle: create/resume (codex_turn), observe, steer, interrupt, respond to pending requests, and persist supervisor state (codex_checkpoint), plus listing/reading past threads. No obvious dead ends or missing operations for the stated purpose.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    B
    maintenance
    Enables ChatGPT to supervise and control local Codex runtime sessions through a secure MCP interface, managing threads, turns, approvals, events, and recovery without exposing direct file, shell, Git, SSH, or model loop access.
    1
    MIT
  • A
    license
    Not graded
    quality
    B
    maintenance
    Enables MCP clients and external AI supervisors to oversee and steer native Codex sessions through a thin local stdio bridge. It exposes eleven codex_* supervisory tools for tasks such as listing threads, starting turns, observing progress, steering, responding to approvals, interrupting, checkpointing, and rolling over work.
    MIT
  • A
    license
    Not graded
    quality
    A
    maintenance
    Enables ChatGPT on Windows to securely access codebases, Git, terminals, language servers, debuggers, SQLite, local HTTP services, adaptive project memory, engineering skills, checkpoints, audit logs, and sandboxed command execution through MCP tools.
    MIT