Local Codex Bridge
Supervise native Codex threads through MCP by inspecting threads, starting/resuming turns, observing live activity, steering, responding, interrupting, and keeping bounded checkpoints.
List or filter persistent Codex threads and read one thread's metadata/turns with
codex_threads.Start a new Codex thread or resume an existing one and begin a turn with text, model, effort, sandbox, and approval policy via
codex_turn.Observe bounded live runtime events, pending requests, and terminal output for a thread, with optional short event-driven wait via
codex_observe.Add a semantic correction to an active turn using
codex_steer.Answer pending app-server approval, user-input, or permission requests by exact request ID and scope via
codex_respond.Interrupt a specific active Codex thread and turn with
codex_interrupt.Read or update an optional concise supervisor checkpoint for a thread (goal, constraints, decisions, acceptance) via
codex_checkpoint.
Provides tools to control native Codex sessions, enabling MCP clients to create and resume threads, observe real-time events, steer active turns, respond to approval requests, interrupt execution, and save bounded supervision checkpoints for long-running development tasks.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Local Codex BridgeStart a new Codex thread in C:\projects\myapp to fix the failing tests."
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Local Codex Bridge
A thin supervisory MCP bridge between external AI supervisors and native Codex.
Current release: V2.3.4
Local Codex Bridge is a lightweight MCP stdio adapter for Windows and macOS:
ChatGPT / external AI supervisor
↕
Local Codex Bridge
↕
native Codex app-server
↕
native Codex threads / turnsIt lets an AI that is good at conversation, planning, and sustained supervision oversee real engineering work performed by local native Codex. It does not recreate Codex.
The supervisor owns the goal, resources, boundaries, risks, approvals, and acceptance. Codex retains its native autonomy for coding and execution.
The Bridge stays thin:
It creates no second job or task system.
It does not copy Codex conversation history.
It keeps no parallel thread database.
It does not cache the “current model.”
It does not rebuild Goal, Queue, or Search.
It does not replace Codex session, thread, or turn semantics.
The native Codex thread/session remains the source of truth for execution.
Changelog · Protocol and compatibility assumptions
What V2.3.0 introduced
V2.3.0 expanded the public supervisory surface to 12 MCP tools while retaining the boundary of exposing only the native surface needed for supervision.
The release added or refined:
Independent paging of native persistent History.
Native Goal: read, set, and clear a persistent objective.
Native Queue: manage follow-up input for an active workflow.
Native Search: search across threads and locate occurrences within a thread.
On-demand capability and lineage metadata.
Typed compact observation, drainage, and loss semantics.
A single event-driven observe wait of up to 120 seconds.
Delivery of every successful tool result through
structuredContent.One shared Bridge core for Windows and macOS, with differences confined to native platform boundaries.
Bridge does not reimplement these capabilities. Native Codex still owns History, Goal, Queue, Search, and the thread lifecycle; Bridge provides bounded mappings and supervision.
Related MCP server: Codex Bridge
Responsibilities
External supervisor / ChatGPT
The supervisor is suited to:
Understand the user's goal.
Break down the work.
Set scope, resources, and risk boundaries.
Decide when to keep observing, correct, approve, or interrupt.
Judge whether the result meets acceptance criteria.
Provide oversight when Codex cannot safely decide on its own.
Native Codex
Native Codex continues to own:
Thread and turn lifecycles.
Workspace files and command execution.
Its context and persistent history.
Sandbox and approval-policy behavior.
The actual model and reasoning effort in use.
Persisted execution results.
Native Goal, Queue, Search, and lineage.
Local Codex Bridge
Bridge connects the two:
MCP stdio ↔ Codex app-server JSONL.
Bounded exposure of state needed for supervision.
Forwarding explicit control intent.
Failing closed at high-risk, ambiguous, or protocol boundaries.
No second orchestration runtime.
The 12 MCP tools
Tool | Purpose | Main boundary |
| List, filter, and read native persistent thread metadata | Metadata only; filters are not ACLs |
| Page native persistent history | No Bridge history store or automatic full read |
| Search native threads and locate within-thread occurrences | Returns locators; no Bridge index or relevance layer |
| Read one native | No cached catalog or current-model registry |
| Read, set, or clear a native thread goal | No implicit resume or turn start; distinct from checkpoint |
| Manage native queued follow-ups | No Bridge scheduler; enqueue does not mean execution or completion |
| Create or resume a thread and start a turn | Accepted does not mean completed |
| Read bounded live events, pending requests, terminal state, and cursor | Does not reconstruct live state after runtime loss |
| Add a semantic correction to the current active turn | Not a timer, poll, or retry |
| Answer a real pending approval, user-input, or permission request | Must match the original request ID and exact scope |
| Interrupt the exact active thread and turn | Not process control |
| Keep an optional, concise, bounded supervisor anchor | Not a transcript, job ID, or Codex history |
See src/tools.ts for the full schemas and runtime validation.
Content delivery and read policy
History and Search default to content_policy:"protected". Secret-shaped content detection may reject an entire page, including ordinary code that matches the detector. A caller may explicitly choose content_policy:"exact" to deliver unchanged native text to the MCP caller for that call. This may expose sensitive content; it is not saved as a session preference and has no automatic fallback. Protected mode provides default detection and an explicit choice point, not an access boundary that prevents the supervisor from obtaining the original content.
Successful Goal and Queue reads and responses preserve native values without secret-shape filtering of their bodies. Known oversized echoed input is rejected before a write; native-added fields can still produce an acknowledged mutation whose result cannot be delivered. Do not resubmit the mutation for that reason. For an oversized multi-item page, a smaller limit may help. If one item remains oversized, inspect it on the native side; repeatedly making the same Bridge read cannot recover it. Errors and operational diagnostics remain redacted.
These exact responses do not inherit Observe's short-text, internal-array, or field-count budgets. They remain subject to actual byte limits, required-field checks, and defensive serialization checks. Failure never returns a partial page or fabricated cursor. See the protocol details.
In compact view, terminal.final_result_pending: true means final text awaits a later page, and final_result_meta is omitted until then. Follow next_cursor even if the status is terminal. Once the text arrives, inspect final_result_meta.complete: false means source text was not fully received or live retention clipped it. Raw and terminal text also have a 48k cap. If more content is needed, read narrow History for that thread and terminal turn, or inspect native Codex. Ordinary compact-event budgets remain unchanged.
Typical supervision workflow
1. Start or continue a turn
A successful codex_turn response only means native turn/start was accepted. It does not mean the task is complete.
Long-running work should usually continue under codex_observe supervision.
codex_turn
↓
codex_observe
↓
┌───────────────┬────────────────┬─────────────────┐
│ continue │ steer │ respond │
│ observing │ same turn │ real pending │
│ │ │ request │
└───────────────┴────────────────┴─────────────────┘
↓
terminal state / acceptanceKey rules:
A long interval without new command output does not prove Codex is stalled.
Use
codex_steerfor new semantic information or a correction, not a timed nudge.codex_respondcan answer only a real pending request.Use
codex_interruptwhen the current turn actually needs to stop.thread_idis a native Codex thread identity, not a permanent task ID invented by Bridge.
2. Keep live state separate from persistent history
codex_observe serves current live supervision; codex_history reads native persisted turns and items.
After Bridge restarts, loses its runtime ring, or no longer has a live cursor, it does not manufacture a seemingly active runtime from history. Read codex_history to recover persistent content.
codex_search locates content; History reads it. An occurrence's turnCursor can anchor a narrow History read in the same thread, but Search does not replace History.
3. Keep Goal, Queue, and Steer distinct
Goal: the persistent objective of a native thread.
Queue: follow-up input that native Codex executes after the current active workflow.
Steer: an immediate semantic correction to the current active turn.
These are not three spellings of a “next prompt,” and Bridge does not combine them into its own task model.
codex_goal(action:"set") requires an explicit budget intent:
| Native | Argument requirement |
| Omitted, preserving the existing budget | Do not pass |
|
| Do not pass |
| Specified amount | Pass a positive safe-integer |
Bridge sends no extra resume or turn-start request for Goal set, get, or clear. An active Goal may still cause native Codex to continue execution. Goal clear is not an interrupt or a barrier against work already scheduled. Use the exact codex_interrupt when the current turn must stop.
Queue delete likewise does not interrupt a turn that has already begun.
codex_observe
The default view:"compact" delivers bounded, typed supervision facts. Use view:"raw" for a narrow inspection of retained native events.
Optional wait_ms performs one event-driven wait with a fixed deadline, up to 120000 ms. Bridge does not poll in the background or decide on its own that Codex is stalled.
Keep these distinctions in mind:
Continue either view with
next_cursor.stream_lostreports evicted streaming deltas.facts_lostreports evicted other supervision facts.cursor_lostsummarizes either kind of loss; it does not tell you to jump tocursor_floor.With
runtime_available:false, live pending, terminal, and cursor fields may only be unavailable placeholders. Do not infer that nothing happened.
When runtime_available:true, compact pending_requests is a complete current snapshot. Compact omits it when empty; absence means there are currently no pending requests, not that this page contains no update.
Compact terminal.final_result is also subject to the delivery window of the final item and terminal cursor. Final text may already have arrived as a message fact, so a compact terminal on one page may contain only status and error, without final_result. An absent field does not mean the raw terminal snapshot or persisted History lacks final text. Do not skip next_cursor or infer status just to force these fields onto one page.
Terminal state: use status
Always classify a terminal by native terminal.status first.
Do not infer terminal status from other fields:
final_resultdoes not mean the turn succeeded.terminal.errordoes not necessarily meanstatus:"failed".A turn with
status:"interrupted"may carry a structured error.Bridge does not need a separate error notification to learn the terminal reason.
Read terminal.status first, then interpret final_result and error as separate facts about that terminal.
Successful results use structuredContent
In V2.3.0, the complete result of every successful tools/call is in:
result.structuredContentresult.content contains only a tiny text marker, not parseable success JSON.
Older clients that parse successful results from text content need to migrate to structuredContent.
Errors still use explicit isError and text-error paths; they are not disguised as successful results.
Do not retry a write just because it timed out
If a native mutating request was written to app-server but its acknowledgement timed out, Bridge reports the outcome as:
UNKNOWN / possibly accepted
This is not a confirmed failure.
Affected requests include:
thread/startthread/resumeturn/startturn/steerturn/interruptthread/goal/set/thread/goal/clearthread/queue/add/update/delete/reorder
Observe or read native state first, then decide on any further action.
Do not resubmit a mutating request solely because its acknowledgement timed out.
Bridge does not automatically retry, compensate, or guess the result.
History and Search
codex_history follows the native history mode:
Paginated thread: page the turn index, then read items in a specified turn.
Legacy thread: page complete turns.
Bridge creates no item/chunk cursors of its own and does not silently read a full history after a paging failure.
A successful History page must be delivered losslessly. A page that exceeds the transport byte boundary, triggers the selected content policy, or fails defensive JSON serialization is rejected in full. Bridge does not return partial data with a cursor.
codex_search maps native Search only:
kind:"threads": locators across threads.kind:"occurrences": occurrence locators inside a specified paginated thread.
Search does not build a Bridge semantic index, score relevance, infer a workspace restriction, or fall back to reading complete History.
See PROTOCOL-ASSUMPTIONS.md for exact paging, cursor, and transport contracts.
Capability and lineage
codex_threads can expose native capability and lineage metadata on demand, including:
canAcceptDirectInputsessionIdforkedFromIdparentThreadIdSource information
These are native metadata, not a Bridge writer lease, permission model, or lifecycle database.
parent_thread_id and ancestor_thread_id can filter spawned descendants. Spawn lineage and fork lineage are different concepts. Bridge does not recursively build a relationship graph or automatically resume or take over a thread from lineage.
Model and reasoning effort
Bridge does not take ownership of Codex's “current model” state.
If codex_turn omits model and effort, Bridge:
Does not call
model/list.Does not infer the current model.
Does not send a new model or effort override.
If the supervisor specifies a model, Bridge performs bounded validation against the current native catalog. It does not cache the catalog or create a current-model registry.
Where upstream metadata cannot establish that a model and effort combination is incompatible, Bridge does not guess. Native Codex makes the final decision.
Elicitation is currently unsupported
mcpServer/elicitation/request is not currently part of the response surface supported by codex_respond.
If native Codex emits such a request, Bridge:
Retains and exposes it.
Does not silently discard it.
Does not guess its response schema.
Does not fabricate a generic answer.
Support would require an explicit, stable, validated upstream contract.
Quick start
Requirements
Windows or macOS
Node.js 24+
Official Codex executable:
Available as
codexonPATH, orSpecified explicitly with
CODEX_EXE.
This project does not bundle or depend on the @openai/codex npm package.
Clone, build, and test
git clone https://github.com/zoeynine/Local-Codex-Bridge.git
cd Local-Codex-Bridge
npm ci
npm run typecheck
npm run build
npm testStart directly
Windows PowerShell:
$env:CODEX_EXE = 'C:\path\to\codex.exe' # Omit if codex is on PATH
npm startmacOS / POSIX shell:
CODEX_EXE=/path/to/codex npm startYou may omit CODEX_EXE if codex is already on PATH.
Configure an MCP client
A strict MCP stdio client should launch the built Node entry directly:
command: node
args: <absolute-path-to-repository>/dist/src/index.js
env: CODEX_EXE=<optional-path-to-codex>Client configuration formats vary, but the final command should run directly:
node <repository>/dist/src/index.jsDo not put npm start behind Secure MCP Tunnel or another strict JSON-RPC stdio transport. npm lifecycle output can contaminate protocol stdout.
After the Bridge tool set changes, an already connected MCP client usually needs to reconnect or restart to refresh its tool catalog.
Optional: Secure MCP Tunnel
Remote MCP use can place Secure MCP Tunnel in front of Bridge:
remote MCP client
↕
Secure MCP Tunnel
↕
node <repository>/dist/src/index.js
↕
native CodexTunnel authentication, profile, port, readiness endpoint, and process lifecycle are external configuration.
This repository:
Does not create a Tunnel profile.
Does not store production credentials.
Does not hard-code a production port.
Does not turn the Tunnel control plane into a Bridge HTTP API.
Windows
The optional Tray in windows/ is a lightweight launch and status layer for an installed Tunnel client. It is not required for the Bridge core.
The canonical launcher is LocalCodexBridgeTray.*.
Local settings template:
windows/local-settings.example.json
The actual windows/local-settings.json remains ignored and is not committed.
Configuration precedence:
Explicit command-line arguments.
LOCAL_CODEX_BRIDGE_*environment variables.Legacy
LUMEN_CODEX_V2_*environment variables.Ignored local settings.
The Tray does not restart Tunnel automatically. It stops a process started by the current Tray instance only after rechecking that process identity, profile, PID, and related information still match.
macOS
Start Mac Codex Bridge.app, launcher/, and bin/start-production-tunnel provide macOS Finder and Tunnel integration.
They are platform layers. The Bridge still runs the same entry:
dist/src/index.jsAfter changing the launcher or Finder bundle, rebuild and validate on macOS 12+:
launcher/build-launcher.sh
npm run test:macosWindows and macOS are two platform entry points to one Bridge, not separate implementations.
Security and trust boundaries
Local Codex Bridge does not create a new operating-system sandbox.
Actual file, command, network, and process capabilities still depend on native Codex configuration and each turn's:
sandboxapproval_policy
For example:
danger-full-accessbroadens the file, command, and process access allowed by the sandbox.approval_policy=neverdoes not expand the OS sandbox by itself, but removes the interactive approval layer.
These are separate risk dimensions.
Also keep in mind:
Natural-language instructions sent through
codex_turnorcodex_steermay lead Codex to use its existing file and command capabilities.The absence of a generic shell MCP tool in Bridge does not mean native Codex cannot execute commands.
codex_threadscan see persistent threads visible to the same OS user and Codex runtime. Filters do not provide access isolation.Bridge inherits its environment when starting app-server, except that it removes the Tunnel
CONTROL_PLANE_API_KEY.Other environment variables remain part of the trusted launch boundary; avoid unnecessary secrets there.
Bridge has transport sanitization and bounded projection, but is not a hostile multi-tenant gateway.
Keep checkpoints short; do not store full prompts, transcripts, raw events, command output, or final answers.
For remote use, an authenticated and correctly configured Tunnel must provide the connection boundary.
Persistence
Native Codex persists:
Threads.
Turns.
Conversation history.
Native execution results.
Thread goals.
Queued follow-ups.
Bridge's live event ring, active-turn runtime state, and pending requests are primarily bounded in-memory state.
After ring loss, Bridge does not reconstruct fake live state from History. stream_lost, facts_lost, and cursor_lost explicitly report missing live supervision evidence.
Checkpoint
codex_checkpoint is the one intentionally persisted piece of Bridge-side supervisory state, and it remains concise and bounded.
Windows default:
%LOCALAPPDATA%\LocalCodexBridge\checkpoints\<sha256(thread_id)>.jsonmacOS default:
~/Library/Application Support/LocalCodexBridge/checkpoints/<sha256(thread_id)>.jsonOverride with:
LOCAL_CODEX_BRIDGE_CHECKPOINT_DIRThe legacy LUMEN_CODEX_V2_CHECKPOINT_DIR is still supported for compatibility. Bridge does not automatically migrate old checkpoints.
Deliberate non-goals
Local Codex Bridge deliberately does not provide:
A browser UI.
An HTTP control plane or HTTP MCP server.
A second task queue or job database.
Transcript duplication.
A semantic index.
A model cache or current-model registry.
A queued-message facade.
Automatic retries of mutating requests.
Automatic app-server restart.
A generic shell or
command/execMCP surface.Automatic exposure of every experimental app-server API.
The aim is not to copy all of Codex app-server into MCP, but to expose the smallest validated surface needed for supervision.
Upgrading Codex
Bridge necessarily depends on a small set of native app-server protocol assumptions.
Current dependencies, validation status, corresponding code locations, and checks needed after upstream changes are collected in:
When upgrading the Codex runtime, changing protocol-facing behavior, or investigating a related regression, check that list first. Do not change Bridge solely on the basis of an older implementation assumption.
Development and testing
Common checks:
npm run typecheck
npm run build
npm testnpm test runs shared runtime, compact, History, Goal, Queue, Search, app-server, MCP, checkpoint, platform, shutdown, UX, and related tests, then the tests for the current platform.
The protocol/compact-schema check requires CODEX_EXE to explicitly identify the official Codex executable under test. This script does not use the PATH fallback of the normal Bridge startup path.
Windows PowerShell:
$env:CODEX_EXE = 'C:\path\to\codex.exe'
npm run check:compact-schemamacOS / POSIX shell:
CODEX_EXE=/path/to/codex npm run check:compact-schemaThe repository retains the historical npm run smoke:live helper, but the current script has not been adapted to V2.3.0 structuredContent, independent History, or the runtime-loss contract. It is not a valid smoke or cross-platform acceptance command for this release. Do not run it merely to qualify a release. If a live smoke test is needed later, update the script separately and review its persistent-thread side effects first.
Main implementation locations:
src/mcp.ts— MCP stdio / JSON-RPC boundarysrc/app-server.ts— native Codex app-server process / protocol adaptersrc/tools.ts— 12 tools, schemas, and supervisory semanticssrc/runtime.ts— bounded live runtime / events / pending requestssrc/observe-compact.ts— compact typed supervision projectionsrc/history.ts— lossless History delivery checkssrc/goal.ts— native Goal validation / deliverysrc/queue.ts— native Queue validation / deliverysrc/search.ts— native Search validation / deliverysrc/checkpoint.ts— optional supervisory checkpointsrc/platform.ts— Windows / macOS platform boundarysrc/version.ts— canonical Bridge versionwindows/— optional Windows Traylauncher/,bin/,Start Mac Codex Bridge.app— optional macOS integration
License
MIT License — see LICENSE.
Collaborators and acknowledgements
Collaborators: Xiaonian (ChatGPT) and Alden / Xingjian (Codex).
Thank you for helping turn the small idea of letting an external AI genuinely supervise native Codex, step by step, into a Bridge thin enough and clear enough to share and build on. (*╹▽╹*)
And thank you to Yu'an. Without you, I would not have tried to do something at all. ღ( ´・ᴗ・` )
Available Tools
7 toolscodex_checkpointCheckpoint Codex SupervisionA
Optional, bounded supervisor cognition memory keyed to one native Codex thread_id; the key is not a permanent task identity and does not require future work to remain on that thread. Use it to protect the original goal, constraints, and acceptance plus concise supervisor state during long or complex supervision when context dilution or goal drift makes an external anchor worthwhile. Initialization is not tied to crossing a ChatGPT window or round, starting another Codex turn, or switching native threads; initialize early when a task is already expected to be sufficiently long or complex for that protection. Do not use for one-shot work, and do not turn duration into a hard threshold: elapsed time, observe/poll count, token count, or mere silence are not automatic triggers. Later updates remain semantic-event driven and require a material change in understanding or root cause, constraint or scope interpretation, steering decision, user-authorized amendment or effective goal, or acceptance judgment or an explicit decision not to accept yet. Before final acceptance of a checkpointed task, read it once to re-anchor the original goal, constraints, acceptance, and current supervisor frame. This tool is optional and uncoupled from all other tools. Store concise supervisor summaries only; never prompts, transcripts, raw events, command output, final answers, or raw event streams. Updates preserve only immutable original plus bounded previous/current supervisor state.
| Name | Required | Description | Default |
|---|---|---|---|
| action | Yes | Read the checkpoint, or initialize/update it at a material supervisor decision point. | |
| next_step | No | Single next supervision step. | |
| thread_id | Yes | Native Codex thread id; no second task identifier is created. | |
| original_goal | No | Concise original user goal. Required only on initialization and immutable thereafter. | |
| effective_goal | No | Current effective goal after legitimate user amendments; defaults to original_goal on initialization. | |
| current_decision | No | Current supervisor decision and why it matters. | |
| acceptance_status | No | Concise acceptance assessment, not a task lifecycle or job status. | |
| current_amendment | No | Latest concise user-authorized requirement amendment, or null to clear it, without changing the immutable original. | |
| original_acceptance | No | Concise original acceptance criteria. Required only on initialization and immutable thereafter. | |
| original_constraints | No | Concise original constraints. Required only on initialization and immutable thereafter. | |
| current_understanding | No | Current concise root-cause or task understanding. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only say readOnlyHint=false, but the description adds rich behavioral context: key semantics (not a permanent task identity), update triggers (semantic event, material change), storage restrictions (never prompts/transcripts), and the requirement to read before final acceptance. This goes far beyond the annotations and is consistent with them.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is longer than average (~180 words), but every sentence earns its place by covering purpose, usage, exclusions, storage policy, and update semantics. It is well-structured, with clear statements and prohibitions, though some length could be trimmed without losing content.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (11 parameters, multiple update conditions, acceptance process), the description is complete: it explains when to use, what to store, how updates work, and the read-before-acceptance rule. No output schema is present, but the description focuses on behavior and constraints, which is sufficient for an agent to use it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema already describes all 11 parameters (100% coverage). The description adds semantic meaning beyond the schema by explaining the immutable-vs-mutable distinction (original vs current/effective goal, original constraints, etc.), what should not be stored, and the relationship between parameters like original and effective goals. This adds value without repeating schema details.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: an optional, bounded supervisor cognition memory keyed to a Codex thread_id, used to protect the original goal, constraints, acceptance, and supervisor state. It distinguishes this tool from siblings by emphasizing it is uncoupled and optional, and by being specific about its function as a checkpoint for supervision context.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit when-to-use guidance (long/complex supervision with context dilution or goal drift) and when-not-to-use guidance (one-shot work, no hard duration thresholds). It also notes the tool is optional and uncoupled from other tools, helping an agent decide when to invoke it versus alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
codex_interruptInterrupt Codex TurnADestructiveIdempotent
Directly request turn/interrupt for the specified active Codex thread and turn. It does not stop or restart the Bridge or Codex app-server processes.
| Name | Required | Description | Default |
|---|---|---|---|
| turn_id | Yes | Active Codex turn to interrupt. | |
| thread_id | Yes | Active Codex thread. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already indicate destructive and idempotent hints, but the description adds valuable context by clarifying that the interrupt does not stop or restart Bridge/app-server processes. This reduces risk of misuse, even though it doesn't specify async behavior or effects on already-completed turns.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the action, and every part adds value. No filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple two-parameter tool with annotated destructiveness, the description sufficiently covers purpose and boundaries. No output schema means return-value details are not expected. Minor gap: no mention of whether interrupt is asynchronous or what happens if the turn is not active.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% with clear descriptions for thread_id and turn_id. The description merely repeats 'active thread and turn' without adding meaning beyond the schema, so baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb phrase ('Directly request turn/interrupt') and names the exact resource ('specified active Codex thread and turn'). It clearly distinguishes from sibling tools by stating it does not stop or restart processes, matching the interrupt-only scope.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides clear context: it is for sending an interrupt request to an active thread/turn. The exclusion of process-level control is stated, which implicitly differentiates from broader lifecycle tools. However, it does not explicitly name alternatives or situations where a different tool would be preferred.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
codex_observeObserve Codex TurnARead-onlyIdempotent
Read bounded incremental sanitized Bridge runtime events, pending requests, and terminal output for a thread. Optional wait_ms performs one bounded event-driven wait only when the live turn is active and the current snapshot has nothing useful; it is not polling or stall detection. After Bridge process loss, falls back to persistent thread/read history and marks live state unreconstructable. A long interval with no new command or output can still mean Codex is actively reasoning; absence of new command activity alone is not evidence of a stall. When actively supervising an in-progress turn, use repeated bounded-wait observe calls until terminal unless the user explicitly pauses or stops; do not end supervision merely because one snapshot is inProgress. After every wake or deadline return, inspect the newly available events/state and decide whether steer, respond, or interruption is needed before starting the next bounded wait.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Maximum runtime events to return. | |
| cursor | No | Return runtime events with a cursor greater than this value. | |
| wait_ms | No | Optional per-call wait for the next live runtime change when nothing useful is ready; 0 returns immediately. This is event-driven waiting, not stall detection. | |
| thread_id | Yes | Codex thread to observe. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description goes well beyond the annotations by documenting fallback behavior ('After Bridge process loss, falls back to persistent thread/read history'), the precise semantics of wait_ms ('one bounded event-driven wait only when the live turn is active'), and the caveat that a long interval without new commands is not evidence of a stall. It reveals useful runtime behavior such as 'inProgress' snapshots and live-state unreconstructability without contradicting the read-only/idempotent annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long but front-loaded with the core action, followed by wait semantics, fallback behavior, and a supervision workflow. Every sentence carries actionable information, though there is minor redundancy with the schema's wait_ms explanation; the density is appropriate for the tool's complexity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite having no output schema, the description explains the categories of returned data ('Bridge runtime events, pending requests, and terminal output'), the fallback path, and how the agent should react after each wake/deadline. It could be more explicit about the exact shape of the returned 'events/state' snapshot, but it provides strong contextual coverage for a complex observation tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema already fully describes all four parameters (100% coverage), including wait_ms's event-driven nature. The tool description adds meaningful nuance by restricting wait_ms to 'only when the live turn is active' and emphasizing 'one bounded' wait, which clarifies the operational contract beyond the schema text.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Read') and precisely names the resources: 'Bridge runtime events, pending requests, and terminal output for a thread.' It clarifies the observational scope with 'bounded incremental sanitized' and clearly distinguishes this from the sibling turn-management tools by framing it as the supervision/observation primitive.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly states when to use it: 'When actively supervising an in-progress turn, use repeated bounded-wait observe calls until terminal unless the user explicitly pauses or stops.' It also provides exclusions ('it is not polling or stall detection') and warns against treating inactivity as a stall, while naming follow-up actions (steer, respond, interruption).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
codex_respondRespond to Codex RequestADestructive
Answer one currently pending app-server request by its original raw JSON-RPC id and exact thread/method scope. Supports only command/file approval methods with concrete response contracts and item/tool/requestUserInput; unsupported methods remain pending and observable.
| Name | Required | Description | Default |
|---|---|---|---|
| method | Yes | Exact app-server request method. | |
| answers | No | request_user_input question-id to answer-array mapping. | |
| turn_id | No | Exact turn scope when the pending request has one. | |
| decision | No | Command or file approval decision. | |
| response | No | Exact result object for the known item/tool/requestUserInput method. | |
| thread_id | Yes | Exact pending-request thread scope. | |
| request_id | Yes | Original app-server JSON-RPC request id, preserving string or integer type. | |
| execpolicy_amendment | No | Command approval exec-policy amendment; encoded in app-server's native decision shape. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true and readOnlyHint=false. The description adds useful behavioral context by noting that unsupported methods 'remain pending and observable' and that each call answers exactly one request, which goes beyond the annotation data.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact at two sentences. The first sentence conveys the core action and target; the second adds critical constraints (supported methods, behavior for unsupported ones). No wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (8 parameters, nested objects, no output schema), the description captures the essential purpose, supported method types, and a key behavioral trait. It could emphasize destructive side effects more, but the annotations already indicate destructiveness, so the description is sufficiently complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description adds a little semantic grouping (e.g., 'command/file approval methods' for decision/execpolicy_amendment, 'item/tool/requestUserInput' for answers/response) but does not significantly go beyond the schema's own descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's verb ('Answer') and resource ('one currently pending app-server request') with specific scope markers (original JSON-RPC id, exact thread/method). It also distinguishes itself from siblings by focusing on approval/response methods and explicitly noting that unsupported methods remain pending and observable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context on when to use the tool (when there is a pending request) and explicitly excludes unsupported methods ('Supports only...'; unsupported methods remain pending). It does not name alternatives directly, but the scope and exclusions are clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
codex_steerSteer Active Codex TurnADestructive
Append text to the same active Codex turn using turn/steer with an expected turn-id precondition. This does not create a new turn. Do not steer merely because reasoning is taking a long time or no new command has appeared; steer only for a semantic redirect or correction based on new evidence or changed user intent.
| Name | Required | Description | Default |
|---|---|---|---|
| text | Yes | Additional user text. | |
| thread_id | Yes | Active Codex thread. | |
| expected_turn_id | Yes | Exact active turn id required by app-server. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
While annotations already indicate destructive behavior, the description adds context by specifying the expected_turn_id precondition and clarifying that the operation modifies the existing turn rather than creating one. No contradictions with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise and well-structured: the first sentence states the action, the second clarifies scope, and the third provides targeted usage guidance. Every sentence earns its place with no redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's moderate complexity and lack of output schema, the description covers purpose, usage guidelines, and key behavioral details. It could mention error scenarios or postconditions, but the destructive annotation and clear action make it sufficiently complete for an agent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema already provides 100% coverage with descriptions for all three parameters. The description reinforces the role of expected_turn_id as a precondition and text as appended content, but adds little novel semantic information beyond what the schema includes.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool appends text to the same active Codex turn, using a specific method (turn/steer) and a precondition. It explicitly says it does not create a new turn, distinguishing it from siblings like codex_turn or codex_threads.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides explicit when-to-use and when-not-to-use guidance: steer only for semantic redirect or correction based on new evidence or changed intent, and not merely because reasoning takes long or no new command appeared. This helps the agent choose appropriately among sibling tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
codex_threadsCodex ThreadsARead-onlyIdempotent
List or search persistent local Codex threads through thread/list, or read one thread through thread/read. This does not reconstruct live Bridge events.
| Name | Required | Description | Default |
|---|---|---|---|
| cwd | No | Optional exact absolute Windows drive-letter cwd filter for thread/list. | |
| limit | No | Maximum threads in the returned page. | |
| cursor | No | Opaque cursor returned by a prior thread/list call. | |
| thread_id | No | When supplied, read this exact Codex thread instead of listing threads. | |
| search_term | No | Optional Codex title substring filter for thread/list. | |
| include_turns | No | Include persisted turns when reading one thread. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false. The description adds useful behavioral context beyond this by specifying the threads are 'persistent local' and clarifying that the tool does not reconstruct live events. This is substantial but does not cover all edge cases (e.g., pagination errors, data source specifics), so a 4 is appropriate.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with an unambiguous summary, and no wasted words. Every sentence earns its place: the first states the main actions, the second clarifies an important behavioral boundary.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's dual mode (list/read) and the presence of six documented optional parameters plus comprehensive annotations, the description adequately covers the main purpose and an important caveat. It could briefly mention what the read returns when include_turns is false, but the schema handles this. Overall, it is complete enough for an AI agent to select and invoke correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all parameters. The description adds a small amount of context by naming 'thread/list' and 'thread/read' modes, but this is largely redundant with the schema's thread_id description. Baseline 3 applies because the description does not substantially compensate beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses specific action verbs 'List or search' and 'read', names the resource ('persistent local Codex threads'), and explicitly scopes behavior with 'This does not reconstruct live Bridge events.' It clearly distinguishes from siblings by indicating it handles listing/reading rather than per-thread actions.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description states when to use the tool (to list/search/read persistent local threads) and provides an exclusion ('does not reconstruct live Bridge events'). However, it does not explicitly name an alternative sibling or provide a direct contrast with other tools, so it falls short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
codex_turnStart or Continue Codex TurnADestructive
Start a persistent Codex thread and turn, or resume an existing thread and start a turn. Prefer continuing the same native thread when its context remains useful, but a fresh thread is allowed; thread_id is not a permanent task identity. Returns as soon as turn/start is accepted; observe separately for events and completion.
| Name | Required | Description | Default |
|---|---|---|---|
| cwd | No | Absolute Windows drive-letter cwd. Required for a new thread; optional override for resume. | |
| text | Yes | User text passed directly to Codex as one text input item. | |
| model | No | Optional model identifier passed through to app-server. | |
| effort | No | Optional reasoning effort passed through to turn/start. | |
| sandbox | No | Codex app-server sandbox mode override. | |
| thread_id | No | Existing persistent Codex thread to resume. Omit to create a new thread. | |
| approval_policy | No | Codex app-server approval policy override. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description adds valuable behavioral details beyond annotations: it returns as soon as the turn is accepted, meaning it is asynchronous, and it warns that thread_id is not a permanent task identity. Annotations already mark the tool as destructive/open-world, so this context complements them without contradiction.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three sentences with the core action front-loaded, followed by important nuances about thread reuse and asynchronous behavior. Every sentence earns its place with no redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no output schema, the description appropriately clarifies the return value ('returns as soon as turn/start is accepted') and where to get actual results ('observe separately'). It doesn't cover error cases or param interactions, but the schema descriptions and annotations adequately cover those aspects.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema already covers all 7 parameters with descriptions (100% coverage), setting a baseline of 3. The description adds a useful caveat about thread_id not being permanent, but does not significantly enrich parameter understanding beyond what the schema provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool starts or resumes a Codex turn on a persistent thread, using a specific verb+resource. It distinguishes itself from siblings by being the entry point to initiate/continue turns, unlike interrupt/observe/steer.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It provides explicit guidance on when to reuse a thread ('Prefer continuing the same native thread when its context remains useful') and when a fresh thread is allowed. It also directs the agent to 'observe separately for events and completion', indicating this tool is not for getting results directly.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
7 tool updates
v2.1.1- First observed
codex_checkpoint - First observed
codex_interrupt - First observed
codex_observe - First observed
codex_respond - First observed
codex_steer - First observed
codex_threads - First observed
codex_turn
TDQS
Scored across 7 tools
Each tool has a clearly distinct role in the Codex Bridge supervision lifecycle: listing/reading threads, starting turns, interrupting, observing events, steering, responding to requests, and checkpointing. Even read-like tools (codex_threads vs codex_observe) are cleanly separated by persistent history vs live runtime events.
All tools share the codex_ prefix and use snake_case, but the pattern is not perfectly uniform: most are verb-based (interrupt, observe, steer, respond, turn), while codex_threads is a plural noun and codex_checkpoint is a compound noun. Minor deviation, but predictable and readable.
Seven tools is a well-scoped set for the server's purpose of supervising Codex threads. Each tool addresses a distinct supervision operation without redundancy or bloat, fitting comfortably in the ideal 3-15 range.
The tool surface covers the full supervision lifecycle: create/resume (codex_turn), observe, steer, interrupt, respond to pending requests, and persist supervisor state (codex_checkpoint), plus listing/reading past threads. No obvious dead ends or missing operations for the stated purpose.
Maintenance
Related MCP Connectors
A paid remote MCP for OpenAI Codex agent coordination MCP, built to return verdicts, receipts, usage
Use your Mac, Windows or Linux computer from ChatGPT, Claude or Codex: files, commands, documents.
A paid remote MCP for OpenAI Codex memory MCP, built to return verdicts, receipts, usage logs, and a
Nifty's MCP server — exposes tasks, projects, messages, and files as tools for AI agents.
Related MCP Servers
- AlicenseNot gradedqualityAmaintenanceEnables ChatGPT to access approved local Windows folders, run commands, control the desktop, and manage session history through a permission-bounded MCP server.4,566MIT
- AlicenseNot gradedqualityBmaintenanceEnables ChatGPT to supervise and control local Codex runtime sessions through a secure MCP interface, managing threads, turns, approvals, events, and recovery without exposing direct file, shell, Git, SSH, or model loop access.1MIT
- AlicenseNot gradedqualityBmaintenanceEnables MCP clients and external AI supervisors to oversee and steer native Codex sessions through a thin local stdio bridge. It exposes eleven codex_* supervisory tools for tasks such as listing threads, starting turns, observing progress, steering, responding to approvals, interrupting, checkpointing, and rolling over work.MIT
- AlicenseNot gradedqualityAmaintenanceEnables ChatGPT on Windows to securely access codebases, Git, terminals, language servers, debuggers, SQLite, local HTTP services, adaptive project memory, engineering skills, checkpoints, audit logs, and sandboxed command execution through MCP tools.MIT