Skip to main content
Glama

Codex Bridge 0.3

Codex Bridge is a local stdio MCP server that lets ChatGPT supervise the real Codex runtime and receive project artifacts without exposing a second shell, Git, SSH, or model loop.

ChatGPT GPT Pro
  -> Secure MCP Tunnel / stdio MCP
  -> Codex Bridge
  -> codex app-server --listen stdio://
  -> local Codex models, tools, projects, and configured MCP servers

The MCP layer manages Codex threads, turns, events, approvals, questions, recovery, and delivery evidence. A read-only artifact layer also lets ChatGPT browse and transfer files from configured projects for inspection, inline preview, and download. Codex remains responsible for project writes, commands, Git, SSH, MCP tools, research work, and its final natural-language answer.

Version 0.3.0 adds a Web-side tool relay: Codex requests a capability through the native bridge_web_tool dynamic tool; ChatGPT uses a tool it actually has, returns text/JSON/sources, and Codex continues the same turn. No browser automation, cookie reuse, connector credentials, or second model loop is involved. Capability names are not restricted to a fixed connector list; actual availability and authorization remain with the ChatGPT host.

Requirements and setup

  • Node.js 20 or newer

  • A locally installed codex CLI

  • A working Codex login for real turns (model/list and protocol smoke checks can still work when a turn login is unavailable)

git clone https://github.com/Wanzhe-Liao/codex-bridge.git
cd codex-bridge
npm ci
npm run build

Copy config.example.toml to ~/.config/codex-supervisor-mcp/config.toml, then register only absolute project directories that ChatGPT may select. ChatGPT supplies project_id and an optional local profile; it cannot supply an arbitrary working directory, provider, or app-server command. The legacy codex-supervisor-mcp configuration and state directory names are intentionally retained for compatibility with existing installations.

max_artifact_bytes = 33554432 # 32 MiB; configurable up to 256 MiB

[projects.default]
cwd = "/absolute/path/to/project"

[profiles.default]
model = ""                    # inherit Codex default
effort = ""                   # inherit Codex default
approval_policy = "on-request"
sandbox_type = "workspace-write"
network_access = false
wait_timeout_seconds = 40

Profile model, effort, and service tier values are checked at task start against this machine's live model/list. Keep them empty to inherit local defaults. workspace-write supplies exactly the registered project as writableRoots.

Run the server:

npm start

The server's stdout is reserved for MCP protocol messages. Diagnostics and app-server stderr go to stderr. Task state uses one WAL-mode SQLite database at ~/.local/share/codex-supervisor-mcp/state.sqlite3; if a sandboxed desktop host cannot create that directory, it falls back to the ignored .supervisor/state.sqlite3 in the launch directory. Override paths only from the local environment with CODEX_SUPERVISOR_CONFIG and CODEX_SUPERVISOR_STATE.

On Windows, the supervisor resolves the native codex.exe on PATH so npm's codex.ps1/codex.cmd shims do not cause a spawn codex ENOENT error. If a machine has multiple Codex installations, set the local (never ChatGPT-provided) CODEX_BIN environment variable to the desired executable path.

Related MCP server: Local Codex Bridge

MCP tools

  • codex_health: configuration, SQLite, app-server, login and simplified live model status.

  • codex_start: create a thread/task and start a natural-language turn.

  • codex_wait: bounded long-poll returning real plans, activity, Codex messages, events, requests and terminal state.

  • codex_status: immediate task status, or active/recent task recovery list.

  • codex_send: turn/steer an active turn or start a new turn on the same completed/lost thread.

  • codex_respond: validate and answer approvals, permissions, user input and MCP elicitations with the original JSON-RPC request ID.

  • codex_submit_tool_result: return a truthful Web-side tool result to the original pending dynamic-tool request; this can advance consequential Codex work.

  • codex_result: return an authoritative terminal delivery, Codex's natural-language final message, and independent evidence.

  • codex_inspect: paginate transcript, plan, diff, commands, bounded output, file changes, MCP calls, warnings, or redacted raw events.

  • codex_cancel: send turn/interrupt and wait for authoritative turn/completed.

  • codex_files: browse or search file metadata inside a configured project; recursive search and pagination are supported.

  • codex_artifact: transfer the original file through MCP and render a ChatGPT preview/download component when supported.

The artifact interface has no extension allowlist. It handles text, source files, PDF, images, audio, video, spreadsheets, archives, and arbitrary binary files. Images and audio use native MCP content blocks; other files use embedded resources plus a dynamic codex-artifact:// resource readable through resources/read. PDFs, images, audio, video, and text get an inline viewer; formats the browser cannot render still get the original-file download button.

Example instructions to GPT Pro:

Use Codex Bridge project default. List PDFs under papers recursively, open the
latest manuscript, show it to me, and make the original file downloadable.
Browse project default for results/summary.csv, open it, inspect the contents,
and attach the original CSV for download. Do not start a Codex task.

turn/start never contains outputSchema; prompts and final answers remain ordinary natural language. No tool response invents percentage progress. Only turn/completed makes the current turn terminal.

Web-side tool relay

codex_start -> codex_wait -> tool_requests[]
  -> ChatGPT uses an available, authorized Web-side tool
  -> codex_submit_tool_result -> Codex continues the original turn
  -> codex_wait (repeat) -> turn/completed -> codex_result

New threads register the fixed bridge_web_tool through thread/start.dynamicTools. Codex supplies capability, request (up to 16,000 characters), optional JSON context, and operation: read|write. Total request arguments are limited to 20,000 characters. Native item/tool/call drives the queue, not commentary parsing. codex_wait wakes immediately and exposes tool_requests; waiting_for_tool is nonterminal. Approvals/questions remain in pending_request (first) and pending_requests (queue). Multiple tasks and out-of-order replies are supported.

Example submission (source identifiers below are illustrative, not real research findings):

{
  "task_id": "<task ID>",
  "request_id": "<relay request UUID from codex_wait>",
  "status": "success",
  "result": {"summary": "Findings actually returned by the host tool"},
  "sources": [{"document_id": "<connector document ID>", "title": "Source document"}],
  "tool_used": "<actual host tool name>"
}

result accepts text or JSON. sources accepts strings (links, document IDs or DOIs) or objects with url, document_id, doi, and optional title. status is success, error, unavailable, or declined. If a tool is absent or lacks permission, return an honest failure explanation and keep waiting; never invent a successful lookup. Responses use the schema-native {contentItems: [{type: "inputText", text: ...}], success} envelope and original JSON-RPC ID. Submission does not start, steer or resume a turn.

Local defaults (no existing configuration changes are required):

[relay]
enabled = true
max_result_bytes = 262144 # 256 KiB, entire submitted status/result/sources/tool_used

Oversize results fail explicitly, without silent truncation. Stored results can be read in full using codex_inspect(kind="tool_requests", item_id="<relay request UUID>", offset=0, limit=8): pages contain character chunks; concatenate text and follow next_offset. Result evidence uses bounded excerpts and labels Web-side results host reports, not independently verified evidence. Authoritative app-server dynamic-tool events are recorded separately.

The SQLite migration retains existing tasks/events and adds a queued interaction ledger plus a per-thread relay flag. Each interaction binds its original typed JSON-RPC ID to a connection UUID, thread, turn and call. Identical repeat submissions return the recorded delivery state without resending; conflicting submissions fail. submitted means accepted by the local pipe, not independent confirmation of a remote write. resolved comes from a server request-resolution or dynamic-tool completion event. Pending requests become stale after cancellation, completion or connection loss; interrupted sends become uncertain. Never automatically redo a Web-side write after uncertain delivery. Restart preserves these records but does not replay old RPC IDs.

Already registered dynamic tools survive thread/resume. Legacy threads without this registration remain readable and continuable but cannot gain the relay by resume: create a new task. This follows the locally verified Codex 0.151.0 experimental schema, where thread/resume has no dynamic-tools registration field. Future Codex updates may require re-verification; see the official app-server documentation.

Example prompts for ChatGPT:

  • Literature search: “Start a new Codex task for project default. Ask Codex to request current cross-centre ECG model literature through bridge_web_tool. Use your available search tool, return sourced summaries and DOIs, and wait for Codex's evidence-backed delivery.”

  • Connector read: “Ask Codex what context it needs from the project planning document. Read the specified document with the available connector, return only relevant non-sensitive passages and the document ID, then continue waiting.”

  • Authorized write: “I authorize adding this exact comment to issue <specified issue>: <approved text>. If Codex requests that operation, use the available GitHub connector and host approval flow, return its actual response/link, then wait for completion. Do not perform other writes.”

A Codex request is not authorization to send email, publish content, delete data or modify remote systems. ChatGPT must obey the original user scope and host approvals. Redaction is heuristic, not a guarantee that arbitrary clinical data is safe: send only necessary non-sensitive context, never raw patient records or credentials.

Doctor, tests, and protocol verification

npm run doctor
npm test

doctor prints the Codex version, login availability without identity or credentials, app-server initialize status, live models, project/profile mappings, SQLite writability, and MCP construction status.

This implementation was verified against schemas generated by the locally installed CLI using:

codex --version
codex app-server generate-ts --experimental --out <temporary-directory>
codex app-server generate-json-schema --out <temporary-directory>

Generated schemas are not copied into the project. Unit tests use a JSONL mock app-server for initialization ordering, IDs, long tasks, plans, commands, diffs, file changes, MCP calls, approvals, questions, persistence, interruption, failure, crash recovery, redaction, tool annotations, wait semantics, artifact MIME handling, PDF byte transfer, viewer registration, size bounds, traversal, and symlink escape.

The optional real integration test is disabled by default because it can consume Codex quota:

CODEX_SUPERVISOR_REAL_INTEGRATION=1 npm run integration

On PowerShell:

$env:CODEX_SUPERVISOR_REAL_INTEGRATION = "1"
npm run integration

The separate opt-in relay fixture consumes a small real Codex turn. It creates a temporary Git project, returns a random marker through a native tool call, and checks that the original turn's final answer contains that marker. The local responder is explicitly a fixture, not Web search:

$env:CODEX_SUPERVISOR_RELAY_INTEGRATION = "1"
npm run integration
Remove-Item Env:CODEX_SUPERVISOR_RELAY_INTEGRATION

Linux: CODEX_SUPERVISOR_RELAY_INTEGRATION=1 npm run integration. Ordinary npm test runs no paid model turns. A real ChatGPT Web connector/search loop is a separate acceptance check: refresh the App, start a new task, have Codex request a genuine search, invoke the host search tool, submit actual citations, and verify the final same-turn delivery. Local fixtures do not verify that host workflow or the host's willingness/ability to follow the wait rule.

Test tool discovery with MCP Inspector from the project directory:

npx @modelcontextprotocol/inspector node /absolute/path/to/codex-bridge/dist/index.js

For a non-UI tool-list smoke check:

npx @modelcontextprotocol/inspector --cli node /absolute/path/to/codex-bridge/dist/index.js --method tools/list

Read-only invocation check: append --method tools/call --tool-name codex_status instead of --method tools/list. To keep diagnostics separate from a running deployment, explicitly pass CODEX_SUPERVISOR_STATE=:memory: to the server in the Inspector session configuration's env object (do not assume it inherits all terminal variables). Unit tests also discover and call the relay over the official MCP in-memory transport. Inspector strict schema review may warn about free-form object values/array elements in result; accepting arbitrary JSON there is intentional, with runtime nesting and byte limits.

Secure MCP Tunnel and ChatGPT

Use absolute paths for both Node and the built script. Keep the control-plane key in the process environment, never in TOML or this repository.

export CONTROL_PLANE_API_KEY="..."

tunnel-client init \
  --sample sample_mcp_stdio_local \
  --profile codex-bridge \
  --tunnel-id <TUNNEL_ID> \
  --mcp-command "/absolute/path/to/node /absolute/path/to/codex-bridge/dist/index.js"

tunnel-client doctor \
  --profile codex-bridge \
  --explain

tunnel-client run \
  --profile codex-bridge

In ChatGPT on the web, enable Developer mode, create a developer App, select the configured Tunnel, refresh both tools and server instructions, and enable the App in the GPT Pro conversation. The embedded server instructions require GPT Pro to stay in the same response and repeatedly call codex_wait, resolve safe requests, wait for authoritative turn/completed, then call codex_result and inspect objective evidence before replying to the user. After upgrading Codex Bridge, stop the old Tunnel with Ctrl+C when its work is finished, restart it and refresh the App's tools and server instructions. Confirm codex_submit_tool_result is listed, then create a new task.

Existing Windows installations with the private .supervisor/start-tunnel.ps1 launcher can keep their one-command startup:

powershell.exe -NoProfile -ExecutionPolicy Bypass -File "C:\absolute\path\to\codex-bridge\.supervisor\start-tunnel.ps1"

That launcher, its project/model/Tunnel mappings and DPAPI-encrypted key are local deployment files, not shipped or changed by this update. Other installations use their existing tunnel-client run --profile <profile> command. Do not run two supervisors against the same state database.

Security boundaries

  • Project paths and profiles are local allowlists; MCP inputs cannot provide arbitrary absolute paths, providers, or process commands.

  • Artifact paths are relative to an allowlisted project. Lexical traversal and symlinks/junctions that resolve outside the project are rejected. Configured sensitive_paths remain unavailable for transfer.

  • Artifact transfer is read-only, has no extension whitelist, and uses the locally configurable max_artifact_bytes transport bound to prevent unbounded JSONL/base64 messages.

  • Start/send/respond/submit-tool-result/cancel are accurately marked mutating and potentially destructive; health/status/wait/result/inspect are read-only.

  • Approval is never globally automatic. High-risk commands, project-external grants, credential-like requests, and arbitrary dynamic tool execution are declined or restricted.

  • Built-in and configurable redaction removes common tokens, private keys, passwords, sensitive paths, and credential fields. Raw reasoning text deltas are not persisted or returned.

  • Command output, artifact size, and pages are bounded. Completed command exit codes, file-change status, MCP status, aggregated diff, and fixed read-only Git checks form the evidence returned by codex_result.

Known operational limitation: a supervisor process cannot prove that an in-flight turn survived an app-server crash when thread/resume no longer reports an active turn. In that case it preserves all IDs/events and reports nonterminal connection_lost; codex_send starts a recovery turn on the same thread rather than claiming completion.

Related MCP Connectors

Related MCP Servers