Skip to main content
Glama
ossmalaysia

codex-local-mcp

by ossmalaysia

codex-local-mcp

An MCP server that lets any LLM drive your local Codex CLI.

The calling model sends a task in plain English. Codex writes and runs code inside an isolated workspace folder on your machine. The server returns Codex's output, the list of files it created, and the images inline.

Codex generates images natively, so codex_generate_image needs no image API key — your existing codex login is enough.

This runs code on your computer. Anything that can call these tools can make Codex execute code as you, inside its sandbox. Read SECURITY.md before using it.

Requirements

  • Node.js 20 or newer

  • Codex CLI installed and logged in

  • Windows, macOS or Linux

Related MCP server: Codex MCP Server

Install

New to this? Follow the step-by-step setup guide instead, which covers finding your paths, configuring each client, verifying it works, and troubleshooting.

npm install -g @openai/codex   # the Codex CLI itself
codex login                    # one time

git clone https://github.com/ossmalaysia/codex-local-mcp.git
cd codex-local-mcp
npm install
npm run build

Connect it to a client

Full walkthrough with per-OS paths and troubleshooting: SETUP.md.

This is a stdio server: the client launches it as a child process. There is no URL and nothing listens on a port.

Use absolute paths for command and for CODEX_BIN. MCP clients launch servers with a minimal environment in which bare node and codex may not resolve.

Claude Desktop

Add to claude_desktop_config.json, then fully quit and reopen the app.

  • Windows: %APPDATA%\Claude\claude_desktop_config.json

  • macOS: ~/Library/Application Support/Claude/claude_desktop_config.json

{
  "mcpServers": {
    "codex": {
      "command": "/absolute/path/to/node",
      "args": ["/absolute/path/to/codex-local-mcp/dist/index.js"],
      "env": {
        "CODEX_BIN": "/absolute/path/to/codex",
        "CODEX_MCP_ROOT": "/absolute/path/to/codex-workspaces",
        "CODEX_TIMEOUT_SEC": "600"
      }
    }
  }
}

On Windows the values look like C:/Program Files/nodejs/node.exe and C:/Users/you/AppData/Roaming/npm/codex.cmd. Forward slashes work fine.

Find the right paths with which node / which codex (macOS, Linux) or (Get-Command node).Source / (Get-Command codex).Source (PowerShell).

Claude Code

claude mcp add codex -s user -- node /absolute/path/to/codex-local-mcp/dist/index.js

Anything else

Any MCP client that can launch a stdio server works. Point it at node /absolute/path/to/dist/index.js.

Tools

Tool

Arguments

Description

codex_run

prompt, workspace?, timeout_sec?, model?, wait_sec?

Run any task.

codex_generate_image

prompt, count?, size?, quality?, workspace?, open?, wait_sec?

Generate images. The prompt is passed through as written.

codex_job_result

job_id, wait_sec?

Collect the result of a task that was still running.

codex_read_artifact

path, max_bytes?

Read a file from a workspace.

Reuse the same workspace name across calls to keep working on the same files.

Long-running tasks

MCP clients built on the official SDK give up on a request after 60 seconds by default, and image generation often takes longer. So a call waits up to wait_sec (default 45, maximum 50) for Codex. If Codex has finished, you get the result as usual. If not, the call returns at once with:

status: running
job_id: a3b1d698-5754-490c-ab5d-bc1ea76149a2
workspace: /path/to/codex-workspaces/my-task

Codex keeps working. Call codex_job_result with the job_id to collect the result; each call waits up to wait_sec again, so poll until it finishes. Finished results stay available for an hour. Jobs are held in memory, so a server restart loses the job_id, but any files produced are still in the workspace shown.

codex_generate_image defaults to quality: "low" and size: "1024x1024" — a fast draft. size must be WIDTHxHEIGHT or auto; a malformed value is rejected before Codex runs. The size is a request to Codex rather than something the server can enforce, so the result reports each image's actual dimensions and flags any that differ from what was asked. Raise quality to high for final assets. open defaults to true and opens each image in your default viewer, because MCP clients deliver tool-result images to the model's context and do not render them into the chat for you.

Configuration

All optional, set through the client's env block.

Variable

Default

Purpose

CODEX_MCP_ROOT

~/.codex-mcp/workspaces

Root for all workspaces. Nothing is written outside it.

CODEX_BIN

codex

Absolute path to the Codex CLI.

CODEX_MODEL

(Codex default)

Model override.

CODEX_TIMEOUT_SEC

300

Default timeout before a run is killed.

CODEX_MAX_TIMEOUT_SEC

1800

Ceiling a caller may request.

CODEX_MAX_OUTPUT_CHARS

40000

Output cap; the head and tail are kept.

CODEX_MAX_INLINE_IMAGE_BYTES

1048576

Above this, a downscaled preview is inlined instead.

CODEX_MAX_REPORTED_FILES

200

Cap on reported changed files.

CODEX_NETWORK_ACCESS

true

Set false to deny Codex the network.

CODEX_WAIT_SEC

45

How long a call waits before returning a job_id. Capped at 50.

CODEX_MAX_RUNNING_JOBS

4

How many Codex jobs may run at once.

CODEX_JOB_RETENTION_SEC

3600

How long a finished job's result stays retrievable.

How it works

Codex runs as --sandbox workspace-write: it can read and write only inside the workspace folder, plus reach the network. Workspace names that would escape the root are refused.

A few decisions worth knowing about:

  • Artifacts are found by filesystem diff. The workspace is snapshotted (path → mtime+size) before and after each run, so whatever Codex does, new files are detected. No output parsing.

  • The prompt travels on stdin (codex exec ... -), so prompt text is never shell-quoted and cannot break out into the command line.

  • Budgets are counted in base64 characters, not file bytes. Encoding inflates data by 4/3, so a budget expressed in file bytes overshoots what the client actually receives by a third. Images are downscaled until their encoded form fits, and a per-response budget caps all of them together. Originals on disk are never modified.

  • Full-resolution files are linked, not embedded. Every reported file also comes back as an MCP resource_link. The server declares the resources capability and serves workspace files through resources/read, so a client can fetch the original on demand instead of having megabytes pushed into every response. Reads outside the workspace root are refused.

  • Generated images open in the OS viewer, since clients don't render tool-result images.

  • Timeouts kill the whole process tree, so a stuck build leaves no orphans.

Development

npm test          # runs against a fake Codex - no CLI or API key needed
npm run build

See CLAUDE.md for architecture and the invariants to preserve, and CONTRIBUTING.md before opening a pull request.

CI runs the build, the test suite and npm audit on Ubuntu, Windows and macOS against Node 20 and 22, plus weekly CodeQL analysis.

Versioning

This project follows Semantic Versioning. Releases are cut automatically from Conventional Commits and published as tags with release notes on the Releases page. To use a specific release, check out its tag:

git checkout v0.1.0   # or any later tag

While the major version is 0, minor releases may still contain breaking changes, which are always called out in the release notes.

Security

Please read SECURITY.md. Report vulnerabilities privately through GitHub's Security → Report a vulnerability, not as a public issue.

License

MIT © OSS Malaysia

Available Tools

4 tools
codex_generate_imageGenerate images via the local Codex CLIA

Generate images with the local Codex CLI. The prompt is passed through as written. Returns the saved file paths and shows the images inline, downscaling a preview when a file is too large to inline. Defaults to a fast low-quality 1024x1024 draft; raise quality for final assets. Image generation often takes longer than wait_sec, in which case this returns status: running and a job_id; call codex_job_result with it to collect the images.

ParametersJSON Schema
NameRequiredDescriptionDefault
openNoOpen the generated images in the user's default image viewer. Default true, because MCP clients do not render tool-result images into the chat.
sizeNoPixel size as WIDTHxHEIGHT, e.g. "1024x1024" (fastest, the default), "1536x1024" landscape, "1024x1536" portrait, or "auto". This is a request to Codex; the actual dimensions are reported in the result.
countNoHow many images. Default 1.
promptYesWhat the image should depict.
qualityNoGeneration quality. Default "low" - fastest, good for drafts. Use "high" for final assets.
wait_secNoSeconds to wait for the result before returning a job_id instead. Default 45, maximum 50, which keeps the call under common 60-second client timeouts.
workspaceNoWorkspace folder name. Defaults to a new folder.
timeout_secNoSeconds before Codex is killed.

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It reveals that the prompt is passed verbatim, that results include file paths and inline previews with downscaling for oversized files, that generation often exceeds wait_sec returning a status and job_id, and how to retrieve results. It also discloses defaults and quality trade-offs. This is exceptionally transparent for a tool with zero annotation support.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is five sentences, each earning its place: the core action, the pass-through behavior, the return format, the default quality guidance, and the asynchronous fallback. It front-loads the primary purpose and packs essential behavioral details without redundancy or fluff.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with 8 parameters, 100% schema coverage, and no output schema, the description covers everything an agent needs: the synchronous/asynchronous split, how to handle long runs, the default quality and size, and the inline display behavior. Sibling tools are correctly referenced for follow-up, and the description is self-sufficient for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so a baseline of 3 applies. The description adds value by explaining the practical impact of quality (draft vs. final), the wait_sec/job_id behavior, and the inline preview downscaling. It doesn't rehash schema definitions but provides usage-level meaning for key parameters, exceeding the baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Generate images with the local Codex CLI.' It clearly differentiates from siblings (codex_job_result, codex_read_artifact, codex_run) by focusing on image generation, while the others handle results, artifacts, and generic execution. The scope is unambiguous and an agent can immediately tell this tool's purpose.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear operational context: it explains the default draft behavior, quality guidance for final assets, and explicitly tells the agent to call codex_job_result when a job_id is returned. However, it does not explicitly state when to avoid this tool or when to prefer a sibling (e.g., 'use codex_run for non-image tasks'), though the name and content make that implicit. This is a minor gap, so not a perfect 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

codex_job_resultCollect the result of a running Codex jobA

Get the result of a codex_run or codex_generate_image call that returned status: running. Waits up to wait_sec for the job to finish; if it is still running, call again. Results stay available for 60 minutes after the job finishes.

ParametersJSON Schema
NameRequiredDescriptionDefault
job_idYesThe job_id returned by codex_run or codex_generate_image.
wait_secNoSeconds to wait for the result before returning a job_id instead. Default 45, maximum 50, which keeps the call under common 60-second client timeouts.

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the behavioral burden and does well: it discloses that the tool waits up to wait_sec, that polling may need to be repeated, and that results persist for 60 minutes. It does not describe response statuses in detail, but the core polling behavior is transparent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences convey purpose, usage condition, polling behavior, and retention in a compact, front-loaded structure. There is no filler or redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description is sufficiently complete for a simple polling tool with two well-documented parameters. However, there is no output schema, so the description could have been more explicit about the response shape or status values returned after waiting. Still, the essential invocation workflow is clear.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema already describes both parameters fully, including that job_id comes from codex_run or codex_generate_image and that wait_sec defaults to 45 with a maximum of 50. The description reinforces these facts but adds little semantic value beyond them, so baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly specifies a distinct action: retrieving the result of a codex_run or codex_generate_image call that previously returned status: running. It names the exact operation, resource type, and state condition, making it unmistakable from siblings that start jobs or read artifacts.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives strong usage context: it is meant for jobs that returned status: running, and it instructs the agent to call again if the job is still running. It does not explicitly contrast with codex_read_artifact or list when not to use this tool, but the operational guidance is clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

codex_read_artifactRead a file produced by CodexA

Read one file from a Codex workspace, for artifacts that were too large to inline. Images come back as image content, everything else as text.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathYesPath to the file, relative to the workspace root or absolute inside it.
max_bytesNoRefuse files larger than this. Default 5000000.

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It adds useful behavior beyond the schema by stating that images are returned as image content and everything else as text. However, it does not disclose error handling, permission requirements, or what happens with nonexistent files—gaps that would matter for a read tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences with no filler. The purpose is front-loaded, and the output-type behavior is stated succinctly. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple tool with two parameters and full schema documentation, the description covers the key context: what file to read and how the content will be returned. It lacks explicit error handling and edge-case details, but does not need to explain return values since there is no output schema and the description handles output type.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both parameters are already documented. The description adds no parameter-specific guidance, but it does not need to because the schema fully covers path and max_bytes semantics. This matches the baseline for full schema coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Read'), a resource ('one file from a Codex workspace'), and the purpose ('artifacts that were too large to inline'). It also differentiates from siblings by clarifying that it reads files, while codex_generate_image generates images and codex_run likely executes tasks.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives the clear context 'for artifacts that were too large to inline,' which tells an agent when to reach for this tool. It does not explicitly name alternative tools or state exclusions, but the use case is specific enough to strongly infer when it applies.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

codex_runRun a task with the local Codex CLIA

Hand a natural-language task to the local Codex CLI coding agent. Codex can write and run code, install packages and call APIs inside an isolated workspace folder, then this returns its output plus any files it created. If Codex takes longer than wait_sec, this returns status: running and a job_id instead; call codex_job_result with it to collect the result.

ParametersJSON Schema
NameRequiredDescriptionDefault
modelNoOverride the Codex model for this run.
promptYesThe task for Codex, written as you would to a developer.
wait_secNoSeconds to wait for the result before returning a job_id instead. Default 45, maximum 50, which keeps the call under common 60-second client timeouts.
workspaceNoRelative workspace folder name under the workspace root. Reuse the same name to continue working on the same files. Defaults to a new timestamped folder.
timeout_secNoSeconds before Codex is killed. Default 300.

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries full behavioral disclosure. It reveals side effects (writing/running code, installing packages), isolation ('inside an isolated workspace folder'), and async behavior with wait_sec and job_id. It does not mention error conditions or persistence of created files beyond the return statement, but covers the main behavioral risks and control flow.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences with no filler. The main action, behavioral scope, and async fallback are all front-loaded and each sentence contributes distinct information needed to select and invoke the tool correctly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a five-parameter, async, no-output-schema tool, the description covers purpose, side effects, async behavior, and continuation semantics. It doesn't specify what the final result payload looks like or how timeout_sec interacts with wait_sec, but what an agent needs to make a correct call and collect results is present.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds meaningful semantics: it explains that wait_sec controls the sync/async split, that a job_id should be passed to codex_job_result, and that reusing workspace continues work on the same files. These enrich beyond the raw schema descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('hand... to'), resource ('local Codex CLI coding agent'), and scope ('can write and run code, install packages and call APIs'). It also distinguishes its output behavior from siblings by explaining it returns output plus created files, and routes async cases to codex_job_result.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear when-to-use context and explicit guidance for the async case: if execution exceeds wait_sec, returns status:running and a job_id, and the agent is instructed to call codex_job_result. It does not explicitly contrast with codex_read_artifact or codex_generate_image, so it falls just short of full alternative differentiation.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 3 tool updatesv0.2.0
    • Changedcodex_generate_image1 field changed
      • addedInput schema / properties / wait_sec
        Added value: +{
        +  "description": "Seconds to wait for the result before returning a job_id instead. Default 45, maximum 50, which keeps the call under common 60-second client timeouts.",
        +  "type": "number"
        +}
    • Addedcodex_job_result
    • Changedcodex_run1 field changed
      • addedInput schema / properties / wait_sec
        Added value: +{
        +  "description": "Seconds to wait for the result before returning a job_id instead. Default 45, maximum 50, which keeps the call under common 60-second client timeouts.",
        +  "type": "number"
        +}
  2. 3 tool updates
    • Changedcodex_generate_image2 fields changed
      • removedInput schema / additionalProperties
        Removed value: -false
      • changedInput schema / properties / size / description
        Previous value: -"Pixel size, e.g. \"1024x1024\" (fastest, the default), \"1536x1024\" landscape, \"1024x1536\" portrait, \"2048x2048\", \"3840x2160\"."New value: +"Pixel size as WIDTHxHEIGHT, e.g. \"1024x1024\" (fastest, the default), \"1536x1024\" landscape, \"1024x1536\" portrait, or \"auto\". This is a request to Codex; the actual dimensions are reported in the result."
    • Changedcodex_read_artifact1 field changed
      • removedInput schema / additionalProperties
        Removed value: -false
    • Changedcodex_run1 field changed
      • removedInput schema / additionalProperties
        Removed value: -false
  3. 3 tool updatesv0.1.0
    • First observedcodex_generate_image
    • First observedcodex_read_artifact
    • First observedcodex_run

TDQS

A4.1/5.0

Scored across 4 tools

Disambiguation5/5

Each tool has a clearly distinct purpose: running a coding task, generating images, polling async job results, and reading workspace artifacts. There is no overlap or ambiguity between them.

Naming Consistency3/5

All tools share the codex_ prefix, but the pattern is inconsistent: codex_run and codex_generate_image are verb-led, codex_read_artifact is verb_noun, while codex_job_result is noun_noun. The naming is readable but not uniform.

Tool Count4/5

Four tools is a compact but reasonable set for a local Codex CLI integration. Each tool serves a necessary function in the workflow, though the surface is slightly minimal.

Completeness4/5

The core lifecycle of submitting a task, polling for results, and reading generated artifacts is covered. Minor gaps exist, such as no way to cancel a running job or list available workspace files, but these are not critical.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    C
    quality
    F
    maintenance
    Connects AI assistants like Claude to the Codex CLI for code analysis, editing, and execution. Supports file references with @ syntax, sandboxed code execution with approval workflows, and structured code changes for automated refactoring and documentation.
    8
    72 npm
    178
    MIT
  • F
    license
    B
    quality
    D
    maintenance
    Connects AI assistants to a local Codex engine for performing deep, project-level code reviews and automated refactoring. It enables context-aware bug fixes and multi-file analysis through a standardized bridge between modern AI clients and local development environments.
    4
    2
    -
  • F
    license
    Not graded
    quality
    B
    maintenance
    Enables external AI (ChatGPT, Claude) to securely access a local workspace via Cloudflare Tunnel, providing file operations, shell commands, Git operations, and more with bearer token authentication and sandbox isolation.
    1
    -