codex-local-mcp
This MCP server lets any LLM drive a local Codex CLI to run coding tasks and generate images inside isolated workspaces, returning output, files, and inline images.
Run coding tasks (
codex_run): send natural-language tasks; Codex writes/runs code in a sandboxed workspace; returns output and created files.Generate images (
codex_generate_image): create images natively via Codex (no separate image API key); choose count, size, quality; images open in the OS viewer by default.Poll long-running jobs (
codex_job_result): collect results for tasks that exceed the wait window, using a returnedjob_id; results retained for up to an hour.Read workspace artifacts (
codex_read_artifact): fetch files that were too large to inline (images as image content, others as text) viaresources/read.Workspace isolation: everything is scoped to a configurable root; only filesystem changes inside the workspace are detected and reported.
Configurable behavior: set timeouts, output caps, network access, model overrides, max concurrent jobs, and job retention via environment variables.
Safe execution: prompt travels via stdin, sandbox uses
workspace-writemode, and timeouts kill the entire process tree.
Enables AI agents to drive the local OpenAI Codex CLI, allowing them to execute coding tasks in an isolated workspace, generate images, and read the resulting artifacts.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@codex-local-mcpWrite a Python script that sorts a list of numbers and save it as sort.py"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
codex-local-mcp
An MCP server that lets any LLM drive your local Codex CLI.
The calling model sends a task in plain English. Codex writes and runs code inside an isolated workspace folder on your machine. The server returns Codex's output, the list of files it created, and the images inline.
Codex generates images natively, so codex_generate_image needs no image API key — your
existing codex login is enough.
This runs code on your computer. Anything that can call these tools can make Codex execute code as you, inside its sandbox. Read SECURITY.md before using it.
Requirements
Node.js 20 or newer
Codex CLI installed and logged in
Windows, macOS or Linux
Related MCP server: Codex MCP Server
Install
New to this? Follow the step-by-step setup guide instead, which covers finding your paths, configuring each client, verifying it works, and troubleshooting.
npm install -g @openai/codex # the Codex CLI itself
codex login # one time
git clone https://github.com/ossmalaysia/codex-local-mcp.git
cd codex-local-mcp
npm install
npm run buildConnect it to a client
Full walkthrough with per-OS paths and troubleshooting: SETUP.md.
This is a stdio server: the client launches it as a child process. There is no URL and nothing listens on a port.
Use absolute paths for command and for CODEX_BIN. MCP clients launch servers with a
minimal environment in which bare node and codex may not resolve.
Claude Desktop
Add to claude_desktop_config.json, then fully quit and reopen the app.
Windows:
%APPDATA%\Claude\claude_desktop_config.jsonmacOS:
~/Library/Application Support/Claude/claude_desktop_config.json
{
"mcpServers": {
"codex": {
"command": "/absolute/path/to/node",
"args": ["/absolute/path/to/codex-local-mcp/dist/index.js"],
"env": {
"CODEX_BIN": "/absolute/path/to/codex",
"CODEX_MCP_ROOT": "/absolute/path/to/codex-workspaces",
"CODEX_TIMEOUT_SEC": "600"
}
}
}
}On Windows the values look like C:/Program Files/nodejs/node.exe and
C:/Users/you/AppData/Roaming/npm/codex.cmd. Forward slashes work fine.
Find the right paths with which node / which codex (macOS, Linux) or
(Get-Command node).Source / (Get-Command codex).Source (PowerShell).
Claude Code
claude mcp add codex -s user -- node /absolute/path/to/codex-local-mcp/dist/index.jsAnything else
Any MCP client that can launch a stdio server works. Point it at
node /absolute/path/to/dist/index.js.
Tools
Tool | Arguments | Description |
|
| Run any task. |
|
| Generate images. The prompt is passed through as written. |
|
| Collect the result of a task that was still running. |
|
| Read a file from a workspace. |
Reuse the same workspace name across calls to keep working on the same files.
Long-running tasks
MCP clients built on the official SDK give up on a request after 60 seconds by default, and
image generation often takes longer. So a call waits up to wait_sec (default 45, maximum 50)
for Codex. If Codex has finished, you get the result as usual. If not, the call returns at once
with:
status: running
job_id: a3b1d698-5754-490c-ab5d-bc1ea76149a2
workspace: /path/to/codex-workspaces/my-taskCodex keeps working. Call codex_job_result with the job_id to collect the result; each
call waits up to wait_sec again, so poll until it finishes. Finished results stay available
for an hour. Jobs are held in memory, so a server restart loses the job_id, but any files
produced are still in the workspace shown.
codex_generate_image defaults to quality: "low" and size: "1024x1024" — a fast draft.
size must be WIDTHxHEIGHT or auto; a malformed value is rejected before Codex runs. The
size is a request to Codex rather than something the server can enforce, so the result
reports each image's actual dimensions and flags any that differ from what was asked.
Raise quality to high for final assets. open defaults to true and opens each image in
your default viewer, because MCP clients deliver tool-result images to the model's context
and do not render them into the chat for you.
Configuration
All optional, set through the client's env block.
Variable | Default | Purpose |
|
| Root for all workspaces. Nothing is written outside it. |
|
| Absolute path to the Codex CLI. |
| (Codex default) | Model override. |
|
| Default timeout before a run is killed. |
|
| Ceiling a caller may request. |
|
| Output cap; the head and tail are kept. |
|
| Above this, a downscaled preview is inlined instead. |
|
| Cap on reported changed files. |
|
| Set |
|
| How long a call waits before returning a |
|
| How many Codex jobs may run at once. |
|
| How long a finished job's result stays retrievable. |
How it works
Codex runs as --sandbox workspace-write: it can read and write only inside the workspace
folder, plus reach the network. Workspace names that would escape the root are refused.
A few decisions worth knowing about:
Artifacts are found by filesystem diff. The workspace is snapshotted (path → mtime+size) before and after each run, so whatever Codex does, new files are detected. No output parsing.
The prompt travels on stdin (
codex exec ... -), so prompt text is never shell-quoted and cannot break out into the command line.Budgets are counted in base64 characters, not file bytes. Encoding inflates data by 4/3, so a budget expressed in file bytes overshoots what the client actually receives by a third. Images are downscaled until their encoded form fits, and a per-response budget caps all of them together. Originals on disk are never modified.
Full-resolution files are linked, not embedded. Every reported file also comes back as an MCP
resource_link. The server declares theresourcescapability and serves workspace files throughresources/read, so a client can fetch the original on demand instead of having megabytes pushed into every response. Reads outside the workspace root are refused.Generated images open in the OS viewer, since clients don't render tool-result images.
Timeouts kill the whole process tree, so a stuck build leaves no orphans.
Development
npm test # runs against a fake Codex - no CLI or API key needed
npm run buildSee CLAUDE.md for architecture and the invariants to preserve, and CONTRIBUTING.md before opening a pull request.
CI runs the build, the test suite and npm audit on Ubuntu, Windows and macOS against
Node 20 and 22, plus weekly CodeQL analysis.
Versioning
This project follows Semantic Versioning. Releases are cut automatically from Conventional Commits and published as tags with release notes on the Releases page. To use a specific release, check out its tag:
git checkout v0.1.0 # or any later tagWhile the major version is 0, minor releases may still contain breaking changes, which are
always called out in the release notes.
Security
Please read SECURITY.md. Report vulnerabilities privately through GitHub's Security → Report a vulnerability, not as a public issue.
License
MIT © OSS Malaysia
Available Tools
4 toolscodex_generate_imageGenerate images via the local Codex CLIA
Generate images with the local Codex CLI. The prompt is passed through as written. Returns the saved file paths and shows the images inline, downscaling a preview when a file is too large to inline. Defaults to a fast low-quality 1024x1024 draft; raise quality for final assets. Image generation often takes longer than wait_sec, in which case this returns status: running and a job_id; call codex_job_result with it to collect the images.
| Name | Required | Description | Default |
|---|---|---|---|
| open | No | Open the generated images in the user's default image viewer. Default true, because MCP clients do not render tool-result images into the chat. | |
| size | No | Pixel size as WIDTHxHEIGHT, e.g. "1024x1024" (fastest, the default), "1536x1024" landscape, "1024x1536" portrait, or "auto". This is a request to Codex; the actual dimensions are reported in the result. | |
| count | No | How many images. Default 1. | |
| prompt | Yes | What the image should depict. | |
| quality | No | Generation quality. Default "low" - fastest, good for drafts. Use "high" for final assets. | |
| wait_sec | No | Seconds to wait for the result before returning a job_id instead. Default 45, maximum 50, which keeps the call under common 60-second client timeouts. | |
| workspace | No | Workspace folder name. Defaults to a new folder. | |
| timeout_sec | No | Seconds before Codex is killed. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It reveals that the prompt is passed verbatim, that results include file paths and inline previews with downscaling for oversized files, that generation often exceeds wait_sec returning a status and job_id, and how to retrieve results. It also discloses defaults and quality trade-offs. This is exceptionally transparent for a tool with zero annotation support.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is five sentences, each earning its place: the core action, the pass-through behavior, the return format, the default quality guidance, and the asynchronous fallback. It front-loads the primary purpose and packs essential behavioral details without redundancy or fluff.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with 8 parameters, 100% schema coverage, and no output schema, the description covers everything an agent needs: the synchronous/asynchronous split, how to handle long runs, the default quality and size, and the inline display behavior. Sibling tools are correctly referenced for follow-up, and the description is self-sufficient for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so a baseline of 3 applies. The description adds value by explaining the practical impact of quality (draft vs. final), the wait_sec/job_id behavior, and the inline preview downscaling. It doesn't rehash schema definitions but provides usage-level meaning for key parameters, exceeding the baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Generate images with the local Codex CLI.' It clearly differentiates from siblings (codex_job_result, codex_read_artifact, codex_run) by focusing on image generation, while the others handle results, artifacts, and generic execution. The scope is unambiguous and an agent can immediately tell this tool's purpose.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear operational context: it explains the default draft behavior, quality guidance for final assets, and explicitly tells the agent to call codex_job_result when a job_id is returned. However, it does not explicitly state when to avoid this tool or when to prefer a sibling (e.g., 'use codex_run for non-image tasks'), though the name and content make that implicit. This is a minor gap, so not a perfect 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
codex_job_resultCollect the result of a running Codex jobA
Get the result of a codex_run or codex_generate_image call that returned status: running. Waits up to wait_sec for the job to finish; if it is still running, call again. Results stay available for 60 minutes after the job finishes.
| Name | Required | Description | Default |
|---|---|---|---|
| job_id | Yes | The job_id returned by codex_run or codex_generate_image. | |
| wait_sec | No | Seconds to wait for the result before returning a job_id instead. Default 45, maximum 50, which keeps the call under common 60-second client timeouts. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral burden and does well: it discloses that the tool waits up to wait_sec, that polling may need to be repeated, and that results persist for 60 minutes. It does not describe response statuses in detail, but the core polling behavior is transparent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences convey purpose, usage condition, polling behavior, and retention in a compact, front-loaded structure. There is no filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description is sufficiently complete for a simple polling tool with two well-documented parameters. However, there is no output schema, so the description could have been more explicit about the response shape or status values returned after waiting. Still, the essential invocation workflow is clear.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema already describes both parameters fully, including that job_id comes from codex_run or codex_generate_image and that wait_sec defaults to 45 with a maximum of 50. The description reinforces these facts but adds little semantic value beyond them, so baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly specifies a distinct action: retrieving the result of a codex_run or codex_generate_image call that previously returned status: running. It names the exact operation, resource type, and state condition, making it unmistakable from siblings that start jobs or read artifacts.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives strong usage context: it is meant for jobs that returned status: running, and it instructs the agent to call again if the job is still running. It does not explicitly contrast with codex_read_artifact or list when not to use this tool, but the operational guidance is clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
codex_read_artifactRead a file produced by CodexA
Read one file from a Codex workspace, for artifacts that were too large to inline. Images come back as image content, everything else as text.
| Name | Required | Description | Default |
|---|---|---|---|
| path | Yes | Path to the file, relative to the workspace root or absolute inside it. | |
| max_bytes | No | Refuse files larger than this. Default 5000000. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It adds useful behavior beyond the schema by stating that images are returned as image content and everything else as text. However, it does not disclose error handling, permission requirements, or what happens with nonexistent files—gaps that would matter for a read tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences with no filler. The purpose is front-loaded, and the output-type behavior is stated succinctly. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple tool with two parameters and full schema documentation, the description covers the key context: what file to read and how the content will be returned. It lacks explicit error handling and edge-case details, but does not need to explain return values since there is no output schema and the description handles output type.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both parameters are already documented. The description adds no parameter-specific guidance, but it does not need to because the schema fully covers path and max_bytes semantics. This matches the baseline for full schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Read'), a resource ('one file from a Codex workspace'), and the purpose ('artifacts that were too large to inline'). It also differentiates from siblings by clarifying that it reads files, while codex_generate_image generates images and codex_run likely executes tasks.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives the clear context 'for artifacts that were too large to inline,' which tells an agent when to reach for this tool. It does not explicitly name alternative tools or state exclusions, but the use case is specific enough to strongly infer when it applies.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
codex_runRun a task with the local Codex CLIA
Hand a natural-language task to the local Codex CLI coding agent. Codex can write and run code, install packages and call APIs inside an isolated workspace folder, then this returns its output plus any files it created. If Codex takes longer than wait_sec, this returns status: running and a job_id instead; call codex_job_result with it to collect the result.
| Name | Required | Description | Default |
|---|---|---|---|
| model | No | Override the Codex model for this run. | |
| prompt | Yes | The task for Codex, written as you would to a developer. | |
| wait_sec | No | Seconds to wait for the result before returning a job_id instead. Default 45, maximum 50, which keeps the call under common 60-second client timeouts. | |
| workspace | No | Relative workspace folder name under the workspace root. Reuse the same name to continue working on the same files. Defaults to a new timestamped folder. | |
| timeout_sec | No | Seconds before Codex is killed. Default 300. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries full behavioral disclosure. It reveals side effects (writing/running code, installing packages), isolation ('inside an isolated workspace folder'), and async behavior with wait_sec and job_id. It does not mention error conditions or persistence of created files beyond the return statement, but covers the main behavioral risks and control flow.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences with no filler. The main action, behavioral scope, and async fallback are all front-loaded and each sentence contributes distinct information needed to select and invoke the tool correctly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a five-parameter, async, no-output-schema tool, the description covers purpose, side effects, async behavior, and continuation semantics. It doesn't specify what the final result payload looks like or how timeout_sec interacts with wait_sec, but what an agent needs to make a correct call and collect results is present.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds meaningful semantics: it explains that wait_sec controls the sync/async split, that a job_id should be passed to codex_job_result, and that reusing workspace continues work on the same files. These enrich beyond the raw schema descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('hand... to'), resource ('local Codex CLI coding agent'), and scope ('can write and run code, install packages and call APIs'). It also distinguishes its output behavior from siblings by explaining it returns output plus created files, and routes async cases to codex_job_result.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear when-to-use context and explicit guidance for the async case: if execution exceeds wait_sec, returns status:running and a job_id, and the agent is instructed to call codex_job_result. It does not explicitly contrast with codex_read_artifact or codex_generate_image, so it falls just short of full alternative differentiation.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
3 tool updates
v0.2.0- Changed
codex_generate_image1 field changed- added
Input schema / properties / wait_secAdded value: +{ + "description": "Seconds to wait for the result before returning a job_id instead. Default 45, maximum 50, which keeps the call under common 60-second client timeouts.", + "type": "number" +}
- Added
codex_job_result - Changed
codex_run1 field changed- added
Input schema / properties / wait_secAdded value: +{ + "description": "Seconds to wait for the result before returning a job_id instead. Default 45, maximum 50, which keeps the call under common 60-second client timeouts.", + "type": "number" +}
3 tool updates
- Changed
codex_generate_image2 fields changed- removed
Input schema / additionalPropertiesRemoved value: -false - changed
Input schema / properties / size / descriptionPrevious value: -"Pixel size, e.g. \"1024x1024\" (fastest, the default), \"1536x1024\" landscape, \"1024x1536\" portrait, \"2048x2048\", \"3840x2160\"."New value: +"Pixel size as WIDTHxHEIGHT, e.g. \"1024x1024\" (fastest, the default), \"1536x1024\" landscape, \"1024x1536\" portrait, or \"auto\". This is a request to Codex; the actual dimensions are reported in the result."
- Changed
codex_read_artifact1 field changed- removed
Input schema / additionalPropertiesRemoved value: -false
- Changed
codex_run1 field changed- removed
Input schema / additionalPropertiesRemoved value: -false
3 tool updates
v0.1.0- First observed
codex_generate_image - First observed
codex_read_artifact - First observed
codex_run
TDQS
Scored across 4 tools
Each tool has a clearly distinct purpose: running a coding task, generating images, polling async job results, and reading workspace artifacts. There is no overlap or ambiguity between them.
All tools share the codex_ prefix, but the pattern is inconsistent: codex_run and codex_generate_image are verb-led, codex_read_artifact is verb_noun, while codex_job_result is noun_noun. The naming is readable but not uniform.
Four tools is a compact but reasonable set for a local Codex CLI integration. Each tool serves a necessary function in the workflow, though the surface is slightly minimal.
The core lifecycle of submitting a task, polling for results, and reading generated artifacts is covered. Minor gaps exist, such as no way to cancel a running job or list available workspace files, but these are not critical.
Maintenance
Related MCP Connectors
Operate Linux, macOS and Windows from your LLM. Every action runs through an auditable allowlist.
Use your own Mac from ChatGPT, Claude or Codex: files, commands, documents, and a browser.
- mcp-serverOAuthai.cdbx
Build Apps and run code in 30 languages — sandboxed, with persistent sessions for agent loops.
Build and supervise fleets of agents from Claude Code, Codex or Cursor. Connects over OAuth.
Related MCP Servers
- AlicenseCqualityFmaintenanceConnects AI assistants like Claude to the Codex CLI for code analysis, editing, and execution. Supports file references with @ syntax, sandboxed code execution with approval workflows, and structured code changes for automated refactoring and documentation.872 npm178MIT
- FlicenseBqualityDmaintenanceConnects AI assistants to a local Codex engine for performing deep, project-level code reviews and automated refactoring. It enables context-aware bug fixes and multi-file analysis through a standardized bridge between modern AI clients and local development environments.42-
- FlicenseNot gradedqualityBmaintenanceEnables external AI (ChatGPT, Claude) to securely access a local workspace via Cloudflare Tunnel, providing file operations, shell commands, Git operations, and more with bearer token authentication and sandbox isolation.1-
- AlicenseAqualityCmaintenanceEnables Codex to delegate coding tasks to an OpenCode CLI locally, returning structured results such as exit codes, session summaries, tool calls, and git diffs.2MIT