Skip to main content
Glama
ossmalaysia

codex-local-mcp

by ossmalaysia

Run a task with the local Codex CLI

codex_run

Run coding tasks via natural language in an isolated workspace using the local Codex CLI, returning output and files or a job ID for async retrieval.

Instructions

Hand a natural-language task to the local Codex CLI coding agent. Codex can write and run code, install packages and call APIs inside an isolated workspace folder, then this returns its output plus any files it created. If Codex takes longer than wait_sec, this returns status: running and a job_id instead; call codex_job_result with it to collect the result.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
modelNoOverride the Codex model for this run.
promptYesThe task for Codex, written as you would to a developer.
wait_secNoSeconds to wait for the result before returning a job_id instead. Default 45, maximum 50, which keeps the call under common 60-second client timeouts.
workspaceNoRelative workspace folder name under the workspace root. Reuse the same name to continue working on the same files. Defaults to a new timestamped folder.
timeout_secNoSeconds before Codex is killed. Default 300.

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed1 schema field changedv0.2.0
    • addedInput schema / properties / wait_sec
      Added value: +{
      +  "description": "Seconds to wait for the result before returning a job_id instead. Default 45, maximum 50, which keeps the call under common 60-second client timeouts.",
      +  "type": "number"
      +}
  2. Changed1 schema field changed
    • removedInput schema / additionalProperties
      Removed value: -false
  3. First observedv0.1.0

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries full behavioral disclosure. It reveals side effects (writing/running code, installing packages), isolation ('inside an isolated workspace folder'), and async behavior with wait_sec and job_id. It does not mention error conditions or persistence of created files beyond the return statement, but covers the main behavioral risks and control flow.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences with no filler. The main action, behavioral scope, and async fallback are all front-loaded and each sentence contributes distinct information needed to select and invoke the tool correctly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a five-parameter, async, no-output-schema tool, the description covers purpose, side effects, async behavior, and continuation semantics. It doesn't specify what the final result payload looks like or how timeout_sec interacts with wait_sec, but what an agent needs to make a correct call and collect results is present.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds meaningful semantics: it explains that wait_sec controls the sync/async split, that a job_id should be passed to codex_job_result, and that reusing workspace continues work on the same files. These enrich beyond the raw schema descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('hand... to'), resource ('local Codex CLI coding agent'), and scope ('can write and run code, install packages and call APIs'). It also distinguishes its output behavior from siblings by explaining it returns output plus created files, and routes async cases to codex_job_result.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear when-to-use context and explicit guidance for the async case: if execution exceeds wait_sec, returns status:running and a job_id, and the agent is instructed to call codex_job_result. It does not explicitly contrast with codex_read_artifact or codex_generate_image, so it falls just short of full alternative differentiation.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.