codex-mcp
Provides a read-only connector to Jira for gathering task/requirement context and evidence during quality review, without mutating Jira data.
Click on "Install Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@codex-mcpreview my candidate bug findings for the login flow"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
codex-mcp
An independent, read-only quality gate for QA artifacts.
codex-mcp is a standalone MCP server that runs Codex as an adversarial second
reviewer over candidate test cases and bug findings — before the authoring
agent writes its final report. Codex inspects the repository itself, forms its
own view of what should be covered or whether a defect is real, and only then
compares that against the candidate it was given.
It returns a review delta. It never writes your artifact.
Authoring agent (Claude, or any MCP client)
│ gathers the requirement, reads the code, drafts candidates
▼
candidate result — in memory, not yet written
│
▼ codex_qualify
codex-mcp ──► Codex (read-only sandbox, rooted at your repo)
│ ├─ reads the code, the diff, the existing tests
│ ├─ reads blast-radius / test-charter if present
│ └─ reads Jira / DB / other MCPs if configured, read-only
▼
review delta: accept · modify · remove · missing · evidence · limitations
│
▼
Authoring agent reconciles, then writes the FINAL artifactContents · Install · Connect to a project · Use it · Configuration · Evidence connectors · Permission boundary · API contract · Troubleshooting · Testing
Why a second model, and why read-only
The failure mode this addresses is not "the agent cannot write test cases". It is that an agent grading its own work agrees with itself. A reviewer that shares the author's context inherits the author's blind spots.
So two properties are load-bearing:
Independence. Codex is prompted to derive expected coverage before it looks closely at the candidate, and to try to falsify each bug claim rather than confirm it. Anchoring it on the candidate first would produce a more agreeable reviewer and a less useful one.
Read-only. The reviewer runs in Codex's read-only sandbox, and every
downstream system it can reach is filtered through a policy layer that classifies
each tool and refuses anything that mutates. A quality gate you cannot safely
point at a live repository is a quality gate nobody runs.
Neither model is authoritative. Source evidence is:
requirement / runtime / code / DB / external evidence > model opinionRelated MCP server: tenth-man-mcp
Install
Requires Node 20+. Four commands, once per machine.
# 1. The Codex CLI. codex-mcp drives it, and it owns your credentials.
npm install -g @openai/codex@latest
# 2. codex-mcp itself.
git clone <this-repo> codex-mcp && cd codex-mcp
npm install && npm run build && npm link
# 3. Sign in. A browser opens once; that is the whole flow.
codex-mcp login
# 4. Write a config, detecting any MCP servers already on this machine.
codex-mcp init --model gpt-5.6-solThen confirm before trusting it:
codex-mcp doctorEvery line should read ok. doctor is read-only and safe against a live
project — see Troubleshooting for what each failure means.
The npm name
codex-mcpbelongs to an unrelated package. Install from source as above, or publish under your own scope.
What init does
It writes ~/.config/codex-mcp/codex-mcp.yaml, and a .env beside it if it
found downstream MCP servers in conventional locations:
$ codex-mcp init --dry-run
Would write into /home/you/.config/codex-mcp
codex-mcp.yaml (new)
.env (new)
Detected:
jira-mcp (jira) -> /home/you/jira-mcp/src/index.js
db-mcp (database) -> /home/you/db-mcp/dist/index.jsDetected servers are written as connector entries you can enable, with their
paths kept in .env so the YAML stays portable across machines. --force
overwrites; without it, existing files are kept.
init is the only command that writes anything, and it runs before any
review exists. Reviews themselves are strictly read-only.
Prefer to write the config yourself? Copy
codex-mcp.example.yaml to
~/.config/codex-mcp/codex-mcp.yaml — every value in it is annotated and is the
built-in default unless its comment says otherwise.
Authentication
One command, once per machine:
codex-mcp login # browser opens; sign in to ChatGPT
codex-mcp auth-status # confirmCredentials land in the Codex CLI's own store (~/.codex/auth.json, mode 0600)
and are refreshed by it. They persist across reboots and terminals — you do not
log in again per project or per session.
codex-mcp never handles the credential itself. It has no OAuth client, no
callback listener, no token storage. It shells out to codex login status and
reads yes/no. Nothing here can open a browser during a review: an
unauthenticated call fails fast instead.
{ "code": "CODEX_AUTH_REQUIRED", "message": "Codex is not authenticated. Run `codex-mcp login`." }Using an API key instead
Mode | Command | Uses |
|
| Browser OAuth, ChatGPT subscription |
|
| An OpenAI API key |
codex-mcp login --mode api # hidden prompt
printenv OPENAI_API_KEY | codex-mcp login --mode apiThe key is read from --api-key, then OPENAI_API_KEY, then a hidden prompt,
and piped to codex login --with-api-key over stdin — never an argv element,
so it stays out of your process table and shell history. The Codex CLI stores it;
codex-mcp does not.
Set auth.mode in your config to match. If the CLI is authenticated in a
different mode than the config claims, reviews fail with a clear error rather
than silently billing the wrong account.
Connect to a project
Pick one of these. Registering twice is the most common setup mistake — see the warning below.
Option A — one project, committed
Create .mcp.json in the project root — not .claude/, which holds a
different set of files and will ignore it:
{
"mcpServers": {
"codex-mcp": { "command": "codex-mcp", "args": ["start"] }
}
}Commit it. Every teammate who has run Install now has the gate, using their own model choice from their own config.
Option B — all your projects, not committed
claude mcp add codex-mcp -- codex-mcp startThis writes to ~/.claude.json and applies wherever you work.
Register in one place only.
claude mcp addwrites at local scope, which outranks.mcp.json. With both present, the project file — including anyenvin it — is silently ignored. Runclaude mcp remove codex-mcpif you are switching to.mcp.json.
Restart Claude Code. /mcp should now list codex-mcp with three tools. If it
does not, see Troubleshooting.
Other MCP clients take the same two fields — command: codex-mcp,
args: ["start"] — in whatever config file they use.
Use it
Make it automatic
Add one rule to the project's CLAUDE.md. This is the entire integration
surface — no project-specific Codex logic anywhere:
## Independent QA qualification
Before finalizing test cases or bug reports, send the complete candidate result,
project root, task/requirement context, and any available blast-radius or
test-charter to codex-mcp for independent qualification.
Reconciling means verifying each objection against the evidence it cites — not
accepting it. Apply what the evidence supports. Reject what it does not, and note
why. codex-mcp is a second opinion, not an approver.
One pass is normal. Run a second only if the first forced substantial high-risk
changes.With that in place you write your normal request and the gate runs itself:
Create test cases for DEV-2951.
The agent gathers the requirement, reads the code, drafts candidates, calls
codex_qualify, reconciles, then writes the report.
Ask for it explicitly
When there is no rule, or you want it on something already drafted:
Before you write the report, send these test cases to codex-mcp with
project.rootset to/path/to/repoandtask.idDEV-2951. Show me what it objects to and whether you agree, then write the final version.
Run the bug findings you just wrote through
codex_qualifywithreviewType: "bugs". For anything it calls a false positive, check the code it cites before you drop the finding.
Qualify these against codex-mcp, but only report objections where the cited evidence actually holds up. Tell me which ones you rejected and why.
Useful variations:
You want | Add to your prompt |
Both tests and bugs in one go |
|
Focus on one risk area |
|
Skip the database |
|
Feed it your artifacts |
|
Reading what comes back
Ask your agent to surface these rather than silently acting on them:
missing— coverage it says you lack. Check the citedfile:lineis real.modify— your expectation contradicts the code. Usually the sharpest finding.remove— redundant. Verify the thing it says supersedes yours actually does.limitations— what it could not verify. A confident review with a long limitations list is a narrow review; read this before trusting the rest.disagreements— it and your agent read the same evidence differently. These need you, not either model.
A good follow-up prompt:
For each objection, tell me the evidence it cited and whether you verified it yourself. List anything you rejected and why.
What it will not do
It never edits files, commits, pushes, writes to Jira or the database, or writes
your report. If your agent claims codex-mcp changed something, it did not — check
git status.
Reconciliation — the part that matters
codex-mcp never tells you to accept Codex. Every response carries:
{
"reconciliation": {
"instruction": "This is an independent second opinion, not a verdict...",
"codexIsNotAuthoritative": true
}
}Codex objection
│
▼
Author verifies the cited evidence
│
├─ evidence supports it → apply
├─ evidence does not → reject, and record why
└─ unclear → investigateThen you write the final artifact.
Loop protection
review.maxPasses (default 2) caps the cycle. Pass 1 is the normal case;
pass 2 exists for reviews that forced substantial high-risk changes. A request
above the limit is refused, and meta.furtherPassesAllowed tells you when the
budget is spent. Iterating until the two models agree is not the goal — agreement
is cheap, and reaching it usually means one of them stopped thinking.
Configuration
Settings live in ~/.config/codex-mcp/codex-mcp.yaml. That is the one file
to edit. codex-mcp.example.yaml is the annotated
version, with the current Codex model ids.
An absent config file is fully supported — the server starts on defaults, which are the safe ones. What you lose is the model pin and every connector, since connectors can only be defined in YAML.
Where the model goes
review:
model: gpt-5.6-sol
requireModel: trueThree layers can set it. Highest wins:
Layer | Scope | Use it when |
| one project, everyone who clones it | the team must all review with one specific model |
| this machine, every project | normal setup — one operator, several projects |
built-in default | — | you accept whatever Codex currently defaults to |
Keep the model in one layer. A key set in two places makes the lower copy
dead — editing it appears to do nothing. doctor warns when the model is set in
both and the two disagree.
To pin a project for a whole team, add env to the .mcp.json from
Option A:
{
"mcpServers": {
"codex-mcp": {
"command": "codex-mcp",
"args": ["start"],
"env": { "CODEX_MODEL": "gpt-5.6-sol", "CODEX_REASONING_EFFORT": "high" }
}
}
}That guarantees everyone gets the same reviewer, which is worth a lot when
findings are compared across a team. The cost: a teammate whose Codex CLI is too
old for that model gets a hard CODEX_MODEL_NOT_AVAILABLE telling them to run
codex update. That failure is deliberate — the alternative is them quietly
getting a weaker reviewer and trusting its verdict.
Which model. Prefer a frontier model. The entire value here is catching what
the authoring agent missed, and a cheaper reviewer mostly agrees with whatever it
is shown. Set requireModel: true on a shared gate so a change to the Codex
default cannot quietly alter review quality. An unavailable model raises
CODEX_MODEL_NOT_AVAILABLE; codex-mcp will not fall back to another one.
Every setting, and its environment variable
Precedence: environment > codex-mcp.yaml > defaults. The variables exist to
override one YAML value without editing the file — in .mcp.json's env to pin
a project, or in the shell for a one-off.
| Environment variable | Default |
|
| (none — Codex decides) |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| (none) |
|
|
|
|
Connector settings invert that precedence: the YAML wins, because it is explicit per-connector intent and these variables are coarse fallbacks for when no YAML says otherwise.
Environment variable | Falls back for |
|
|
|
|
|
|
|
|
Those toggles match the connector's name, not its kind — connectors called
jira-mcp and db-mcp match neither jira nor database, so both fall under
CUSTOM_MCPS_ENABLED. Setting enabled: in the YAML avoids the question
entirely.
Two variables have no YAML equivalent, since they are read before any config file
is located: CODEX_MCP_CONFIG (path to the config file) and XDG_CONFIG_HOME
(where ~/.config/codex-mcp/ is looked for).
Where the config file is found
First hit wins:
--config <path> → $CODEX_MCP_CONFIG → ./codex-mcp.yaml → ~/.config/codex-mcp/codex-mcp.yamldoctor prints which one it loaded. A .env beside the chosen file is read if
present; nothing requires one. It is worth having only for values that differ per
machine — connector paths the YAML references as ${JIRA_MCP_PATH} — and even
those can carry a ${VAR:-fallback} default instead.
Never put credentials in .env or the YAML. CHATGPT_TOKEN,
SESSION_TOKEN, ACCESS_TOKEN, and REFRESH_TOKEN are ignored outright, and
their presence is reported as a configuration warning. Codex authentication
belongs to the Codex CLI and your OS credential store.
Evidence connectors
Codex never talks to Jira or your database directly. It connects to the codex-mcp evidence broker, a separate read-only process that discovers each downstream server's tools, classifies them, and forwards only what passes policy — re-checking on every call, not just at discovery.
# ~/.config/codex-mcp/codex-mcp.yaml
connectors:
jira-mcp:
enabled: true
kind: jira
approval: once
transport: stdio
command: node
args: ['/path/to/jira-mcp/src/index.js']
cwd: /path/to/jira-mcp
db-mcp:
enabled: true
kind: database
approval: once
transport: stdio
command: node
args: ['/path/to/db-mcp/dist/index.js']
cwd: /path/to/db-mcp
allowTools: ['execute_query']
denyTools: ['update_query']
maxRows: 500
timeoutMs: 10000kind drives normalization onto a stable vocabulary — requirement.read,
database.query_readonly, testmanagement.search, external_file.read — so the
reviewer prompt can ask for "the requirement" without knowing whether your
connector calls it getJiraIssue or get_jira_ticket. Unmapped tools are still
exposed under their own names; adding a new read-oriented MCP requires no code
change.
A downstream server receives only PATH, HOME, and the env its own config
declares — never the codex-mcp process environment.
An unreachable connector degrades the review to a recorded limitation rather than failing it. Missing evidence is a fact about the review, and the response says so.
Run codex-mcp doctor after adding one. Each connector line reports how many
tools were exposed and how many were withheld by policy:
[ ok ] Connector: jira-mcp
4 read-only tool(s) exposed, 0 withheld by policy.
[ ok ] Connector: db-mcp
6 read-only tool(s) exposed, 1 withheld by policy.Asking permission — the approval field
Reading the project you were handed needs no permission: you supplied
project.root, so reading it is the request. Reaching outside it — a
ticket tracker, a production database, a file server — is a separate decision,
and enabled: true in a config file written weeks ago is not informed consent
for today's review.
| Behavior |
| Ask before every review |
| Ask once per server session — the default |
| Never ask |
The prompt is delivered through MCP elicitation, so it reaches the human in
your MCP client. If your client cannot show prompts, the connector is skipped
and recorded in limitations — not silently allowed. A prompt nobody can see is
not consent. Set approval: trusted on connectors you have already vetted.
Requirements
When a jira-kind connector is configured and task.id is set, Codex reads the
ticket itself and treats any requirement text you passed as the authoring
agent's interpretation — a claim to reconcile, not a source. Without a
connector it falls back to the text you supplied and records that it could not
verify it independently.
Database
Consulted only where it can change a verdict: persistence, relationships, tenant ownership, state transitions, migrations, data integrity, verifying a reported defect. The prompt says so explicitly, and the policy layer enforces the rest.
Use a read-only database account. codex-mcp refuses every mutating statement, but a read-only grant is the boundary that does not depend on this server being correct.
The permission boundary
The central rule: Codex may inspect broadly and mutate nothing.
Read "broadly" literally — see read scope below before pointing this at a machine holding secrets you care about.
Local | |
Read files, search, list, inspect tests, read artifacts | allow |
| allow |
Edit, create, delete files | deny |
| deny |
Shell wrappers, metacharacters, redirection, unknown binaries | deny |
Jira | |
Read issue, search, comments, linked issues, acceptance criteria | allow |
Create, edit, comment, transition, delete | deny |
Database | |
Read schema, | allow |
| deny |
Multi-statement payloads, | deny |
Enforcement, in layers:
Codex's own
read-onlysandbox — the primary boundary.Command policy — argv-based, default-deny. Unknown binaries are refused; shell wrappers are refused because their payload cannot be classified.
SQL policy — comments and string literals are stripped before keyword scanning, so a mutation cannot hide inside a quoted value. One statement per call, row cap injected when the query has none.
Tool policy — every downstream MCP tool is classified
read/write/destructive/unknown; onlyreadis exposed.unknownis denied unless explicitly allowlisted, and no allowlist can rescue a mutating tool — a boundary you can argue your way past is not a boundary.
The classifier is deliberately asymmetric: any hint of mutation beats any hint of reading, and a tool must look positively read-only to be exposed. A tool that sounds unsafe but is not costs you one line of config; a tool that sounds safe but is not costs you data.
tests/security/ asserts all of this, including that a refused call never
reaches the downstream server and that a fixture repository is byte-identical
after a review.
Read scope is wider than the project
Codex's read-only sandbox constrains writes, not reads. Inside it, Codex
can read any file your user account can read — not only files under
project.root. Verified directly:
$ codex exec --sandbox read-only -C ./proj "read ../outside.txt"
exec sed -n '1,$p' ../outside.txt in .../proj
succeeded: SECRET_OUTSIDE=canary-9f3a2bThe Codex CLI offers no option to narrow read scope; sandbox_permissions only
grants further access. So the honest statement of the guarantee is:
Nothing is modified, anywhere. Reads are bounded by your OS file permissions, not by
project.root.
project.root steers where the reviewer looks — it is the working directory
and the subject of the prompt — but it is not a read jail.
What this means in practice:
A
.env, private key, or credentials file anywhere readable by your user is reachable by the reviewer, and its contents may be sent to OpenAI as part of the model's context.codex-mcp's own artifact-path containment (
assertArtifactPathAllowed) stops codex-mcp from reading files outside the project into the prompt. It does not and cannot constrain what Codex reads inside its own sandbox.Findings are redacted before they are logged, but that is a logging control, not a containment one.
If that matters for your environment, run codex-mcp inside a container or VM with only the project mounted. That is the only reliable way to bound reads today.
What the reviewer reads by design
Within the project root it reads everything, including dot-directories.
Hiding .claude, .cursor, .github, or a team's own .qa from the reviewer
is how it ends up ignoring the very rules the project wrote down for it. Known
tool caches (.venv, .pytest_cache, .next, and the like) are still listed
but are not recommended to it as reading material.
Convention files — CLAUDE.md, AGENTS.md, CONTRIBUTING.md, TESTING.md,
.cursorrules, CODEOWNERS — are surfaced to the prompt as "read these first".
The contract
codex_qualify
Required: reviewType, project.root, and a candidate set matching the review
type. Everything else is optional and never blocks a review.
{
"reviewType": "test-design",
"project": { "root": "/absolute/path/to/project", "branch": "feature/DEV-123" },
"task": {
"id": "DEV-123",
"source": "jira",
"title": "Archive a resource",
"description": "A user may archive a resource belonging to their own tenant.",
"acceptanceCriteria": ["Archiving an active resource sets status to archived."]
},
"artifacts": {
"blastRadiusPath": "docs/blast-radius.md",
"testCharterPath": "docs/test-charter.md"
},
"candidate": {
"testCases": [{ "id": "TC-001", "title": "Archive an active resource", "priority": "high" }],
"bugs": []
},
"options": { "useJira": true, "useDatabase": true, "useExternalMcps": true }
}Candidates travel in the payload. They have not been written anywhere yet, and requiring a temporary report file would defeat the point.
Artifact paths are resolved inside project.root; a path escaping it is refused.
Review types
Type | Reviews |
| Coverage, redundancy, weak assertions, missing high-value scenarios |
| Whether each finding is real, a false positive, a duplicate, or unproven |
| Both, as two separate Codex runs — fusing the prompts degrades both |
Test-design result
{
"status": "CHANGES_REQUIRED",
"summary": { "accepted": 18, "modify": 2, "remove": 1, "missing": 3 },
"accepted": ["TC-001", "TC-002"],
"modify": [{
"candidateId": "TC-014",
"reason": "Expected state contradicts persistence logic.",
"evidence": [{ "source": "code", "location": "src/session/service.ts:143" }],
"recommendation": "Queue should remain persisted after this transition."
}],
"remove": [{ "candidateId": "TC-022", "reason": "Duplicates TC-018.", "supersededBy": "TC-018" }],
"missing": [{
"title": "Verify cross-tenant access is rejected",
"priority": "high",
"dimension": "authorization",
"reason": "Target lookup accepts an externally supplied identifier.",
"evidence": [{ "source": "code", "location": "src/resource/controller.ts:82" }]
}],
"disagreements": [],
"limitations": []
}Bug result
{
"status": "CHANGES_REQUIRED",
"summary": { "verified": 1, "falsePositive": 1, "needsMoreEvidence": 0, "other": 0 },
"findings": [{
"candidateId": "BUG-003",
"verdict": "FALSE_POSITIVE",
"confidence": "high",
"severityAssessment": null,
"reason": "Ownership validation occurs in router-level middleware.",
"evidence": [
{ "source": "code", "location": "src/routes/users.ts:42" },
{ "source": "code", "location": "src/middleware/access.ts:91" }
],
"recommendation": "Remove the finding unless runtime evidence contradicts the middleware."
}],
"limitations": []
}status: PASS · CHANGES_REQUIRED · INCONCLUSIVE · ERROR
verdict: VERIFIED · FALSE_POSITIVE · NEEDS_MORE_EVIDENCE ·
SEVERITY_DISAGREEMENT · DUPLICATE_OR_ALREADY_COVERED · INCONCLUSIVE
What the envelope guarantees
codex-mcp normalizes the reviewer's output before returning it, because a model
grading a list will sometimes drift:
ids the reviewer invented are dropped, with a note — you cannot act on a reference to a test case that does not exist;
a candidate the reviewer never mentioned is recorded as unreviewed, never promoted to accepted, because silence is not approval;
a bug with no verdict becomes an explicit
INCONCLUSIVE;summarycounts are recomputed from the arrays;statusis derived from the delta, not from the reviewer's self-assessment.
meta.evidence reports what the review was actually based on — whether git,
blast-radius, test-charter, and requirement access were available, and which
connectors were reachable. The project path itself is never logged or returned;
meta.evidence.projectRootId is a hash.
codex_auth_status
Whether Codex is authenticated, in which mode, and whether that matches your
configured auth.mode. Never returns a credential.
codex_capabilities
Diagnostic. What evidence this instance can reach, which downstream tools were withheld and why, and an explicit list of what the reviewer is forbidden to do.
CLI
codex-mcp init # write ~/.config/codex-mcp/, detecting local MCP servers
codex-mcp start # run the MCP server on stdio (what a client launches)
codex-mcp login # authenticate (--mode chatgpt|api)
codex-mcp auth-status # report auth state, never credentials
codex-mcp doctor # diagnose everything; mutates nothinginit takes --model <id>, --force, and --dry-run. start and doctor
take --config <path>. doctor also takes --project <path> and --json.
codex-mcp broker is internal — the evidence broker that Codex launches. You do
not run it by hand.
Troubleshooting
Symptom | Cause | Fix |
| Client not restarted, or | Restart; move the file to the project root |
| A |
|
| Not signed in |
|
| Codex CLI too old, or the model is not on your account |
|
| Codex CLI missing from |
|
Auth-mode mismatch error |
| Change one to match; do not silently bill the wrong account |
Connector missing from |
| Check the YAML; |
Connector skipped mid-review | Your client cannot show elicitation prompts | Set |
|
| Re-run |
Config edits do nothing | An environment variable outranks the file |
|
Everything doctor reports is either ok, warn (works, but looser than it
should be), or FAIL (reviews cannot work).
Errors
Stable codes, safe to branch on. Payloads are redacted before they leave the process.
CODEX_AUTH_REQUIRED CODEX_NOT_INSTALLED
CODEX_MODEL_NOT_CONFIGURED CODEX_MODEL_NOT_AVAILABLE
INVALID_PROJECT_ROOT PROJECT_ACCESS_DENIED
INVALID_REVIEW_REQUEST INVALID_REVIEW_TYPE
DOWNSTREAM_MCP_UNAVAILABLE DOWNSTREAM_MCP_PERMISSION_DENIED
DB_QUERY_DENIED DB_QUERY_TIMEOUT
CODEX_EXECUTION_FAILED CODEX_OUTPUT_INVALID
REVIEW_TIMEOUT INTERNAL_ERRORIf Codex returns output that does not match the schema, codex-mcp retries
once with an explicit correction that forbids re-analysis. If that also
fails it returns CODEX_OUTPUT_INVALID. It does not return a partially parsed
review — you would act on it.
Observability
Structured JSON to stderr (stdout belongs to the MCP transport). Logged: review id and type, hashed project id, model, timings, connector availability, candidate counts, Codex exit status, schema-validation status.
Never logged: tokens, passwords, DB credentials, cookies, secrets found in
source. Redaction runs at every level, including debug.
Testing
Five layers, cheapest first. Work down them — a failure at one layer makes the next layer's result meaningless.
1. Automated suite — free, offline, ~7s
npm install
npm run build
npm test
npm run typecheck400+ tests against a fake Codex CLI and a deliberately hostile fake MCP server. No network, no model calls, deterministic. This is what you run on every change and in CI.
tests/security/ is the part worth reading: it asserts that file edits,
commits, pushes, issue writes, and DB mutations are refused — and that a refused
call never reaches the downstream server.
2. doctor — is this install wired up correctly
codex-mcp doctor
codex-mcp doctor --project /path/to/repoRead-only, safe against a live project. Checks Node, the Codex CLI, auth, auth mode agreement, model, sandbox, config file, and every configured connector.
3. codex_capabilities — which evidence can it actually reach
doctor gives counts; this gives the per-tool breakdown, including why each
withheld tool was withheld. Call it from your MCP client, or:
node -e "
import('./dist/src/config/config.js').then(async ({loadConfig}) => {
const {CodexMcpServer} = await import('./dist/src/server.js');
const {Logger} = await import('./dist/src/util/logger.js');
const s = new CodexMcpServer({config: loadConfig(), logger: new Logger('error', {}, {write(){}})});
console.log(JSON.stringify(await s.callToolForTesting('codex_capabilities', {}), null, 2));
process.exit(0);
});"Check that the tools you expect are in allowedTools, and that every entry in
deniedTools is one you want denied. A read-only tool with an unusual name
lands in deniedTools as unknown — add it to that connector's allowTools.
4. npm run try — a real review, real model, real cost
This is the only layer that spends budget. It proves the whole path: auth, model, sandbox, evidence collection, connectors, prompt, structured output.
npm run try -- --project /path/to/repo
npm run try -- --project /path/to/repo --type bugs
npm run try -- --project /path/to/repo --type combined --task DEV-123
npm run try -- --project /path/to/repo --candidates ./candidates.json --jsonWith no --candidates it sends a set seeded with known flaws — two
duplicates, one assertion the code contradicts, and several obvious gaps. That
is the point: you are testing the reviewer, so use input whose correct answer
you already know.
Judge it on:
did it put the duplicate in
remove?did it put the contradicted assertion in
modify, citing the code?does every
missingentry have a realfile:line, not a vague area?is the repository unchanged afterwards (
git status)?
A PASS on the seeded set means something is wrong, not that your code is
clean.
Supply --candidates with your own JSON to rehearse a real workflow:
{ "testCases": [{ "id": "TC-1", "title": "..." }], "bugs": [] }5. End-to-end fixture — opt-in
CODEX_MCP_E2E=1 npm test -- tests/e2eBuilds a fixture repository containing a real coverage gap (idempotency) and a bug report that router middleware already refutes, runs a full qualification against the real Codex CLI, and asserts the fixture is byte-identical afterwards. Takes a few minutes.
Driving it as an MCP server
Once the layers above pass, drive it the way a client will:
printf '%s\n%s\n%s\n' \
'{"jsonrpc":"2.0","id":1,"method":"initialize","params":{"protocolVersion":"2025-06-18","capabilities":{},"clientInfo":{"name":"smoke","version":"1"}}}' \
'{"jsonrpc":"2.0","method":"notifications/initialized"}' \
'{"jsonrpc":"2.0","id":2,"method":"tools/list"}' \
| codex-mcp startThen register it in Claude Code and use it on a real ticket.
Development
src/
config/ resolution, precedence, validation
auth/ Codex CLI delegation for both auth modes
codex/ process spawning, argv construction, output parsing
review/ orchestration, per-type reviewers, output normalization
evidence/ repository, git, artifacts, requirement, database, external
mcp-broker/ downstream clients, discovery, classification, the broker server
policy/ command, SQL, MCP-tool, permission, and consent decisions
prompts/ base reviewer, test-design, bug-review
schemas/ public request and result contracts
tools/ the three MCP toolsLicense
MIT
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Servers
AlicenseNot gradedqualityAmaintenanceEnables multi-agent code review with cross-verification of findings against source code, catching hallucinations and improving agent accuracy over time.36438MIT- AlicenseAqualityDmaintenanceAdversarial review system that spawns three independent contrarian reviewers to catch issues before AI coding agents execute critical changes.39MIT
- AlicenseNot gradedqualityBmaintenanceProvides tools for agents to manage a local review graph, tracking acceptance behaviors, evidence, review passes, and human waivers to decouple review convergence from shipping readiness.3MIT
- AlicenseNot gradedqualityBmaintenanceA local-first, auditable code review MCP server that freezes Git changes, creates immutable ReviewBundles, provides role-isolated contexts for correctness, security, architecture, and test reviewers, validates structured findings, and generates deterministic JSON/Markdown reports.7Apache 2.0
Related MCP Connectors
Deterministic pre-execution audit for trading agents. PASS/WAIT/FAIL, reproducible verdict_hash.
Deterministic AI code review, with an audit record. Governance inside the agent loop.
Agentic code review, no signup to try: reality gates + frontier-model review, with veto.
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/salmansrabon/codex-mcp'
If you have feedback or need assistance with the MCP directory API, please join our Discord server