mcp-server-codex
This server lets an MCP client drive the Codex CLI as a local autonomous coding agent, delegating tasks, reviews, diffs, and image generation through structured tools.
Run Codex agents: start new sessions with a prompt, working directory, sandbox mode, model, config overrides, images, and timeouts.
Resume or fork sessions: continue prior work by session id/thread id or the most recent session, or branch a new session while leaving the original untouched.
Review code: run Codex reviews over uncommitted changes, a base branch diff, or a specific commit, with optional custom instructions.
Apply diffs: apply a Codex task's latest diff to the working tree via
git apply.List sessions: read saved sessions from disk to find ids for resume/fork.
Generate images: create images through Codex's built-in image generation, with prompt composition, output paths, sizes, transparency, reference images, and named presets.
Manage image presets: list, reload, set, update, and delete reusable preset configurations persisted to a JSON file.
Handle long runs: background execution with job ids, status checks, paginated event logs, and cancellation.
Exchange messages with Codex: read questions/findings from running agents, reply to blocked questions, and send unsolicited messages to one or all runs.
Enforce safety boundaries: path allowlists, default sandbox restrictions, and gating of dangerous full-access modes.
Provides tools for driving the OpenAI Codex CLI locally, including running and resuming agent sessions, forking threads, reviewing code, applying diffs, and generating images.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@mcp-server-codexReview my uncommitted changes and flag any issues"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
mcp-server-codex
Serveur MCP (Model Context Protocol) en TypeScript pour piloter le CLI Codex localement : lancer des sessions d'agent, les reprendre, les bifurquer, faire des revues de code, appliquer des diffs et générer des images — le tout depuis n'importe quel client MCP (Claude Code, Claude Desktop, Codex lui-même…).
Le code, les identifiants et les messages d'erreur sont en anglais : ils sont lus par des agents. La documentation est en français.
Ce que ça fait
Le CLI Codex est conçu pour un humain devant un terminal. Ce serveur le rend pilotable par un agent :
Il force
--jsonpartout et parse le flux JSONL en résultats structurés.Il gère les runs longs sans faire tomber l'appel d'outil (voir Exécution hybride).
Il restreint ce que Codex peut toucher sur le disque.
Il expose la génération d'images, que Codex n'offre pas comme service appelable.
Related MCP server: codex-mcp-server
Prérequis
Outil | Rôle | Installation Windows |
Codex CLI ≥ 0.154 | le binaire piloté |
|
Node.js ≥ 22.0 | exécution du serveur (≥ 22.18 pour développer) |
|
git |
|
|
Vous devez être authentifié côté Codex (codex login). Le serveur n'a besoin d'aucune clé API, y compris pour les images.
⚠️ La version de Codex publiée sur winget est en retard sur celle de npm. Ce serveur est écrit contre le comportement de la 0.154 (voir Notes d'implémentation). Préférez npm. Un Codex installé par npm n'apparaît pas dans
winget list: évitez de mélanger les deux voies, sous peine d'avoir deux binaires concurrents dans lePATH. En cas de doute, pointezCODEX_BINsur le bon exécutable.
Installation
npm install
npm run buildConfiguration client
Claude Code
claude mcp add codex -- node C:/chemin/vers/mcp-server-codex/dist/index.jsClaude Desktop / configuration JSON générique
{
"mcpServers": {
"codex": {
"command": "node",
"args": ["C:/chemin/vers/mcp-server-codex/dist/index.js"],
"env": {
"CODEX_MCP_ALLOWED_ROOTS": "C:/projets/mon-app",
"CODEX_MCP_DEFAULT_SANDBOX": "workspace-write"
}
}
}
}Le serveur parle stdio. Tous ses diagnostics vont sur stderr : stdout transporte le protocole et écrire dedans corromprait la session.
Ce que le client voit à l'initialisation
Le serveur publie une description — ce qu'il pilote et à quoi ça sert, lu par l'humain qui décide de l'installer — et des instructions lues par l'agent appelant, qui doit arbitrer entre faire le travail lui-même et le déléguer.
Drives the Codex CLI locally, so an agent can hand a coding task to a second autonomous agent instead of doing it turn by turn. Codex is at its best on work that is long, mechanical and verifiable: a refactor across many files, making a failing suite pass, a code review, tracing a bug through an unfamiliar codebase. […]
Les instructions couvrent ce que le schéma des outils ne dit pas : quand déléguer à Codex plutôt que d'éditer soi-même, donner un objectif vérifiable plutôt qu'une procédure, le fait qu'un dépassement de délai rende un job_id au lieu d'échouer, la reprise par thread_id, les deux garde-fous qui rejettent un appel avant tout lancement, et le pont.
Le pont publie les siennes, lues par Codex : à quel moment poser une question vaut mieux que deviner, et qu'une question expirée n'est pas une impasse.
Variables d'environnement
Variable | Défaut | Rôle |
|
| Chemin ou nom du binaire Codex. |
|
| Racine Codex : sessions et images générées y sont lues. |
| cwd du serveur | Répertoires autorisés, séparés par |
|
| À |
|
| Sandbox par défaut : |
|
| Attente avant bascule en arrière-plan. |
|
| Taille du tampon d'événements par job. |
|
| Durée de consultation d'un job terminé. |
|
| Boîte aux lettres partagée avec les processus pont. |
|
| Attente d'un |
| Node courant | Exécutable que Codex lance pour le pont. |
|
| Script du pont passé à cet exécutable. |
| (aucun) | Fichier JSON de presets d'image nommés. C'est par là qu'une instance se spécialise. |
Une valeur invalide fait échouer le démarrage (code 78) plutôt que de retomber silencieusement sur un défaut : une faute de frappe ne doit pas devenir une politique de sécurité différente de celle demandée.
Exécution hybride
Un run Codex dure de quelques secondes à plusieurs dizaines de minutes, alors que les clients MCP coupent les appels d'outil bien avant. Chaque outil d'exécution fait donc la course contre son propre timeout_seconds :
il finit à temps → résultat complet en un aller-retour ;
le délai expire → le processus continue, l'appel rend un
job_idimmédiatement.
Le délai n'annule jamais le run : perdre dix minutes de travail du modèle à cause d'une échéance arbitraire du client est précisément ce que ce design évite.
codex_exec { prompt: "…", timeout_seconds: 60 }
└─ dépassement → { job_id: "job-3-a1b2", mode: "background", thread_id: "…" }
├─ codex_job_status { job_id }
├─ codex_job_logs { job_id, since: 42 } ← pagination par curseur
└─ codex_job_cancel { job_id } ← SIGTERM puis SIGKILLLes jobs vivent le temps de la session MCP.
Outils
Outil | Rôle |
| Nouvelle session Codex sur un prompt. Rend le message final, les commandes exécutées et un |
| Reprend une session avec tout son historique ( |
| Bifurque une session existante, l'originale reste intacte. |
| Revue de code : |
| Applique le dernier diff d'une tâche Codex ( |
| Liste les sessions enregistrées. Lecture disque, aucun processus lancé. |
| Génère une image et l'écrit sur disque. |
| État d'un run passé en arrière-plan. |
| Événements JSONL paginés d'un job, filtrables par type. |
| Arrête un run en cours. |
| Relève ce que Codex a envoyé : questions bloquantes, constats, alertes. |
| Répond à une question de Codex ; le run en attente repart aussitôt. |
| Envoie à Codex un message qu'il n'a pas demandé — correction, changement de cap, arrêt. |
| Détail complet des presets d'image de l'instance, et le fichier d'où ils viennent. |
| Relit le fichier de presets depuis le disque, sans redémarrage. |
| Définit ou remplace un preset, en mémoire et dans le fichier. |
| Modifie certains champs d'un preset ; |
| Supprime un preset de l'instance et du fichier. |
Les outils d'exécution acceptent en commun : cwd, model, sandbox, images, config, enable, disable, output_schema, worktree, ephemeral, skip_git_repo_check, timeout_seconds.
Génération d'images
{
"prompt": "un robot bleu, style plat minimaliste",
"output_path": "assets/robot.png",
"use_case": "logo-brand",
"size": "1024x1024",
"transparent": true,
"constraints": "pas de texte, pas de watermark"
}Spécialiser une instance : les presets
Le paquet ne connaît aucun sujet en particulier. C'est le déploiement qui le spécialise : une instance pointe CODEX_MCP_IMAGE_PRESETS vers un fichier JSON décrivant les sujets que son équipe dessine en boucle — une mascotte, un produit, un style maison — et les appelants nomment un preset au lieu de tout redécrire à chaque fois.
// ~/mon-equipe/presets.json — hors du dépôt, propre à l'instance
{
"mascotte": {
"subject": "personnage en costume, corps ovoïde vert, antennes en feuille",
"style": "pixel art 16-bit, palette limitée, contours nets",
"constraints": "pas de watermark, fond simple",
"use_case": "stylized-concept",
"reference_images": ["./ref/mascotte.jpg"]
}
}{ "preset": "mascotte", "prompt": "de profil, qui salue", "output_path": "sprites/salut.png" }Le preset fournit des défauts, jamais un verrou : tout champ donné à l'appel gagne, si bien qu'une image peut s'écarter du style maison sans le redéfinir.
Trois choix qui méritent d'être dits :
L'instance annonce ses presets. Leurs noms sont ajoutés à la description de
codex_generate_image, sinon l'agent appelant n'aurait aucun moyen d'apprendre qu'ils existent.Les chemins relatifs se résolvent depuis le fichier de presets, pas depuis le répertoire courant du serveur : le preset et son image de référence voyagent ensemble, alors que le répertoire de lancement est accidentel.
Un fichier illisible, un JSON invalide ou un champ inconnu font échouer le démarrage. Un preset silencieusement ignoré produirait des images génériques qui ressemblent à un raté du modèle, et personne n'irait regarder la configuration.
Les images de référence d'un preset passent par l'allowlist comme les autres : être de la configuration ne vaut pas dérogation.
Modifier les presets sans redémarrer
Le fichier est lu au démarrage, mais il n'y est pas figé. Trois outils couvrent la boucle de celui qui affine un sujet :
{ "name": "mascotte", "subject": "…", "style": "pixel art 16-bit" } // codex_preset_set
{} // codex_preset_reload
{} // codex_preset_listLes outils MCP sont le chemin normal, et c'est ce qui garantit le format : use_case est validé contre la taxonomie, size contre une dimension en pixels. Un ui_mockup au lieu de ui-mockup était auparavant accepté et partait tel quel dans le prompt, dégradant le rendu sans que rien ne le signale. La même validation vit dans le chargeur, donc une édition du fichier à la main ne contourne pas la garantie.
codex_preset_set définit ou remplace intégralement — pour qu'un champ puisse être retiré. codex_preset_update change certains champs et laisse les autres, null effaçant un champ : c'est ce qu'on veut pour ajuster un style sans redire un long sujet, puisqu'un sujet redit est un sujet qui dérive. codex_preset_delete supprime, et échoue sur un nom inconnu plutôt que de faire passer une faute de frappe pour un nettoyage réussi. codex_preset_reload sert quand le fichier a été édité en dehors du serveur.
Deux garanties tiennent des deux côtés :
Un changement refusé ne dégrade rien. La validation porte sur l'ensemble avant de remplacer quoi que ce soit : un JSON cassé ou un champ inconnu laisse en place les presets qui marchaient, et n'écrit pas dans le fichier.
L'annonce suit. La description de
codex_generate_imagenomme les presets, et elle est construite à l'enregistrement de l'outil ; chaque changement la republie et émetnotifications/tools/list_changed. Sans ça, l'instance connaîtrait un preset que l'agent appelant n'a aucun moyen de découvrir.
Ce qui n'a pas été assoupli : un fichier de presets absent ou illisible arrête toujours le serveur au démarrage. Un chemin qui n'existe pas est presque toujours une faute de frappe, et démarrer sans presets est le mode d'échec que tout ceci existe pour empêcher. Pour laisser l'agent remplir le fichier, créez-le avec {} — c'est un acte explicite.
L'outil renvoie le chemin du fichier, pas les octets : un PNG de 850 Ko pèse ~1,1 Mo en base64 et saturerait le contexte de l'agent appelant.
Codex n'expose aucun service de génération d'images : le protocole app-server contient ImageGenerationThreadItem comme type d'événement mais aucune méthode RPC image/*, et il n'existe pas de sous-commande codex image. Le seul accès est agentique — le modèle décide d'appeler son outil interne image_gen, guidé par la skill système imagegen. Ce serveur en tire deux conséquences :
Le prompt est composé, pas transmis tel quel. La skill attend une spécification étiquetée (
Use case:,Primary request:,Constraints:…) ; lui donner du texte brut dégrade nettement le résultat.Le fichier doit être retrouvé.
image_genn'émet aucun item JSONL : le flux d'événements ne dit jamais où l'image a atterri. Le serveur vérifie doncoutput_path, puis se rabat sur$CODEX_HOME/generated_images/<thread_id>/et y copie le fichier le plus récent. Si les deux échouent, il le dit explicitement plutôt que de renvoyer un chemin fantôme.
Messagerie bidirectionnelle
Le pont est toujours actif : chaque run lancé par ce serveur est démarré avec un serveur MCP supplémentaire, claude_bridge, que Codex lance lui-même. Du point de vue de Codex, l'agent Claude qui le supervise est simplement trois outils de plus.
agent Claude boîte aux lettres run codex exec
──────────── ───────────────── ──────────────
codex_inbox ──── lit ────▶ un fichier JSON ◀── écrit ── ask_claude (bloque)
codex_reply ─── écrit ───▶ par message ─── lit ───▶ check_claude
codex_tell ─── écrit ───▶ (écriture ─── lit ───▶ check_claude
atomique)Côté Codex (claude_bridge) :
Outil | Rôle |
| Pose une question et attend la réponse. Au-delà du délai, rend un |
| Envoie un message sans attendre de réponse. |
| Relève ce que Claude a envoyé depuis le dernier passage, y compris de sa propre initiative, et collecte la réponse à un |
Ce qui rend l'asynchrone supportable des deux côtés. Claude a un rythme naturel — ses tours d'outils — donc de son côté rien ne bloque : codex_inbox rend ce qui est arrivé et coûte presque rien. Codex, lui, ne peut pas reprendre un tour plus tard sans le perdre : c'est donc de ce côté que l'attente est faite, avec une échéance et un identifiant de repli. Une question qui expire n'est jamais perdue.
Attribution des runs. Un superviseur peut piloter plusieurs runs simultanément. Chaque message porte le job_id du run qui l'a produit, et codex_reply renvoie la réponse au run qui a posé la question. Un codex_tell sans job_id est une diffusion : tous les runs le voient — ce qu'on veut précisément pour un « arrête tout ».
Le pont ne desserre pas le bac à sable. Il s'attache par des surcharges -c portées par la ligne de commande du run, sans toucher à $CODEX_HOME ni à un autre run. Appeler un outil MCP depuis un run sous bac à sable demande normalement une approbation que personne n'est là pour donner ; le pont résout ça par approvals_reviewer="auto_review" plus une politique granulaire qui autorise les seules sollicitations MCP :
approval_policy={granular={mcp_elicitations=true,rules=false,sandbox_approval=false}}Vérifié de bout en bout contre le vrai CLI : sous -s read-only, l'appel ask_claude aboutit et reçoit sa réponse, tandis que l'écriture d'un fichier est refusée (writing is blocked by read-only sandbox). L'élargissement du bac à sable aurait fonctionné aussi, et aurait été le mauvais choix.
Sécurité
L'installation d'un serveur MCP donne à un agent la capacité d'exécuter du code sur votre machine. Les défauts sont donc restrictifs :
Allowlist de répertoires. Tout
cwd,add_dir,images,output_schemaetoutput_pathest résolu en chemin réel — liens symboliques compris — puis vérifié comme descendant d'une racine autorisée. Un chemin refusé l'est avant tout lancement de processus : un appel rejeté n'a aucun effet de bord.Sandbox par défaut
workspace-write, approbations surnever(aucun humain n'est là pour répondre ; un refus revient au modèle comme un échec exploitable au lieu de bloquer le run).danger-full-accesset--dangerously-bypass-approvals-and-sandboxsont refusés saufCODEX_MCP_ALLOW_DANGEROUS=1.
Le garde-fou de chemins ne protège pas contre un Codex lancé en danger-full-access : ce mode retire les limites côté Codex lui-même.
Notes d'implémentation
Trois comportements de Codex 0.154, vérifiés empiriquement, façonnent le code :
Seul
codex execaccepte-s/--sandbox,-C/--cd,--add-diret-p/--profile.exec resume,exec forketexec reviewne les ont pas : le sandbox y passe par-c sandbox_mode="…".Le
codex reviewde premier niveau n'a pas--json— seulcodex exec reviewl'a. Toutes les revues passent donc parexec review.Codex lit stdin dès qu'il n'est pas sur un TTY (« Reading additional input from stdin… »). Le prompt est toujours passé via
-sur stdin, puis stdin est refermé. Cela contourne aussi la limite de 8191 caractères de la ligne de commande Windows et tout l'échappement de quotes.
codex resume sans identifiant ouvre un sélecteur TUI, impilotable en MCP : codex_list_sessions lit donc directement $CODEX_HOME/sessions/**/rollout-*.jsonl (source de vérité) et enrichit avec $CODEX_HOME/session_index.jsonl, qui ne contient que les threads nommés. Seule la première ligne de chaque rollout est lue — ces fichiers atteignent couramment des dizaines de méga-octets.
Le parseur JSONL est délibérément tolérant : un item.type inconnu est conservé tel quel plutôt que rejeté, pour qu'une montée de version de Codex dégrade le résumé au lieu de casser le serveur.
CI/CD et publication
Deux workflows GitHub Actions, sans secret à configurer : le GITHUB_TOKEN fourni automatiquement suffit.
ci.yml — à chaque push et pull request sur main
Matrice Node 22 et 24 × Ubuntu et Windows : typecheck, tests, build. Windows n'est pas du zèle — l'allowlist de chemins, la gestion des lettres de lecteur et le contournement de la limite de 8191 caractères sont des comportements spécifiquement Windows.
Un job supplémentaire vérifie le plancher d'exécution : package.json annonce node >= 22.0, ce job construit avec une chaîne récente puis charge le dist/ sous Node 22.0. Les tests ne peuvent pas y tourner (le type stripping exige 22.18), mais la promesse est prouvée au lieu d'être supposée.
release.yml — sur un tag v*
npm version patch # ou minor / major : met à jour package.json et crée le tag
git push --follow-tagsLe workflow rejoue typecheck, tests et build sur l'arbre exact qui va être publié — CI prouve qu'un commit est sain, la release prouve que le tag l'est —, puis publie sur npm.pkg.github.com et crée la GitHub Release avec des notes générées. Un workflow_dispatch permet de rejouer une release à partir d'un tag existant.
Garde-fou de version : si le tag et package.json divergent, la publication échoue avant le npm publish. Une version npm ne pouvant jamais être republiée, une release mal étiquetée serait définitive.
Scope dérivé à la publication. GitHub Packages n'accepte qu'un paquet scopé au compte propriétaire. Plutôt que de figer un owner dans le dépôt — ce qui casserait tout fork —, scripts/github-package.mjs réécrit le nom en @owner/mcp-server-codex au moment de publier, à partir de github.repository_owner, en le passant en minuscules (GitHub conserve la casse des comptes, npm la refuse).
Installer depuis GitHub Packages
Le registre GitHub exige une authentification, même en lecture. Dans le .npmrc du projet consommateur :
@owner:registry=https://npm.pkg.github.com
//npm.pkg.github.com/:_authToken=${GITHUB_TOKEN}avec un token portant le scope read:packages, puis :
npm install @owner/mcp-server-codexDéveloppement
npm test # 194 tests, sans lancer Codex ni consommer de tokens
npm run typecheck
npm run buildLe développement demande Node ≥ 22.18, première version où le type stripping est actif sans drapeau : la suite exécute les .ts directement. Le paquet publié, lui, n'est que du JavaScript compilé et tourne dès Node 22.0 — plancher imposé par execa, qui déclare node >=22 et utilise Set.prototype.union.
Node exécute TypeScript nativement : ni tsx ni ts-node.
L'architecture tient à une couture : toute interaction avec le système passe par CodexRunner (src/codex/runner.ts). Les tests unitaires injectent soit un faux binaire Codex scriptable (test/helpers/fake-codex.mjs, qui exerce le vrai chemin spawn/stdin/streaming/annulation), soit un runner stub piloté à la main pour les scénarios de timeout et d'annulation.
src/
index.ts binaire, transport stdio
runtime.ts racine de composition
server.ts enregistrement MCP des outils
schemas.ts schémas zod (les descriptions sont lues par l'agent appelant)
config.ts environnement et garde-fous
codex/argv.ts pur : options → argv
codex/events.ts pur : JSONL → événements typés → résumé
codex/runner.ts seule couture avec le système
codex/sessions.ts lecture des sessions sur disque
jobs/store.ts registre, tampon circulaire, TTL
jobs/hybrid.ts course run / timeout
security/paths.ts allowlist de racines
tools/ un fichier par outilLicence
MIT
Available Tools
18 toolscodex_applyApply a Codex diffADestructive
Apply the latest diff produced by a Codex task to the local working tree, as a git apply. Modifies files on disk.
| Name | Required | Description | Default |
|---|---|---|---|
| cwd | No | Repository to apply into. Must sit inside the server allowlist. | |
| task_id | Yes | Codex task id whose latest diff should be applied with git apply. | |
| timeout_seconds | No | Inline wait before backgrounding. 0 returns immediately. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already carry destructiveHint=true and readOnlyHint=false, so the description's 'Modifies files on disk' mostly echoes structured metadata. It does add useful operational context ('as a git apply' and 'local working tree'), but it does not disclose possible failure modes like apply conflicts, dirty working trees, or whether changes are staged.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two short sentences with no filler. The core action and target are front-loaded, and the side-effect note is kept to one clause. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutating tool with no output schema, the description gives the core operation but is thin on surrounding context. It does not mention what happens after apply, how failures or timeouts are surfaced, whether a missing diff is an error, or how an agent should verify the result. Structured fields cover parameters and safety, but an agent still lacks enough operational context to call this confidently.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all parameters (cwd, task_id, timeout_seconds) are already documented in the input schema. The description does not add extra meaning beyond 'latest diff' aligning with task_id, so it stays at the baseline rather than adding new parameter-level insight.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific action ('Apply'), a specific resource ('the latest diff produced by a Codex task'), and a specific target ('the local working tree... as a git apply'). It clearly distinguishes this from sibling tools like codex_exec or codex_review by focusing on applying an existing diff rather than creating or reviewing one.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives context by indicating this applies a task's latest diff to the local tree, but it never explicitly says when to use this tool versus alternatives. With siblings like codex_exec, codex_resume, and codex_review available, an agent gets no direct guidance on sequencing, such as 'use after codex_exec' or 'consider codex_review before applying.'
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
codex_execRun CodexADestructive
Start a new Codex agent session against a prompt. Codex can read and edit files and run commands in the sandbox. Returns the final message, the commands it ran and a thread_id you can pass to codex_resume. Long runs move to the background and return a job_id.
| Name | Required | Description | Default |
|---|---|---|---|
| cwd | No | Working directory for the run. Must sit inside the server allowlist. Defaults to the first allowed root. | |
| model | No | Model slug, e.g. "gpt-5.5". Defaults to the Codex configuration. | |
| config | No | Raw Codex config overrides as key=value, TOML-parsed, e.g. 'model_reasoning_effort="high"'. | |
| enable | No | Codex feature flags to enable for this run. | |
| images | No | Image files to attach to the prompt. Each must be inside the allowlist. | |
| prompt | Yes | Instructions for the Codex agent. Sent over stdin, so length is unconstrained. | |
| add_dir | No | Extra directories Codex may write to, beyond the working directory. Each must be inside the allowlist. | |
| disable | No | Codex feature flags to disable for this run. | |
| profile | No | Codex config profile to layer on top of the base configuration. | |
| sandbox | No | Sandbox for commands Codex runs. "read-only" forbids writes, "workspace-write" (default) allows writes inside the workspace, "danger-full-access" removes all limits and is refused unless the server was started with CODEX_MCP_ALLOW_DANGEROUS=1. | |
| worktree | No | Run in a fresh managed git worktree instead of the working directory. | |
| ephemeral | No | Do not persist the session to disk. It cannot be resumed afterwards. | |
| output_schema | No | Path to a JSON Schema file constraining the shape of the agent final response. | |
| timeout_seconds | No | How long to wait inline before handing back a job_id and continuing in the background. 0 means return immediately. Defaults to the server setting (120s). | |
| skip_git_repo_check | No | Allow running outside a git repository. | |
| dangerously_bypass_approvals_and_sandbox | No | Remove every approval and sandbox check. Refused unless CODEX_MCP_ALLOW_DANGEROUS=1. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotations already mark the tool as mutating, open-world, and destructive, so the bar is lower. The description adds useful behavioral context by explaining that Codex can read and edit files and run commands in the sandbox, and that long runs move to the background and return a job_id. It also discloses the return surface (final message, commands, thread_id), which goes beyond what annotations provide.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded: it opens with the core purpose, then covers capabilities, return values, and backgrounding behavior in just a few sentences. Every clause earns its place and there is no filler or restating of the tool name.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex 16-parameter tool with no output schema, the description supplies the essential missing pieces: what the agent can do, what it returns, and how continuation and backgrounding work. Sandbox modes, timeouts, and safety flags are left to the fully covered schema, which is a reasonable division. A short explicit caution about destructive potential would push this to 5, but the destructiveHint annotation already covers that.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all 16 parameters and the baseline is 3. The description adds little parameter-level meaning beyond framing the prompt as the input to the agent and mentioning the thread_id/job_id outputs. It does not need to compensate, but it also does not enrich individual parameter understanding.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action ('Start a new Codex agent session') against a clearly identified resource ('a prompt'), and immediately distinguishes the tool from siblings by naming its two output tokens: thread_id (for codex_resume) and job_id (for background runs). This gives an agent a clear mental model of codex_exec as the entry-point tool rather than resume, review, or apply.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly implies when to use the tool: for a brand-new session against a prompt. It also points to the follow-up tool by saying the returned thread_id can be passed to codex_resume, which helps an agent understand the workflow. It does not explicitly state when not to use it or name excluded siblings, so it falls just short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
codex_forkFork a Codex sessionADestructive
Branch an existing Codex session into a new one, leaving the original untouched. Useful for trying a different approach from a shared starting point.
| Name | Required | Description | Default |
|---|---|---|---|
| cwd | No | Working directory for the run. Must sit inside the server allowlist. Defaults to the first allowed root. | |
| model | No | Model slug, e.g. "gpt-5.5". Defaults to the Codex configuration. | |
| config | No | Raw Codex config overrides as key=value, TOML-parsed, e.g. 'model_reasoning_effort="high"'. | |
| enable | No | Codex feature flags to enable for this run. | |
| images | No | Image files to attach to the prompt. Each must be inside the allowlist. | |
| prompt | No | Message to send in the forked session. | |
| disable | No | Codex feature flags to disable for this run. | |
| sandbox | No | Sandbox for commands Codex runs. "read-only" forbids writes, "workspace-write" (default) allows writes inside the workspace, "danger-full-access" removes all limits and is refused unless the server was started with CODEX_MCP_ALLOW_DANGEROUS=1. | |
| worktree | No | Run in a fresh managed git worktree instead of the working directory. | |
| ephemeral | No | Do not persist the session to disk. It cannot be resumed afterwards. | |
| thread_id | No | Alias for session_id, matching the thread_id returned by codex_exec. | |
| session_id | No | Session UUID or thread name to branch from. The original is left untouched. | |
| output_schema | No | Path to a JSON Schema file constraining the shape of the agent final response. | |
| timeout_seconds | No | How long to wait inline before handing back a job_id and continuing in the background. 0 means return immediately. Defaults to the server setting (120s). | |
| skip_git_repo_check | No | Allow running outside a git repository. | |
| dangerously_bypass_approvals_and_sandbox | No | Remove every approval and sandbox check. Refused unless CODEX_MCP_ALLOW_DANGEROUS=1. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description usefully adds the key behavioral detail that the original session is left untouched, which is valuable and goes beyond the annotations. It does not disclose approval requirements, background execution, or the lifecycle of the new session, though the annotations already signal a non-read-only, potentially destructive operation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences with no filler. The primary action is stated immediately, and the important safety trait 'leaving the original untouched' is front-loaded. Every word earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex 16-parameter mutation tool with no output schema and destructive annotations, the description is thin. It does not explain what is returned, how the forked session runs, or how the tool relates to codex_exec and codex_resume, though the rich schema and sibling list partially compensate.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the input schema fully documents all 16 parameters. The tool description adds no parameter-level meaning beyond what the schema already provides, so the baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states a specific action and resource: branch an existing Codex session into a new one while leaving the original untouched. This distinguishes it from related operations like resume or exec, though it never names or contrasts with any sibling tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Useful for trying a different approach from a shared starting point' gives a clear use case but no explicit when-to-use versus alternatives. The agent must infer that codex_fork is the right choice over codex_exec or codex_resume from context rather than being told.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
codex_generate_imageGenerate an image with CodexADestructive
Generate an image using the image_gen tool built into Codex, and save it to output_path. No API key is needed: it runs through your Codex session. Returns the path of the file written, not the image bytes.
| Name | Required | Description | Default |
|---|---|---|---|
| cwd | No | Working directory for the run. Must sit inside the server allowlist. Defaults to the first allowed root. | |
| size | No | Requested size as WIDTHxHEIGHT, e.g. "1024x1024", "1536x1024", "3840x2160". | |
| model | No | Model slug, e.g. "gpt-5.5". Defaults to the Codex configuration. | |
| style | No | Style or medium, e.g. "flat minimal vector", "studio product photography". | |
| config | No | Raw Codex config overrides as key=value, TOML-parsed, e.g. 'model_reasoning_effort="high"'. | |
| enable | No | Codex feature flags to enable for this run. | |
| images | No | Image files to attach to the prompt. Each must be inside the allowlist. | |
| preset | No | Named preset configured on this server instance. It supplies the subject, style, constraints and reference art, so the prompt only has to say what differs. The tool description lists the presets this instance knows; omit it on an instance that has none. | |
| prompt | Yes | What the image should show. Plain language; the server wraps it in the spec Codex expects. With a preset, describe only the variation — the preset already carries the subject. | |
| disable | No | Codex feature flags to disable for this run. | |
| sandbox | No | Sandbox for commands Codex runs. "read-only" forbids writes, "workspace-write" (default) allows writes inside the workspace, "danger-full-access" removes all limits and is refused unless the server was started with CODEX_MCP_ALLOW_DANGEROUS=1. | |
| use_case | No | Taxonomy slug steering the style. One of: product-mockup, ui-mockup, logo-brand, illustration-story, infographic-diagram, photorealistic-natural, stylized-concept, ads-marketing. | |
| worktree | No | Run in a fresh managed git worktree instead of the working directory. | |
| ephemeral | No | Do not persist the session to disk. It cannot be resumed afterwards. | |
| constraints | No | Things the image must avoid or preserve, e.g. "no text, no watermark". | |
| output_path | Yes | Where to write the image, e.g. "assets/hero.png". Must sit inside the server allowlist. | |
| transparent | No | Ask for a genuinely transparent background and preserve the alpha channel. | |
| output_schema | No | Path to a JSON Schema file constraining the shape of the agent final response. | |
| timeout_seconds | No | How long to wait inline before handing back a job_id and continuing in the background. 0 means return immediately. Defaults to the server setting (120s). | |
| reference_images | No | Reference images for style or composition. Each must be inside the allowlist. | |
| skip_git_repo_check | No | Allow running outside a git repository. | |
| dangerously_bypass_approvals_and_sandbox | No | Remove every approval and sandbox check. Refused unless CODEX_MCP_ALLOW_DANGEROUS=1. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already carry readOnlyHint=false and destructiveHint=true, so the write/destructive profile is covered. The description adds genuine value beyond annotations by disclosing the return format (path, not image bytes) and the no-API-key requirement. No contradiction with annotations — the file-writing behavior aligns with the non-read-only hint.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences with zero waste. The core action and output path are front-loaded, followed by the auth note and return-format clarification. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex 22-parameter tool with no output schema, the description clarifies the return value and the execution mechanism, which is essential. The rich per-parameter schema descriptions cover the remaining configuration surface. Slightly more safety/limitation context would help given the destructive annotation, but the description is largely adequate.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so all 22 parameters are documented in the input schema, setting the baseline at 3. The description adds marginal semantic value by tying output_path to the save behavior and clarifying the return, but it does not elaborate on individual parameters beyond what the schema already provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Generate an image'), a concrete resource, and the output behavior (saves to output_path). It also clarifies the mechanism (Codex's built-in image_gen) and the return type (path, not bytes), making it clearly distinct from siblings like codex_exec, codex_apply, and the job/session tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives useful context about how the tool runs (through the Codex session, no API key needed), but offers no explicit when-to-use vs. alternatives guidance and no exclusions. The purpose is self-evident relative to siblings, but the description relies on inference rather than stating usage boundaries.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
codex_inboxRead messages from CodexARead-only
Collect anything a running Codex agent has sent — a question it is blocked on, a finding, a warning. Cheap and non-blocking, so it is worth calling between other tools while a run is in flight. Pass the next_cursor from the previous call to see only what is new; awaiting_answer counts the questions still waiting on codex_reply.
| Name | Required | Description | Default |
|---|---|---|---|
| job_id | No | Only messages from this run. Omit to watch every run at once. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already signal read-only/non-destructive behavior; the description adds non-blocking/cost characteristics, cursor-based incremental reads, and the meaning of awaiting_answer. It could go slightly deeper on message retention or consumption behavior, but the key traits are disclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences lead with purpose and examples, then move to usage and behavioral guidance. There is no repetition of schema/annotations and no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read-only tool with one optional parameter, it covers purpose, cost, non-blocking behavior, pagination, and a key return field. It does not fully describe the message envelope shape or when to use codex_job_logs instead, but the overall guidance is sufficient for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema already documents job_id with 100% coverage, so the description does not need to compensate. It adds no parameter-level detail, and the baseline of 3 applies because the schema carries the semantic load.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses concrete verbs ('Collect'/'Read') and a specific resource ('messages from Codex'), with examples of content types (questions, findings, warnings). This clearly distinguishes it from log/status/reply sibling tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives clear context: the tool is cheap, non-blocking, and worth calling between other tools while a run is in flight. It also provides pagination guidance via next_cursor, though it does not explicitly contrast with siblings or state when not to use it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
codex_job_cancelCancel a background Codex jobADestructive
Stop a run that is still going: SIGTERM first, then SIGKILL after a grace period. Reports honestly whether this call is what cancelled it.
| Name | Required | Description | Default |
|---|---|---|---|
| job_id | Yes | Job id to cancel. Sends SIGTERM, then SIGKILL after a grace period. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description goes well beyond the annotations by revealing the exact kill mechanism (SIGTERM, then SIGKILL after a grace period) and the return semantics ('Reports honestly whether this call is what cancelled it'). This is valuable behavioral context that the destructiveHint annotation alone does not provide, and nothing contradicts the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences convey purpose, mechanism, and outcome without waste. The core action is front-loaded, and each sentence earns its place by adding essential operational detail.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter destructive tool with no output schema, this description is complete: it explains what is cancelled, how (signal sequence), when the tool is appropriate, and what result to expect. An agent has enough context to call it correctly and interpret the outcome.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%: the job_id parameter already has a full description including signal behavior. The tool description repeats this rather than adding new parameter-level meaning, so the baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states an unambiguous action ('Stop a run that is still going') with a specific resource and even details the shutdown sequence (SIGTERM then SIGKILL). It clearly differentiates from siblings like codex_job_status and codex_job_logs by focusing on cancellation rather than inspection.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'Stop a run that is still going' conveys the primary use case: cancel an active background job. It does not explicitly name alternatives or say when not to use it, but the context is clear enough for an agent to route to this tool instead of job_status or job_logs.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
codex_job_logsRead background Codex job eventsARead-only
Page through the JSONL events of a background run. Pass the returned next_cursor back as "since" to read only what is new. Filter with "types" to cut the noise.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Maximum events per page. Defaults to 200. | |
| since | No | Cursor from a previous call. Omit to read from the start. | |
| types | No | Keep only these event types, e.g. ["item.completed", "turn.completed"]. | |
| job_id | Yes | Job id returned by a tool call that went to the background. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and destructiveHint=falsecars, so safety is covered. The description adds meaningful behavioral detail beyond that: it explains cursor-based pagination semantics and the incremental-read pattern ('Pass the returned next_cursor back as "since"'), which is valuable for correct use.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, each earning its place: the first states the core purpose, the second teaches cursor resumption, and the third explains filtering. No filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only paginated log tool, the description covers what the agent needs: how to page, how to continue from a cursor, and how to filter. The annotations cover safety, and the schema covers parameters. The absence of an output schema is mitigated by mentioning JSONL events and next_cursor.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3. The description adds value by clarifying that 'since' is a cursor returned by a previous call and that 'types' can reduce noise, going slightly beyond the schema's bare definitions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action and resource: 'Page through the JSONL events of a background run.' This clearly positions the tool as an event-log reader and distinguishes it from siblings like codex_job_status and codex_job_cancel, which cover status and cancellation rather than raw event streams.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides concrete usage context: paginate with next_cursor, resume with 'since' for new data, and reduce volume via 'types'. It does not explicitly name alternatives or state when not to use the tool, but the event-log framing makes its intended role clear against the sibling tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
codex_job_statusCheck a background Codex jobARead-only
Report the state of a run that moved to the background: whether it is still running, its exit code, how long it has taken and how many events it has produced.
| Name | Required | Description | Default |
|---|---|---|---|
| job_id | Yes | Job id returned by a tool call that went to the background. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish read-only, non-destructive behavior. The description adds useful behavioral detail about what the status check returns: running state, exit code, elapsed time, and event count. It does not contradict annotations and supplements them with meaningful output expectations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single focused sentence that front-loads the core action and then lists the returned status details. Every clause adds useful information, with no redundant or filler text.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's simplicity, one required parameter, and no output schema, the description adequately covers input, behavior, and return contents. It tells the agent what data the call will provide, which is sufficient for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema covers the single parameter job_id with a clear description ('Job id returned by a tool call that went to the background'), and schema coverage is 100%. The tool description does not add additional parameter-level meaning beyond this, so the baseline score of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: 'Report the state of a run that moved to the background' and lists the concrete details it reports (still running, exit code, elapsed time, event count). This distinguishes it from sibling tools like codex_job_logs and codex_job_cancel, which serve different purposes.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It provides clear context by explaining that the tool applies to backgrounded runs rather than active or cancellable jobs. It does not explicitly name alternatives or exclusion conditions, but the domain context from the sibling tools makes the intended use reasonably clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
codex_list_sessionsList Codex sessionsARead-only
List recorded Codex sessions with their id, name, working directory and last update, read straight from disk. Use it to find a session_id for codex_resume or codex_fork.
| Name | Required | Description | Default |
|---|---|---|---|
| cwd | No | Only sessions whose working directory matches this path. | |
| limit | No | Maximum number of sessions to return. Defaults to 20. | |
| query | No | Case-insensitive substring match on the thread name or id. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and destructiveHint=false. The description adds meaningful behavior beyond annotations by stating sessions are 'read straight from disk,' clarifying the source and implying no remote calls or state changes. This is useful context for an agent deciding whether invocation is safe and appropriate.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two compact sentences, with the core action and return fields front-loaded and the usage guidance in the second sentence. No filler words and no redundancy with the schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple, read-only list tool with three optional and fully documented parametershol, the description covers what the tool returns composites and why to use it. No output schema exists, but the description names the key return fields, making the tool safely callable without needing more context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all three parameters. The description adds little beyond the schema, except that the returned data includes fields like working directory, which maps to the cwd filter. Baseline 3 is appropriate since the schema carries the full parameter burden.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'List recorded Codex sessions' and enumerates the returned fields (id, name, working directory, last update). It clearly differentiates itself from session-consuming tools like codex_resume and codex_fork by positioning itself as the way to discover a session_id.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides explicit usage guidance: 'Use it to find a session_id for codex_resume or codex_fork.' This names the dependent siblings and the exact purpose, leaving no ambiguity about when to select this tool over alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
codex_preset_deleteDelete an image presetADestructive
Remove a preset from this instance and from its presets file. Fails if no preset of that name exists, rather than reporting success for a name that was never there.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | Preset to delete. Fails if no preset of that name exists. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already indicate destructiveHint=true, but the description adds significant behavior beyond that: the preset is removed from a persistent file, and the operation fails rather than falsely succeeding if the name does not exist. This failure mode is critical for an agent to predict outcomes correctly.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two short sentences, front-loaded with the action, and contains no filler. The first sentence states scope and persistence; the second clarifies the error behavior, which earns its place by preventing a common false-success expectation.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter delete tool, the description covers the essential context: what is deleted, the persistence effect, and failure behavior. Annotations already convey destructive intent, so an agent has everything it needs to invoke the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already fully documents the single 'name' parameter. The description repeats the failure condition but does not add any new parameter-level meaning, matching the baseline of 3.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses the specific verb 'Remove' and identifies the resource ('a preset') and the scope ('from this instance and from its presets file'). This distinguishes it from sibling tools like codex_preset_set, codex_preset_update, and codex_preset_list without needing to open them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description makes clear that the tool is for deleting a preset and that the preset must already exist, which is a clear context for use. It does not explicitly name siblings or say when not to use it, but no exclusions are needed for a straightforward delete operation.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
codex_preset_listList image presetsARead-only
Show the image presets this instance is configured with, in full — subject, style, constraints and reference art — plus the file they come from. The description of codex_generate_image only names them; use this to see what a preset actually contains before using or changing it.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and destructiveHint=false, so the safety profile is covered. The description adds that it shows full details and the file source, which is useful output context. No contradictions with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences, front-loaded with the core function ('Show the image presets'), and every phrase adds value. It is concise and well-structured.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter, read-only list tool, the description fully covers what it returns (subject, style, constraints, reference art, file) and when to use it. No output schema exists, but the description provides sufficient context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so schema coverage is trivially 100%. With no parameters to document, the description doesn't need to add parameter semantics, and the baseline of 4 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool 'Show the image presets this instance is configured with', listing specific fields (subject, style, constraints, reference art) and the source file. It explicitly distinguishes from codex_generate_image, which only names presets, making the purpose unambiguous and differentiating it from siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit usage guidance: 'use this to see what a preset actually contains before using or changing it'. It also contrasts with codex_generate_image, telling the agent when to use this tool instead. This is clear and actionable.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
codex_preset_reloadReload image presetsA
Re-read the presets file from disk, picking up edits made outside this server without a restart. If the file has become unusable the call fails and the presets already loaded are kept, so a bad edit never leaves the instance with none.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description adds meaningful behavior beyond the annotations: it discloses that a bad file will cause the call to fail and that currently loaded presets are preserved. This gives the agent important safety information about failure modes, though it doesn't describe the success return value or side effects in detail.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences with no filler. The primary action is front-loaded, and the failure-safety behavior is added in a compact second sentence. Every word contributes value.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter, no-output-schema tool, the description covers the essential context: what the tool does, when it is useful, and what happens on failure. It is complete enough for an agent to invoke it correctly, even if it omits a precise success response description.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so there are no parameter semantics to document. The baseline of 4 applies because no parameter explanation is needed and the description focuses on the operation itself rather than inventing unnecessary parameter details.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('re-read') and resource ('presets file from disk'), making the action immediately clear. It also distinguishes itself from sibling tools like codex_preset_set, codex_preset_update, and codex_preset_list by focusing on reloading from disk rather than mutating or listing presets.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly implies when to use the tool: after presets were edited outside the server, to apply changes without a restart. It does not explicitly name alternatives or exclusions, but the context is strong enough that an agent would not confuse it with the sibling preset tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
codex_preset_setAdd or replace an image presetA
Define a named preset for a subject drawn repeatedly, so later calls name it instead of restating the description. Written to the presets file, so it survives a restart, and usable immediately. Replaces a preset of the same name outright rather than merging, so a field can be removed.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | Name callers will pass as "preset". Replaces an existing preset of the same name. | |
| size | No | Requested size as WIDTHxHEIGHT, e.g. "1024x1024", "1536x1024", "3840x2160". | |
| style | No | Style or medium, e.g. "pixel art 16-bit, limited palette". | |
| subject | No | The recurring subject, prepended to every prompt using this preset. This is the field that earns a preset. | |
| use_case | No | Taxonomy slug steering the style. One of: product-mockup, ui-mockup, logo-brand, illustration-story, infographic-diagram, photorealistic-natural, stylized-concept, ads-marketing. | |
| constraints | No | What images from this preset must avoid, e.g. "no watermark". | |
| transparent | No | Ask for a transparent background by default. | |
| reference_images | No | Reference art for this subject. Each must sit inside the server allowlist. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description adds material behavioral detail beyond the annotations: it persists to the presets file, survives restart, is usable immediately, and replaces same-name presets outright so fields can be removed. This replacement/removal behavior is exactly the kind of side effect an agent needs to know.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences with no filler: purpose first, then persistence and overwrite semantics. It does not repeat schema details and every sentence contributes critical selection or invocation information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an 8-parameter tool with complete schema coverage and no output schema, the description covers purpose, file persistence, immediate usability, and replace-not-merge behavior. An agent has enough information to decide whether to call it and what to expect when it does.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
All eight parameters already have schema descriptions, so the baseline is 3. The description adds value by tying the two central parameters together—'name' is the handle used in later calls, and the 'subject' is the recurring content being preset—clarifying their relationship beyond the individual schema entries.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific action—'Define a named preset for a subject drawn repeatedly'—and names the resource and purpose: to let later calls use the preset name instead of restating the subject. It also distinguishes this from update-style siblings by stating it 'Replaces a preset of the same name outright rather than merging.'
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives a clear use context: use this when a subject is drawn repeatedly and you want a reusable named preset. It also conveys when not to rely on it, since replacement is outright rather than a merge, implying partial updates are not performed here.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
codex_preset_updateModify an image presetA
Change some fields of an existing preset and leave the rest alone. Use this rather than codex_preset_set when adjusting one thing: restating a long subject just to tweak a style is how a subject drifts. Pass null for a field to clear it. Fails if the preset does not exist.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | Preset to modify. It must already exist; use codex_preset_set to define one. | |
| size | No | New default size. Pass null to clear it. | |
| style | No | New style. Pass null to clear it. | |
| subject | No | New subject. Pass null to clear it. | |
| use_case | No | New use case. Pass null to clear it. | |
| constraints | No | New constraints. Pass null to clear them. | |
| transparent | No | New transparency default. Pass null to clear it. | |
| reference_images | No | New reference art. Pass null to clear it. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, destructiveHint=false, so the mutation nature is known. The description adds crucial behavioral context beyond annotations: the null-to-clear semantics and the failure condition when the preset is missing. It does not contradict annotations. While it doesn't mention permissions or side effects, given the annotations, this is a solid disclosure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with no filler. The core purpose and differentiation are front-loaded in the first sentence, followed by the null-clearing and failure notes in the second. Every word earns its place, and the structure is scannable for an agent.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With 8 parameters, all documented in the schema, the description covers the essential decision-making: when to use, null behavior, and failure mode. It doesn't mention the return value, but since there is no output schema, that might be expected. Still, the description is complete enough for correct invocation; a slightly richer mention of what happens after success (e.g., in-memory or persisted) would push it to 5.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so each parameter is documented in the schema. The description adds a global semantic: 'Pass null for a field to clear it' and 'leave the rest alone', which clarifies the behavior of omitted and null fields—valuable beyond what the schema individually states. This elevates it above the baseline for high-coverage schemas.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb ('Change'), the resource ('existing preset'), and the scoping ('some fields... leave the rest alone'). It explicitly differentiates from the sibling codex_preset_set by naming it and describing the exact condition for choosing this tool over it. No ambiguity remains about what this tool does.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives explicit when-to-use guidance: 'Use this rather than codex_preset_set when adjusting one thing', and explains the rationale ('restating a long subject just to tweak a style is how a subject drifts'). It also warns about a failure condition ('Fails if the preset does not exist') and clarifies that omitted fields remain unchanged. This fully directs an agent on when to invoke this tool vs alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
codex_replyAnswer a Codex questionADestructive
Answer a question codex_inbox reported. The Codex run is blocked waiting for exactly this, so answering promptly is what keeps it moving; if it has already timed out the answer is still delivered and collected at its next turn. Fails if no such question exists, rather than posting an answer nobody is waiting for.
| Name | Required | Description | Default |
|---|---|---|---|
| text | Yes | The answer. Codex is blocked waiting for it, so be direct. | |
| message_id | Yes | Id of the Codex message being answered, from codex_inbox. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description discloses useful timing semantics: answering promptly unblocks the run, late answers are still delivered at the next turn, and the call fails rather than posting orphaned answers. This goes beyond the annotations by explaining failure modes and delivery behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three focused sentences, with the core purpose first and each sentence adding necessary runtime context. No repetition of schema details or filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter tool with full schema coverage and annotations declaring side effects, the description covers invocation, timing, and failure behavior completely. No additional return-format details are necessary for correct use.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, and both required parameters already have clear descriptions in the schema. The description reinforces that message_id comes from codex_inbox and that text should be direct, but adds little parameter-level meaning beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States the action and target clearly: 'Answer a question codex_inbox reported.' This is distinct from sibling tools by describing a blocking inbox exchange, not a general codex instruction.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Clearly explains when to use it: a Codex run is blocked waiting for this answer, and it fails if no such question exists, which serves as a when-not. It does not explicitly name an alternative sibling for non-reply communication.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
codex_resumeResume a Codex sessionADestructive
Continue a previous Codex session with its full history, either by session_id or with last=true for the most recent one. Use this instead of codex_exec to follow up on earlier work without re-explaining the context.
| Name | Required | Description | Default |
|---|---|---|---|
| cwd | No | Working directory for the run. Must sit inside the server allowlist. Defaults to the first allowed root. | |
| last | No | Continue the most recent recorded session. Mutually exclusive with "session_id". | |
| model | No | Model slug, e.g. "gpt-5.5". Defaults to the Codex configuration. | |
| config | No | Raw Codex config overrides as key=value, TOML-parsed, e.g. 'model_reasoning_effort="high"'. | |
| enable | No | Codex feature flags to enable for this run. | |
| images | No | Image files to attach to the prompt. Each must be inside the allowlist. | |
| prompt | No | Message to send after resuming. Omit to just replay the session. | |
| disable | No | Codex feature flags to disable for this run. | |
| sandbox | No | Sandbox for commands Codex runs. "read-only" forbids writes, "workspace-write" (default) allows writes inside the workspace, "danger-full-access" removes all limits and is refused unless the server was started with CODEX_MCP_ALLOW_DANGEROUS=1. | |
| worktree | No | Run in a fresh managed git worktree instead of the working directory. | |
| ephemeral | No | Do not persist the session to disk. It cannot be resumed afterwards. | |
| thread_id | No | Alias for session_id, matching the thread_id returned by codex_exec. | |
| session_id | No | Session UUID or thread name to continue. Mutually exclusive with "last". | |
| output_schema | No | Path to a JSON Schema file constraining the shape of the agent final response. | |
| timeout_seconds | No | How long to wait inline before handing back a job_id and continuing in the background. 0 means return immediately. Defaults to the server setting (120s). | |
| skip_git_repo_check | No | Allow running outside a git repository. | |
| dangerously_bypass_approvals_and_sandbox | No | Remove every approval and sandbox check. Refused unless CODEX_MCP_ALLOW_DANGEROUS=1. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description adds meaningful behavioral context by noting the session includes 'full history' and that no re-explanation is needed, which characterizes stateful continuation beyond what annotations provide. Annotations already signal destructive and open-world behavior, so the tool's mutation risk is covered even though the description does not elaborate on side effects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two tight sentences with no filler. The core action and selection mechanism are front-loaded, and the alternative-tool guidance is condensed into the second sentence.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite having 17 parameters and no output schema, the description plus the fully documented schema and annotations provide enough context for correct invocation. The main gap is that it does not describe the return/background-job behavior, but the schema and sibling tools like codex_job_status cover much of that workflow.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3 and the schema already documents all 17 parameters. The description reinforces the role of session_id and last=true, but it does not add semantic details beyond what the parameter descriptions already specify.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly names the verb 'Continue', the resource 'previous Codex session', and the two selection modes (session_id or last=true). It also explicitly differentiates itself from codex_exec, so an agent can quickly tell when to pick this tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It says directly to 'use this instead of codex_exec to follow up on earlier work without re-explaining the context'. This gives a concrete condition for using the tool and names the relevant alternative, leaving little to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
codex_reviewReview code with CodexA
Run a Codex code review over uncommitted changes (default), a diff against a base branch, or a specific commit. Optionally steer it with custom instructions.
| Name | Required | Description | Default |
|---|---|---|---|
| cwd | No | Working directory for the run. Must sit inside the server allowlist. Defaults to the first allowed root. | |
| base | No | Review the diff against this base branch. | |
| model | No | Model slug, e.g. "gpt-5.5". Defaults to the Codex configuration. | |
| title | No | Title shown in the review summary. | |
| commit | No | Review the changes introduced by this commit SHA. | |
| config | No | Raw Codex config overrides as key=value, TOML-parsed, e.g. 'model_reasoning_effort="high"'. | |
| enable | No | Codex feature flags to enable for this run. | |
| images | No | Image files to attach to the prompt. Each must be inside the allowlist. | |
| prompt | No | Custom review instructions, e.g. "focus on race conditions". | |
| disable | No | Codex feature flags to disable for this run. | |
| sandbox | No | Sandbox for commands Codex runs. "read-only" forbids writes, "workspace-write" (default) allows writes inside the workspace, "danger-full-access" removes all limits and is refused unless the server was started with CODEX_MCP_ALLOW_DANGEROUS=1. | |
| worktree | No | Run in a fresh managed git worktree instead of the working directory. | |
| ephemeral | No | Do not persist the session to disk. It cannot be resumed afterwards. | |
| uncommitted | No | Review staged, unstaged and untracked changes. This is the default. | |
| output_schema | No | Path to a JSON Schema file constraining the shape of the agent final response. | |
| timeout_seconds | No | How long to wait inline before handing back a job_id and continuing in the background. 0 means return immediately. Defaults to the server setting (120s). | |
| skip_git_repo_check | No | Allow running outside a git repository. | |
| dangerously_bypass_approvals_and_sandbox | No | Remove every approval and sandbox check. Refused unless CODEX_MCP_ALLOW_DANGEROUS=1. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description does not disclose important behavioral traits: a Codex review can run commands under a configurable sandbox, write to the workspace, and continue in the background after the inline timeout, but none of that appears here. The annotations do signal readOnlyHint=false and openWorldHint=true, but the description itself adds no behavioral context beyond the basic review action.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with no filler; the main purpose and default behavior are front-loaded, and optional customization comes second. This is appropriately concise given that the schema already carries the detailed parameter documentation.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given 18 parameters, no output schema, and asynchronous behavior implied by timeout_seconds and job_id, the short description is thin: it does not tell the agent what to expect back (inline result vs. job_id) or how the review interacts with sandbox and background execution. The rich parameter schema partially compensates, making this minimally viable but with clear gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3 even without parameter info in the description. The description usefully summarizes the main review targets (uncommitted, base, commit) and the prompt parameter, but those map closely to the already-detailed schema fields and add grouping rather than genuinely new semantics.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific action ('Run a Codex code review') and clearly names the input scopes: uncommitted changes (default), a diff against a base branch, or a specific commit. This makes its purpose distinct from siblings like codex_exec and codex_apply even though no sibling is explicitly named.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It clearly explains when to use the tool by describing the three review scopes and marking uncommitted changes as the default, so an agent knows what will happen if no target is specified. It also mentions optional custom instructions. However, it does not give explicit when-not-to-use guidance or name alternatives like codex_exec for broader tasks.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
codex_tellSend Codex a messageADestructive
Send a running Codex agent something it did not ask for — a correction, a change of direction, a stop. It arrives at the next tool turn of that run rather than interrupting it mid-command. Omit job_id to reach every run at once.
| Name | Required | Description | Default |
|---|---|---|---|
| job_id | No | Only tell this run. Omit to reach every run, which is what you want for "stop". | |
| message | Yes | What to tell Codex. It reads this at its next tool turn. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already carry destructive and non-read-only hints. The description adds useful behavioral detail: delivery happens at next tool turn (non-interrupting) and omitting job_id broadcasts to all runs. This exceeds annotation baseline.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with purpose, no filler. Every phrase contributes useful information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter tool with no output schema, the description covers the message content, targeting, and delivery timing. It is complete enough for correct invocation, though it does not address edge cases like missing runs.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema covers 100% of parameters. The description adds semantic nuance by explaining that omitting job_id reaches all runs and is intended for stop, which goes beyond the schema's phrasing.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific action (send a message to a running Codex agent) and enumerates use cases (correction, direction change, stop). It distinguishes itself from siblings by emphasizing unsolicited messages and delivery timing, so an agent can separate it from codex_reply or codex_job_cancel.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives concrete scenarios (correction, direction change, stop) and parameter guidance (omit job_id to reach every run, especially for stop). It does not explicitly name alternative tools, but the context is clear enough for selection.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
9 tool updates
v0.0.7- Changed
codex_generate_image6 fields changed- added
Input schema / properties / presetAdded value: +{ + "description": "Named preset configured on this server instance. It supplies the subject, style, constraints and reference art, so the prompt only has to say what differs. The tool description lists the presets this instance knows; omit it on an instance that has none.", + "type": "string" +} - changed
Input schema / properties / prompt / descriptionPrevious value: -"What the image should show. Plain language; the server wraps it in the spec Codex expects."New value: +"What the image should show. Plain language; the server wraps it in the spec Codex expects. With a preset, describe only the variation — the preset already carries the subject." - changed
Input schema / properties / size / descriptionPrevious value: -"Requested size, e.g. \"1024x1024\", \"1536x1024\", \"3840x2160\"."New value: +"Requested size as WIDTHxHEIGHT, e.g. \"1024x1024\", \"1536x1024\", \"3840x2160\"." - added
Input schema / properties / size / patternAdded value: +"^\\d{2,5}x\\d{2,5}$" - changed
Input schema / properties / use_case / descriptionPrevious value: -"Taxonomy slug steering the style: product-mockup, ui-mockup, logo-brand, illustration-story, infographic-diagram, photorealistic-natural, stylized-concept, ads-marketing."New value: +"Taxonomy slug steering the style. One of: product-mockup, ui-mockup, logo-brand, illustration-story, infographic-diagram, photorealistic-natural, stylized-concept, ads-marketing." - added
Input schema / properties / use_case / enumAdded value: +[ + "product-mockup", + "ui-mockup", + "logo-brand", + "illustration-story", + "infographic-diagram", + "photorealistic-natural", + "stylized-concept", + "ads-marketing" +]
- Added
codex_inbox - Added
codex_preset_delete - Added
codex_preset_list - Added
codex_preset_reload - Added
codex_preset_set - Added
codex_preset_update - Added
codex_reply - Added
codex_tell
10 tool updates
v0.0.4- First observed
codex_apply - First observed
codex_exec - First observed
codex_fork - First observed
codex_generate_image - First observed
codex_job_cancel - First observed
codex_job_logs - First observed
codex_job_status - First observed
codex_list_sessions - First observed
codex_resume - First observed
codex_review
TDQS
Scored across 18 tools
Each tool has a clearly distinct purpose: session lifecycle, job control, agent messaging, and preset management are all cleanly separated. Even the closely related codex_inbox, codex_reply, and codex_tell are easy to tell apart from their descriptions.
All tools share a consistent codex_ prefix and snake_case format, with clear verb or noun segments following it. The naming is predictable enough that an agent can infer the role of an unfamiliar tool from the pattern.
18 tools is slightly above the ideal compact range, but the count is justified by the multiple coherent clusters: sessions, jobs, inbox/reply messaging, presets, and code review. It feels organized rather than bloated.
The core lifecycle is well covered: start, resume, fork, list, monitor, cancel, and communicate with runs, plus full preset CRUD. Minor gaps exist, such as no explicit session deletion or a way to list all background jobs, but they do not block the main workflows.
Maintenance
Related MCP Connectors
A paid remote MCP for OpenAI Codex agent coordination MCP, built to return verdicts, receipts, usage
Remote MCP server for supportsheep: run AI interviews and manage support content for your blog.
Develop, manage, and debug Railway projects, services, and deployments from within agents.
Personal assistant MCP server with search, execute, packages, jobs, secrets, and integrations.
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceEnables MCP clients to spawn and control Codex CLI and Claude Code sessions on the host machine, with session management and filesystem access.4MIT
- AlicenseCqualityDmaintenanceBridges MCP clients with local Codex CLI to execute autonomous coding tasks, manage threads, and inspect history via SQLite state.131,102 npm4Apache 2.0
- AlicenseAqualityBmaintenanceLocal MCP bridge that lets Codex operate local Claude Code sessions, including listing, starting, resuming, forking, prompting, and stopping conversations via the Remote Control CLI.14MIT
- AlicenseNot gradedqualityCmaintenanceEnables Grok Build to orchestrate the local Codex CLI for code reviews, adversarial reviews, task rescue, session transfer, and background job management through MCP tools.1Apache 2.0