Skip to main content
Glama

Delegate a task to Codex

codex_delegate

Delegate coding tasks to the local Codex CLI, picking model and reasoning effort per run. Use it for second opinions, long investigations, or file edits without flooding your own conversation.

Instructions

Delegate a task to the local Codex CLI (OpenAI's coding agent), choosing model and reasoning effort. Use it when the user asks for Codex, or when handing work off clearly serves their request: a second opinion from a different model family, or an investigation that would otherwise flood this conversation. When the user has not named a model, call codex_recommend first and present its suggested model and effort to the user in the same message in which you say you are going to delegate, then pass both explicitly here. That recommendation is advice for an already-authorised delegation, not a replacement for the user's own preference. Everything passed in prompt, context and target_files is sent to OpenAI, and every run spends the user's own Codex usage, so do not delegate what you can answer directly, and tell the user when you delegate. Codex runs read-only unless a different default sandbox is configured. Set sandbox to workspace-write to let it edit files. Codex cannot see this conversation, so pass everything it needs in prompt, context, and target_files.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
modeNoblocking (default) waits and streams progress; background returns a job_id immediately.
modelNoCatalog slug from list_codex_models. If omitted, the configured default is used; with no default configured the call is refused and the recommended model is returned.
promptYesThe task for Codex. Be specific and self-contained: Codex cannot see this conversation.
contextNoBackground Codex needs: prior findings, constraints, relevant excerpts.
sandboxNoSandbox policy. Uses the configured default when omitted; without one, Codex runs read-only.
add_dirsNoAdditional absolute directories that should be writable alongside working_dir.
web_searchNoEnable Codex's API-backed live web-search tool for this run. In a read-only sandbox, shell commands have no network access, so this is the route to current external information. When omitted, Codex's own configured web_search mode applies.
working_dirNoAbsolute path Codex uses as its working root.
auto_approveNoAdds --approve-for-me so Codex auto-approves its own commands. Only applies when sandbox is workspace-write.
target_filesNoPaths Codex should focus on, relative to working_dir.
use_worktreeNoRun in a managed git worktree. Writes outside it remain subject to the sandbox policy and add_dirs.
output_schemaNoA JSON Schema object for this turn's final message, which then comes back as JSON in a delimited block. OpenAI's structured outputs apply: every object needs "additionalProperties": false and every property listed in "required". At most 64 KiB serialised.
timeout_secondsNoWall-clock budget. Defaults to 1800s.
reasoning_effortNoReasoning depth, independent of model choice. Uses the configured default or the model's default when omitted. Clamped to supported levels within the configured ceiling; refused if none qualify.
acceptance_criteriaNoConcrete conditions that must hold for the task to be considered done.
skip_git_repo_checkNoAllow running outside a git repository.
system_instructionsNoPersona or extra rules inherited from the orchestrator, layered on the built-in quality contract.

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed1 schema field changedv0.4.0
    • addedInput schema / properties / output_schema
      Added value: +{
      +  "additionalProperties": {},
      +  "description": "A JSON Schema object for this turn's final message, which then comes back as JSON in a delimited block. OpenAI's structured outputs apply: every object needs \"additionalProperties\": false and every property listed in \"required\". At most 64 KiB serialised.",
      +  "propertyNames": {
      +    "type": "string"
      +  },
      +  "type": "object"
      +}
  2. Changed2 schema fields changedv0.3.0
    • changedInput schema / properties / sandbox / description
      Previous value: -"Sandbox policy. Defaults to read-only: Codex analyses and reports but cannot modify files."New value: +"Sandbox policy. Uses the configured default when omitted; without one, Codex runs read-only."
    • changedInput schema / properties / web_search / description
      Previous value: -"Enable live web search for this run, through Codex's web_search = \"live\" setting. When omitted, Codex's own configured web_search mode applies."New value: +"Enable Codex's API-backed live web-search tool for this run. In a read-only sandbox, shell commands have no network access, so this is the route to current external information. When omitted, Codex's own configured web_search mode applies."
  3. Changed5 schema fields changedv0.2.0
    • changedInput schema / properties / auto_approve / description
      Previous value: -"Adds --approve-for-me so Codex auto-approves its own commands. Only applies when sandbox allows writes."New value: +"Adds --approve-for-me so Codex auto-approves its own commands. Only applies when sandbox is workspace-write."
    • changedInput schema / properties / model / description
      Previous value: -"Catalog slug from list_codex_models. Omitted means the recommendation matrix picks one."New value: +"Catalog slug from list_codex_models. If omitted, the configured default is used; with no default configured the call is refused and the recommended model is returned."
    • changedInput schema / properties / reasoning_effort / description
      Previous value: -"Reasoning depth, independent of model choice. Clamped to what the chosen model supports."New value: +"Reasoning depth, independent of model choice. Uses the configured default or the model's default when omitted. Clamped to supported levels within the configured ceiling; refused if none qualify."
    • changedInput schema / properties / use_worktree / description
      Previous value: -"Run in a managed git worktree so changes never touch the current working tree."New value: +"Run in a managed git worktree. Writes outside it remain subject to the sandbox policy and add_dirs."
    • changedInput schema / properties / web_search / description
      Previous value: -"Enable Codex's native web search tool."New value: +"Enable live web search for this run, through Codex's web_search = \"live\" setting. When omitted, Codex's own configured web_search mode applies."
  4. First observedv0.1.0

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint=false and openWorldHint=true, but the description adds substantial non-obvious context: prompt/context/target_files are sent to OpenAI, each run spends the user's own Codex usage, the user must be told when delegating, and the default sandbox is read-only unless reconfigured. This is exactly the kind of disclosure structured fields cannot carry.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The core purpose and usage trigger are front-loaded, and for a 17-parameter tool the length is defensible. It is a dense single block with some redundancy (the 'Codex cannot see this conversation' constraint is echoed by the schema), but each sentence carries operational weight.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a high-complexity tool with 17 parameters, no output schema, and a nested output_schema parameter, the description covers delegation conditions, privacy, cost, sandbox behavior, and the recommendation handoff. Return behavior (job_id for background, structured output block) is left to schema descriptions rather than explained, which is the main remaining gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds real meaning beyond the schema: the model parameter is tied to a codex_recommend workflow, sandbox read-only is the default and workspace-write is required for edits, and target_files/context/prompt are flagged as externally transmitted. This lifts it above the schema-only baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (delegate) and resource (a task to the local Codex CLI, described as OpenAI's coding agent), with additional scoping that distinguishes it from siblings by naming codex_recommend as the prerequisite recommendation tool. An agent can tell exactly what this does and how it differs from list_codex_models and codex_recommend.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit when-to-use triggers ('when the user asks for Codex, or when handing work off clearly serves their request: a second opinion from a different model family, or an investigation that would otherwise flood this conversation') and an explicit when-not ('do not delegate what you can answer directly'). It also names the alternative workflow (codex_recommend first when no model is chosen), leaving nothing to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.