Skip to main content
Glama

Pi Conductor for Codex

The Codex sister project of omp-conductor, adapted from its pi-backend branch. Codex orchestrates persistent Pi RPC workers through a local MCP server. Workers implement, investigate or review; Codex inspects reports and real diffs before merging.

The projects are maintained separately: omp-conductor serves Claude Code; omp-conductor-codex serves Codex. Shared Pi worker concepts and fixes can move between them under the MIT license. This is a standalone repository, rather than a GitHub fork tied to the original branch history. This package has no Claude runtime, React pane, Haiku call, shell stdin wrapper or Claude hook dependency.

Install

Requires Node 22.19+, Git, and Pi on PATH. The optional companion launcher also requires tmux and a current Codex CLI with --no-daemon. The toolbox runner uses Linux bash, timeout and flock (matching the upstream Linux workflow).

Install the prebuilt plugin directly from GitHub:

codex plugin marketplace add SamiulH25/omp-conductor-codex --ref main
codex plugin add omp-conductor-codex@omp-conductor-codex

Start a new Codex session to load the tools and skill. The committed dist/server.mjs bundles the server dependencies; installation does not require npm or a build step.

If Pi is not installed yet:

npm install -g --ignore-scripts @earendil-works/pi-coding-agent
pi --version

To install from a local checkout instead:

git clone https://github.com/SamiulH25/omp-conductor-codex.git
cd omp-conductor-codex
codex plugin marketplace add .
codex plugin add omp-conductor-codex@omp-conductor-codex

A minimal plugin ZIP is also available from GitHub Releases. Extract it, run codex plugin marketplace add /absolute/path/to/omp-conductor-codex, then use the same codex plugin add command above. Keep API keys in your environment or local worker data directory; do not add them to this checkout.

Pi workers use OpenCode Go by default. Export OPENCODE_GO_API_KEY in the environment Codex starts from, or put OPENCODE_GO_API_KEY=... in ~/.codex/plugin-data/omp-conductor-codex/env with mode 600. The legacy ~/.pi-workers/env file is also read if the new file is absent. The key is passed only to Pi; it is never returned in status or reports. No OpenAI API key is needed.

The server creates its worker settings, model definition and session directories on first use. OMP_CONDUCTOR_DATA_DIR overrides the data directory. Its startup check probes Pi and credentials without making an inference request.

Related MCP server: Vibechemy

Companion split

Launch Codex with the live dashboard beside it:

./bin/conductor -- -C /path/to/your/project

From an installed GitHub plugin, use the bundled launcher:

~/.codex/plugins/cache/omp-conductor-codex/omp-conductor-codex/1.1.0/bin/conductor -- -C /path/to/your/project

No build is needed. If tmux is missing, install it first (sudo apt install tmux on Ubuntu/Debian). The launcher creates a tmux session, or adds a sidebar to the current pane when already inside tmux. Codex receives any arguments after --. It uses --no-daemon so the MCP server inherits the panel ID and data directory even if another Codex daemon is already running.

Use Ctrl-b then Left/Right to switch panes with default tmux bindings. Press q in the dashboard to close that pane; j/k or Up/Down scroll its cards. The sidebar also closes when Codex exits. Ctrl-b then d detaches a newly created session; the launcher prints an attach command when used with launch --detach --.

The dashboard shows animated worker faces and spinners, states, activity, files touched, elapsed time, a time-budget bar, token rates and sparklines, cost, errors and verification warnings. Flat-rate usage is labeled plan. Snapshots are refreshed without model calls, are scoped to each launch, and stay private under the worker data directory. If the server stops updating, active workers are shown as interrupted rather than as live work.

Preview the UI or inspect existing snapshots:

./bin/conductor dashboard --demo
./bin/conductor dashboard --once
./bin/conductor dashboard --data-dir /path/to/worker-data

A standalone dashboard shows recent sessions across launch groups. The split launcher selects only its own group. The plugin installed in Codex must be v1.1.0 or newer to publish dashboard snapshots; update it using the commands below and start a new session. The launcher can also be run directly from the extracted release ZIP.

Use

Ask Codex: “Use Pi Conductor to split this implementation into independent worker tasks, review their changes, and merge the accepted work.” The bundled orchestrate-pi skill covers briefing, checks and the review loop.

Tools

Purpose

pi_spawn, pi_send

Start a worker or recall it with context retained

pi_status, pi_wait, pi_digest

Activity, token/cost metrics, completion and reports

pi_diff, pi_merge

Review real changes and merge an accepted branch

pi_kill, pi_cleanup

Stop work; remove an idle worker and its worktree

pi_dict, pi_tools

Project glossary and runnable verification checks

pi_model, pi_effort

Show/change settings with optional value; reset restores defaults

explore and review workers have read-only tools. dev and general workers default to isolated worktrees in git repositories. Non-git tasks run in place with snapshots for file-tool edits. A maximum of four workers runs per server session.

Workers get a soft time warning, a partial-report request and a hard stop after a grace period. Expected-file and verification warnings remain visible in digests. Toolbox checks can serialize across workers; supervisor verification supports bounded fix rounds. Detailed investigation/review reports are preserved rather than compressed by another model.

Merges refuse tracked changes in the main checkout, abort main-checkout conflicts and hand conflicts to the worker worktree. Cleanup refuses unmerged changes unless force: true is explicitly used. Worktrees and the path guard are accident-prevention measures, not a security sandbox: editing workers can run shell commands as the current user. Workers do not automatically load project instruction files; include the applicable instructions in their briefs.

State and host differences

State is stored separately from the Claude version in the Codex data directory. With CODEX_THREAD_ID, a restart reloads that thread's worker records and Pi session IDs. Without it, each MCP process uses a fresh UUID so separate clients do not share workers. An interrupted worker can resume through pi_send when its saved Pi session is available. Model, effort, dictionaries and toolboxes persist across sessions.

The Claude embedded pane is replaced by the companion tmux dashboard; Codex still uses pi_status and bounded pi_wait to review workers. No unsupported Codex UI hooks are installed. The dashboard is read-only: worker corrections, merging and cleanup remain MCP operations inside Codex.

Development and validation

npm ci --ignore-scripts
npm test
npm run dashboard -- --demo
npm run package  # clean ZIP under artifacts/, with no node_modules or development files

Tests use a local fake Pi executable and real Git repositories. They cover isolated edits, recall, reviewed new-file diffs, merge/cleanup, concurrent limits, crashes, verification failures, merge conflicts, and MCP initialization/input validation. They also check live snapshot updates, session isolation, stale-server indicators, terminal escape sanitization, and real tmux splits both outside and inside an existing session. They do not spend model credits. npm run build type-checks TypeScript and bundles the server with esbuild.

The host-independent worker engine and Pi guard/toolbox are adapted from upstream commit 8bca6ad under the included MIT license. Codex packaging follows OpenAI's plugin documentation.

Updating

Refresh the GitHub marketplace and reinstall the plugin to pick up new files, then start a new session:

codex plugin marketplace upgrade omp-conductor-codex
codex plugin add omp-conductor-codex@omp-conductor-codex

Sister-project maintenance

Report Codex packaging, MCP, or runtime issues here. Report issues in the original Claude host integration to omp-conductor. When porting a shared worker-engine fix, cite the source commit and preserve the MIT copyright notice. Neither project's working directory or release process depends on the other repository.

Available Tools

13 tools
pi_cleanupA

Remove a finished worker's worktree and branch and forget the worker. Refuses if it has unmerged work unless force is true.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesWorker id from pi_spawn
forceNo

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does disclose the critical behavior: it is destructive (removes worktree and branch, forgets the worker) and it has a safety guard that refuses on unmerged work unless force is true. It omits irreversibility details, required permissions, and whether the worker must be stopped first.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, zero waste, with the effect front-loaded and the precondition immediately after. Every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a two-parameter cleanup tool with no annotations and no output schema, the description covers what it does and its main failure mode. It stops short of describing the return value or defining precisely what "forget the worker" entails, leaving minor gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 50%: id is documented in the schema ("Worker id from pi_spawn") but force is not. The description compensates by explaining force's semantics (override the unmerged-work refusal), which is the key meaning an agent needs.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives specific verbs (remove, forget) and concrete resources (worktree, branch, worker), making the effect clear. It does not explicitly differentiate from siblings like pi_kill or pi_merge, so the agent must infer the boundary, but the core purpose is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied by "a finished worker's worktree," signalling this runs after completion, and the unmerged-work refusal states a precondition. However, no alternative tool is named (e.g. whether to pi_merge first), so the agent must infer the workflow ordering.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

pi_dictA

The project dictionary: short term -> definition entries (what a system is called, what it does, where it lives, conventions) that are injected into every worker for that project. Seed it from what you already understand BEFORE the first pi_spawn; after reviewing a worker, add facts you verified (never unverified worker claims). Not for task notes or history. action: show | set (upsert entries) | remove (terms).

ParametersJSON Schema
NameRequiredDescriptionDefault
dirNoAny directory inside the project (default: session cwd)
termsNoFor remove
actionYes
entriesNoFor set: [{term, definition}]. Definitions are one or two sentences (max 300 chars); name file paths.

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are present, so the description carries the full burden, and it does disclose the key non-obvious behavior: entries are injected into every worker for that project, and set is an upsert. The verification rule is valuable operational context. It stops short of describing what show returns, persistence scope, or auth, so it is strong but not exhaustive.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the resource definition, then usage, then exclusion, then action legend. Dense and mostly waste-free, though the first sentence is long and the action legend is somewhat compressed.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema and no annotations, the description covers purpose, timing, write policy, and action semantics adequately for a 4-param tool. The remaining gap is the return shape of show and any limits on entry count, which an agent might want.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 75%, so the schema already documents dir, terms, and entries with per-field descriptions. The description's 'action: show | set (upsert entries) | remove (terms)' reinforces the mapping and adds upsert semantics, but adds little beyond the existing schema text.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific resource ('project dictionary: short term -> definition entries') and what it does (show/set/remove entries that are injected into every worker for that project). It also draws an explicit boundary against siblings ('Not for task notes or history'), so an agent can separate it from pi_digest or task-note tools without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit when-to-use guidance with sequencing ('Seed it from what you already understand BEFORE the first pi_spawn') and a clear post-condition ('after reviewing a worker, add facts you verified (never unverified worker claims)'). It also names an exclusion ('Not for task notes or history'), which is the when-not half.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

pi_diffB

git diff (with stat) of a worker's directory so you can review the real changes.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesWorker id from pi_spawn
maxCharsNo

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It discloses that the tool runs a git diff with stat scoped to a worker directory, implying a read-only review operation, but it omits permissions, side effects, rate limits, and output characteristics.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single front-loaded sentence with no filler. The parenthetical '(with stat)' is slightly dense but earns its place by specifying output content.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema and no annotations. The description gives enough to know it is a diff review tool, but it does not explain the return format, the effect of maxChars, or how it fits into the worker workflow, leaving gaps for an agent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 50%. The id parameter is already documented in the schema as 'Worker id from pi_spawn,' and the description adds no further meaning. The maxChars parameter is undocumented in both the schema and the description, so the description does not compensate for the coverage gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'git diff (with stat) of a worker's directory.' It clearly distinguishes this from most siblings, though it does not explicitly name an alternative. The intent 'so you can review the real changes' reinforces the purpose.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied by 'so you can review the real changes,' suggesting it is used after a worker has produced changes. However, there is no explicit when-to-use, when-not-to-use, prerequisites, or reference to alternatives such as pi_status or pi_merge.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

pi_digestA

Compact report of one worker: files changed, commands, errors, final message, git status. detail "full" adds recent raw events. Treat claims as unverified; check with pi_diff.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesWorker id from pi_spawn
detailNo

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden, and it does disclose meaningful traits: the report is a summary of worker self-reported claims that should be treated as unverified, and detail=full exposes raw events. It does not explicitly state read-only nature or any access constraints, but the trust caveat is genuinely valuable behavioral context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three dense sentences with no filler; the report contents come first, followed by the detail modifier and the verification caveat. Every clause carries information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description must describe returns, and it does so by enumerating the report fields. The main residual gap is the detail parameter's default/brief behavior and any pagination or size limits on raw events.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 50%: id is documented as "Worker id from pi_spawn" but the detail enum has no schema description. The description compensates partially by explaining that detail="full" adds recent raw events, but never clarifies the default or what "brief" omits, so the low-coverage gap is only half filled.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific resource (one worker) and enumerates exactly what the report contains: files changed, commands, errors, final message, git status. It also distinguishes itself from the sibling pi_diff by directing verification there, so an agent can tell the two apart without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides real guidance on interpreting output ("Treat claims as unverified; check with pi_diff") and names the alternative tool for verification. It does not, however, state when to call pi_digest versus pi_status or pi_wait, so the primary selection condition is left implicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

pi_effortB

Show or change the reasoning effort every Pi worker uses

ParametersJSON Schema
NameRequiredDescriptionDefault
valueNoOmit to show; set a new value or reset.

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries full burden. It does not disclose whether the setting applies globally or per-worker, whether the change is immediate or persisted, what happens on reset, or what the show mode returns. For a read/write configuration tool with zero annotation coverage, this is a significant gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single, front-loaded sentence that communicates the operation and target without waste. Nothing is redundant or buried.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With one optional parameter, full schema coverage, and no output schema, the definition is minimally viable. However, for a tool that mutates a shared configuration, the absence of behavioral details (scope, persistence, reset semantics) leaves it short of complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, and the schema description already explains that omitting value shows the setting and setting it applies a value or reset. The description adds no syntax, valid values, or examples beyond the schema, so baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (show/change) and resource (reasoning effort for Pi workers). The dual show/change nature is captured. It does not explicitly differentiate from siblings like pi_model, which could also seem to configure Pi workers, leaving some ambiguity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool versus alternatives such as pi_model or pi_status. The description implies it's a configuration getter/setter but provides no context about prerequisites or when a caller should reach for it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

pi_killC

Abort a running worker and stop its Pi process.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesWorker id from pi_spawn

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It conveys that this is destructive but says nothing about reversibility, what happens to in-flight work or outputs, error behavior for an invalid/already-dead worker, or permissions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no wasted words. It is efficient, though the brevity contributes to the missing behavioral context rather than compensating for it.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-parameter tool this is minimally adequate: the action and target are clear and the parameter is fully documented. However, for a destructive operation with no annotations and no output schema, more context on irreversibility and post-conditions would be needed to reach a higher score.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the sole parameter is documented as 'Worker id from pi_spawn', which even ties it to the sibling that produces the id. The description adds nothing beyond the schema, so the baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Abort') and resource ('a running worker'), plus the concrete effect of stopping its Pi process. It is clearly distinct from pi_spawn and other siblings, though it does not explicitly name an alternative such as pi_cleanup.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no indication of when to use this versus pi_cleanup or pi_status, nor any preconditions (e.g. worker must be running). The agent must infer usage entirely from the name and the required id.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

pi_mergeA

Commit a finished worker's worktree changes and merge its branch (--no-ff) into the repo it was spawned from. Refuses while the worker runs or the main tree has tracked changes; aborts and reports on conflict.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesWorker id from pi_spawn
messageNoMerge commit message

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does well: it discloses the merge strategy (--no-ff), the two refusal preconditions, and conflict behavior (aborts and reports). It omits what state the worktree/branch is left in after abort and whether the commit is unconditional, leaving a small gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences with zero filler; the action is front-loaded and the failure semantics follow immediately. Every clause carries information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation tool with no annotations and no output schema, the description covers preconditions and failure modes adequately. It could say more about post-merge state and what the success response contains, but an agent has what it needs to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and both parameters are documented there ('worker id from pi_spawn', 'merge commit message'). The description adds no format or behavioral detail about id or message, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a precise compound action: commit a finished worker's worktree changes and merge its branch (--no-ff) into the spawning repo. The verb+resource is unambiguous and the 'worker' framing separates it from generic git tools and from siblings like pi_cleanup or pi_diff.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Clearly scopes usage to a 'finished worker' and gives explicit preconditions under which it refuses (worker still running, main tree has tracked changes). It does not name a sibling alternative (e.g. pi_cleanup for discarding work), so it falls just short of the explicit when-not/alternative bar.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

pi_modelC

Show or change the model every Pi worker uses

ParametersJSON Schema
NameRequiredDescriptionDefault
valueNoOmit to show; set a new value or reset.

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It discloses scope (affects "every Pi worker") which is genuinely useful, but says nothing about whether existing workers are updated, whether a restart is needed, permissions required, or side effects of changing the model.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single, front-loaded sentence with no wasted words. It is appropriately sized for a one-parameter toggle tool, though it is terse enough to leave gaps.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple one-parameter tool this is minimally adequate: the core show/change behavior is conveyed. However, with no annotations and no output schema, the description could have disclosed the reset semantics and the global-effect behavior that currently live only in the schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the single parameter is documented in the schema ("Omit to show; set a new value or reset."). The description adds no meaning beyond that, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific dual-purpose verb ("show or change") and a clear resource ("the model every Pi worker uses"). It is distinguishable from siblings like pi_status or pi_effort, though it doesn't explicitly name a contrasting alternative.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It signals that the tool both reads and writes, but offers no guidance on when to call it for showing versus changing, nor any prerequisites or timing constraints relative to the sibling tools. The reader must infer the intended workflow.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

pi_sendA

Send a follow-up message to a finished worker. Its Pi process (or, after 30 idle minutes, its saved session) keeps everything it already read, so use this for fixes and for the next task in the same area instead of re-spawning (no re-reading, no re-learning the layout). Send only the new instruction; it already has the background. The worker must not be running.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesWorker id from pi_spawn
effortNoChange this worker's reasoning effort from this message on.
messageYes
maxMinutesNoNew time limit for this run (use it to give a worker that hit its limit more time).

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden, and it delivers real value: the worker's Pi process — or its saved session after 30 idle minutes — retains prior context, and the caller should send only the new instruction. It does not disclose the consequence of calling while the worker is running, nor whether the send blocks or returns immediately.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three front-loaded sentences with no filler; the purpose and the re-spawn contrast come first. The parenthetical about the 30-minute saved session is dense but earns its place by explaining the retention guarantee.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema and no annotations, so the description must stand alone — it covers purpose, precondition, context retention, and message framing. It omits the failure behavior when the worker is running and what the call returns, which are the remaining operational gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 75% and already documents id, effort, and maxMinutes. The description adds genuine meaning for the undocumented 'message' parameter — 'Send only the new instruction; it already has the background' — telling the caller not to restate context.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (send a follow-up message) and resource (a finished worker), and explicitly contrasts with the sibling pi_spawn ('instead of re-spawning'). An agent can route between pi_send and pi_spawn without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Names the alternative ('instead of re-spawning') and the conditions that select this tool: 'for fixes and for the next task in the same area'. It also states the precondition plainly — 'The worker must not be running.'

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

pi_spawnA

Start a Pi worker in the background on a self-contained task and return its id at once. Workers always use the model the user set with pi_model (not selectable here). Give it a complete standalone brief (it has no access to this conversation). Use disjoint dirs or let worktree isolation separate parallel workers. Then use pi_wait / pi_digest to read compact results and judge them. Workers are long-lived RPC processes that keep their context: for a fix or a follow-up in the same area, pi_send an existing worker instead of spawning a new one.

ParametersJSON Schema
NameRequiredDescriptionDefault
dirNoWorking directory (default: session cwd)
taskYesComplete standalone instructions for the worker
agentNoAgent type (default general). general: unrestricted, short SUMMARY (the pre-agent-types behaviour); dev: implements changes: read, edit, write, shell; short SUMMARY of what changed. Worktree by default; explore: read-only investigator: finds and reads code, returns a long evidence-backed FINDINGS report (never summarized). No worktree; review: read-only reviewer: returns every issue found with file:line, severity and a fix (never summarized). No worktree
titleNoShort label
checksNoToolbox checks (pi_tools) the worker must run and pass after its last edit. Default: the toolbox entries marked required. [] = none for this task. The worker runs them itself with `check <name>`; serial ones never run twice at once.
effortNoReasoning effort for this worker only (default: the pi_effort setting). Raise it for tasks that need real reasoning.
expectNoPaths (relative to dir) the task must change. If the worker finishes without changing one, the digest, status and wake-up message carry a warning.
noDictNoSkip the project-dictionary requirement for a dev/general worker (the first one in a project is refused until pi_dict has entries).
verifyNoA shell command the PLUGIN runs itself in the worker's directory after the worker finishes (one at a time across workers, up to 10 min). Prefer toolbox checks (pi_tools), which the worker runs and fixes itself; use verify for a final gate the worker must not run. The result goes in the digest; a failure counts as a warning.
worktreeNoIsolate in a new git worktree (default depends on agent type; dev and general: true when dir is a git repo, explore and review: false)
fixRoundsNoWith verify: how many times (0-3, default 0) a failing verify is sent back to the worker to fix automatically.
maxMinutesNotime limit for the run (default depends on agent type, 15-20). At 85% the worker is told to wrap up; at 100% it is asked for a partial report and gets 2 more minutes before it is stopped. pi_send can resume it.
verifyTimeoutSecNoTime limit for the verify command (default 300, max 600).

TDQS

A4.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does well: it discloses the async/non-blocking nature (returns id at once), that the model is fixed by the pi_model setting and not selectable, that workers have no access to this conversation, and that they are long-lived RPC processes that retain context. It does not discuss failure/timeout behavior (covered only in schema descriptions) or concurrency limits, keeping it short of a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Six sentences, each carrying distinct information: what it does, the model constraint, the standalone-brief requirement, isolation guidance, result-reading path, and the pi_send alternative. The core action and return value are front-loaded, and there is no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 13-parameter tool with no output schema and no annotations, the description covers the whole lifecycle an agent needs: spawning semantics, isolation, model binding, follow-up routing, and result inspection, while the rich schema descriptions handle per-parameter detail. Nothing essential to a correct invocation is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3; the description adds genuine meaning beyond the schema by explaining that the task brief must be fully standalone because the worker cannot see this conversation, and by noting the model is not a parameter here. That said, the per-parameter semantics (checks vs verify, effort, maxMinutes) live almost entirely in the schema, not the description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Start a Pi worker') plus the scope ('in the background on a self-contained task') and the immediate return value ('return its id at once'). It explicitly distinguishes itself from pi_send (for follow-ups), pi_wait/pi_digest (for reading results), so an agent can route correctly without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit when-to-use and when-not-to-use guidance: use pi_send for a fix or follow-up in the same area instead of spawning a new worker, use disjoint dirs or worktree isolation for parallel workers, and use pi_wait/pi_digest to judge results. Alternatives are named with the condition that selects them.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

pi_statusB

One line per worker: state, elapsed, last action, files touched.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the burden. It does disclose the returned fields (state, elapsed, last action, files touched), which is genuinely useful behavioral context given there is no output schema, but it says nothing about read-only safety, whether it blocks, or freshness/caching of the status data.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single tight sentence with zero filler, and the most important information (per-worker line format) is front-loaded. It is perhaps too terse to earn a 5, but every word earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description needs to describe return values, and it does so adequately by enumerating the per-worker fields. With zero parameters and a simple read operation, the main missing piece is when to prefer this over pi_wait/pi_digest.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, which is the baseline-4 case. Nothing in the description contradicts or misleads about invocation, so no penalty applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description conveys that this tool reports status, but does so indirectly through the return format ('one line per worker: state, elapsed, last action, files touched') rather than stating a verb+resource like 'report worker status'. It does not differentiate itself from siblings such as pi_wait, pi_digest, or pi_dict, leaving the agent to infer the distinction.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to call this versus the many sibling tools (pi_wait, pi_digest, pi_dict, pi_diff). No prerequisites, timing, or exclusion conditions are stated; the agent must guess the use case from the name alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

pi_toolsA

The project toolbox: the checks (tests, compile, lint) workers are expected to run THEMSELVES with check <name> before reporting. Shown in every dev/general worker's prompt; required checks are enforced (the worker is reminded, and the digest warns if one was not run or failed). serial checks hold a project-wide lock, so a tool that cannot run twice at once (a Unity editor, a Gradle daemon) is safe with parallel workers: do not forbid workers from running it. A Unity project with no toolbox set gets built-in unity-editmode (required) and unity-playmode checks. action: show | set (upsert checks) | remove (names) | reset (drop all, back to the built-in default).

ParametersJSON Schema
NameRequiredDescriptionDefault
dirNoAny directory inside the project (default: session cwd)
namesNoFor remove
actionYes
checksNoFor set: checks to add or replace. The first set on a project starts from the built-in default if there is one.

TDQS

A4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and discloses rich behavior: required checks are enforced with worker reminders and digest warnings, serial checks hold a project-wide lock and are safe with parallel workers, and Unity projects receive built-in default checks. It also explains that `set` upserts and `reset` drops all back to the built-in default. This goes well beyond the schema and gives the agent a clear operational model.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but front-loaded with the core concept: it opens by defining the toolbox and check-running contract before moving to locking behavior, defaults, and action semantics. It is longer than average but most sentences carry operational information, with only minor redundancy in the enforcement and Unity-default notes.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex, mutation-capable tool with no annotations and no output schema, the description covers the essential behavioral context: required checks, serial locking, built-in Unity defaults, and action-specific effects. It does not explain return values or error cases, but those gaps are minor given the focus on invocation correctness.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is already 75%, so the baseline is 3. The description adds meaningful semantics for the `action` parameter by mapping `show`, `set` (upsert checks), `remove` (names), and `reset` (drop all, back to built-in default), which clarifies how to use the enum beyond the bare schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states that this is the project toolbox for checks and enumerates the four actions (show, set, remove, reset), so the verb and resource are clear. Sibling tools are unrelated (pi_spawn, pi_dict, etc.), and no other tool in the list manages checks, but the description does not explicitly differentiate itself from siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explains that workers are expected to run checks themselves with `check <name>` and describes enforcement and locking behavior, but it never states when an agent should call pi_tools versus alternatives. Usage is implied through the action list and the note about not forbidding workers from running serial checks, with no explicit when-not guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

pi_waitA

Wait until the given workers (default: all running) finish or timeoutSec passes (max 60 per call; call again to keep waiting), then return their digests.

ParametersJSON Schema
NameRequiredDescriptionDefault
idsNo
modeNodefault all
timeoutSecNo

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations exist, so the description carries the full burden. It does disclose the significant blocking behavior and the 60-second per-call cap, but says nothing about timeout return semantics, error conditions, or whether waiting on already-completed workers is a no-op.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One dense sentence that front-loads the blocking action, then the default, the timeout cap, and the return value. No filler and nothing repeated from the schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-annotation, no-output-schema tool with three parameters, the description covers the blocking behavior, default scope, timeout cap, and return payload. The unexplained 'any' mode and the absence of timeout/error behavior keep it from being complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 33%, so the description must compensate. It clarifies that 'ids' defaults to all running workers and that 'timeoutSec' is capped at 60, but it never explains the 'any' vs 'all' semantics of 'mode', leaving a key enum undocumented.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('wait until the given workers ... finish') plus the observable outcome ('return their digests'), which is far more than the name alone. It does not, however, distinguish itself from siblings like pi_status or pi_digest, so an agent must infer the boundary.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives useful operating context ('default: all running', 'max 60 per call; call again to keep waiting'), which tells the agent how to invoke it repeatedly. It never says when to prefer this over pi_status/pi_digest or what happens on timeout, so the usage guidance is implied rather than explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 13 tool updatesv1.1.0
    • First observedpi_cleanup
    • First observedpi_dict
    • First observedpi_diff
    • First observedpi_digest
    • First observedpi_effort
    • First observedpi_kill
    • First observedpi_merge
    • First observedpi_model
    • First observedpi_send
    • First observedpi_spawn
    • First observedpi_status
    • First observedpi_tools
    • First observedpi_wait

TDQS

A3.8/5.0

Scored across 13 tools

Disambiguation4/5

Most tools have a clearly distinct lifecycle role (spawn, kill, merge, cleanup, send). There is mild overlap in reporting: pi_wait, pi_digest, and pi_status all surface worker state/digests, though the blocking-vs-snapshot distinction is described well enough to disambiguate.

Naming Consistency5/5

Every tool uses a uniform pi_ prefix with a short, readable suffix. Action verbs (spawn, wait, send, kill, merge, cleanup) and config nouns (dict, tools, model, effort, status) are used consistently within their categories.

Tool Count5/5

13 tools is well-scoped for a worker-orchestration server, covering the full lifecycle without redundant entries. Each tool earns its place (lifecycle ops, config, and two project-context stores).

Completeness5/5

The surface covers the whole worker lifecycle: spawn, monitor (wait/digest/status), review (diff), follow-up (send), abort (kill), integrate (merge), and teardown (cleanup), plus model/effort config and shared project context. No obvious gaps or dead ends.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    C
    maintenance
    Enables Codex to delegate bounded coding tasks to MiMo Code through a shared local daemon, supporting task boundaries, Git Worktrees, and a collaborative review workflow.
    3
    MIT
  • A
    license
    Not graded
    quality
    A
    maintenance
    An MCP server for orchestrating a fleet of CLI coding agents in isolated git worktrees. It exposes tools for spawning workers, sending instructions, reviewing diffs, and merging changes, with full terminal visibility.
    5 npm
    4
    MIT
  • F
    license
    Not graded
    quality
    B
    maintenance
    Enables MCP hosts like Claude Code and Codex to spawn, manage, and interact with persistent, reusable Pi coding-agent sessions, supporting task dispatch, status checks, and session lifecycle control.
    2 npm
    -