CoS Codex Bridge
CoS Codex Bridge is a local MCP server that lets a Chief of Staff client orchestrate Codex and Claude Code CLI tasks: create/find projects, submit prompts, monitor receipts, and continue sessions.
Project management:
bridge_projectslists, creates, registers, or inspects allowed project directories;bridge_session_managerenames, pins/unpins, and assigns sessions to projects.Session discovery:
bridge_sessionsfinds and reads allowed Codex or Claude Code sessions, including queued items and output.Work submission:
bridge_submitstarts or continues async Codex/Claude Code work with stable request IDs, whole prompts, attached artifacts, and direct or desktop-queue delivery.Progress tracking:
bridge_receiptreads durable receipts with hashes, state, and bounded output;bridge_steerretries/recovery for queued items;bridge_cancelrequests cancellation.Clarification handling:
bridge_answeranswers pending Codex clarification questions without approving permission expansion.Artifacts:
bridge_artifactwrites versioned UTF-8 text artifacts without overwriting, or reads them.Capability checks:
bridge_doctorreports installed CLI availability, versions, mode, sandbox, and scope.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@CoS Codex BridgeFind my project, submit this prompt to Codex, then track it."
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
CoS Codex Bridge
Your Chief of Staff. Now in charge of Codex, too.

A local MCP server that lets a Chief of Staff client find Codex tasks, create project work, deliver whole prompts, follow progress and continue the same conversation. An optional Claude Code CLI route in the same MCP does this with local project folders and saved CLI sessions. Free MIT core. No bridge subscription or checkout. Your existing Codex or Claude Code access is required for real execution.
Choose your release: 0.1.4 on npm is the stable Codex-only route (@latest). 0.2.0-beta.1 adds Claude Code CLI support (@beta) in the same MCP. Install locally, connect your client, then use the bridge_* tools. No daemon or Desktop-owner adapter installation is required. The Codex core workflow and the Grok Bot → Claude Code CLI handoff have passed owner field testing on macOS. Desktop-owned paused-queue recovery remains a known Codex limitation. Sidebar rendering on newer Codex Desktop versions is not certified; check the compatibility table before relying on it. A queued receipt is never proof that work started.
The bridge is indexed in the official MCP Registry and Glama. These are discovery listings; the bridge still installs and runs locally.
What your Chief of Staff can do
Turn research or a product idea into a new Codex project and task.
Send the complete prompt and attached text without manual copying and pasting.
Find an existing Codex task, read its progress and continue the same conversation.
Track delivery with durable receipts instead of assuming an accepted prompt ran.
Rename, assign, pin and cancel scoped work across approved projects.
Keep local project access inside explicit allowlisted directories.
Example: “Create a project for this idea, send my research to Codex, monitor the build, and follow up in the same task with the review findings.”
Claude Code route (beta)
The same MCP can now start and continue Claude Code CLI work. Pass provider:"claude-code" to bridge_projects, bridge_sessions and bridge_submit; the existing Codex route remains the default. A local folder is the Claude Code project context. The bridge creates a saved CLI session there, returns a durable receipt, reads its result and follows up in the same session. This was locally tested with Claude Code 2.1.269, including a real file edit in a disposable folder.
In an owner field test, the Grok Bot Chief of Staff used the installed MCP to create an allowlisted folder, send a prompt to Claude Code, read its completion, and follow up in the same saved session. Both bridge receipts and Claude Code's native session history were checked. This proves the local CLI handoff, not Claude Desktop or account Project control.
Claude account Projects, ordinary chats, Cowork/Dispatch and Desktop-owned sessions are separate surfaces. This route does not create or control them, and a CLI session does not automatically appear in Claude Desktop's sidebar. Claude setup, exact workflow and limits.
See the handoff
The animation illustrates the local handoff. It uses no private Desktop data or project screenshots.
Related MCP server: deepseek_harness
Try the safe demo
New here? Start with the no-account quickstart. It uses the stable npm package, needs no Git clone, and shows what a completed receipt looks like. Report your install result, including which local MCP client you use.
This exercises the MCP workflow locally without a Codex account, model call or project changes:
This clone uses GitHub main, currently the Claude Code CLI preview (0.2.0-beta.1). To demo the stable Codex-only source instead, run git checkout v0.1.4 after cd cos-codex-bridge and before npm ci.
git clone https://github.com/AV-Labs-Co/cos-codex-bridge.git
cd cos-codex-bridge
npm ci
npm run demoThe demo prints a completed receipt, payload hash and task metadata so you can see the bridge contract before granting access to a real project.
Easiest setup: give this link to your local assistant
Open the repository and installation instructions. An assistant with local terminal access can install it for you. A browser-only or cloud-only chat cannot install software on your computer.
Copy this installation request to your Chief of Staff:
Install CoS Codex Bridge from https://github.com/AV-Labs-Co/cos-codex-bridge. Read the README and SECURITY.md first. Use only a project folder I approve, keep read-only defaults, and preserve existing client configuration. Follow the installer instructions, run doctor, and connect the generated stdio MCP entry to my local client. Tell me what passed and what still needs setup. Wait for my first task before submitting any work. Do not publish or deploy anything.
You can also download the source archive from Releases, extract it and follow the steps below. Git cloning makes later updates easier. A public npm package is available, but no hosted endpoint is required.
Install
Requires Node.js 22+, npm, a locally authenticated Codex CLI for Codex work and/or Claude Code CLI for Claude work, and a local client supporting stdio MCP. Codex Desktop registration additionally requires Codex Desktop on macOS.
The Git clone below installs the current main branch, which is the Claude Code CLI preview (0.2.0-beta.1). For the stable Codex-only source, run git checkout v0.1.4 immediately after cd cos-codex-bridge, before npm ci. The pinned npm @0.1.4 command below is the stable package route.
git clone https://github.com/AV-Labs-Co/cos-codex-bridge.git
cd cos-codex-bridge
npm ci
npm test
node scripts/install.mjs --root /absolute/path/to/your/projectsThe installer writes a private config, launcher and MCP snippet under ~/.local/share/cos-codex-bridge. Keep the checkout in place. Default execution is read-only; use --write only for approved project edits. Choose specific project roots, never your entire home directory. Add --codex /absolute/path/to/codex or --claude /absolute/path/to/claude if either CLI is not on the MCP client's PATH. See installer and upgrade steps.
If you prefer npm to Git, install a pinned release into a dedicated folder, then run its same local installer. Use @0.1.4 for stable Codex only or @0.2.0-beta.1 for Codex plus the Claude Code CLI preview:
mkdir -p "$HOME/.local/share/cos-codex-bridge-package"
npm install --prefix "$HOME/.local/share/cos-codex-bridge-package" cos-codex-bridge@0.1.4
node "$HOME/.local/share/cos-codex-bridge-package/node_modules/cos-codex-bridge/scripts/install.mjs" --root "/absolute/path/to/your/projects"Keep that package folder: the generated launcher points to it. The npm path was checked with clean stable and beta installs, isolated demo configs and doctor. The installer does not edit any MCP client settings for you.
~/.local/share/cos-codex-bridge/cos-codex-bridge doctorCheck codexAvailable, claudeCodeAvailable, claudeCodeAuthenticated, mode, sandbox and roots. The Claude sign-in check is made from the MCP host's process and may differ from a sandboxed terminal; doctor does not verify Desktop sidebar rendering. For a model-free demo, install into a separate prefix with --demo.
Paste the generated mcp-client.json into your client's MCP configuration. Equivalent shape:
{"mcpServers":{"cos-codex-bridge":{"command":"/absolute/path/to/node","args":["/absolute/path/to/cos-codex-bridge/dist/cli.js","--config","/absolute/path/to/config.json","mcp"]}}}A complete handoff
Tell your Chief of Staff: “Find my app project, send this entire implementation brief to Codex, monitor it, then continue that same task with the review findings.”
Resolve the exact project and task with
bridge_projectsandbridge_sessions.Create a directory if needed, then register it. A directory alone is not a Desktop project.
Call
bridge_submitwithproject, the wholeprompt, and a stablerequestId. IncludethreadIdfor follow-ups.Poll
bridge_receipt. Verify hashes, task ID and eventual completion. Handle clarification withbridge_answer.Use
bridge_session_manageto rename, assign or pin the task. Stored metadata and visible Desktop rendering are separate proofs.
A successful model completion does not independently prove that generated code works. Review and test the result.
Ten tools, twelve features and one known limitation
Tool | Purpose |
| List aliases, create a folder, register/open or inspect a Desktop project |
| Find and read existing allowed tasks, including externally created tasks |
| Start or follow up; stable request IDs and complete UTF-8 payloads |
| Durable progress, hashes, bounded output and explicit uncertainty |
| Retry recovery of an existing bridge queue item, without resending it |
| Answer pending clarification; never approve permission expansion |
| Request cancellation of bridge-owned direct work |
| Read/write versioned text artifacts without overwrite |
| Rename, pin/unpin and assign to a registered project |
| Report mode, Codex version, scope and honest capability limits |
Known limitation (uncommon): If a session already has an active writer and a steering prompt is sent, the prompt waits for a natural pause/stopping point. On a long autonomous run, the only human intervention needed is pressing Steer in that case.
See the v1 capability contract and verification matrix.
Busy tasks, steering and interruption recovery
Direct submission defaults to SESSION_BUSY when another writer owns the task. To opt into the existing Desktop execution policy, use onBusy:"queue" and acceptDesktopPolicy:true. delivery:"desktop-queue" explicitly queues to an existing task. These paths use Codex's first-party queue API, the equivalent of codex queue, and preserve separate text inputs and stable client message IDs.
Receipts distinguish busy, queued, steered, delivered, completed, blocked and uncertain. thread/queue/start recovery has passed a local live test with the writer available. Desktop-owned paused recovery is deferred for v0.1; it may require a human Steer click. bridge_steer retries the exact saved item when the writer is available, including an existing CLI-created item adopted by its queue ID. bridge_sessions with includeQueue:true exposes pending IDs and inferred needs-steer state. It never silently forks or treats queue disappearance as delivery. See queue semantics.
Security defaults
Source preview, not in npm 0.1.4 or 0.2.0-beta.1: direct submissions now accept per-handoff writeIntent. A read-only handoff narrows a write-enabled configuration; a write handoff cannot expand configured authority. Desktop queue rejects explicit per-handoff intent because its existing permissions cannot be narrowed by this route. Contract and examples.
Default-deny realpath allowlists, read-only direct execution, private local receipts, bounded UTF-8 input, explicit project/task matching and no implicit cloud endpoint. Prompts and artifacts are stored locally in plaintext for receipt integrity; do not treat them as encrypted storage.
Direct workers disable inherited connectors and deny permission approvals. Desktop queue is a separate, explicit policy boundary: it uses the existing task's permissions and tools. The bridge cannot enforce a narrower sandbox inside that already-running task. No automatic store submission, social posting or publication is authorized. See SECURITY.md.
Client matrix
Environment | Evidence |
Grok Bot / CoS on owner's Mac | Codex orchestration and a Claude Code CLI create → complete → same-session follow-up field-tested; installation-specific integration |
Standard stdio MCP client | Protocol handshake, schemas, errors, installer and demo tested automatically on macOS and Ubuntu |
Codex Desktop macOS / CLI 0.153.4 | Owner-tested registration, assignment, pinning, continuity and queue delivery |
Codex Desktop 0.155.0-alpha.9.2 | Native metadata observed in field; sidebar rendering not certified; legacy adapter disabled |
Other MCP clients | Expected protocol compatibility; not individually field-certified |
Claude Code CLI 2.1.269 | Local create, complete, read and same-session follow-up verified; restricted file write verified in a disposable folder |
Claude Desktop/Claude account Projects | No creation, sidebar registration, pinning or live-session control claim; saved local transcript may be readable but not resumable through this route |
Windows runtime / Linux Desktop integration | Not verified; no macOS Desktop parity claim |
Ordinary ChatGPT chats | Not supported |
Hosted service / Composio cloud | Not provided or listed |
Native project assignment is supported through the installed experimental App Server API. An optional, version-gated legacy Desktop assignment adapter has backup and race checks; it is off by default and remains experimental.
Community
This project focuses on reliable Chief of Staff handoffs rather than a universal superiority claim. Other Codex MCP projects solve useful adjacent workflows. We do not claim “most advanced,” all-account control or blanket autonomy.
Ask an installation question or share a client recipe in GitHub Discussions. Use Issues for reproducible bugs, with redacted version, state and error details. Never post full private prompts or credentials.
Star it to follow development and fork it for your client. MIT License.
Available Tools
10 toolsbridge_answerA
Answer a pending Codex clarification using its receipt and question IDs. Cannot approve permissions or widen sandbox.
| Name | Required | Description | Default |
|---|---|---|---|
| answers | Yes | ||
| receiptId | Yes | ||
| questionId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already indicate this is not read-only and not destructive, but the description adds non-obvious behavioral limits: only pending clarifications can be answered, and permissions/sandbox changes are out of scope. This is valuable context beyond the annotations, though it does not describe side effects or outcome details.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with no filler: the first states the purpose, the second states a key limitation. Both sentences earn their place, and the most important information is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers the core operation and an important limitation, but given the nested `answers` schema, absent output schema, and no parameter descriptions, it leaves key workflow details unclear: how to construct the answers payload, where the IDs come from, and what a successful response looks like. Adequate but with clear gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It names receipt and question IDs as identifiers, which adds some meaning, but it does not explain the nested `answers` object structure, its key semantics, or its relationship to the top-level `questionId`. This is a significant gap for a three-parameter tool with nested objects.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action ('Answer') applied to a specific resource ('a pending Codex clarification') and names the identifying inputs ('receipt and question IDs'). This clearly distinguishes bridge_answer from siblings like bridge_submit and bridge_receipt, and it is not a tautology.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It clearly identifies the intended context: answering pending clarifications identified by receipt and question IDs. It also gives explicit exclusions ('Cannot approve permissions or widen sandbox'), but it does not name an alternative tool for those excluded actions, so it falls short of full routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
bridge_artifactA
Create a versioned UTF-8 artifact without overwriting, or read it. Artifact creation does not install skills or schedule routines.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | ||
| text | No | ||
| action | Yes | ||
| project | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the annotations (readOnlyHint=false, destructiveHint=false), the description adds meaningful behavioral context: writes are versioned and non-overwriting, and creation has no side effects on skills or routines. This clarifies the tool's safety profile and side-effect scope. It does not mention potential error conditions or what happens on duplicate writes, but it goes beyond the minimal annotation coverage.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences with no fluff. The core purpose is front-loaded in the first sentence, and the second adds a critical side-effect clarification. Every word contributes to the agent's understanding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 4-parameter tool with no output schema and sparse annotations, the description is incomplete. It omits definitions for 'name', 'project', and 'text', does not describe return values or error behavior (e.g., what happens if the artifact exists), and gives no context on how these parameters relate to other bridge tools. An agent would need additional information to call it correctly in all cases.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has 0% description coverage, so the description must compensate. However, it only clarifies the 'action' parameter (write/read) through the enum and the phrase 'without overwriting'. It does not explain the purpose of 'name', 'project', or 'text', leaving the agent to guess these are artifact identifiers and content. The description adds little value for parameter meaning.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states a specific verb('Create' or 'read') and a resource ('versioned UTF-8 artifact'), and adds a distinctive property ('without overwriting') that differentiates it from ordinary writes. It also explicitly says what it does not do ('does not install skills or schedule routines'), which helps set it apart from sibling tools like bridge_submit or bridge_steer without needing their schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides a negative usage hint: artifact creation does not install skills or schedule routines, implying those actions belong to other tools. However, it does not explicitly name alternatives or state concrete conditions for when to use this tool versus a sibling. The guidance is present but implicit, leaving the agent to infer the full decision.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
bridge_cancelA
Request cancellation of a bridge-owned running job. Poll receipt to verify. Does not undo file changes.
| Name | Required | Description | Default |
|---|---|---|---|
| receiptId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already mark the tool as non-read-only and non-destructive, so the bar is lower. The description adds useful behavioral details beyond that: cancellation is a request requiring receipt polling, and it does not revert file changes. This gives the agent a realistic model of the tool's behavior without contradicting the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, each carrying distinct information: the action, the verification workflow, and the non-undo limitation. The most important procedural detail is front-loaded, and there is no redundant filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with one required parameter and no output schema, the description captures the essential workflow and limitation. It could explain how the receiptId is obtained or what the cancellation receipt looks like, but the combination of annotations and concise description makes it sufficiently complete for a simple cancel operation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has 0% description coverage, and the description does not explain the receiptId parameter or where it comes from. 'Poll receipt to verify' hints at the role of a receipt, but it never explicitly ties receiptId to the cancellation request or to a prior bridge_submit call, so the agent must infer the parameter's meaning.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action ('Request cancellation'), a specific target ('bridge-owned running job'), and adds a key distinction with 'Does not undo file changes.' It does not explicitly name sibling tools, but the resource and action are clear enough to separate it from bridge_submit or bridge_steer.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear operational context: cancellation is requested, not performed instantly, and the agent should poll a receipt to verify. It also warns that file changes are not undone, which helps the agent avoid using this tool as a rollback. It stops short of explicitly naming alternatives or when-not-to-use conditions, but the guidance is still actionable.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
bridge_doctorARead-only
Read local bridge capabilities and Codex/Claude Code versions. Does not authenticate or change settings.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotations already establish readOnlyHint=true and destructiveHint=false, and the description adds valuable specificity: it does not authenticate and does not change settings, and it is scoped to local information. This goes beyond the annotation flags by clarifying authentication behavior and local scope, which helps an agent know what side effects to expect (none). No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences with no filler. The primary purpose is front-loaded in the first sentence, and the negative scope in the second adds value without redundancy. Every word contributes to understanding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple, parameterless, read-only diagnostic tool, the description covers what it reads (local capabilities and versions) and what it excludes (auth and settings). There is no output schema, but the description adequately tells an agent what information to expect. Given the annotations and zero parameters, nothing essential is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so the schema has no properties to describe. The description correctly focuses on behavior rather than parameters. This matches the baseline for a zero-parameter tool, where no parameter-level meaning is needed.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb and resource: it 'Read[s] local bridge capabilities and Codex/Claude Code versions,' which clearly defines the action and object. It also explicitly states what it does not do (authenticate or change settings), further distinguishing it from sibling tools that likely perform mutations or other reads. Verdict: clearly identifiable and differentiated.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies a diagnostic/preflight use case: checking local capabilities and versions, likely before using other bridge tools. It explicitly notes it is not for authentication or settings changes, giving a when-not, but it does not name alternative tools or provide explicit conditions for selecting it over siblings. This is clear context with an exclusion, but not fully explicit alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
bridge_projectsC
List aliases or create an allowed directory. Codex can register with Desktop; Claude Code uses the folder as local CLI project context, without Desktop registration.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | ||
| action | No | list | |
| parent | No | ||
| project | No | ||
| provider | No | codex |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description adds one useful behavioral detail about Codex Desktop registration versus Claude Code local folder context. However, it omits the 'register' and 'inspect' actions, doesn't disclose side effects or prerequisites, and the annotations only say the operation isn't read-only without adding depth.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences with no filler, and the primary actions are front-loaded. It is concise, though the brevity contributes to missing information rather than being a model of efficiency.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given five parameters, no output schema, and only weak annotations, this description is not complete enough for an agent to call the tool correctly. It covers only a subset of the action enum and gives no information about return values, path semantics, or when 'create' vs 'register' is appropriate.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for parameter meaning. It explains provider behavior somewhat ('Codex can register with Desktop; Claude Code uses the folder as local CLI project context'), which maps loosely to the 'provider' and 'project' parameters, but 'name', 'action', and 'parent' are left undocumented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description mentions specific actions ('List aliases', 'create an allowed directory') but the resource and intended scope are vague ('aliases', 'allowed directory'). It doesn't distinguish the tool from sibling bridge tools or account for the 'register'/'inspect' actions present in the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is given on when to use bridge_projects instead of related tools. The only contextual note differentiates Codex and Claude Code behavior, but it doesn't state when this tool should or shouldn't be selected.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
bridge_receiptARead-only
Read durable delivery receipt, hashes, state and bounded output. uncertain means inspect the session before retrying. completed means the agent turn finished, not that its claims were independently verified.
| Name | Required | Description | Default |
|---|---|---|---|
| receiptId | Yes | ||
| includeOutput | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, which the description aligns with by saying 'Read'. It adds meaningful behavioral nuance by explaining that 'completed' only means the agent turn finished, not independent verification, and notes 'bounded output' hinting at size limits. This goes beyond the annotations and adds value.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise and well-structured: one sentence for the core action and one for state interpretation. It front-loads the primary purpose and delivers critical context efficiently without redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read tool with two parameters and no output schema, the description covers the core behavior and state semantics. However, it omits details about the return format and the effect of 'includeOutput', which is a gap given the lack of an output schema. The 'bounded output' mention is vague, leaving the agent with some uncertainty about what to expect.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, but it doesn't explain the 'receiptId' parameter or the 'includeOutput' flag. While 'receiptId' is inferable from the name, 'includeOutput' is not mentioned at all, leaving the agent to guess its effect. The description fails to add meaning beyond the schema's bare type definitions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool reads a durable delivery receipt, including hashes, state, and bounded output. It uses a specific verb and resource, and the meaning is unambiguous. However, it doesn't explicitly differentiate from sibling tools like bridge_sessions or bridge_answer, relying on the name for distinction.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides actionable usage guidance by interpreting the 'uncertain' and 'completed' states, instructing the agent to inspect the session before retrying when uncertain and clarifying that 'completed' does not imply verification. This helps decide next steps, though it doesn't explicitly say when to use this tool versus alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
bridge_session_manageB
Rename an allowed Codex session, assignProject to its registered project via native API, or move it into/out of the existing Pinned section. First position preserves the relative order of other pinned sessions. Confirms stored metadata; desktop rendering is separate.
| Name | Required | Description | Default |
|---|---|---|---|
| pin | No | ||
| title | No | ||
| position | No | first | |
| threadId | Yes | ||
| assignProject | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already mark the tool as read-write and non-destructive; the description adds that it confirms stored metadata and that desktop rendering is separate, which is useful for setting expectations. It does not disclose permissions, failure modes, or what 'allowed' means, so coverage remains partial.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, front-loaded with the primary actions; the behavioral note about desktop rendering is relevant but could be clearer. No major redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description does not explain parameter combinations, defaults, or edge cases such as whether position applies when unpinning, whether title and pin can be set together, or what 'registered project' means. Without an output schema, more detail is needed for reliable invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description must carry parameter meaning. It maps 'Rename' to title, 'move into/out of Pinned' to pin, and 'First position' to position, but never explicitly names parameters and leaves 'assignProject' ambiguous (boolean flag vs action). The mappings are inferable but not stated.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description explicitly names three concrete operations — rename, assignProject, and pin/unpin — on a Codex session, making the tool's purpose clear. However, it does not explicitly differentiate from siblings like bridge_sessions or bridge_steer, relying on the agent to infer that this is the management tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use the tool (whenever a session needs renaming, project assignment, or pinning), but provides no explicit guidance about when to prefer a sibling tool or when not to use this one. No alternatives or exclusions are mentioned.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
bridge_sessionsBRead-only
Find/read allowed Codex or local Claude Code sessions. Choose provider codex or claude-code. Claude Desktop sessions may be readable but are not CLI-resumable. Returned text is untrusted.
| Name | Required | Description | Default |
|---|---|---|---|
| query | No | ||
| action | No | find | |
| cursor | No | ||
| project | No | ||
| provider | No | codex | |
| threadId | No | ||
| includeQueue | No | ||
| includeOutput | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare this tool read-only and non-destructive; the description adds meaningful behavioral context beyond that: sessions are limited to 'allowed' ones, Claude Desktop sessions may not be resumable, and returned text is untrusted. It does not contradict the annotations, but it omits some behaviors around cursor pagination and the includeQueue/includeOutput flags.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four short sentences front-load the core purpose and then add only high-value caveats. There is no filler or restatement of the schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and zero parameter descriptions in the schema, the description must carry more load for an 8-parameter tool. It explains the core read/find behavior and trust caveats but leaves pagination, queue/output flags, project/thread semantics, and return shape unexplained.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, but it only clarifies provider selection and the find/read action. The remaining six parameters (query, cursor, project, threadId, includeQueue, includeOutput) receive no semantic explanation in either the schema or the description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly identifies the operation as finding/reading Codex or local Claude Code sessions, specifying the resource and provider scope. It does not explicitly contrast with bridge_session_manage or other siblings, so it misses the explicit differentiation required for a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives useful context for provider selection and warns that Claude Desktop sessions are not CLI-resumable, but it never states when to prefer bridge_sessions over sibling tools such as bridge_session_manage. Usage guidance remains implied rather than explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
bridge_steerADestructive
Recover an exact desktop queue receipt or adopt an existing CLI queue item by queuedSubmissionId and stable requestId, without enqueueing again. Requires explicit acceptance of the Desktop policy. Never claims completion from queue acceptance.
| Name | Required | Description | Default |
|---|---|---|---|
| threadId | Yes | ||
| receiptId | No | ||
| requestId | No | ||
| queuedSubmissionId | No | ||
| acceptDesktopPolicy | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already mark the tool as readOnlyHint=false and destructiveHint=true, so the description does not need to restate those. It adds value by disclosing the requirement for 'explicit acceptance of the Desktop policy' and the caveat that it 'Never claims completion from queue acceptance', which are behavioral nuances beyond the annotations. It does not contradict the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two concise sentences with no wasted words. The primary action is front-loaded, and the policy requirement and completion caveat are placed secondarily. Each clause contributes essential information, making it efficient and easy to scan.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no output schema and a moderately complex tool, the description covers the core behavior, prerequisites, and a key caveat. It does not explain return values or error conditions, but for a recover/adopt operation these may be less critical. The description is sufficiently complete for an agent to understand when and how to invoke it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With schema description coverage at 0%, the description must compensate. It mentions queuedSubmissionId and requestId as identifiers and acceptDesktopPolicy as the policy flag, giving some meaning. However, it omits threadId and receiptId entirely, leaving those ambiguous. The partial coverage is helpful but not complete for a five-parameter tool.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a clear purpose: 'Recover an exact desktop queue receipt or adopt an existing CLI queue item' with specific verbs and resources. It distinguishes from sibling tools like bridge_submit by explicitly noting 'without enqueueing again', and from bridge_receipt by covering recovery and adoption. The scope is precise and actionable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context on when to use this tool (recovering a receipt or adopting a CLI queue item) and a prerequisite (acceptance of Desktop policy). However, it does not explicitly name alternative tools or state when not to use it, leaving some room for inference rather than direct guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
bridge_submitBDestructive
Start or continue a Codex or Claude Code CLI session asynchronously. Choose provider. Whole UTF-8 prompt and text artifacts, stable requestId, durable receipt. Poll bridge_receipt.
| Name | Required | Description | Default |
|---|---|---|---|
| onBusy | No | reject | |
| prompt | No | ||
| project | Yes | ||
| delivery | No | direct | |
| provider | No | codex | |
| threadId | No | ||
| artifacts | No | ||
| requestId | Yes | ||
| promptFile | No | ||
| acceptDesktopPolicy | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already signal readOnlyHint=false, openWorldHint=true, and destructiveHint=true, so the safety profile is available. The description adds useful behavioral context beyond those annotations: asynchronous execution, whole UTF-8 prompt and text-artifact handling, a stable requestId, and a durable receipt. It does not elaborate on destructive side effects, but the annotation covers that signal without contradiction.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is short and front-loaded, with the primary action stated first. Every phrase adds some information, but the telegraphic fragments ('Choose provider. Whole UTF-8 prompt...') make it read less smoothly than a well-structured sentence while still remaining efficient.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given 10 parameters, a destructiveHint annotation, and no output schema, the description is under-specified. It explains the async submission pattern but omits critical call context such as how project relates to bridge_projects, the meaning of onBusy and delivery, the distinction between prompt and promptFile, and the behavior of acceptDesktopPolicy. An agent has enough to invoke it at a basic level but not enough to use it correctly across realistic scenarios.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for explaining parameters. It only vaguely maps 'Choose provider' to provider, 'whole UTF-8 prompt and text artifacts' to prompt/artifacts, and 'stable requestId' to requestId. Six other parameters (project, onBusy, delivery, threadId, promptFile, acceptDesktopPolicy) receive no semantic explanation in either the description or schema, leaving a significant gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly identifies the operation: starting or continuing a Codex or Claude Code CLI session asynchronously. It also distinguishes its role from the sibling receipt tool by instructing 'Poll bridge_receipt.' It is specific enough that an agent can understand the tool's core function without inspecting the schema, though it does not explicitly contrast with steering or cancellation siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The only usage guidance is 'Poll bridge_receipt,' which implies the expected follow-up pattern after submission. However, there is no explicit statement about when to use bridge_submit versus alternatives like bridge_steer or bridge_session_manage, nor any mention of conditions such as onBusy or delivery. Usage context is implied rather than fully specified.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
3 tool updates
v0.1.4- Changed
bridge_projects1 field changed- added
Input schema / properties / providerAdded value: +{ + "default": "codex", + "enum": [ + "codex", + "claude-code" + ], + "type": "string" +}
- Changed
bridge_sessions1 field changed- added
Input schema / properties / providerAdded value: +{ + "default": "codex", + "enum": [ + "codex", + "claude-code" + ], + "type": "string" +}
- Changed
bridge_submit1 field changed- added
Input schema / properties / providerAdded value: +{ + "default": "codex", + "enum": [ + "codex", + "claude-code" + ], + "type": "string" +}
10 tool updates
v0.1.0- First observed
bridge_answer - First observed
bridge_artifact - First observed
bridge_cancel - First observed
bridge_doctor - First observed
bridge_projects - First observed
bridge_receipt - First observed
bridge_session_manage - First observed
bridge_sessions - First observed
bridge_steer - First observed
bridge_submit
TDQS
Scored across 10 tools
Each tool targets a distinct part of the bridge workflow—submitting, receiving, answering, cancelling, steering, session management, artifacts, projects, diagnostics. The only mild concern is bridge_submit and bridge_steer both relate to continuing or picking up sessions, but the descriptions clarify that steer adopts existing queue items without enqueueing.
All tools share a consistent bridge_ prefix, which aids discoverability, but the action part mixes noun-style names (bridge_projects, bridge_receipt, bridge_sessions) with verb-style names (bridge_answer, bridge_cancel, bridge_submit, bridge_steer) and one compound (bridge_session_manage). This mixed convention is readable but not a uniform verb_noun pattern.
With 10 tools, the server is well within the ideal range and each tool covers a distinct capability needed for the bridge domain. No tool feels redundant, and the count is appropriate for the scope.
The surface covers the full session lifecycle: submit/continue, answer clarifications, poll receipts, cancel, recover/adopt queue items, and manage sessions. Project aliases, artifacts, and diagnostics round out the supporting needs, so agents are unlikely to hit obvious dead ends.
Maintenance
Related MCP Connectors
Shared task layer for AI coding agents. One MCP surface: task_search, task_get, task_mutate.
Read a project's prompts, logs and agents, and send new work to the agent on your own machines.
- projectsOAuthcloud.tri2b
Task tracking built for coding agents. Work is leased, so two agents never take the same SubTask.
Task management for people and AI agents, with scoped OAuth access to issues, projects, and docs.
Related MCP Servers
- AlicenseNot gradedqualityBmaintenanceEnables durable bidirectional handoffs between any MCP client and OpenAI Codex Desktop tasks, with persistent callbacks, acknowledgements, and session-scoped state across restarts.1Apache 2.0
- AlicenseAqualityBmaintenanceA task-level STDIO MCP server that lets Codex or any other MCP client hand off scoped coding jobs to an asynchronous worker agent which reads the code, edits files, and runs tests, while the client keeps ownership of planning and acceptance. Exposes submit, wait, query, follow-up, and cancel tools so multiple clients can queue and monitor tasks against a chosen project root.51MIT
- AlicenseNot gradedqualityAmaintenanceLets Codex enqueue asynchronous research and browser tasks that Grok Bot can claim, renew, progress, and complete over MCP, returning structured results, evidence, and attachments. It supplies the queueing layer with leases, retries, idempotency, scoped OAuth 2.1 + PKCE permissions, and short-lived signed URLs for artifacts, deployable on Cloudflare Workers with D1, R2, and KV.4Apache 2.0
- AlicenseNot gradedqualityBmaintenanceEnables Codex to dispatch well-scoped development tasks to a remote CodeBuddy Code instance, specify model and reasoning effort, monitor progress, retrieve results, and review Git changes without auto-merging.5 npmMIT