codex-thread-bridge
This MCP server lets you create, inspect, and message Codex threads/tasks on a running Codex App Server, with deduplicated mutations and Git worktree support.
Check connection and report capabilities/compatibility limits (
get_capabilities).Create a retained Codex thread/session in an existing directory, optionally with title, prompt, model, sandbox, and App Server project ID (
create_thread).Create a retained, locked Git worktree and task at a full commit ID, with expected sandbox policy and bridge-managed retained mode (
create_worktree_thread).Send one message to an explicitly selected idle thread, resuming it if needed (
send_message_to_thread).List unarchived backend threads without loading them, with paging (
list_threads).Read thread metadata and newest-first turn history without resuming (
read_thread).Wait up to 50 seconds for a specific recent turn without resuming or interrupting (
wait_thread).Read persistent Goal state for a thread (
get_goal).Read a mutation receipt by
request_idto inspect known IDs after partial or uncertain delivery (get_operation).
Lets agents create, resume, and message OpenAI Codex sessions through the local Codex App Server, including isolated Git worktree threads, listing/reading thread history, waiting on turns, and inspecting goals and operation receipts.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@codex-thread-bridgeCreate a thread in /home/user/project with the prompt "Summarize the current git changes.""
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
codex-thread-bridge
An MCP server for creating, messaging, and steering Codex tasks through a running App Server on the same host.
Install
Requires Python 3.11+, uv, and an authenticated Codex App Server with a Unix WebSocket control socket. Git is required for worktree launches. The bridge runs as the same user and on the same host as App Server.
From a clone of this repository:
uv sync --frozen --no-dev
.venv/bin/codex-thread-bridge --help
codex mcp add codex-thread-bridge -- /absolute/path/codex-thread-bridge/.venv/bin/codex-thread-bridgeThe default socket is $CODEX_HOME/app-server-control/app-server-control.sock;
CODEX_HOME defaults to ~/.codex. Use --socket /path/to.sock to override it.
Authentication and model usage belong to the existing App Server.
On Windows, the bridge uses codex app-server proxy --sock as a raw WebSocket
byte relay. On Unix, it connects directly to the socket. Use --transport proxy
to select the relay explicitly, and --codex-binary /path/to/codex if the CLI
is not on PATH. The equivalent environment variables are
CODEX_THREAD_BRIDGE_TRANSPORT and CODEX_THREAD_BRIDGE_CODEX.
After changing bridge code or MCP configuration, request a refresh with:
uv run --locked codex-thread-bridge-reloadThe command accepts --socket, asks for y/N confirmation, and queues a refresh
for loaded tasks. Its response does not confirm that every task has refreshed.
Related MCP server: hato
Tools
Tool | Behavior |
| Report server identity and bridge capabilities |
| Create a retained task in an existing directory, with an optional title and initial prompt |
| Create a retained, locked Git worktree and a task at a specified commit |
| Apply and verify permissions for an idle task with an expected identity |
| Resume an idle task and start a turn without settings overrides |
| Append a message to the active turn identified by |
| List unarchived backend tasks without loading them |
| Read task metadata and paginated history without resuming |
| Wait up to 50 seconds for a specified recent turn |
| Read persistent Goal state |
| Read the retained receipt for a mutation request |
Tool schemas are exposed through MCP. Their definitions are in server.py.
Example
Call create_thread with these MCP arguments:
{
"request_id": "repository-overview-001",
"cwd": "/absolute/path/to/project",
"title": "Repository overview",
"prompt": "Summarize the project structure."
}Pass the returned threadId and turnId to wait_thread. Use
send_message_to_thread for an idle task or steer_thread for an active turn.
Each intentional new message requires a new request_id.
Behavior
create_threaddefaults to sandboxread-onlyand approval policynever. Expliciton-requestapproval requires App Server Auto-review. Omitted model and reasoning effort use server defaults. Creation rejects explicit network-enabled read-only policies; permission updates accept them.Permission updates require the task's expected identity and an idle task. The idle check and update are separate operations, so concurrent clients can race them. Workspace writable roots cannot be existing files, sockets, or devices.
Each mutation uses a stable
request_id. Repeating it with matching arguments returns the recorded receipt without repeating or continuing the operation.acceptedreports API acceptance; turn completion is reported bywait_thread. Failed or uncertain operations can leave tasks, turns, or worktrees behind;get_operationreturns the recorded outcome and known IDs.Receipts persist in
$XDG_STATE_HOME/codex-thread-bridge, defaulting to~/.local/state/codex-thread-bridge.--state-diroverrides this location. Deleting this state discards request deduplication history.Worktree launches require a full local commit ID and an absent, canonical absolute destination outside existing repositories. Worktrees are detached, locked, and retained until manual cleanup. Dirty files are not copied, and Git hooks and checkout filters are disabled. Approval policy is
never;expected_sandbox_policyverifies returned settings rather than applying overrides.Project IDs belong to App Server's registry. Desktop controls its own project association and task listing. Worktrees created by the bridge have a manual lifecycle rather than a Desktop-managed lifecycle.
New and resumed turns carry bridge instructions as
toolOutputwith the bridge tool's name. The bridge leaves client-side tool calls and approval requests unanswered so it cannot consume another client's shared callback. A capable client such as Desktop must be subscribed to the thread to handle them; the bridge does not establish that subscription or provide interactive approvals. Dispatch keeps the shared connection open so concurrent reads can finish.
Development
uv sync --frozen --group dev
uv run pytest
uv run ty check src
uv run ruff check .
uv run ruff format --check .
uv buildTests use a fake App Server and temporary Git repositories. Test temporary
directories must be outside existing repositories; pytest accepts --basetemp.
The source and tests define the detailed behavior. Contribution requirements are in CONTRIBUTING.md.
References
MIT licensed. Independent project, not affiliated with or endorsed by OpenAI.
Available Tools
9 toolscreate_threadADestructive
Create a retained session in an existing cwd, optionally with an initial prompt.
Requires approval of this action and sandbox. No worktree or persistent Goal is created. Approval policy is never. Omitted model/reasoning use configured defaults. Supply only an App Server project ID, never assume a Desktop saved-project ID is interchangeable. Returns actual settings and IDs; verify Desktop association separately. Reusing request_id returns its receipt without resending. A failed/unknown operation may have created a thread.
| Name | Required | Description | Default |
|---|---|---|---|
| cwd | Yes | ||
| model | No | ||
| title | No | ||
| prompt | No | ||
| sandbox | No | read-only | |
| request_id | Yes | ||
| app_server_project_id | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Goes well beyond the annotations by disclosing approval requirements, idempotent behavior on request_id reuse, and the possibility that a failed/unknown operation may still have created a thread. It also clarifies model/sandbox defaults and the need to verify Desktop association, all without contradicting readOnlyHint=false or destructiveHint=true.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The key purpose is front-loaded and every sentence adds a meaningful caveat, but a few phrases are cryptic (e.g., 'Approval policy is never') and could be clearer without adding length.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 7-parameter mutation tool with rich side effects, the description covers when to use it, what it does not create, idempotency, failure ambiguity, approval, and return-value caveats. The output schema supplies the return details, so the remaining gaps are minor.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, but the description compensates for most parameters: cwd must exist, prompt is optional/initial, model/reasoning fall back to defaults, request_id is reusable/idempotent, and app_server_project_id must not be a Desktop saved-project ID. A few parameter details (title and request_id origin/format) remain implicit, preventing a 5.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Opens with a specific verb-plus-resource: 'Create a retained session in an existing cwd, optionally with an initial prompt.' It distinguishes itself from the sibling worktree tool by explicitly stating 'No worktree or persistent Goal is created.'
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear context for use (existing cwd, retained session, no worktree/goal) and important invocation constraints (App Server project ID only, never Desktop saved-project ID). It does not explicitly name when to prefer create_worktree_thread or other siblings, relying instead on the exclusion clause.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_worktree_threadADestructive
Create a retained, locked Git worktree and a task at an approved full commit ID.
Requires approval of bridge-managed-retained ownership, source checkout root, immutable commit, absent absolute destination (existing parent), permissions and exact prompt. Creates a detached checkout and Git metadata; disables hooks/filters, copies no dirty files, runs no setup, sets no Goal. No automatic cleanup, archive or Desktop binding. Approval policy is never. expected_sandbox_policy is the complete expected response: e.g. {"type":"readOnly","networkAccess":false}. Actual settings and workspace roots must match before prompt dispatch. Omit prompt for readiness-only creation; caller owns further readiness and Goal policy. Model/reasoning defaults are preserved when omitted. Reuse request_id after uncertainty: receipts replay without continuing partial work. Known artifacts and recovery requirements are retained even on failure/cancellation.
| Name | Required | Description | Default |
|---|---|---|---|
| model | No | ||
| title | No | ||
| prompt | No | ||
| sandbox | Yes | ||
| request_id | Yes | ||
| destination | Yes | ||
| worktree_mode | Yes | ||
| reasoning_effort | No | ||
| source_repository | Yes | ||
| starting_revision | Yes | ||
| app_server_project_id | No | ||
| expected_sandbox_policy | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description goes far beyond the annotations, disclosing that hooks/filters are disabled, no dirty files are copied, no setup or Goal is set, no automatic cleanup/archive/Desktop binding occurs, and that request_id replays receipts without continuing partial work. This is rich behavioral context and fully consistent with the destructiveHint and openWorldHint annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense and information-rich rather than wasted. The main purpose is front-loaded, and each sentence adds meaningful detail about preconditions, side effects, idempotency, or parameter behavior. It could be slightly better organized into preconditions versus behaviors, but it is not bloated.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity—12 parameters, required approval gates, destructive annotations, and no schema descriptions—the description covers the critical ground: approval policy, expected sandbox policy matching, readiness-only usage, idempotent request_id behavior, artifact retention on failure, and what is intentionally not done. An output schema also exists, so return values do not need to be described.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It does add meaning for several parameters: starting_revision must be an approved full commit ID, destination must be absent with an existing parent, expected_sandbox_policy must match the full expected response, prompt can be omitted for readiness-only creation, and model/reasoning defaults are preserved when omitted. A few parameters like source_repository, sandbox, worktree_mode, and app_server_project_id are not individually explained, but most critical semantics are covered.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Create a retained, locked Git worktree and a task at an approved full commit ID.' This clearly distinguishes the tool from siblings like create_thread, which would not create a worktree. The scope is concrete and actionable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides strong usage context: it lists required approvals, prerequisites, the need for an exact prompt, and readiness-only creation by omitting the prompt. However, it does not explicitly name alternatives or state when not to use this tool versus create_thread or other siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_capabilitiesARead-only
Check connection and report implemented capabilities and compatibility limits.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already indicate a read-only, non-destructive operation, and the description adds useful behavioral detail about what the tool actually does: it performs a connection check and reports implemented capabilities and compatibility limits. This is consistent with annotations and adds context beyond them, though it does not disclose error handling or timeout behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, tightly written sentence with no filler. Every word contributes: 'check connection' states the action, and 'report implemented capabilities and compatibility limits' states the output scope.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter, read-only introspection tool with an output schema present, the description is complete. It tells the agent what the tool checks and what it reports, and the annotations cover the safety profile. No essential information is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so parameter semantics are trivial. With an empty schema and 100% coverage, there is nothing the description needs to clarify; the baseline for zero-parameter tools applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses specific verbs ('Check', 'report') and a clear resource: connection status, implemented capabilities, and compatibility limits. This makes the tool's purpose immediately obvious and distinguishes it naturally from the thread, goal, and operation siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly implies when to use this tool: when an agent needs to verify connectivity or discover the server's capabilities and compatibility limits. It does not explicitly list exclusions or alternatives, but none of the sibling tools serve this introspection role, so the context is clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_goalBRead-only
Read persistent Goal state without modifying it; text over 4000 characters is marked.
| Name | Required | Description | Default |
|---|---|---|---|
| thread_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description adds meaningful behavioral context beyond the annotations: 'text over 4000 characters is marked'. This is a useful detail that affects how an agent interprets results. The statement 'without modifying it' is consistent with the readOnlyHint annotation, and no contradiction exists.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The entire description is one concise sentence that front-loads the core purpose and then adds a single important behavioral detail. There is no filler or repetition of what annotations already state.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
This is a simple tool with one required parameter and a read-only safety profile already covered by annotations. An output schema exists, so the return format does not need to be described. The description covers the key behavioral nuance (4000-character marking) and leaves little missing for correct invocation, though it could be more explicit about how thread_id maps to a Goal.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 0% description coverage, and the description does not explain the thread_id parameter at all. The parameter name is self-explanatory to some degree, but the description provides no additional semantic meaning, such as what goal state is associated with the thread or whether thread_id refers to a specific goal thread. With low schema coverage, the description should compensate but does not.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'Read persistent Goal state'. It clearly identifies the tool as a read operation on Goal state, which distinguishes it from write-oriented siblings like create_thread and send_message_to_thread. However, it does not explicitly differentiate it from read_thread or get_operation, so the distinction is implied rather than stated.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no guidance on when to use this tool versus alternatives. It provides the purpose but no context about when read_thread, get_capabilities, or get_operation would be more appropriate. The usage is only implied by the tool name and description.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_operationARead-only
Read a mutation receipt, including known IDs after partial or uncertain delivery.
| Name | Required | Description | Default |
|---|---|---|---|
| request_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotations already mark this as read-only and non-destructive, so the safety profile is covered. The description adds useful context about the receipt containing known IDs under uncertain delivery, but it does not disclose lifecycle details, error behavior, or what happens when the receipt is unavailable.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single front-loaded sentence with no filler. Every clause earns its place by adding either the core action or the key use case.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-required-parameter read operation with an output schema and read-only annotations, the description gives sufficient context. It could be more explicit about how request_id relates to the mutation receipt, but the overall picture is complete enough for an agent to select and invoke the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has 0% description coverage and the description does not explain request_id beyond what the schema title already implies. It never clarifies that request_id is the mutation receipt identifier or how it is obtained, so the description fails to compensate for the schema gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Read') and a specific resource ('mutation receipt'), and the qualifier 'including known IDs after partial or uncertain delivery' makes the tool's scope clear. It is also distinguishable from sibling tools like read_thread and get_capabilities without needing to inspect their schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear invocation context: it is meant for reading a mutation receipt after partial or uncertain delivery. It does not explicitly name alternative tools or state when not to use it, so it falls just short of full guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_threadsARead-only
List unarchived backend threads without loading them. Project IDs are backend IDs.
| Name | Required | Description | Default |
|---|---|---|---|
| cwd | No | ||
| limit | No | ||
| cursor | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already mark this as read-only and non-destructive. The description adds meaningful context beyond annotations: it does not load threads, only lists unarchived ones, and warns that project IDs refer to backend IDs. This helps an agent understand the operation's lightweight nature and potential ID interpretation pitfalls.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two short sentences with no filler. The core action and scope are front-loaded, and the additional ID clarification is kept to a single relevant sentence. Every word earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists and annotations cover the safety profile, so those are not gaps. Still, the description omits guidance on cursor-based pagination, limit behavior, and explicit selection between this tool and read_thread. Adequate for a simple list tool but with clear gaps in parameter semantics and usage routing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for explaining cwd, limit, and cursor. It adds only the note that 'Project IDs are backend IDs,' which may clarify cwd but leaves limit and cursor completely unexplained. This is insufficient for a paginated list tool.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb ('List'), a specific resource ('unarchived backend threads'), and a behavioral qualifier ('without loading them'). This clearly distinguishes it from siblings like read_thread, send_message_to_thread, and create_thread.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'without loading them' implies this tool is for lightweight enumeration rather than fetching full thread content, which gives some usage context relative to read_thread. However, it does not explicitly name alternatives, state when not to use it, or explain when the unarchived filter matters.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
read_threadARead-only
Read metadata and one newest-first turn page, without resuming; truncation is marked.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| cursor | No | ||
| thread_id | Yes | ||
| max_text_chars | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint and non-destructive behavior, so the safety profile is covered. The description adds meaningful behavior beyond annotations: exactly one page, newest-first ordering, no resumption, and explicit truncation marking.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single dense sentence front-loads the verb and resource, then packs in ordering, paging behavior, resumption behavior, and truncation handling without filler. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema and read-only annotations, the description covers the core behavior well. However, the meaning of cursor and the exact relationship between 'without resuming' and pagination are left ambiguous, which is a noticeable gap for a 4-parameter tool with zero schema descriptions.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the description never names thread_id, limit, cursor, or max_text_chars. Phrases like 'turn page' and 'truncation is marked' weakly imply pagination and character limits, but the agent is left to infer the exact parameter roles from titles and defaults.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb ('Read') and resource ('metadata and one newest-first turn page'), and the 'without resuming' clause separates it from wait/resume-style siblings while 'Read' contrasts with create/send tools. It is compact but unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives a clear sense of scope ('one newest-first turn page, without resuming') but does not explicitly state when to prefer this over list_threads, wait_thread, or send_message_to_thread, nor does it name alternatives. Usage is implied rather than spelled out.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
send_message_to_threadADestructive
Resume the explicitly selected idle session without overrides and send one message.
Requires user authorization. Refuses an active thread and an interactive approval policy. Resume may load the session; its actual settings are returned. Does not steer, interrupt, set Goals, or retry delivery. Use a stable request_id; inspect get_operation on uncertainty.
| Name | Required | Description | Default |
|---|---|---|---|
| message | Yes | ||
| thread_id | Yes | ||
| request_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the annotations, the description discloses substantial behavioral detail: it requires user authorization, refuses active threads and interactive approval policies, may load the session, returns actual settings, does not retry delivery, and advises using a stable request_id. These details add real operational context without contradicting readOnlyHint=false, destructiveHint=true, or openWorldHint=true.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded with the primary purpose, followed by high-signal constraints and usage guidance. Every sentence carries operational weight, with no filler or repetition of schema information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with required parameters and an output schema, the description covers the full calling context: what the tool does, what preconditions it requires, what it refuses, what side effects it may have, and how to handle uncertainty. An agent has enough information to select and invoke this tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 0% description coverage, so the description must compensate. It does: thread_id maps to the 'explicitly selected idle session,' request_id is explained as needing to be stable and tied to get_operation for uncertainty, and message is understood as the single message being sent. A bit more per-parameter specificity would be ideal, but the key invocation semantics are present.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific action: 'Resume the explicitly selected idle session without overrides and send one message.' It clearly identifies the resource (thread session) and the operation (send message), and goes further to distinguish itself by stating what it does not do: 'Does not steer, interrupt, set Goals, or retry delivery.' This lets an agent separate it from sibling tools such as create_worktree_thread or wait_thread.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context for when to use the tool: resuming an explicitly selected idle session. It also provides exclusions, such as 'Refuses an active thread and an interactive approval policy,' and suggests falling back to get_operation when uncertain. It does not explicitly name alternative sibling tools, but the when-not conditions are strong enough to guide selection.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
wait_threadARead-only
Wait up to 50 seconds for a specific recent turn, without resuming or interrupting it.
Zero returns one snapshot. A timeout leaves the turn running. Only the latest 100 turns are inspected; use read_thread pagination for older turns. Completion can mean failure or interruption: inspect turn.status. The response never substitutes another turn.
| Name | Required | Description | Default |
|---|---|---|---|
| turn_id | Yes | ||
| thread_id | Yes | ||
| timeout_seconds | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Even with readOnlyHint=true and destructiveHint=false, the description adds rich behavioral detail: a timeout leaves the turn running, completion can mean failure or interruption, and the response never substitutes another turn. These details go well beyond what the annotations or schema convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded with the primary action, then adds precise caveats. Every sentence adds useful information, with no fluff or repetition of schema details.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The output schema exists, so return values need no explanation. The description covers the important edge cases: no resumption/interruption, timeout behavior, the 100-turn limit, status interpretation, and the guarantee against substituting another turn. This is complete for an agent to invoke the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema provides no descriptions for its 3 parameters, but the description compensates by explaining timeout behavior ('up to 50 seconds', 'Zero returns one snapshot', 'A timeout leaves the turn running') and scoping turn_id to the latest 100 turns. It does not explicitly walk through thread_id and turn_id, but their roles are reasonably clear from context.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb and resource: 'Wait up to 50 seconds for a specific recent turn.' It also clarifies what the tool does not do ('without resuming or interrupting it'), distinguishing it from sibling tools that operate on turns. The scope is further sharpened by noting it only inspects the latest 100 turns.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context and an explicit alternative: 'use read_thread pagination for older turns.' It does not enumerate all when-not-to-use cases, but the recent-turn limitation and read-only waiting behavior make the intended use clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
9 tool updates
v0.1.0- First observed
create_thread - First observed
create_worktree_thread - First observed
get_capabilities - First observed
get_goal - First observed
get_operation - First observed
list_threads - First observed
read_thread - First observed
send_message_to_thread - First observed
wait_thread
TDQS
Scored across 9 tools
Each tool targets a clearly distinct action: capabilities, thread creation variants, listing/reading/waiting, goal inspection, and operation receipt inspection. Even the two creation tools are sharply separated by the worktree behavior, and send_message_to_thread has no overlap with the create or read tools.
All tool names are lowercase snake_case and verb-led, following a consistent verb_noun structure such as create_thread, list_threads, read_thread, get_goal, and get_operation. The one compound name, send_message_to_thread, is still semantically aligned and does not introduce a different convention.
Nine tools is well within the ideal range for a bridge server. Each tool covers a distinct lifecycle or inspection concern without redundancy, and the count feels proportionate to the domain.
The core workflow is well covered: create threads, send messages, read/wait on turns, and inspect goals or operation receipts. Minor gaps exist around thread lifecycle management (no archive/delete) and goal mutation, but these appear intentionally left to external ownership and can be worked around.
Maintenance
Related MCP Connectors
Agent communication platform for agent to agent messaging via MCP. Messages, channels, skills.
Agent-to-agent messaging: directory, public lobby, DMs, channels, search. Stateless MCP + REST.
Agent-native collaboration network: orchestrate a team of long-running agents from any MCP client.
Join durable public agent discussions and invite-only private group rooms through MCP.
Related MCP Servers
- AlicenseAqualityDmaintenanceEnables Claude Code sessions to communicate with each other, allowing discovery, messaging, and synchronous queries across sessions.647 npmMIT
- AlicenseNot gradedqualityBmaintenanceEnables inter-session messaging for Claude Code, allowing sessions on different machines to send messages to each other, with delivery as user turns and support for offline queuing.209 npmMIT
- AlicenseNot gradedqualityBmaintenanceEnables managing native Codex and Claude Code sessions through a shared interface, allowing agents in one harness to create, read, continue, update, search, and delete sessions in the other.118 npmMIT
- AlicenseNot gradedqualityAmaintenanceEnables agents to create and drive Claude Code sessions inside a running Remote Control bridge, allowing a caller to spawn a session with a prompt and then survey or steer it from other sessions.96 npmMIT