workbuddy-subagent-bridge
This server is an MCP bridge that lets a main agent delegate coding, file modification, and diagnostic tasks to WorkBuddy sessions while managing permissions, reports, and feedback.
Create WorkBuddy tasks with explicit workspace directory, prompt, permission mode, model, and optional acceptance criteria (
workbuddy_session_start).Send follow-up messages to an existing session for iterative feedback and revisions (
workbuddy_session_send).Query task/session status and inspect pending ACP permission requests (
workbuddy_session_status).Reply to pending permission requests with the raw ACP response payload (
workbuddy_session_permission_reply).Read the in-memory final report of a completed task (
workbuddy_session_report).Cancel an active session (
workbuddy_session_cancel).Close the local ACP process and mark the task completed (
workbuddy_session_close).List all persisted tasks and their lifecycle status (
workbuddy_status).List available models from the WorkBuddy ACP session (
workbuddy_models).Run environment diagnostics for Node, CodeBuddy CLI, and ACP stdio (
workbuddy_doctor).Support multiple independent tasks with separate task IDs, sessions, working directories, permission modes, and reports.
Offer configurable permission modes ranging from read-only
plantobypassPermissions/fullAccess(dangerous, not default).Generate MCP configuration for Codex via
init --client codex(not a runtime MCP tool but a CLI capability).
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@workbuddy-subagent-bridgestart a plan-mode task to analyze this repo and report risks"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
WorkBuddy Subagent Bridge
WorkBuddy Subagent Bridge 是一个本地 MCP 桥接器,让主 Agent 通过 MCP 调用 WorkBuddy 执行任务。
主 Agent 负责理解目标、拆解任务、审批权限和验收结果;WorkBuddy 负责具体执行。任务完成后,主 Agent 可以读取报告,并继续向同一会话发送修改意见。
本项目是独立的社区工具,与 WorkBuddy 官方没有隶属关系。
适用场景
让主 Agent 把编码、文件修改或诊断任务交给 WorkBuddy。
管理多个独立的 WorkBuddy 任务,并分别查询状态和报告。
在 WorkBuddy 请求修改文件或执行操作时,由主 Agent 审批权限。
让同一 WorkBuddy 会话持续接收反馈,完成“分派—执行—验收—修订”流程。
当前公开版本为 0.1.0,已在 Windows 上完成真实验证。
Related MCP server: pi-delegate
安装
前置条件
Windows
Node.js 20 或更高版本
已安装并登录 WorkBuddy / CodeBuddy
当前验证环境:WorkBuddy 5.5.6、CodeBuddy CLI 2.137.1、Node.js 22.22.3。
交由 Agent 安装
复制下面的 Prompt 发送给 Agent 即可:
安装并配置 WorkBuddy Subagent Bridge。
项目仓库:https://github.com/BenjaminNH/workbuddy-subagent-bridge
先阅读并执行仓库中的 docs/agent-install.md,完成后返回安装和验证结果。从 npm 安装
npm install -g workbuddy-subagent-bridge
workbuddy-subagent doctordoctor 会检查 Node.js、npm、WorkBuddy、CodeBuddy CLI、ACP 连接和模型目录。开始使用前,请确认关键检查项均为 ok。
从本地源码安装
npm install
npm run build
npm pack
npm install -g .\workbuddy-subagent-bridge-0.1.0.tgz
workbuddy-subagent doctor配置 MCP
生成 Codex 配置文件:
workbuddy-subagent init --client codex --output .\workbuddy-mcp.jsoninit 只生成配置文件,不会自动修改 Codex 的全局配置。请将生成的 workbuddy-subagent-bridge 条目加入 Codex 的 MCP 配置,然后重启 Codex。
如果目标文件已存在,init 默认拒绝覆盖;确认需要替换时,显式添加 --force。
基本使用
配置完成后,主 Agent 可以创建任务:
请分析当前项目,实现指定功能,并在完成后报告修改内容和测试结果。典型流程:
使用
workbuddy_session_start创建任务。使用
workbuddy_session_status查询任务状态和待处理权限请求。使用
workbuddy_session_report读取最终报告。使用
workbuddy_session_send向同一会话发送反馈或修改要求。完成后使用
workbuddy_session_close关闭本地会话进程。
一个 MCP Bridge 可以管理多个独立任务。每个任务都有自己的 taskId、WorkBuddy session、工作目录、权限模式和报告。
权限模式
权限模式在创建任务时指定,默认是 plan。主 Agent 应根据任务范围选择模式。
模式 | 说明 |
| 只读分析和规划,默认模式 |
| 需要执行受限操作时逐项请求权限 |
| 自动接受文件编辑,其他操作仍受限制 |
| 由 WorkBuddy 自动判断风险 |
| 无法交互时拒绝需要询问的操作 |
| 由主 Agent 管理权限请求 |
| 跳过权限提示,风险较高 |
| 跳过全部权限检查,包括危险命令,风险极高 |
bypassPermissions 和 fullAccess 不会被默认使用,也不会被后续 session_send 自动提升。主 Agent 应在使用前确认工作目录和任务范围。
MCP 工具
工具 | 用途 |
| 创建 WorkBuddy 任务 |
| 向同一会话发送后续反馈 |
| 查询任务状态和权限请求 |
| 读取当前任务的最终报告 |
| 回复权限请求 |
| 取消任务 |
| 关闭本地会话进程 |
| 列出任务及其状态 |
| 查询模型目录、来源和更新时间 |
| 检查本机运行环境 |
安全边界和当前限制
工作目录必须由调用方明确提供,Bridge 不自动创建 worktree 或合并 diff。
默认使用
plan,高风险权限必须在任务级别显式指定。日志只记录基础生命周期信息,不记录 prompt、结果正文、文件内容或凭据。
ACP 是多轮任务的主路径;CLI 只用于一次性任务或 ACP 启动前失败的兼容场景。
ACP prompt 已提交但结果不确定时,不会自动切换 CLI 重放。
completed和cancelled是任务状态标签;如果 WorkBuddy 仍支持加载 session,可以继续发送反馈。报告保存在当前 Bridge 进程内。Bridge 进程重启后的无缝任务恢复不属于当前 MVP 的保证范围。
当前版本不提供远程公网 Agent 服务、A2A、多账号池、多供应商调度、自动 worktree、GUI 自动化或 OpenAI/Anthropic API 兼容层。
开发与验证
npm install
npm test
npm run typecheck
npm run build
npm run doctor问题反馈
提交问题时,请附带:
WorkBuddy 和 CodeBuddy CLI 版本
workbuddy-subagent doctor输出最小复现步骤
使用的权限模式
相关任务状态和原始错误字符串
许可证
本项目采用 MIT License。
Available Tools
10 toolsworkbuddy_doctorB
Check Node, CodeBuddy CLI, and ACP stdio initialization.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral disclosure burden. The verb 'Check' implies a read-only diagnostic actionchers, which is a meaningful behavioral signal. However, it does not state what happens on success/failure, whether any system state can be changed, or what output the agent should expect, leaving important behavior implicit.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, direct sentence with no filler or redundant wording. It front-loads the action and lists the specific resources being checked, making it easy for an agent to parse quickly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-argument diagnostic tool, invocation is trivial and the description is sufficient for initial selection. However, with no output schema, the agent is not told what kind of result to expect or how to interpret the check outcome. The phrase 'initialization' is somewhat vague, so the definition is adequate but not fully complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parametersched, so schema coverage is trivially 100% and there is no parameter meaning to add. The baseline of 4 for no-parameter tools applies here because the description and schema are consistent and no parameter documentation is needed.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb, 'Check', and names distinct resources: Node, CodeBuddy CLI, and ACP stdio initialization. This clearly communicates a diagnostic/health-check tool and differentiates it from the session-management siblings. It could be slightly stronger by stating the expected outcome (e.g., 'report environment readiness'), but the core purpose is understandable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no guidance on when to use this tool instead of siblings like workbuddy_status or workbuddy_session_status. There is no mention of troubleshooting contexts, prerequisites, or scenarios where this should not be used. The usage intent must be inferred entirely from the tool name and the word 'Check'.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workbuddy_modelsB
List models returned by the current WorkBuddy ACP session/new response.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. It only restates the basic listing action and does not clarify whether the tool is read-only, whether it requires an active session, whether it triggers new work, or what happens when no session exists.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single concise sentence with no filler or redundancy. It front-loads the primary action and resource, making it easy to parse quickly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the zero-parameter schema and simple listing intent, the description is close to sufficient, but it remains vague about session lifecycle and what 'new response' means. An agent would need to infer timing and valid call contexts from sibling tool names rather than from the description itself.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has zero parameters, so parameter semantics have minimal impact. The description adds a little context by indicating the models are tied to the current session/new response, but since there are no parameters to document, the baseline of 4 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb ('List') and resource ('models returned by the current WorkBuddy ACP session/new response'), which is more than a tautology and distinguishes it from session-action siblings. However, the phrase 'session/new response' is slightly ambiguous about whether it refers to an existing session's models or a newly returned response.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit guidance on when to call this tool versus alternatives such as workbuddy_session_status or workbuddy_session_send. The context is implied by sibling names, but the description does not state prerequisites like requiring an active session or that this should be called after a response is generated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workbuddy_session_cancelC
Cancel the active WorkBuddy session.
| Name | Required | Description | Default |
|---|---|---|---|
| taskId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It implies a destructive action ('cancel') but does not specify side effects, reversibility, authentication requirements, or what happens to the session. This is a significant gap for a mutation tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, short sentence with no filler words. It is front-loaded with the action. While it lacks detail, that is a completeness issue, not a conciseness issue. Every word earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple tool with one parameter but zero annotation coverage and no output schema, the description is severely incomplete. It fails to explain the taskId parameter, any preconditions, or what 'active' means. An agent cannot reliably call this tool with the provided information.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has one required parameter, taskId, with no description. Schema description coverage is 0%, so the description must explain what taskId refers to. It does not, leaving the agent to guess what value to provide, which is critical for correct invocation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb 'Cancel' and the resource 'active WorkBuddy session', which is specific. However, it does not distinguish from the sibling 'workbuddy_session_close', which likely has a similar or overlapping purpose. An agent might not know when to use this tool vs the close tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus alternatives like 'close' or 'send'. There are no conditions, prerequisites, or exclusions mentioned. The description only states the action without any context on when it is appropriate.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workbuddy_session_closeB
Close the local ACP process and mark the task completed.
| Name | Required | Description | Default |
|---|---|---|---|
| taskId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Describes the core effect (closing the process and marking the task completed) but does not disclose side effects, reversibility, or error states. With no annotations provided, the description carries the burden and only partially covers it.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Single sentence with no redundant wording. The verb and object are front-loaded, and every word earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Adequate for a simple close operation but lacks usage context and parameter clarification. It does not explain how it relates to other session tools or what state the process should be in, leaving some ambiguity for an agent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Does not explain the taskId parameter beyond its existence. Since schema coverage is 0%, the description should clarify that taskId identifies the task to close, but it omits this entirely, leaving the agent to infer its role.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States exactly what it does: closes the local ACP process and marks the task completed. This clearly distinguishes it from sibling tools like session_start or session_send, and the verb-resource pair is specific.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Offers no guidance on when to use this tool versus alternatives like session_cancel or session_send. It does not mention prerequisites, ordering, or conditions that would make this the correct choice.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workbuddy_session_permission_replyA
Reply to a pending ACP permission request. Inspect workbuddy_session_status first and pass the raw response payload expected by the current WorkBuddy version.
| Name | Required | Description | Default |
|---|---|---|---|
| taskId | Yes | ||
| response | Yes | Raw ACP permission response payload; do not invent a format if the current WorkBuddy request provides an explicit choice schema. | |
| requestId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It discloses that a raw response payload is required and advises not to invent a format if an explicit schema is given. It also mandates a status check. However, it does not mention side effects (e.g., whether the reply resolves the pending request), error conditions, or return behavior. This is adequate but not fully transparent for a mutation-like action.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with no redundancy. The first sentence states the core purpose, and the second provides essential guidance (check status, use raw payload). It is front-loaded and every word earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With 3 required params, no output schema, and no annotations, the description should provide more context. It offers a prerequisite and payload guidance but omits where to obtain taskId and requestId, what constitutes a valid response, how errors are surfaced, and the effect of a successful reply. It is sufficient for a simple call but leaves an agent inferring missing details.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 33% (only 'response' has a description). The description adds context for 'response' (raw payload, version-specific) but does not elaborate on 'taskId' or 'requestId' beyond their names, which are likely identifiers from the status. The description partially compensates for the low schema coverage but does not fully explain how these IDs relate to the pending request.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Reply') and a specific resource ('a pending ACP permission request'). It clearly distinguishes this from sibling tools like workbuddy_session_send (for general messages) and workbuddy_session_status (for inspection) by specifying the permission-reply context. The purpose is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly instructs to 'Inspect workbuddy_session_status first', providing a clear prerequisite and ordering. It implies this tool is only for pending permission requests, but does not explicitly state when not to use it or name alternative tools. Still, the context is clear enough for an agent to know when to invoke it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workbuddy_session_reportA
Read the in-memory report produced by a completed prompt. Reports are not persisted in the task registry.
| Name | Required | Description | Default |
|---|---|---|---|
| taskId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It usefully discloses that the report is in-memory and not persisted, and 'Read' implies non-mutation. However, it does not cover failure cases (e.g., incomplete prompt or unknown taskId) or whether reading affects the report.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tightly written sentences, with the core purpose front-loaded and the critical limitation in the second sentence. No filler or redundant restatement of the tool name.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no annotations and no output schema, the description should say a bit more about report availability/format. It covers purpose and a key storage caveat, but an agent must infer what a successful response contains and when the report may be absent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, and the description never explicitly defines taskId; it only implies via 'completed prompt' that the taskId identifies that prompt. The single parameter's name is conventional and self-explanatory, so an agent can infer it, but the description should have stated the mapping directly.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Read') and resource ('in-memory report produced by a completed prompt'), and is clearly distinct from the sibling session lifecycle tools (start/send/status/cancel/close/permission). The non-persistence note further sharpens what this tool does.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides clear usage context: the tool is for reports after a prompt has completed, and 'not persisted in the task registry' warns against expecting later retrieval. It does not explicitly name alternatives or when-not-to-use conditions, but the context is sufficient.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workbuddy_session_sendC
Send a follow-up message to the same WorkBuddy session.
| Name | Required | Description | Default |
|---|---|---|---|
| taskId | Yes | ||
| message | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries full responsibility for disclosing behavioral traits. It only states the action without revealing side effects (e.g., message delivery, session state changes), requirements (active session), or error behavior. The agent is left guessing about potential mutations or constraints.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single concise sentence with no filler, front-loaded with the action. It is appropriately short, though it could add a bit more context without becoming verbose, hence not a 5.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with no annotations and no output schema, the description is insufficiently complete. It lacks parameter explanations, usage context, and behavioral expectations, making it inadequate for an agent to reliably select and invoke the tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate by explaining taskId and message. It does not mention either parameter, leaving their meaning and purpose completely undocumented. This is a critical gap for a tool with only two required parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('send') and a resource ('follow-up message to the same WorkBuddy session'), which clearly indicates the action and distinguishes it from start, cancel, close, and status. However, it does not elaborate on what 'follow-up' implies or contrast with permission_reply, leaving some ambiguity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage within an existing session ('same WorkBuddy session') but provides no explicit guidance on when to use it versus alternatives, nor when not to use it (e.g., when no session exists). No exclusions or prerequisites are mentioned.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workbuddy_session_startA
Start a current-version WorkBuddy session. ACP mode returns a taskId immediately by default; poll status and request the report after completion. Set waitForCompletion=true when a blocking report is preferred.
| Name | Required | Description | Default |
|---|---|---|---|
| cwd | Yes | Explicit existing workspace directory. | |
| prompt | Yes | ||
| backend | No | auto prefers ACP and may use one-shot CLI fallback only before a prompt is submitted. | auto |
| modelId | No | Optional model ID. Local validation fallback tries hy4-preview then deepseek-v4.1-flash when omitted. | |
| permissionMode | No | WorkBuddy permission mode. Default plan: read/analyze only. acceptEdits: auto-accept file edits. bypassPermissions: skip permission prompts (danger). fullAccess: skip all permission checks including dangerous commands (extreme danger). Other values pass through to the current WorkBuddy version. | plan |
| waitForCompletion | No | If true, wait for the first prompt report; otherwise return the taskId immediately. | |
| acceptanceCriteria | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description must carry behavioral disclosure. It discloses the asynchronous default (taskId returned immediately, poll status for report) and the blocking alternative. However, it does not mention potential side effects, permission implications, or the fact that it initiates a potentially long-running process. The schema mentions permissionMode dangers but the description does not reinforce this.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences with no filler. It front-loads the core purpose and then provides essential behavioral guidance. Every word adds value.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a session-start tool with 7 params and no output schema, the description covers the key workflow aspects (immediate vs blocking, polling). It does not explain the return format beyond taskId, but that is implied. It also does not mention prerequisites like cwd existence, though that is in the schema. Overall, it is adequately complete for an agent to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 71%, above the 50% threshold, so baseline is 3. The description adds workflow context (poll status, request report) but does not provide extra meaning for individual parameters beyond what the schema already documents. The schema already covers backend, permissionMode, and waitForCompletion well.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb 'Start' and the resource 'WorkBuddy session', making the purpose unambiguous. It also distinguishes itself from sibling tools like send/status/cancel by describing the session initiation behavior.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It provides clear guidance on the two usage modes (default immediate return vs. waitForCompletion=true for blocking) and the associated polling workflow. However, it does not explicitly contrast with sibling tools (e.g., when to use send instead of start), so it lacks explicit exclusions or alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workbuddy_session_statusA
Read task/session status and any pending ACP permission requests without sending a prompt.
| Name | Required | Description | Default |
|---|---|---|---|
| taskId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral burden and does a good job: it explicitly marks the operation as a read and discloses the key side-effect-free trait of not sending a prompt. It does not elaborate on return shape or error cases, but for a status read that is a minor gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, front-loaded sentence that states the operation, the object, and the key qualifier with zero filler. Every word earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a low-complexity, one-parameter read tool, the description covers the essential invocation facts: what is read, what is returned conceptually, and that no prompt is sent. It omits an explicit return format and prerequisites, but the absence of an output schema and simple scope keep this from being a major gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema provides no description for taskId (0% coverage), and the tool description never explicitly defines it. 'task/session status' implies taskId identifies a session, but no format, ownership, or lifecycle context is given, so the description does not compensate for the empty schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb ('Read'), a concrete resource ('task/session status and any pending ACP permission requests'), and adds a distinguishing qualifier ('without sending a prompt'). This clearly separates it from action-oriented siblings like workbuddy_session_send or workbuddy_session_cancel.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'without sending a prompt' implies this is the safe status-checking alternative to session_send, but the description never states when to prefer it over workbuddy_status or the session action tools. No explicit when-to-use or when-not-to-use guidance is provided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workbuddy_statusA
List persisted WorkBuddy tasks and their current lifecycle status.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden of behavioral disclosure. 'List' conveys a read-only query and 'persisted' clarifies that it reads durable state rather than live sessions, but the description does not mention return format, ordering, or whether any filtering applies.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single front-loaded sentence with no filler. Every word contributes to identifying the resource and the kind of result returned.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a parameterless, read-only list tool with no output schema, the description is largely sufficient: the agent knows what action to take and what resource is involved. It loses one point because it does not explicitly distinguish task status from session status among the closely related sibling tools.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so there is nothing for the description to add beyond what the schema already conveys. The baseline of 4 applies because parameter semantics are not a meaningful gap here.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('List'), a defined resource ('persisted WorkBuddy tasks'), and a clear output focus ('current lifecycle status'). It is unambiguous and can be distinguished from siblings by resource type, though it does not explicitly name workbuddy_session_status as the session-focused alternative.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is given about when to use this tool versus the sibling workbuddy_session_status or other session tools. The description implies it is for querying persisted tasks, but it does not state exclusions or point to alternatives for session-level status.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
10 tool updates
v0.1.0- First observed
workbuddy_doctor - First observed
workbuddy_models - First observed
workbuddy_session_cancel - First observed
workbuddy_session_close - First observed
workbuddy_session_permission_reply - First observed
workbuddy_session_report - First observed
workbuddy_session_send - First observed
workbuddy_session_start - First observed
workbuddy_session_status - First observed
workbuddy_status
TDQS
Scored across 10 tools
Most tools target distinct session actions (start, send, cancel, close, permission reply, report) and are easy to tell apart. The only potential confusion is between workbuddy_session_status (per-session state) and workbuddy_status (persisted task registry), which are related but serve different purposes.
All tools share the workbuddy_ prefix and mostly follow a predictable pattern: workbuddy_session_<verb> for session operations, plus standalone nouns (status, models, doctor). The mix of verb-based session actions with noun-only utility tools is a minor deviation, but the overall convention is clear.
Ten tools is well within the ideal range for a session-bridge server and each one maps to a distinct lifecycle stage or utility concern. No tool feels redundant or unnecessary for the stated purpose.
The session lifecycle is well covered: start, send, status, cancel, close, permission reply, and report. The persisted task listing, model enumeration, and environment diagnostic round out the surface. Minor gaps exist, such as no explicit session resume or task-triggered report fetch, but agents can generally work around these.
Maintenance
Related MCP Connectors
- DazbenchOAuthapp.dazbench
Task management your AI agents can actually run. One line becomes a context-ready task over MCP.
Agent-first task marketplace MCP — discover, claim, and deliver paid workspace tasks.
Discover and call AI agents via MCP. Supports A2A agents and platform agents with async tasks.
Work management where AI agents are first-class members: tasks, projects, memory over hosted MCP
Related MCP Servers
- AlicenseNot gradedqualityBmaintenanceEnables MCP clients like Claude Code to delegate coding tasks to the local Cursor Agent CLI, with persistent per-workspace sessions that resume across calls.12 npmMIT
- AlicenseNot gradedqualityBmaintenanceEnables Claude Code to delegate coding implementation and testing to a local pi agent via MCP tools, including dispatching task books, monitoring status, steering or aborting runs mid-execution, and retrieving results and transcripts.37 npm4MIT
- AlicenseAqualityBmaintenanceEnables AI coding agents like Claude Code, Codex, Cursor, and OpenCode to delegate tasks to the WorkBuddy CLI as a sub-agent via an MCP tool, with one-command setup and configurable model/permission options.193 npm1MIT
- AlicenseAqualityAmaintenanceEnables orchestrator agents to route coding tasks to native CLI worker agents via MCP, providing tools for agent discovery, task submission, waiting, inspection, and cancellation.53MIT