workbuddy-delegate
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@workbuddy-delegateSummarize these meeting notes and flag action items for review."
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
WorkBuddy Delegate
Community / experimental · Windows-first · 2026-09-23 reliability update
WorkBuddy Delegate is a local Codex plugin that sends bounded, low-risk text work to an existing WorkBuddy installation, then returns a small, inspectable draft for Codex to verify. It is designed for extraction, summarization, classification, translation, rewriting, and code drafts—not autonomous decisions or external actions.
中文:这是一个 Windows 优先的 Codex 社区插件,将边界清晰、低风险的文本工作交给本机已有的 WorkBuddy。主任务仍负责核验与最终判断。
工作流程与 Token 目标
flowchart LR
U["用户任务:多份长文本"] --> C{"Codex 判断任务"}
C -->|"低风险、边界明确的文本工作"| P["插件检查文件范围、权限与调用限制"]
C -->|"需要关键判断或执行操作"| M["Codex 直接处理"]
F["明确指定的 UTF-8 文件"] --> P
P -->|"检查通过"| W["WorkBuddy 阅读原文并生成草稿"]
P -->|"检查失败"| M
W --> R["精简答案 + 原文引文 + 不确定项"]
R --> V["Codex 按需核验引文并作最终判断"]
F -.->|"必要时抽查"| V
V --> O["交付结果"]
M --> O**设计目标:**让 WorkBuddy 承担简单但耗费大量阅读上下文的提取、摘要、分类、翻译和改写工作。这样,Codex 主任务通常只需接收任务说明、精简草稿和核验所需的引文,而不必把全部文件内容放入自己的上下文。WorkBuddy 本身仍会消耗模型用量;实际节省多少 Codex Token、总成本和质量,需要用同一批任务做配对测量,不能由流程图保证。
Related MCP server: lmstudio-mcp
What it does
Packages a Codex skill and a dependency-free local MCP server.
Reads only explicit UTF-8 files beneath user-managed
allowed_roots.Blocks common credential/config paths and strips secret-like inherited environment variables.
Applies daily call, input-size, timeout, cache, and retention limits.
Uses no copied API key: authentication stays in WorkBuddy.
Stores small local result artifacts for inspection and labels them
needs_main_agent_review.
It does not intercept every prompt before Codex sees it, guarantee savings or equal quality, edit files, run worker tools, or publish/send anything.
Requirements
Windows 10/11
A current local Codex build with Agent Plugins v1 support (tested on
codex-cli 0.155.0-alpha.2.6)Python 3.10+
Node.js available on
PATHWorkBuddy installed and signed in
The adapter uses WorkBuddy's bundled CodeBuddy CLI, which is not a documented stable public API. A WorkBuddy update may require a compatibility update here.
Install from GitHub
codex plugin marketplace add renyuxin6666-tech/codex-workbuddy-delegate --ref main
codex plugin add workbuddy-delegate@renyuxin-toolsRestart the Codex desktop app and open a new task. Then ask:
Show WorkBuddy Delegate status and its manager command.
The first file-based delegation needs an allowed root. Preserve the complete returned manager command (including configuration and state paths), then append the subcommand:
python "<installed-plugin>\scripts\manage.py" --config "<status.config_path>" --state-dir "<status.state_path>" roots add "D:\path\to\your\project"
python "<installed-plugin>\scripts\manage.py" --config "<status.config_path>" --state-dir "<status.state_path>" self-testProvided text can be dry-run without an allowed file root. No manager command reads or prints API keys.
Example prompts
“Check whether summarizing these five notes is safe to delegate. If yes, use WorkBuddy and verify the cited excerpts.”
“Show WorkBuddy Delegate status and this week's local invocation count.”
“Use WorkBuddy for a first-pass translation of these allowed Markdown files; keep the final edit in Codex.”
Management
python plugins/workbuddy-delegate/scripts/manage.py doctor
python plugins/workbuddy-delegate/scripts/manage.py config init
python plugins/workbuddy-delegate/scripts/manage.py config set model auto
python plugins/workbuddy-delegate/scripts/manage.py config set daily_call_limit 10
python plugins/workbuddy-delegate/scripts/manage.py roots add "D:\project"
python plugins/workbuddy-delegate/scripts/manage.py usage --days 7
python plugins/workbuddy-delegate/scripts/manage.py cache pruneBy default, portable installs place config and runtime state under the plugin's managed ${PLUGIN_DATA} directory. Direct script use falls back to %LOCALAPPDATA%\WorkBuddyDelegate.
The legacy Codex MCP launcher now requires either PLUGIN_DATA or both explicit
WORKBUDDY_DELEGATE_CONFIG and WORKBUDDY_DELEGATE_STATE_DIR environment
variables. It will stop instead of silently opening the direct-script fallback.
For a headless run, point both explicit variables at the paths returned by the
native workbuddy_status; do not copy secrets or replace the allowed-root list.
After setup, call the native workbuddy_status and verify that its allowed_roots
and model match the changes. If the manager and MCP disagree, they are reading
different configuration paths; use the explicit path arguments above. Do not
remove the allowlist or disable the sandbox to fix this. Shell fallback commands
may need scoped host approval to access WorkBuddy's installation and login.
To require approval for real delegation while auto-approving read-only status tools, use Codex's plugin-scoped MCP policy. See the current OpenAI plugin packaging documentation because config keys may evolve.
Routing boundary
Good candidates:
First-pass summaries of several non-sensitive notes
Exact-field extraction with source quotations
Low-stakes classification or translation batches
Rewrite variants and unapplied code drafts
Keep in Codex:
Final research claims, evidence adjudication, or literature novelty
Medical, legal, financial, admissions, safety, or employment decisions
Secrets, credentials, hidden config, or unrestricted filesystem reading
File changes, deployments, messages, purchases, or other external actions
Tasks whose authority or currentness is ambiguous
Evaluation snapshot
To test actual Token efficiency and quality, see the preregistered paired experiment plan. It separates Codex usage, WorkBuddy usage, blind quality scores, failures, and optional lower-tier Codex subagents. The 2026-09-23 preflight stopped before scored A/B runs; the repair follow-up fixed path routing and hardened JSON handling, but a live check now reports a possible WorkBuddy login issue. Status remains inconclusive, with no savings result claimed.
The bounded prototype scored 93.18/100 on one synthetic 22,049-character authority/version extraction fixture, but failed the strict completeness gate because it captured only 6/11 superseded historical relations. It captured current constraints 5/5, unresolved issues 1/1, and exact evidence 6/6, with no accepted prompt-injection instruction. This supports current-state extraction with Codex review—not complete historical auditing or generalized performance. See BENCHMARK.md.
Security and privacy
Read PRIVACY.md and SECURITY.md before adding sensitive projects. Results and quoted excerpts are stored locally until retention or pruning removes them. Source content is sent to the model/provider configured in WorkBuddy.
Development
python -m unittest discover -s tests -v
python plugins/workbuddy-delegate/scripts/manage.py self-test
python C:\Users\<you>\.codex\skills\.system\plugin-creator\scripts\validate_plugin.py plugins\workbuddy-delegate
python C:\Users\<you>\.codex\skills\.system\skill-creator\scripts\quick_validate.py plugins\workbuddy-delegate\skills\workbuddy-delegateStatus
This GitHub release is a local/repo marketplace plugin. It is not an official WorkBuddy integration and is not listed in the universal OpenAI Plugins Directory; public directory submission normally expects a public HTTPS MCP endpoint, while this plugin must access the user's local WorkBuddy installation.
MIT licensed. WorkBuddy, CodeBuddy, Codex, OpenAI, and their marks belong to their respective owners. This community project is not affiliated with or endorsed by them.
Local dispatch reliability update
The manager command returned by native status includes the exact config and state paths;
keep these arguments when adding an approved project root. Terminal defaults may refer
to a different configuration. dry_run now verifies CLI discovery, runtime-directory
write access and the usage ledger without calling a model or consuming invocation quota.
It does not verify login or network connectivity. Permission errors return scoped repair
instructions; CLI failures return sanitized diagnostic categories, never raw output.
After reinstalling, start a new Codex task to load the updated MCP process.
Available Tools
5 toolsdelegate_to_workbuddyA
Send one bounded low-risk text task to the user's installed WorkBuddy. Use explicit relative files or short provided text. Supported kinds: summarize, extract, classify, translate, rewrite, code_draft. Only risk=low is routed. The main agent must verify the returned result.
| Name | Required | Description | Default |
|---|---|---|---|
| kind | Yes | ||
| risk | Yes | ||
| text | No | ||
| files | No | ||
| dry_run | No | ||
| workspace | Yes | ||
| instruction | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already indicate readOnly=false and destructive=false. The description adds valuable behavioral context: it sends a task to an external WorkBuddy, only routes risk=low, and requires the main agent to verify the returned result. This goes beyond the annotations and discloses a key operational requirement.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is tight and well-structured. It front-loads the core purpose, then enumerates kinds, constraints, and the verification note. Each sentence adds necessary information without fluff, and the list of supported kinds is presented clearly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given 7 parameters, no output schema, and annotations that only cover safety hints, the description is incomplete. It does not explain the required workspace parameter, the dry_run behavior, or the format of the returned result. The existence of a sibling workbuddy_read_result suggests the result retrieval mechanism should be mentioned, but it is not.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description must compensate. It explains kind via the supported list, risk via 'Only risk=low', and text/files via 'explicit relative files or short provided text'. However, it does not explain the workspace (required), dry_run, or the instruction length constraint beyond 'bounded', leaving gaps for those parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action ('Send one bounded low-risk text task to the user's installed WorkBuddy') with a clear resource and scope, and enumerates supported kinds. This clearly distinguishes it from siblings like workbuddy_status or workbuddy_plan, which handle different concerns.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives clear context: use only for low-risk, bounded text tasks, and it notes the main agent must verify the result. However, it does not explicitly reference sibling tools or state when NOT to use it, so exclusions are absent but the scope is well-defined.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workbuddy_planBRead-onlyIdempotent
Classify eligibility for bounded delegation without reading files or sending a model request.
| Name | Required | Description | Default |
|---|---|---|---|
| kind | Yes | ||
| risk | Yes | ||
| instruction | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already indicate readOnlyHint, idempotentHint, and non-destructive behavior. The description adds a more specific behavioral constraint: it performs no file reads and no model request. This goes beyond the generic annotations and helps the agent understand the tool's cost and side-effect profile. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence with no filler. It immediately conveys the core action and key constraints, making it easy for an agent to parse quickly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and no parameter documentation, the description leaves the agent unable to determine what kind or risk values to provide or what the classification result looks like. It is a useful high-level summary but insufficient for correct invocation of a three-required-parameter tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate by explaining instruction, kind, and risk. It does not define any of these parameters, their valid values, or how they map to eligibility. The phrase 'bounded delegation' provides only a vague hint about their purpose.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action: classify eligibility for bounded delegation. It distinguishes itself from siblings by explicitly noting it does not read files or send a model request, separating it from workbuddy_read_result and delegate_to_workbuddy. However, 'bounded delegation' is not defined, leaving some ambiguity about the exact scope.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies this is a pre-delegation eligibility check, especially given the sibling delegate_to_workbuddy. The phrase 'without reading files or sending a model request' suggests it is a lightweight planning step. Yet it never explicitly states when to use this tool versus alternatives or what conditions would make it inappropriate.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workbuddy_read_resultB
Read a plugin-created result artifact when its short preview was truncated.
| Name | Required | Description | Default |
|---|---|---|---|
| artifact | Yes | ||
| max_chars | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
There are no annotations to cover safety or side-effect expectations, so the description carries the full burden. It says the tool 'reads' an artifact, but it does not disclose what is returned, whether the output is truncated, how max_chars affects behavior, or any error conditions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence with no filler. Every word contributes to the core purpose and the triggering condition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no annotations and no output schema, the description should provide enough context to call the tool correctly. It explains when to call it but leaves key operational details unstated: what an artifact identifier looks like, what the tool returns, and how max_chars should be used.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description does not compensate. It does not explain what 'artifact' refers to, how to obtain its value, or what 'max_chars' controls beyond the schema's default and bounds.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description identifies a specific verb ('Read'), a specific resource ('plugin-created result artifact'), and a precise trigger ('when its short preview was truncated'). This clearly distinguishes it from the sibling tools, which are about status, planning, delegation, and usage rather than reading artifacts.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides a clear usage context: use this tool when a short preview was truncated. However, it does not explicitly mention alternatives or state when not to use it, so it falls short of full exclusionary guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workbuddy_statusARead-onlyIdempotent
Check the local WorkBuddy bridge, configured limits, and installation without sending a model request.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already mark the operation as read-only, idempotent, and non-destructive; the description adds important non-obvious behavior by stating that no model request is sent and that it only inspects local state. This goes beyond what the annotations alone convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single sentence that front-loads the action and object, then adds the key behavioral caveat. No filler, no repetition of the schema or annotations.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter status tool, this is complete: it says what is checked (bridge, limits, installation), what is guaranteed (no model request), and annotations cover the safety profile. The absence of an output schema is acceptable because the named check items imply what the result will report.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so parameter-level semantics are not a concern. The baseline of 4 applies because the description cannot add more meaning in this dimension.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Check') and identifies concrete resources: the local WorkBuddy bridge, configured limits, and installation. The closing constraint 'without sending a model request' further scopes it and distinguishes it from delegation tools like delegate_to_workbuddy.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use it: whenever you need local status, limits, or installation info without incurring a model request. It does not explicitly name sibling alternatives or state exclusions, so it stops short of the strongest guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workbuddy_usageA
Show local invocation counts. This does not call WorkBuddy or estimate money.
| Name | Required | Description | Default |
|---|---|---|---|
| days | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Since no annotations are provided, the description carries the behavioral burden. It transparently states that the tool does not call WorkBuddy or estimate money, which are important non-obvious behavioral constraints. It does not elaborate on output format or side effects, but for a local read-style display tool this is reasonably transparent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, with the main action first and the key boundary ('does not call WorkBuddy or estimate money') immediately after. No wasted words and no redundancy with the schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple one-parameter tool, the core purpose and key non-behaviors are covered. However, the role of 'days' is left to inference and no output shape is described, so an agent may not fully understand what it will receive or how to adjust the time window.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description should compensate for the undocumented 'days' parameter, but it does not mention it at all. The schema's default/min/max provide some constraints, yet the description fails to explain whether 'days' is a lookback window or how it affects the counts.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Show') and resource ('local invocation counts'), making the tool's function immediately clear. It also implicitly distinguishes itself from siblings like delegate_to_workbuddy by clarifying it does not call WorkBuddy.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description conveys that this is for local usage counts and not for executing WorkBuddy or estimating costs, but it never explicitly names alternatives or states clear when-to-use versus them. Usage context is implied rather than fully specified.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
5 tool updates
v0.1.0- First observed
delegate_to_workbuddy - First observed
workbuddy_plan - First observed
workbuddy_read_result - First observed
workbuddy_status - First observed
workbuddy_usage
TDQS
Scored across 5 tools
The tools are mostly distinct: status, plan, delegate, read_result, and usage each serve a clear purpose. However, 'workbuddy_plan' (classify eligibility) and 'workbuddy_status' (check limits) could be slightly confused as both are pre-flight checks, but descriptions clarify.
All tools share the 'workbuddy_' prefix and use a consistent verb_noun pattern (status, plan, delegate_to, read_result, usage). 'delegate_to_workbuddy' is slightly more verbose but still fits the pattern, so minor deviation.
With 5 tools, the server is well-scoped for its purpose of managing delegation to WorkBuddy. Each tool covers a distinct aspect: status, planning, delegation, result reading, and usage tracking, which is appropriate.
The surface covers the main lifecycle: check status (status), eligibility (plan), delegate, read result, and usage. A minor gap is the lack of a cancel/abort tool for pending delegations, but the bounded tasks are short-lived, so this is not critical.
Maintenance
Related MCP Connectors
AI work orchestration for plans, tasks, teams, and coding-agent dispatch.
Lets AI agents use a real human as a tool: visual checks, taste, phone calls, unblocking, approvals
1Operate Linux, macOS and Windows from your LLM. Every action runs through an auditable allowlist.
API for AI agents to delegate tasks to real humans.
Related MCP Servers
- AlicenseBqualityDmaintenanceEnables coding agents like Claude Code and Codex to offload boilerplate generation, summarization, and other bounded text tasks to local or cheap cloud LLMs, keeping the frontier agent in charge of judgment and code edits.93MIT
- AlicenseNot gradedqualityCmaintenanceEnables Claude Code to delegate mechanical tasks (summaries, boilerplate, reformatting) to local models running in LM Studio.1MIT
- AlicenseAqualityBmaintenanceLets a frontier coding agent delegate research, cataloguing, and long-running computation to a local LLM with guarded filesystem, web, and Python execution tools, preserving the agent's context and tokens.6MIT
- AlicenseNot gradedqualityBmaintenanceEnables Codex to dispatch well-scoped development tasks to a remote CodeBuddy Code instance, specify model and reasoning effort, monitor progress, retrieve results, and review Git changes without auto-merging.5 npmMIT