Skip to main content
Glama

WorkBuddy Delegate

Community / experimental · Windows-first · 2026-09-23 reliability update

WorkBuddy Delegate is a local Codex plugin that sends bounded, low-risk text work to an existing WorkBuddy installation, then returns a small, inspectable draft for Codex to verify. It is designed for extraction, summarization, classification, translation, rewriting, and code drafts—not autonomous decisions or external actions.

中文:这是一个 Windows 优先的 Codex 社区插件,将边界清晰、低风险的文本工作交给本机已有的 WorkBuddy。主任务仍负责核验与最终判断。

工作流程与 Token 目标

flowchart LR
    U["用户任务:多份长文本"] --> C{"Codex 判断任务"}
    C -->|"低风险、边界明确的文本工作"| P["插件检查文件范围、权限与调用限制"]
    C -->|"需要关键判断或执行操作"| M["Codex 直接处理"]
    F["明确指定的 UTF-8 文件"] --> P
    P -->|"检查通过"| W["WorkBuddy 阅读原文并生成草稿"]
    P -->|"检查失败"| M
    W --> R["精简答案 + 原文引文 + 不确定项"]
    R --> V["Codex 按需核验引文并作最终判断"]
    F -.->|"必要时抽查"| V
    V --> O["交付结果"]
    M --> O

**设计目标:**让 WorkBuddy 承担简单但耗费大量阅读上下文的提取、摘要、分类、翻译和改写工作。这样,Codex 主任务通常只需接收任务说明、精简草稿和核验所需的引文,而不必把全部文件内容放入自己的上下文。WorkBuddy 本身仍会消耗模型用量;实际节省多少 Codex Token、总成本和质量,需要用同一批任务做配对测量,不能由流程图保证。

Related MCP server: lmstudio-mcp

What it does

  • Packages a Codex skill and a dependency-free local MCP server.

  • Reads only explicit UTF-8 files beneath user-managed allowed_roots.

  • Blocks common credential/config paths and strips secret-like inherited environment variables.

  • Applies daily call, input-size, timeout, cache, and retention limits.

  • Uses no copied API key: authentication stays in WorkBuddy.

  • Stores small local result artifacts for inspection and labels them needs_main_agent_review.

It does not intercept every prompt before Codex sees it, guarantee savings or equal quality, edit files, run worker tools, or publish/send anything.

Requirements

  • Windows 10/11

  • A current local Codex build with Agent Plugins v1 support (tested on codex-cli 0.155.0-alpha.2.6)

  • Python 3.10+

  • Node.js available on PATH

  • WorkBuddy installed and signed in

The adapter uses WorkBuddy's bundled CodeBuddy CLI, which is not a documented stable public API. A WorkBuddy update may require a compatibility update here.

Install from GitHub

codex plugin marketplace add renyuxin6666-tech/codex-workbuddy-delegate --ref main
codex plugin add workbuddy-delegate@renyuxin-tools

Restart the Codex desktop app and open a new task. Then ask:

Show WorkBuddy Delegate status and its manager command.

The first file-based delegation needs an allowed root. Preserve the complete returned manager command (including configuration and state paths), then append the subcommand:

python "<installed-plugin>\scripts\manage.py" --config "<status.config_path>" --state-dir "<status.state_path>" roots add "D:\path\to\your\project"
python "<installed-plugin>\scripts\manage.py" --config "<status.config_path>" --state-dir "<status.state_path>" self-test

Provided text can be dry-run without an allowed file root. No manager command reads or prints API keys.

Example prompts

  • “Check whether summarizing these five notes is safe to delegate. If yes, use WorkBuddy and verify the cited excerpts.”

  • “Show WorkBuddy Delegate status and this week's local invocation count.”

  • “Use WorkBuddy for a first-pass translation of these allowed Markdown files; keep the final edit in Codex.”

Management

python plugins/workbuddy-delegate/scripts/manage.py doctor
python plugins/workbuddy-delegate/scripts/manage.py config init
python plugins/workbuddy-delegate/scripts/manage.py config set model auto
python plugins/workbuddy-delegate/scripts/manage.py config set daily_call_limit 10
python plugins/workbuddy-delegate/scripts/manage.py roots add "D:\project"
python plugins/workbuddy-delegate/scripts/manage.py usage --days 7
python plugins/workbuddy-delegate/scripts/manage.py cache prune

By default, portable installs place config and runtime state under the plugin's managed ${PLUGIN_DATA} directory. Direct script use falls back to %LOCALAPPDATA%\WorkBuddyDelegate.

The legacy Codex MCP launcher now requires either PLUGIN_DATA or both explicit WORKBUDDY_DELEGATE_CONFIG and WORKBUDDY_DELEGATE_STATE_DIR environment variables. It will stop instead of silently opening the direct-script fallback. For a headless run, point both explicit variables at the paths returned by the native workbuddy_status; do not copy secrets or replace the allowed-root list.

After setup, call the native workbuddy_status and verify that its allowed_roots and model match the changes. If the manager and MCP disagree, they are reading different configuration paths; use the explicit path arguments above. Do not remove the allowlist or disable the sandbox to fix this. Shell fallback commands may need scoped host approval to access WorkBuddy's installation and login.

To require approval for real delegation while auto-approving read-only status tools, use Codex's plugin-scoped MCP policy. See the current OpenAI plugin packaging documentation because config keys may evolve.

Routing boundary

Good candidates:

  • First-pass summaries of several non-sensitive notes

  • Exact-field extraction with source quotations

  • Low-stakes classification or translation batches

  • Rewrite variants and unapplied code drafts

Keep in Codex:

  • Final research claims, evidence adjudication, or literature novelty

  • Medical, legal, financial, admissions, safety, or employment decisions

  • Secrets, credentials, hidden config, or unrestricted filesystem reading

  • File changes, deployments, messages, purchases, or other external actions

  • Tasks whose authority or currentness is ambiguous

Evaluation snapshot

To test actual Token efficiency and quality, see the preregistered paired experiment plan. It separates Codex usage, WorkBuddy usage, blind quality scores, failures, and optional lower-tier Codex subagents. The 2026-09-23 preflight stopped before scored A/B runs; the repair follow-up fixed path routing and hardened JSON handling, but a live check now reports a possible WorkBuddy login issue. Status remains inconclusive, with no savings result claimed.

The bounded prototype scored 93.18/100 on one synthetic 22,049-character authority/version extraction fixture, but failed the strict completeness gate because it captured only 6/11 superseded historical relations. It captured current constraints 5/5, unresolved issues 1/1, and exact evidence 6/6, with no accepted prompt-injection instruction. This supports current-state extraction with Codex review—not complete historical auditing or generalized performance. See BENCHMARK.md.

Security and privacy

Read PRIVACY.md and SECURITY.md before adding sensitive projects. Results and quoted excerpts are stored locally until retention or pruning removes them. Source content is sent to the model/provider configured in WorkBuddy.

Development

python -m unittest discover -s tests -v
python plugins/workbuddy-delegate/scripts/manage.py self-test
python C:\Users\<you>\.codex\skills\.system\plugin-creator\scripts\validate_plugin.py plugins\workbuddy-delegate
python C:\Users\<you>\.codex\skills\.system\skill-creator\scripts\quick_validate.py plugins\workbuddy-delegate\skills\workbuddy-delegate

Status

This GitHub release is a local/repo marketplace plugin. It is not an official WorkBuddy integration and is not listed in the universal OpenAI Plugins Directory; public directory submission normally expects a public HTTPS MCP endpoint, while this plugin must access the user's local WorkBuddy installation.

MIT licensed. WorkBuddy, CodeBuddy, Codex, OpenAI, and their marks belong to their respective owners. This community project is not affiliated with or endorsed by them.

Local dispatch reliability update

The manager command returned by native status includes the exact config and state paths; keep these arguments when adding an approved project root. Terminal defaults may refer to a different configuration. dry_run now verifies CLI discovery, runtime-directory write access and the usage ledger without calling a model or consuming invocation quota. It does not verify login or network connectivity. Permission errors return scoped repair instructions; CLI failures return sanitized diagnostic categories, never raw output. After reinstalling, start a new Codex task to load the updated MCP process.

Available Tools

5 tools
delegate_to_workbuddyA

Send one bounded low-risk text task to the user's installed WorkBuddy. Use explicit relative files or short provided text. Supported kinds: summarize, extract, classify, translate, rewrite, code_draft. Only risk=low is routed. The main agent must verify the returned result.

ParametersJSON Schema
NameRequiredDescriptionDefault
kindYes
riskYes
textNo
filesNo
dry_runNo
workspaceYes
instructionYes

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate readOnly=false and destructive=false. The description adds valuable behavioral context: it sends a task to an external WorkBuddy, only routes risk=low, and requires the main agent to verify the returned result. This goes beyond the annotations and discloses a key operational requirement.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is tight and well-structured. It front-loads the core purpose, then enumerates kinds, constraints, and the verification note. Each sentence adds necessary information without fluff, and the list of supported kinds is presented clearly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given 7 parameters, no output schema, and annotations that only cover safety hints, the description is incomplete. It does not explain the required workspace parameter, the dry_run behavior, or the format of the returned result. The existence of a sibling workbuddy_read_result suggests the result retrieval mechanism should be mentioned, but it is not.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 0% schema description coverage, the description must compensate. It explains kind via the supported list, risk via 'Only risk=low', and text/files via 'explicit relative files or short provided text'. However, it does not explain the workspace (required), dry_run, or the instruction length constraint beyond 'bounded', leaving gaps for those parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific action ('Send one bounded low-risk text task to the user's installed WorkBuddy') with a clear resource and scope, and enumerates supported kinds. This clearly distinguishes it from siblings like workbuddy_status or workbuddy_plan, which handle different concerns.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives clear context: use only for low-risk, bounded text tasks, and it notes the main agent must verify the result. However, it does not explicitly reference sibling tools or state when NOT to use it, so exclusions are absent but the scope is well-defined.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

workbuddy_planB
Read-onlyIdempotent

Classify eligibility for bounded delegation without reading files or sending a model request.

ParametersJSON Schema
NameRequiredDescriptionDefault
kindYes
riskYes
instructionYes

TDQS

B3.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate readOnlyHint, idempotentHint, and non-destructive behavior. The description adds a more specific behavioral constraint: it performs no file reads and no model request. This goes beyond the generic annotations and helps the agent understand the tool's cost and side-effect profile. No contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence with no filler. It immediately conveys the core action and key constraints, making it easy for an agent to parse quickly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema and no parameter documentation, the description leaves the agent unable to determine what kind or risk values to provide or what the classification result looks like. It is a useful high-level summary but insufficient for correct invocation of a three-required-parameter tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate by explaining instruction, kind, and risk. It does not define any of these parameters, their valid values, or how they map to eligibility. The phrase 'bounded delegation' provides only a vague hint about their purpose.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific action: classify eligibility for bounded delegation. It distinguishes itself from siblings by explicitly noting it does not read files or send a model request, separating it from workbuddy_read_result and delegate_to_workbuddy. However, 'bounded delegation' is not defined, leaving some ambiguity about the exact scope.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies this is a pre-delegation eligibility check, especially given the sibling delegate_to_workbuddy. The phrase 'without reading files or sending a model request' suggests it is a lightweight planning step. Yet it never explicitly states when to use this tool versus alternatives or what conditions would make it inappropriate.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

workbuddy_read_resultB

Read a plugin-created result artifact when its short preview was truncated.

ParametersJSON Schema
NameRequiredDescriptionDefault
artifactYes
max_charsNo

TDQS

B3.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

There are no annotations to cover safety or side-effect expectations, so the description carries the full burden. It says the tool 'reads' an artifact, but it does not disclose what is returned, whether the output is truncated, how max_chars affects behavior, or any error conditions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence with no filler. Every word contributes to the core purpose and the triggering condition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no annotations and no output schema, the description should provide enough context to call the tool correctly. It explains when to call it but leaves key operational details unstated: what an artifact identifier looks like, what the tool returns, and how max_chars should be used.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, and the description does not compensate. It does not explain what 'artifact' refers to, how to obtain its value, or what 'max_chars' controls beyond the schema's default and bounds.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description identifies a specific verb ('Read'), a specific resource ('plugin-created result artifact'), and a precise trigger ('when its short preview was truncated'). This clearly distinguishes it from the sibling tools, which are about status, planning, delegation, and usage rather than reading artifacts.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides a clear usage context: use this tool when a short preview was truncated. However, it does not explicitly mention alternatives or state when not to use it, so it falls short of full exclusionary guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

workbuddy_statusA
Read-onlyIdempotent

Check the local WorkBuddy bridge, configured limits, and installation without sending a model request.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already mark the operation as read-only, idempotent, and non-destructive; the description adds important non-obvious behavior by stating that no model request is sent and that it only inspects local state. This goes beyond what the annotations alone convey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single sentence that front-loads the action and object, then adds the key behavioral caveat. No filler, no repetition of the schema or annotations.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter status tool, this is complete: it says what is checked (bridge, limits, installation), what is guaranteed (no model request), and annotations cover the safety profile. The absence of an output schema is acceptable because the named check items imply what the result will report.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, so parameter-level semantics are not a concern. The baseline of 4 applies because the description cannot add more meaning in this dimension.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Check') and identifies concrete resources: the local WorkBuddy bridge, configured limits, and installation. The closing constraint 'without sending a model request' further scopes it and distinguishes it from delegation tools like delegate_to_workbuddy.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies when to use it: whenever you need local status, limits, or installation info without incurring a model request. It does not explicitly name sibling alternatives or state exclusions, so it stops short of the strongest guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

workbuddy_usageA

Show local invocation counts. This does not call WorkBuddy or estimate money.

ParametersJSON Schema
NameRequiredDescriptionDefault
daysNo

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Since no annotations are provided, the description carries the behavioral burden. It transparently states that the tool does not call WorkBuddy or estimate money, which are important non-obvious behavioral constraints. It does not elaborate on output format or side effects, but for a local read-style display tool this is reasonably transparent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, with the main action first and the key boundary ('does not call WorkBuddy or estimate money') immediately after. No wasted words and no redundancy with the schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple one-parameter tool, the core purpose and key non-behaviors are covered. However, the role of 'days' is left to inference and no output shape is described, so an agent may not fully understand what it will receive or how to adjust the time window.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description should compensate for the undocumented 'days' parameter, but it does not mention it at all. The schema's default/min/max provide some constraints, yet the description fails to explain whether 'days' is a lookback window or how it affects the counts.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Show') and resource ('local invocation counts'), making the tool's function immediately clear. It also implicitly distinguishes itself from siblings like delegate_to_workbuddy by clarifying it does not call WorkBuddy.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description conveys that this is for local usage counts and not for executing WorkBuddy or estimating costs, but it never explicitly names alternatives or states clear when-to-use versus them. Usage context is implied rather than fully specified.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 5 tool updatesv0.1.0
    • First observeddelegate_to_workbuddy
    • First observedworkbuddy_plan
    • First observedworkbuddy_read_result
    • First observedworkbuddy_status
    • First observedworkbuddy_usage

TDQS

A3.8/5.0

Scored across 5 tools

Disambiguation4/5

The tools are mostly distinct: status, plan, delegate, read_result, and usage each serve a clear purpose. However, 'workbuddy_plan' (classify eligibility) and 'workbuddy_status' (check limits) could be slightly confused as both are pre-flight checks, but descriptions clarify.

Naming Consistency4/5

All tools share the 'workbuddy_' prefix and use a consistent verb_noun pattern (status, plan, delegate_to, read_result, usage). 'delegate_to_workbuddy' is slightly more verbose but still fits the pattern, so minor deviation.

Tool Count5/5

With 5 tools, the server is well-scoped for its purpose of managing delegation to WorkBuddy. Each tool covers a distinct aspect: status, planning, delegation, result reading, and usage tracking, which is appropriate.

Completeness4/5

The surface covers the main lifecycle: check status (status), eligibility (plan), delegate, read result, and usage. A minor gap is the lack of a cancel/abort tool for pending delegations, but the bounded tasks are short-lived, so this is not critical.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers