claude-worker-delegation
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@claude-worker-delegationSummarize the errors in debug.log and save the report to summary.md"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
claude-worker-delegation
An MCP server that turns Claude Code into an architect that hands bulk work to cheap worker models on OpenRouter.
Premium models are great at planning and review, but paying premium rates for them to read a 300 KB log or write boilerplate is wasteful. This server gives Claude Code three tools that offload that work. The big file contents go straight from disk to the worker model, so they never enter Claude's context.
┌──────────────────────────┐
you ──────▶│ Claude Code (architect) │ plans, decides, reviews, runs commands
└────────────┬─────────────┘
│ MCP (stdio)
┌────────────▼─────────────┐
│ workers MCP server │ tiers · budget cap · path sandbox · cost log
└────────────┬─────────────┘
│ OpenRouter API
┌──────────┬────────┴──┬────────────┐
▼ ▼ ▼ ▼
fast code reason long
(summaries) (diffs/tests) (2nd opinion) (1M-token files)Tools
Tool | What it does |
| Run one task on a worker tier. Pass |
| Run several independent jobs in parallel (concurrency set in config). |
| Token and cost totals for today and all time, broken down by tier. |
Related MCP server: opencode-mcp
Features
Tiered models in
models.json:fast,code,reason, andlong. Each tier has a description the architect uses to pick one.Daily budget cap. Calls are refused once the day's spend reaches
daily_budget_usd.Cost logging. Every call is appended to
usage.jsonlwith its tokens, dollar cost, and latency.Path sandbox. Workers can only read and write files under the session's working directory, plus any extra roots set in
WORKERS_ALLOWED_ROOTS.Guard-railed worker prompt. Workers are told to answer tersely, return unified diffs for code, and never invent file contents they weren't shown.
Provider settings. OpenRouter routing sorts by price, allows fallbacks, and denies data collection.
Large-read hook (
hooks/delegate-large-reads.js). A Claude CodePreToolUsehook that blocks whole-file reads over 50 KB, whether throughRead,cat, orGet-Content, and points Claude todelegateinstead.
Real usage
The first 12 calls cost $0.09 total, and three of them fed about 363K input tokens of logs and files to the long tier. All of that input would otherwise have gone through the premium model's context.
Setup
npm install
export OPENROUTER_API_KEY=sk-or-... # Windows: setx OPENROUTER_API_KEY "sk-or-..."
claude mcp add workers -- node /path/to/claude-worker-delegation/server.jsThen:
Add the rules in
examples/CLAUDE.mdto your~/.claude/CLAUDE.mdso Claude knows what to delegate.Optionally add the hook and permissions from
examples/settings.jsonto~/.claude/settings.json.Run
npm run smoketo check that the tools, every tier, and the path sandbox all work.
Config (models.json)
"tiers": {
"fast": { "model": "deepseek/deepseek-v4-flash", "use": "summaries, extraction, boilerplate" },
"code": { "model": "qwen/qwen3-coder-next", "use": "code, tests, diffs" },
"reason": { "model": "deepseek/deepseek-v4-pro", "use": "debugging, second-opinion reviews" },
"long": { "model": "google/gemini-3.1-flash-lite", "use": "very large files (1M context)" }
},
"daily_budget_usd": 3.0,
"concurrency": 4To change models, edit this file. You don't need to change any code.
Stack
Node.js · Model Context Protocol SDK · zod · OpenRouter API · Claude Code hooks
License
MIT
Available Tools
3 toolsdelegateA
Hand a self-contained subtask to a cheap worker model (OpenRouter). Use for bulk reading/summarizing, boilerplate, tests, mechanical edits as diffs, first drafts, or a second opinion. Review the result before acting on it.
| Name | Required | Description | Default |
|---|---|---|---|
| task | Yes | Precise instructions for the worker. Include acceptance criteria and output format. | |
| tier | No | fast: summaries, extraction, boilerplate, simple edits; code: writing/refactoring code, tests, diffs; reason: tricky logic, debugging hypotheses, second-opinion reviews; long: very large logs/files (1M context) | fast |
| files | No | Paths the server loads into the worker prompt — do NOT read them yourself first. | |
| write_to | No | Save full output here and return only a 20-line preview. Use for anything > ~50 lines. | |
| max_tokens | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description discloses that it uses a 'cheap worker model' and advises to 'Review the result before acting on it', indicating the output may need verification. It also implies self-contained subtasks, suggesting no dependencies. It doesn't cover rate limits or auth, but for a delegation tool this is adequate.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with purpose, and a clear list of use cases. No fluff. Every word earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema, but the description gives enough to know what the tool does and cautions to review results. It doesn't describe return format or error handling, but given the tool's simplicity, it's sufficient.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 80%, so most parameters are documented. The description does not add additional parameter semantics beyond the schema, so it doesn't enhance understanding of the parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states 'Hand a self-contained subtask to a cheap worker model' which is a clear verb+resource. It lists specific use cases like 'bulk reading/summarizing, boilerplate, tests' which helps differentiate from worker_stats and delegate_batch, though it doesn't explicitly name the siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly lists appropriate use cases ('bulk reading/summarizing, boilerplate, tests, mechanical edits as diffs, first drafts, or a second opinion') and advises to 'Review the result before acting on it.' It does not mention alternatives or exclusions, but the context is clear enough to guide an agent.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
delegate_batchA
Run several independent worker jobs in parallel. Returns one result block per job.
| Name | Required | Description | Default |
|---|---|---|---|
| jobs | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It usefully discloses parallel execution and the per-job result shape, but it does not mention failure behavior, partial results, timeouts, side effects of running worker jobs, or resource/cost implications.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences with no filler. The action is front-loaded, and the return behavior is stated in the second sentence. Every word earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is moderately complex due to nested job fields, yet the description is thin on edge-case behavior and sibling differentiation. The schema fills in field-level details, but the absence of annotations and output schema leaves gaps around failure handling and when to prefer this over 'delegate.'
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description does little to compensate: it only restates that jobs are 'independent worker jobs.' It does not explain the job structure, tier semantics, file loading, write_to behavior, or max_tokens, even though the nested schema does describe these.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb and resource: 'Run several independent worker jobs in parallel' and states the return behavior: 'Returns one result block per job.' The plural and 'parallel' wording clearly distinguishes it from the sibling tool 'delegate' and from 'worker_stats'.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context: this is for several independent jobs that can safely run in parallel. It does not explicitly name 'delegate' as the single-job alternative or list when not to use this tool, so it stops short of full exclusion guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
worker_statsA
Worker token/cost totals: today and all-time, by tier.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It implies a read-only operation by stating 'totals,' but does not explicitly disclose side-effect-free behavior, permissions, or response format. For a simple stats tool, this is minimally acceptable but leaves gaps.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, tightly worded sentence that front-loads the core purpose ('Worker token/cost totals') and follows with scope details. Every word earns its place, with no fluff.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the lack of parameters and output schema, the description covers the main purpose and scope. It does not specify the exact output format or timezone for 'today,' but for a simple stats tool with no inputs, this is a minor omission.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so the description does not need to clarify parameter meaning. Baseline for 0 params is 4, and the description appropriately avoids redundant parameter details.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool provides worker token/cost totals, covering today and all-time, broken down by tier. This is specific and distinguishes it from sibling tools (delegate, delegate_batch) which are action-oriented rather than informational.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus the sibling delegate tools. While the purpose implies it is for reading stats rather than delegating, explicit usage context is absent.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
3 tool updates
v1.0.0- First observed
delegate - First observed
delegate_batch - First observed
worker_stats
TDQS
Scored across 3 tools
worker_stats is clearly a metrics-only tool, while delegate and delegate_batch are distinguished by single vs. parallel execution. There is no meaningful overlap between the tools.
delegate and delegate_batch share a clear prefix and action-oriented style, but worker_stats breaks the pattern by using a noun-based name. The inconsistency is minor given the small tool count.
Three tools is appropriate for a focused worker-delegation server. Each tool serves a distinct purpose without unnecessary redundancy.
The core delegation workflow is fully covered: single delegation, batch delegation, and usage statistics. Since results are returned directly, no additional retrieval or status tools appear necessary.
Maintenance
Related MCP Connectors
- mcpOAuthcom.fivexer
Route, roster, and track work in a Fivexer workspace from Claude Code and other agentic tools
Token guard and rate limiter preventing runaway API cost spikes for OpenAI and Anthropic.
AI routing, memory, guardrails, and governance. Routes across Claude, GPT, Gemini.
Exact Claude API cost calc with real cache economics, plus a tiktoken-misuse scanner.
Related MCP Servers
- AlicenseCqualityAmaintenanceRoutes coding tasks across multiple AI CLIs (Copilot, Claude Code, Gemini, etc.) with cost-aware tier routing and parallel wave orchestration.552Apache 2.0
- AlicenseNot gradedqualityBmaintenanceLets Claude Code delegate tasks to the OpenCode CLI, choosing cost-effective models by intelligence tier and tracking usage.23,279 npm1MIT
- AlicenseAqualityCmaintenanceLets Claude Code delegate tasks to OpenRouter-backed Claude Code sessions in isolated child processes, keeping Anthropic and OpenRouter credentials separate. It adds a model catalog, per-job cost tracking, and API-key management for running tasks on 400+ OpenRouter models.710 npmMIT
- AlicenseNot gradedqualityBmaintenanceEnables Claude Code to delegate tasks to free OpenRouter models as callable tools while keeping the Anthropic endpoint for the orchestrator, with optional sandboxed file and shell access.MIT