delegations-mcp
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@delegations-mcprun devteam:implement to add error handling to the login function"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
delegations-mcp
An MCP server that exposes a library of delegation prompts as tools. A smart orchestrator LLM uses these tools to hand off bounded tasks to a reduced LLM or coding agent, which receives a fully-constructed, self-contained prompt requiring no broader context.
Configuration
Config is discovered by walking up from the working directory, then merged with
~/.delegations.toml. Project config extends global; library blocks merge
field-by-field (project overrides global per field).
# .delegations.toml
[agent]
executable = "copilot"
args = ["--model", "gpt-4o-mini", "--prompt", "Instructions in: {prompt_path}"]
output_dir = "/tmp" # where prompts and transcripts are written
[library.devteam]
path = "./library/devteam"
test_executable = "/usr/bin/python3"
test_args = ["-m", "pytest"]{prompt_path} in agent args is replaced with the path to the rendered prompt file.
Related MCP server: multi-model-mcp
Running
The server must run on the host filesystem — the agent it spawns edits files directly and needs access to your project.
stdio mode — recommended, zero config
The MCP client spawns the server automatically, inheriting its working directory. Config is discovered from there. No setup required beyond installing the package.
MCP client config (e.g. Claude Desktop):
{
"mcpServers": {
"delegations": {
"command": "uv",
"args": ["run", "--directory", "/path/to/delegations-mcp", "delegations-mcp"]
}
}
}HTTP mode — persistent session server
Useful when multiple agents share a session: the registry is loaded once, and
the async lock for lock: true delegations is shared across all connections.
Run from your project directory (so config discovery finds .delegations.toml):
cd /your/project
delegations-mcp --transport http --port 8000
# or: uv run --directory /path/to/delegations-mcp delegations-mcp --transport httpConnect your MCP client to http://localhost:8000/mcp. One server instance per
project — the working directory at startup determines which config and libraries
are used.
Tools exposed
Tool | Description |
| Lists available delegations; refreshes registry from disk |
| Returns full details and merged input schema for a delegation |
| Runs a delegation; returns |
Delegation names are library:delegation (e.g. devteam:implement).
Libraries
See library/devteam/README.md for the bundled
devteam library. See docs/library-implementation.md
to author your own.
Available Tools
3 toolsget_delegationC
Get full details for a delegation by its library:name.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. It states this is a 'Get' operation, implying it's likely read-only, but doesn't confirm safety aspects like whether it requires authentication, has rate limits, or what happens on errors. For a tool with zero annotation coverage, this is a significant gap in transparency.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence with zero waste. It's front-loaded with the core purpose and includes the key parameter detail. Every word earns its place, making it highly concise and well-structured.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's low complexity (1 parameter, no nested objects) and the presence of an output schema (which likely covers return values), the description is somewhat complete but has gaps. It lacks behavioral context due to no annotations and minimal parameter semantics. For a simple lookup tool, it's adequate but not fully helpful.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema description coverage is 0%, meaning the input schema provides no descriptions for the single parameter 'name'. The description adds some meaning by specifying it's a 'library:name' identifier, but this is minimal—it doesn't explain the format, examples, or constraints of the 'name' parameter. With low coverage, the description doesn't adequately compensate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose with a specific verb ('Get') and resource ('full details for a delegation'), and it specifies the lookup method ('by its library:name'). However, it doesn't explicitly distinguish this from its sibling 'list_delegations' (which likely lists multiple delegations) or 'run_delegation' (which likely executes a delegation), so it misses full sibling differentiation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus its siblings. It doesn't mention alternatives like 'list_delegations' for listing multiple delegations or 'run_delegation' for execution, nor does it specify prerequisites or contexts for usage. This leaves the agent without clear usage direction.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_delegationsB
List all available delegations. Also refreshes the registry from disk.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It discloses a behavioral trait: 'refreshes the registry from disk', indicating a side effect that might impact performance or data freshness. However, it lacks details on permissions, rate limits, or what 'refreshes' entails (e.g., overwrites cache).
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise with two clear sentences. However, the second sentence about refreshing the registry could be better integrated or explained, slightly reducing efficiency. It's front-loaded with the primary purpose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given 0 parameters, 100% schema coverage, and an output schema exists, the description is reasonably complete. It covers the main action and a side effect. However, for a tool with behavioral implications (refreshing), more context on when and why to use it would enhance completeness.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There are 0 parameters, and schema description coverage is 100%, so no parameter documentation is needed. The description doesn't add param info, but with no params, this is acceptable. Baseline is 4 as per rules for 0 parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb 'list' and resource 'delegations', making the purpose specific. However, it doesn't explicitly distinguish this from sibling tools like 'get_delegation' (which likely retrieves a single delegation) or 'run_delegation' (which likely executes one).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus alternatives like 'get_delegation' or 'run_delegation'. The description mentions 'refreshes the registry from disk', which hints at a side effect but doesn't clarify if this is a primary use case or when it's appropriate compared to other listing methods.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
run_delegationC
Run a delegation by its library:name. Returns summary, prompt_path, transcript_path.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | ||
| inputs | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries full burden. It discloses return values ('summary, prompt_path, transcript_path'), which adds context beyond the input schema, but fails to describe behavioral traits like whether it's read-only, destructive, requires authentication, has side effects, or rate limits. For a tool named 'run', this is a significant gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the main action and resource, followed by return values. It's concise with two clauses, but the second clause could be more integrated; overall, it's efficient with minimal waste.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given complexity (2 parameters, nested object, no annotations) and an output schema exists, the description is moderately complete. It covers the purpose and return values, but lacks behavioral context and full parameter details. The output schema reduces the need to explain returns, but gaps in usage and transparency remain.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It adds meaning by specifying the 'name' parameter as 'library:name', clarifying its format, but doesn't explain 'inputs' (a nested object with no details). This partial compensation meets the baseline for low coverage, but leaves 'inputs' undocumented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states the action ('Run a delegation') and resource ('by its library:name'), but it's vague about what 'run' entails—does it execute, simulate, or test? It distinguishes from siblings 'get_delegation' and 'list_delegations' by implying execution rather than retrieval, but the purpose lacks specificity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this tool versus alternatives like 'get_delegation' or 'list_delegations'. It mentions 'by its library:name', hinting at a prerequisite (a delegation must exist), but no explicit when/when-not rules or context for selection are provided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
3 tool updates
v0.1.0- First observed
get_delegation - First observed
list_delegations - First observed
run_delegation
TDQS
Scored across 3 tools
Each tool has a clearly distinct purpose: get_delegation retrieves details for a specific delegation, list_delegations lists all delegations and refreshes the registry, and run_delegation executes a delegation and returns results. There is no overlap in functionality, making tool selection unambiguous.
All tool names follow a consistent verb_noun pattern with snake_case: get_delegation, list_delegations, run_delegation. The naming is predictable and readable, with no deviations in style or convention.
With only 3 tools, the set feels thin for a delegations management server, as it lacks operations like create, update, or delete delegations. However, the tools cover basic retrieval, listing, and execution, which might be sufficient for a minimal scope.
The tools provide read and execute capabilities (get, list, run), but there are notable gaps in lifecycle management, such as creating, updating, or deleting delegations. This could limit agent workflows that require full CRUD operations, though the existing tools support core usage.
Maintenance
Related MCP Connectors
Nifty's MCP server — exposes tasks, projects, messages, and files as tools for AI agents.
LLM Orchestration Agent (Mcp)
MCP server for AI agents to plan, verify, and deploy Cloudflare-native apps.
MCP server for building and testing AI agents with multi-model experimentation and insights.
Related MCP Servers
- AlicenseNot gradedqualityBmaintenanceMCP server that enables AI agents to run a deterministic orchestration loop with decomposition, subagent execution, and review feedback across multiple LLM backends.59MIT
- FlicenseAqualityCmaintenanceAn MCP server that exposes tools for sub-agent style reasoning across multiple LLM providers, enabling delegation of prompts to various models and running critique loops, debates, red-teaming, and answer ranking.6-
- AlicenseAqualityBmaintenanceLocal MCP server that exposes delegation tools for Codex, Claude, and Antigravity CLI, enabling an orchestrator agent to assign tasks to these sub-agents via non-interactive CLI commands.3MIT
- AlicenseNot gradedqualityCmaintenanceMCP server that enables delegation of small coding implementations to a cheaper language model with propose_patch and apply_patch tools, while the primary agent retains architecture, review, and approval.MIT