Skip to main content
Glama
kaijfox

delegations-mcp

by kaijfox

delegations-mcp

An MCP server that exposes a library of delegation prompts as tools. A smart orchestrator LLM uses these tools to hand off bounded tasks to a reduced LLM or coding agent, which receives a fully-constructed, self-contained prompt requiring no broader context.

Configuration

Config is discovered by walking up from the working directory, then merged with ~/.delegations.toml. Project config extends global; library blocks merge field-by-field (project overrides global per field).

# .delegations.toml

[agent]
executable = "copilot"
args = ["--model", "gpt-4o-mini", "--prompt", "Instructions in: {prompt_path}"]

output_dir = "/tmp"   # where prompts and transcripts are written

[library.devteam]
path = "./library/devteam"
test_executable = "/usr/bin/python3"
test_args = ["-m", "pytest"]

{prompt_path} in agent args is replaced with the path to the rendered prompt file.

Related MCP server: multi-model-mcp

Running

The server must run on the host filesystem — the agent it spawns edits files directly and needs access to your project.

The MCP client spawns the server automatically, inheriting its working directory. Config is discovered from there. No setup required beyond installing the package.

MCP client config (e.g. Claude Desktop):

{
  "mcpServers": {
    "delegations": {
      "command": "uv",
      "args": ["run", "--directory", "/path/to/delegations-mcp", "delegations-mcp"]
    }
  }
}

HTTP mode — persistent session server

Useful when multiple agents share a session: the registry is loaded once, and the async lock for lock: true delegations is shared across all connections.

Run from your project directory (so config discovery finds .delegations.toml):

cd /your/project
delegations-mcp --transport http --port 8000
# or: uv run --directory /path/to/delegations-mcp delegations-mcp --transport http

Connect your MCP client to http://localhost:8000/mcp. One server instance per project — the working directory at startup determines which config and libraries are used.

Tools exposed

Tool

Description

list_delegations()

Lists available delegations; refreshes registry from disk

get_delegation(name)

Returns full details and merged input schema for a delegation

run_delegation(name, inputs)

Runs a delegation; returns summary, prompt_path, transcript_path

Delegation names are library:delegation (e.g. devteam:implement).

Libraries

See library/devteam/README.md for the bundled devteam library. See docs/library-implementation.md to author your own.

Available Tools

3 tools
get_delegationC

Get full details for a delegation by its library:name.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. It states this is a 'Get' operation, implying it's likely read-only, but doesn't confirm safety aspects like whether it requires authentication, has rate limits, or what happens on errors. For a tool with zero annotation coverage, this is a significant gap in transparency.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, efficient sentence with zero waste. It's front-loaded with the core purpose and includes the key parameter detail. Every word earns its place, making it highly concise and well-structured.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's low complexity (1 parameter, no nested objects) and the presence of an output schema (which likely covers return values), the description is somewhat complete but has gaps. It lacks behavioral context due to no annotations and minimal parameter semantics. For a simple lookup tool, it's adequate but not fully helpful.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema description coverage is 0%, meaning the input schema provides no descriptions for the single parameter 'name'. The description adds some meaning by specifying it's a 'library:name' identifier, but this is minimal—it doesn't explain the format, examples, or constraints of the 'name' parameter. With low coverage, the description doesn't adequately compensate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose with a specific verb ('Get') and resource ('full details for a delegation'), and it specifies the lookup method ('by its library:name'). However, it doesn't explicitly distinguish this from its sibling 'list_delegations' (which likely lists multiple delegations) or 'run_delegation' (which likely executes a delegation), so it misses full sibling differentiation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus its siblings. It doesn't mention alternatives like 'list_delegations' for listing multiple delegations or 'run_delegation' for execution, nor does it specify prerequisites or contexts for usage. This leaves the agent without clear usage direction.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_delegationsB

List all available delegations. Also refreshes the registry from disk.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It discloses a behavioral trait: 'refreshes the registry from disk', indicating a side effect that might impact performance or data freshness. However, it lacks details on permissions, rate limits, or what 'refreshes' entails (e.g., overwrites cache).

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise with two clear sentences. However, the second sentence about refreshing the registry could be better integrated or explained, slightly reducing efficiency. It's front-loaded with the primary purpose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given 0 parameters, 100% schema coverage, and an output schema exists, the description is reasonably complete. It covers the main action and a side effect. However, for a tool with behavioral implications (refreshing), more context on when and why to use it would enhance completeness.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There are 0 parameters, and schema description coverage is 100%, so no parameter documentation is needed. The description doesn't add param info, but with no params, this is acceptable. Baseline is 4 as per rules for 0 parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb 'list' and resource 'delegations', making the purpose specific. However, it doesn't explicitly distinguish this from sibling tools like 'get_delegation' (which likely retrieves a single delegation) or 'run_delegation' (which likely executes one).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives like 'get_delegation' or 'run_delegation'. The description mentions 'refreshes the registry from disk', which hints at a side effect but doesn't clarify if this is a primary use case or when it's appropriate compared to other listing methods.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

run_delegationC

Run a delegation by its library:name. Returns summary, prompt_path, transcript_path.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYes
inputsNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries full burden. It discloses return values ('summary, prompt_path, transcript_path'), which adds context beyond the input schema, but fails to describe behavioral traits like whether it's read-only, destructive, requires authentication, has side effects, or rate limits. For a tool named 'run', this is a significant gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with the main action and resource, followed by return values. It's concise with two clauses, but the second clause could be more integrated; overall, it's efficient with minimal waste.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given complexity (2 parameters, nested object, no annotations) and an output schema exists, the description is moderately complete. It covers the purpose and return values, but lacks behavioral context and full parameter details. The output schema reduces the need to explain returns, but gaps in usage and transparency remain.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It adds meaning by specifying the 'name' parameter as 'library:name', clarifying its format, but doesn't explain 'inputs' (a nested object with no details). This partial compensation meets the baseline for low coverage, but leaves 'inputs' undocumented.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states the action ('Run a delegation') and resource ('by its library:name'), but it's vague about what 'run' entails—does it execute, simulate, or test? It distinguishes from siblings 'get_delegation' and 'list_delegations' by implying execution rather than retrieval, but the purpose lacks specificity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool versus alternatives like 'get_delegation' or 'list_delegations'. It mentions 'by its library:name', hinting at a prerequisite (a delegation must exist), but no explicit when/when-not rules or context for selection are provided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 3 tool updatesv0.1.0
    • First observedget_delegation
    • First observedlist_delegations
    • First observedrun_delegation

TDQS

B3.2/5.0

Scored across 3 tools

Disambiguation5/5

Each tool has a clearly distinct purpose: get_delegation retrieves details for a specific delegation, list_delegations lists all delegations and refreshes the registry, and run_delegation executes a delegation and returns results. There is no overlap in functionality, making tool selection unambiguous.

Naming Consistency5/5

All tool names follow a consistent verb_noun pattern with snake_case: get_delegation, list_delegations, run_delegation. The naming is predictable and readable, with no deviations in style or convention.

Tool Count3/5

With only 3 tools, the set feels thin for a delegations management server, as it lacks operations like create, update, or delete delegations. However, the tools cover basic retrieval, listing, and execution, which might be sufficient for a minimal scope.

Completeness3/5

The tools provide read and execute capabilities (get, list, run), but there are notable gaps in lifecycle management, such as creating, updating, or deleting delegations. This could limit agent workflows that require full CRUD operations, though the existing tools support core usage.

Maintenance

ActivityInactive
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers