Skip to main content
Glama

mcp-delegate

Un servidor MCP que ofrece a Claude Code (como orquestador) una herramienta para delegar una tarea a un bucle agéntico completo y separado que se ejecuta en un modelo distinto (local vía Ollama, o remoto vía OpenRouter), con su propio acceso a herramientas (archivos, bash, etc.), y que devuelve únicamente un resultado final — funcionalmente equivalente a un subagente nativo, pero independiente del modelo.

Consulta mcp-subagent-delegation-plan.md para el plan de construcción completo, estructurado en fases como commits/checkpoints independientes.

Estado

Fases 1, 2, 3 y 4 completadas.

  • delegate_task — una sola llamada de completado de chat contra un endpoint compatible con OpenAI y configurado (Ollama, LM Studio, vLLM, OpenRouter, ...).

  • delegate_agentic_task — le da al modelo delegado su propio bucle de uso de herramientas (read_file, write_file, run_bash) en el directorio de trabajo especificado por quien llama; se ejecuta hasta que deja de llamar a herramientas, alcanza max_iterations o supera timeout_seconds.

  • list_recent_delegations — permite inspeccionar qué hicieron realmente las delegaciones anteriores (con cualquiera de las dos herramientas), sin tener que bucear en los registros ni volver a ejecutar nada.

  • get_delegation_transcript — transcripción completa del intercambio de mensajes y llamadas a herramientas de una delegación, cuando se ejecutó con capture_transcript=True (p. ej., para ejecuciones de comparación o evaluación de modelos).

Desviación del plan original: la fase 2 pedía envolver agent-loop como un subproceso. agent-loop solo admite Linux/macOS/WSL, y este servidor necesita ejecutarse de forma nativa en Windows, así que hemos construido en su lugar el bucle en el propio proceso descrito como alternativa en la fase 5: la misma interfaz de herramientas, sin la complejidad de subprocesos ni de eliminación de códigos ANSI, y evita por completo la licencia AGPL/sin uso comercial de agent-loop. Ver delegate/agentic.py.

Nota de seguridad: working_dir lo especifica el llamante; no es un sandbox fijo: el modelo delegado obtiene acceso no supervisado a archivos y bash del directorio al que se le apunte. Las herramientas de archivo (read_file/write_file) están limitadas para permanecer dentro de working_dir; run_bash se ejecuta con ese directorio como cwd, pero los comandos de shell no están supervisados del todo y podrían escapar de él (p. ej., cd ..). Apunta esto a un directorio en el que te sientas cómodo con que un modelo autónomo pueda leer, escribir y ejecutar comandos.

Nota sobre guardarraíles: el plan original, en su fase 4, pedía confirmar que las protecciones propias de agent-loop (tope de iteraciones, detección de repeticiones) estuvieran activas. Como no estamos usando agent-loop, la validación no es directamente aplicable: nuestro bucle tiene sus propios topes max_iterations y timeout_seconds (verificado en pruebas), pero no tiene detección de repeticiones. Un modelo que se quede atascado alternando entre dos llamadas a herramientas se ejecutará hasta alcanzar max_iterations en lugar de ser detectado a tiempo. Merece la pena añadirlo si eso llega a pasar en la práctica.

Related MCP server: deepseek-subagent-mcp

Configuración

uv sync
cp .env.example .env             # fill in DELEGATE_BASE_URL / DELEGATE_API_KEY / DELEGATE_MODEL
cp models.json.example models.json   # optional: named backends, see below

Varios backends

Ambas herramientas aceptan un parámetro opcional backend que consulta base_url/model/api_key desde models.json en lugar de las variables de entorno DELEGATE_* por defecto — p. ej., backend="ollama-local" para una llamada y backend="openrouter-free" para otra dentro del mismo turno, cada una ejecutándose de forma concurrente. model, si se proporciona también, sobrescribe únicamente la cadena del modelo dentro de ese backend.

Para referenciar una variable de entorno para una clave en lugar de escribirla en models.json:

{
  "openrouter-free": {
    "base_url": "https://openrouter.ai/api/v1",
    "model": "nvidia/nemotron-nano-9b-v2:free",
    "api_key_env": "OPENROUTER_API_KEY"
  }
}

models.json está incluido en .gitignore, igual que .env.

Concurrencia

Las llamadas a herramientas MCP ya se ejecutan en hilos de trabajo independientes, por lo que las delegaciones concurrentes se ejecutan en paralelo sin ningún tipo de cableado adicional. DELEGATE_MAX_CONCURRENCY (por defecto 4, ver .env.example) limita cuántas delegaciones — mediante ambas herramientas, para cualquier backend — se ejecutan a la vez, para evitar que una gran reventa sobrecargue un servidor de modelos local o los límites de frecuencia de una API de pago.

Ejecete el servidor directamente (sobre todo útil para comprobar que arranca sin errores; después espera en stdio a un cliente MCP):

uv run server.py

Registro (logging)

Cada llamada a delegate_task/delegate_agentic_task — tanto de éxito como fallida — se registra en un archivo SQLite local, conclusiones.db (gitignored, se crea en el primer uso): herramienta, backend, model, texto de la tarea, hora de inicio/fin, número de iteraciones, éxito/fallo, un avance truncado del resultado/error y el consumo de tokens si el backend lo proporcionó. Puedes consultarlo mediante la herramienta list_recent_delegations o directamente con sqlite3 delegations.db "select * from delegations order by id desc limit 20". El registro es con los mejores esfuerzos: en fallo del registro no dará al traste con una delegación que por lo demás haya tenido éxito.

Ambas herramientas añaden también una línea final [tokens: N prompt / N completion / N total ($cost)] a su propio valor de retorno cuando el backend informa del uso, para que el agente que llama lo vea de inmediato, sin una llamada adicional a list_recent_delegations.

Seguimiento del costo

pricing.json asigna a cada cadena de modelo una tarifa en USD {input_per_million, output_per_million}. Cuando el modelo resuelto de una llamada tiene entrada, el coste se calcula a partir del uso real de tokens, se registra en delegations.db (columna cost_usd) y se incluye en el sufijo [tokens: ...]. Un modelo sin entrada registra cost_usd = NULL: desconocido, no se asume como gratis, de modo que una entrada que falte no puede infravalorar el gasto en silencio. Los modelos locales no suelen tener entrada por esa razón; los modelos son realmente gratuitos (p. ej., :free de OpenRouter) reciben una entrada explícita {"input_per_million": 0, "output_per_million": 0} en lugar de quedar fuera.

A diferencia de .env/models.json, pricing.json no es un secreto ni específico del entorno, por lo que se incluye con seguimiento de versiones directamente, en lugar de estar republicado en .gitignore. Los precios cambiana: el archivo implementado se obtuvo de /api/v1/models de OpenRouter el 2026-08-21 para los modelos mencionados en una competición de modelos para la que se creó esto; vuelve a consultarlo y edítalo para añadir o actualizar modelos según haga falta.

Captura de transcripción (comparación/evaluación de modelos)

Ambas herramientas tienen capture_transcript: bool = False. Cuando está configurado, se registra el intercambio completo de mensajes — cada mensaje del modelo, cada llamada a herramientas y su resultado, no solo la respuesta final —, y el valor de retorno recibe un sufijo [delegation_id: N]. Puedes recuperarlo con get_delegation_transcriptn(delegation_id).

Esto existe para ejecutar la misma tarea con varios modelos/backends y comparar no solo las respuesta final sino cómo ha llegado cada uno hasta ella (la selección de herramientas, llamadas a herramientas con formato incorrecto, reintentos) — p. ej., un examen comparado entre modelos candidatos antes de elegir uno para producción. Está desactivado por defecto porque es un sobrecoste de registro que no quieres en una delegación rutinaria.

Registro con Claude Code

Un .mcp.json a nivel de proyecto ya está incluido en el repositorio (uv run server.py). Reinicia Claude Code en este directorio, o ejecuta claude mcp list para confirmar que ha cargado el servidor delegate, y luego pídele que llame a la herramienta delegate_task con un prompt trivial para confirmar el ruta de ida y vuelta.

Herramientas

  • delegate_task(prompt, model=None, system_prompt=None, backend=None, capture_transcript=False) -> str — llamada de chat de una sola pasada contra el backend configurado.

  • delegate_agentic_task(task, working_dir, model=None, max_iterations=20, timeout_seconds=600, backend=None, cache_transcript=False) -> str — delegación en varios pasos con las herramientas read_file/write_file/run_bash, limitadas a working_dir. Devuelve solo la respuesta final, no la transcripción completa, salvo que capture_transcript=True.

  • list_recent_delegations(limit=20) -> list[dict] — delegaciones registradas más recientes, primero las más nuevas.

  • get_delegation_transcript(delegation_id) -> list[dict] — transcripción completa de una delegación registrada con capture_transcript=True.

delegate_task/delegate_agentic_task devuelven los errores (configuración incorrecta, endpoint inaccesible, timeouts, límite de iteraciones) como cadenas "Error: ..." en lugar de lanzar una excepción, para que un agente que llama pueda ver qué ha pasado.

Available Tools

4 tools
delegate_agentic_taskA

Delegate a multi-step task to a model with its own tool-use loop (read_file, write_file, run_bash) scoped to working_dir. Runs until the model stops calling tools, hits max_iterations, or exceeds timeout_seconds. Returns only the final answer, not the full transcript.

The delegated model gets unattended file/bash access within working_dir for the duration of the call - point it at a directory you're comfortable it can read, write, and execute commands in.

Args: task: The task instruction to give the delegated model. working_dir: Directory the model's tools are scoped to. model: Override just the model string for this call. max_iterations: Stop after this many tool-call rounds. timeout_seconds: Wall-clock budget for the whole task. backend: Named backend from models.json (base_url/model/api_key) to use instead of the default DELEGATE_* env vars. model, if also given, overrides the model within that backend. capture_transcript: Log every model message and tool call/result for later retrieval via get_delegation_transcript, instead of just the final answer. Off by default; useful when comparing models (e.g. a bake-off) rather than for routine use.

ParametersJSON Schema
NameRequiredDescriptionDefault
taskYes
modelNo
backendNo
working_dirYes
max_iterationsNo
timeout_secondsNo
capture_transcriptNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description fully carries the behavioral disclosure burden. It clearly states that the delegated model gets unattended read/write/execute access within working_dir, that only the final answer is returned, that there are termination conditions, and that transcript capture is opt-in. This is comprehensive and honest about side effects and limits.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Despite being long, the description is tightly structured: a core behavior paragraph, a safety warning, then a bulleted Args list. Every sentence earns its place, and the most important info (what it does, termination, permissions) is front-loaded. No fluff or redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex delegation tool with 7 parameters, no annotations, and a dangerous access profile, the description covers all critical aspects: scope, termination, access level, return value, optional transcript capture, and backend override. The existence of an output schema is acknowledged but not required to detailed since it says returns only the final answer. Nothing an agent needs to invoke it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description is the only source of parameter meaning. It explains every parameter in the Args block, including the nuanced interplay between model and backend (backend as a base_url/model/api_key bundle, and that `model` overrides within that backend). This fully compensates for the schema's lack of descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a precise verb and resource: 'Delegate a multi-step task to a model with its own tool-use loop...'. It clearly states the operation's scope (working_dir) and distinguishes itself from tools like get_delegation_transcript by explaining that it returns only the final answer, not the full transcript. This is a specific, unambiguous definition that lets an agent know exactly what it does.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explains the conditions under which the delegated model stops (no more tool calls, max_iterations, timeout_seconds) and warns about unattended file/bash access. It also suggests capture_transcript for comparison scenarios, indirectly routing to get_delegation_transcript. However, it does not explicitly contrast with delegate_task or state when to choose this tool over that sibling, leaving some inference to the agent.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delegate_taskA

Delegate a single-shot task to a configured OpenAI-compatible model (e.g. local Ollama or OpenRouter) and return its text response verbatim.

Args: prompt: The task/question to send to the delegated model. model: Override just the model string for this call. system_prompt: Optional system prompt to steer the delegated model. backend: Named backend from models.json (base_url/model/api_key) to use instead of the default DELEGATE_* env vars. model, if also given, overrides the model within that backend. capture_transcript: Log the full message exchange for later retrieval via get_delegation_transcript. Off by default; useful when comparing models (e.g. a bake-off) rather than for routine use.

ParametersJSON Schema
NameRequiredDescriptionDefault
modelNo
promptYes
backendNo
system_promptNo
capture_transcriptNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are present, so the description carries the behavioral burden. It discloses the side-effect of transcript capture, the verbatim return behavior, and backend/model override semantics. It does not discuss latency, cost, or authentication, but those are not critical for selecting or invoking this tool correctly.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is organized with a front-loaded summary followed by a clear Args block. Every parameter is explained in one or two lines, and there is no redundant or filler content.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-shot delegation tool, the description covers purpose, parameter semantics, backend resolution, and the return behavior. With an output schema present and sibling context available, no critical invocation detail is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema has 0% description coverage, but the description fully documents all five parameters, including the relationship between backend and model, overriding behavior, and the opt-in nature of capture_transcript. This completely compensates for the schema gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool 'Delegate a single-shot task to a configured OpenAI-compatible model' and 'return its text response verbatim.' The 'single-shot' qualifier distinguishes it from the sibling delegate_agentic_task, though it does not explicitly name that sibling.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives concrete guidance on when to use capture_transcript ('when comparing models, e.g. a bake-off') and when not ('rather than for routine use'), and explains backend selection versus DELEGATE_* env vars. It does not explicitly describe when to choose delegate_task over delegate_agentic_task, but context signals and the 'single-shot' phrasing provide reasonable guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_delegation_transcriptA

Full message transcript (every model message and tool call/result) for one delegation, if it was run with capture_transcript=True. Get the id from list_recent_delegations. Returns an error string if no transcript was captured for that id.

Args: delegation_id: The id field from a list_recent_delegations row.

ParametersJSON Schema
NameRequiredDescriptionDefault
delegation_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the burden. It discloses the error condition for missing transcripts, which is the key behavioral nuance. It does not explicitly state read-only semantics, but that is reasonably implied for a retrieval tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise, with two clear sentences and a brief args section. No redundant or filler content; it efficiently conveys all necessary information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given there is an output schema (as indicated in context), the description need not explain return formats. It covers the essential context: the source of the id, the capture condition, and error behavior. This makes it complete for a single-parameter retrieval tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The parameter delegation_id is explained beyond the schema: it is the id from a list_recent_delegations row. This provides actionable meaning on how to obtain the correct value, enhancing the bare integer type definition.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool returns the full transcript for a delegation, using a specific verb ('get') and resource ('transcript'). It is distinct from siblings (list_recent_delegations lists, delegate_task delegates), so no ambiguity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly notes the precondition (capture_transcript=True), the error behavior when no transcript exists, and instructs to obtain the delegation_id from list_recent_delegations. This gives clear when-to-use guidance and differentiates it from alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_recent_delegationsA

List the most recent delegate_task / delegate_agentic_task calls (backend, model, task, duration, iterations, success, token usage, USD cost if the model has a pricing.json entry, truncated result), most recent first. Answers "what did the delegated model actually do" without re-running anything.

Args: limit: Max number of records to return (default 20).

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full behavioral disclosure burden. It discloses the read-only nature (without re-running), sorting (most recent first), truncation of results, and conditional cost reporting. It does not mention pagination or error behavior, but for a simple read-only listing tool these are minor omissions; the disclosed traits exceed typical descriptions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is moderately concise, listing the returned fields in a parenthetical that is useful but slightly dense. The core purpose is stated upfront, and the parameter doc is separated. It could be tightened by moving the field list to a separate line, but it remains efficient and well-organized.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has one optional parameter, no annotations, and an output schema (not provided). The description covers the return semantics (fields, ordering, truncation, cost condition) and the read-only intent. Given the simplicity, nothing essential for correct invocation is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate for the single parameter 'limit'. It does so explicitly: 'Max number of records to return (default 20).' This adds full semantic meaning beyond the bare schema field, making the tool usable without additional inference.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool lists recent delegate_task / delegate_agentic_task calls, enumerates the returned fields (backend, model, task, duration, iterations, success, token usage, USD cost, truncated result), and specifies ordering (most recent first). It also states the intended purpose—answering what a delegated model actually did—which distinguishes it from sibling tools that create delegations or fetch full transcripts.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies a clear use case for inspecting prior delegations without re-running them, but it does not explicitly contrast with siblings like get_delegation_transcript or delegate_task. It lacks explicit when-not-to-use guidance, though the mention of 'without re-running anything' strongly suggests a read-only inspection context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 4 tool updatesv0.1.0
    • First observeddelegate_agentic_task
    • First observeddelegate_task
    • First observedget_delegation_transcript
    • First observedlist_recent_delegations

TDQS

A4.7/5.0

Scored across 4 tools

Disambiguation5/5

Each tool has a distinct purpose: delegate_task for single-turn, delegate_agentic_task for multi-step with tool use, list_recent_delegations for querying history, and get_delegation_transcript for retrieving full logs. No overlap.

Naming Consistency5/5

All tool names follow a consistent snake_case verb_noun pattern (delegate_task, delegate_agentic_task, list_recent_delegations, get_delegation_transcript), with clear action prefixes.

Tool Count5/5

Four tools precisely cover the core delegation workflow: create a delegation (two variants), list delegations, and inspect a transcript. No unnecessary extras.

Completeness5/5

The tool set covers creating delegations, retrieving summaries, and fetching full transcripts. No update/delete is needed for delegation records, so the surface is complete for its purpose.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    A
    quality
    C
    maintenance
    Enables Claude to delegate tasks to external coding agents (Codex or Antigravity) for independent reviews, separate quota usage, and async processing.
    6
    MIT
  • A
    license
    A
    quality
    B
    maintenance
    Enables AI coding agents like Claude Code or Codex to delegate tasks to a DeepSeek Harness subagent with its own context window, providing tools for task delegation, result waiting, continuation, and supervision with sandboxed execution.
    6
    MIT
  • A
    license
    A
    quality
    B
    maintenance
    Enables delegating coding tasks to a pi agent as a steerable background worker, allowing mid-run redirection, follow-ups, and keeping the delegate's context isolated from your main conversation.
    12
    12
    16
    MIT