Skip to main content
Glama

Agent Team Lab

Mini plataforma reutilizable para cargar un equipo de agentes expertos en cualquier proyecto. El flujo conecta Product Manager, Arquitectura, Programación, QA y Seguridad; cada resultado alimenta las etapas siguientes y PM consolida la decisión final.

Puede consumirse de tres formas:

  • como servidor MCP desde Codex, Hermes u otro cliente compatible;

  • mediante API REST desde una aplicación;

  • con una CLI local para pruebas.

El núcleo no depende de un único modelo. Incluye adaptadores para OpenAI Responses API, endpoints compatibles con OpenAI, Hermes Agent CLI y un proveedor determinista para tests.

Flujo

flowchart TD
    PM1[PM: alcance] --> ARQ[Arquitectura]
    ARQ --> DEV[Programación]
    DEV --> QA[QA]
    DEV --> SEC[Seguridad]
    QA --> PM2[PM: decisión]
    SEC --> PM2

La versión 0.1.0 es de solo lectura: recoge archivos de texto del proyecto dentro de una raíz permitida, limita el contexto y devuelve recomendaciones. No ejecuta código del repositorio analizado ni escribe cambios.

Related MCP server: all-agents-mcp

Inicio rápido

Requisitos: Python 3.11+ y uv.

git clone https://github.com/Benemox/agent-team-lab.git
cd agent-team-lab
cp .env.example .env
uv sync --extra dev

Configura OPENAI_API_KEY en tu entorno y prueba:

uv run agent-team roles
uv run agent-team run "Revisa el flujo de pagos" --project /ruta/al/proyecto

La ruta debe estar dentro de AGENT_TEAM_ALLOWED_ROOT. Para una prueba sin gastar tokens:

AGENT_TEAM_PROVIDER=mock AGENT_TEAM_ALLOWED_ROOT="$PWD" \
  uv run agent-team run "Comprueba el equipo" --project "$PWD"

Proveedores

OpenAI / Codex

Usa Responses API. OPENAI_MODEL admite el identificador de modelo disponible en tu proyecto de OpenAI.

AGENT_TEAM_PROVIDER=openai
OPENAI_API_KEY=...
OPENAI_MODEL=gpt-5.4-mini

Hermes Agent

Instala y configura Hermes por separado, comprueba que hermes funciona y selecciona el adaptador:

hermes setup
AGENT_TEAM_PROVIDER=hermes uv run agent-team run "Analiza este módulo" --project .

El adaptador usa una ejecución finita hermes chat --oneshot -q; conserva la configuración de proveedor/modelo que hayas elegido en Hermes.

Ollama, vLLM, LocalAI u otro compatible

AGENT_TEAM_PROVIDER=openai-compatible
OPENAI_COMPATIBLE_BASE_URL=http://127.0.0.1:11434/v1
OPENAI_COMPATIBLE_API_KEY=local
OPENAI_COMPATIBLE_MODEL=qwen3-coder

API REST

AGENT_TEAM_ALLOWED_ROOT=/ruta/a/proyectos uv run agent-team serve

Endpoints:

  • GET /health

  • GET /v1/roles

  • POST /v1/experts/{role_id}

  • POST /v1/team/run

  • documentación OpenAPI en http://127.0.0.1:8008/docs

Ejemplo:

curl -X POST http://127.0.0.1:8008/v1/team/run \
  -H 'Content-Type: application/json' \
  -H "Authorization: Bearer $AGENT_TEAM_API_TOKEN" \
  -d '{"task":"Revisa el login", "project_path":"/ruta/a/proyectos/mi-app"}'

Define AGENT_TEAM_API_TOKEN antes de exponer el servicio fuera de localhost. El contenedor monta los proyectos en modo solo lectura.

MCP y plugin

El repositorio ya incluye .codex-plugin/plugin.json, .mcp.json y cinco skills. Tras instalar dependencias, el servidor puede arrancarse con:

uv run agent-team-mcp

Expone las herramientas:

  • list_experts

  • ask_expert

  • run_team

Para incorporar el equipo a otro repositorio sin instalar el plugin completo, añade este servidor a la configuración MCP del cliente y usa el directorio clonado como directorio de trabajo:

{
  "mcpServers": {
    "agent-team-lab": {
      "command": "uv",
      "args": ["run", "--directory", "/ruta/agent-team-lab", "agent-team-mcp"],
      "env": {
        "AGENT_TEAM_PROVIDER": "openai",
        "AGENT_TEAM_ALLOWED_ROOT": "/ruta/a/proyectos"
      }
    }
  }
}

No escribas la API key dentro del JSON si el archivo se va a versionar; pásala mediante variables del sistema o un almacén de secretos.

Personalizar el equipo

  • Edita perfiles en skills/*/SKILL.md.

  • Cambia orden, paralelismo y objetivos en config/team.yaml.

  • Añade un proveedor implementando AgentProvider y registrándolo en providers/factory.py.

  • Ajusta extensiones, límites y carpetas ignoradas en context.py.

Verificación

uv run ruff check src tests
uv run pytest

Diseño de seguridad

  • raíz de proyectos permitida y bloqueo de path traversal;

  • límites de bytes y archivos enviados al modelo;

  • carpetas de dependencias, caché y control de versiones excluidas;

  • contexto marcado como no confiable frente a prompt injection;

  • sin shell para analizar proyectos y sin escrituras automáticas;

  • Hermes se invoca con argumentos, no mediante shell=True;

  • token Bearer opcional para REST y escucha local por defecto.

Licencia

MIT.

Available Tools

3 tools
ask_expertC

Ask one specialist to review a task and the readable source files in a project.

ParametersJSON Schema
NameRequiredDescriptionDefault
taskYes
role_idYes
providerNo
project_pathNo.

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of disclosing behavior. It only says the tool 'asks' a specialist to review, without stating whether this launches an async request, returns a response, requires permissions, or has side effects.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single efficient sentence with the core action front-loaded and no filler. However, it is so brief that important context is omitted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With four parameters, no annotations, and no output schema, the description is too thin. It lacks guidance on how to select role_id, when to use this tool versus run_team, and what the agent should expect as a result.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It loosely maps 'specialist' to role_id and 'project' to project_path, but it does not explain provider, the meaning of task, or how role_id values are obtained.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific action ('Ask one specialist') and a clear resource (a task and readable source files in a project). The phrase 'one specialist' weakly contrasts with run_team, but it does not fully clarify that role_id identifies the specialist, so it stops short of 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The wording implies a single-specialist review, which hints at a distinction from run_team, but it never explicitly states when to use this tool versus list_experts or run_team. It also does not mention that list_experts should be used to find a valid role_id.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_expertsA

List the expert roles available in this agent team.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations exist, so the description carries the burden of behavioral disclosure. 'List' clearly implies a read-only operation, but the description does not explicitly confirm non-mutation, ordering, or whether available roles can change by team context. This is acceptable for a simple list tool but leaves some behavioral details implicit.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence with no filler. It states the action and object efficiently and is appropriately sized for a tool with no parameters.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter list tool with an output schema, the description is complete enough for correct invocation and expected result. It could further mention that listing roles is useful before ask_expert, but nothing essential to calling the tool is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters and the input schema is empty, so there is no parameter information for the description to add. Baseline for a zero-parameter tool is 4, and no additional semantics are required.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('List') and a specific resource ('expert roles available in this agent team'). It distinguishes the tool from siblings ask_expert (which consults a role) and run_team (which executes the team), so an agent can tell them apart immediately.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives no guidance on when to use this tool versus ask_expert or run_team. There is no mention of prerequisites, exclusions, or a suggested sequence, so the agent must infer the usage context from the tool name alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

run_teamC

Run the connected PM, architecture, programming, QA and security workflow.

ParametersJSON Schema
NameRequiredDescriptionDefault
taskYes
providerNo
project_pathNo.

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of disclosing side effects and behavioral traits, but 'Run ... workflow' only suggests broad orchestration. It does not state whether the tool mutates files, invokes sub-agents, requires permissions, or returns any result.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence with no filler, and the core action is front-loaded. It is concise to the point of under-specification, but the wording is efficient and readable.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool that orchestrates a multi-role workflow, this definition is far too thin. With no annotations, no output schema, and no parameter explanations, an agent lacks essential context about task format, project scope, provider behavior, and expected outcomes.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description needed to explain what 'task', 'provider', and 'project_path' mean. It mentions none of these parameters, leaving the agent without any semantic grounding for the required input.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a clear action ('Run') on a defined resource: the connected PM, architecture, programming, QA, and security workflow. It is not a tautology and gives an agent a reasonable idea of what the tool orchestrates, though it does not explicitly distinguish it from the sibling expert tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided about when to use this tool versus list_experts or ask_expert. The description mentions no prerequisites, exclusions, or decision criteria, so an agent cannot determine when run_team is the right choice.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 3 tool updatesv0.1.0
    • First observedask_expert
    • First observedlist_experts
    • First observedrun_team

TDQS

B3.4/5.0

Scored across 3 tools

Disambiguation5/5

Each tool has a clear, distinct purpose: listing expert roles, consulting a single expert, and running the full team workflow. There is no ambiguity or overlap between them.

Naming Consistency5/5

All tool names follow a consistent verb_noun pattern using snake_case: list_experts, ask_expert, run_team. The naming is uniform and predictable.

Tool Count5/5

Three tools is a well-scoped set for an agent team lab. Each tool addresses a core interaction—discover, consult, execute—without unnecessary bloat.

Completeness4/5

The core workflows are covered: listing experts, asking one expert, and running the full team. Minor gaps exist such as no explicit status/result polling or workflow cancellation, but the surface is reasonably complete for its scope.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    C
    maintenance
    A reusable AI software development team built on MCP. 13 specialized agents (Project Manager, Backend, Frontend, QA, Security, DevOps, UX, and more) collaborate via shared SQLite state. Exposes 44 MCP tools across 12 domains. All discussions and decisions are stored in the database so agents get back up to speed immediately when re-loaded.
    MIT
  • A
    license
    B
    quality
    F
    maintenance
    Enables orchestrating multiple AI CLI agents (Claude Code, Codex, Gemini CLI, Copilot CLI) through a unified MCP interface for task delegation, cross-agent comparison, and specialized tools like code review and debugging.
    14
    7 npm
    14
    MIT
  • F
    license
    Not graded
    quality
    D
    maintenance
    Enables collaborative software development through a team of specialized AI agents (Architect, Developer, Reviewer, Tester) communicating via the Model Context Protocol, with tools for file management, task tracking, and code review.
    -