Skip to main content
Glama

Ejecutor de pruebas MCP

Un servidor de Protocolo de Contexto de Modelo (MCP) para ejecutar y analizar resultados de pruebas de múltiples entornos de prueba. Este servidor proporciona una interfaz unificada para ejecutar pruebas y procesar sus resultados, y admite:

  • Bats (Sistema de pruebas automatizadas Bash)

  • Pytest (marco de pruebas de Python)

  • Pruebas de Flutter

  • Jest (Marco de pruebas de JavaScript)

  • Pruebas Go

  • Pruebas de óxido (prueba de carga)

  • Genérico (para ejecución de comandos arbitrarios)

Instalación

npm install test-runner-mcp

Related MCP server: Tailscale MCP Server

Prerrequisitos

Es necesario instalar los siguientes marcos de prueba para sus respectivos tipos de prueba:

Uso

Configuración

Agregue el ejecutor de pruebas a su configuración de MCP (por ejemplo, en claude_desktop_config.json o cline_mcp_settings.json ):

{
  "mcpServers": {
    "test-runner": {
      "command": "node",
      "args": ["/path/to/test-runner-mcp/build/index.js"],
      "env": {
        "NODE_PATH": "/path/to/test-runner-mcp/node_modules",
        // Flutter-specific environment (required for Flutter tests)
        "FLUTTER_ROOT": "/opt/homebrew/Caskroom/flutter/3.27.2/flutter",
        "PUB_CACHE": "/Users/username/.pub-cache",
        "PATH": "/opt/homebrew/Caskroom/flutter/3.27.2/flutter/bin:/usr/local/bin:/usr/bin:/bin"
      }
    }
  }
}

Nota: Para las pruebas de Flutter, asegúrese de reemplazar:

  • /opt/homebrew/Caskroom/flutter/3.27.2/flutter con su ruta de instalación real de Flutter

  • /Users/username/.pub-cache con su ruta de caché de publicación actual

  • Actualice PATH para incluir las rutas reales de su sistema

Puede encontrar estos valores ejecutando:

# Get Flutter root
flutter --version

# Get pub cache path
echo $PUB_CACHE   # or default to $HOME/.pub-cache

# Get Flutter binary path
which flutter

Ejecución de pruebas

Utilice la herramienta run_tests con los siguientes parámetros:

{
  "command": "test command to execute",
  "workingDir": "working directory for test execution",
  "framework": "bats|pytest|flutter|jest|go|rust|generic",
  "outputDir": "directory for test results",
  "timeout": "test execution timeout in milliseconds (default: 300000)",
  "env": "optional environment variables",
  "securityOptions": "optional security options for command execution"
}

Ejemplo para cada marco:

// Bats
{
  "command": "bats test/*.bats",
  "workingDir": "/path/to/project",
  "framework": "bats",
  "outputDir": "test_reports"
}

// Pytest
{
  "command": "pytest test_file.py -v",
  "workingDir": "/path/to/project",
  "framework": "pytest",
  "outputDir": "test_reports"
}

// Flutter
{
  "command": "flutter test test/widget_test.dart",
  "workingDir": "/path/to/project",
  "framework": "flutter",
  "outputDir": "test_reports",
  "FLUTTER_ROOT": "/opt/homebrew/Caskroom/flutter/3.27.2/flutter",
  "PUB_CACHE": "/Users/username/.pub-cache",
  "PATH": "/opt/homebrew/Caskroom/flutter/3.27.2/flutter/bin:/usr/local/bin:/usr/bin:/bin"
}

// Jest
{
  "command": "jest test/*.test.js",
  "workingDir": "/path/to/project",
  "framework": "jest",
  "outputDir": "test_reports"
}

// Go
{
  "command": "go test ./...",
  "workingDir": "/path/to/project",
  "framework": "go",
  "outputDir": "test_reports"
}

// Rust
{
  "command": "cargo test",
  "workingDir": "/path/to/project",
  "framework": "rust",
  "outputDir": "test_reports"
}

// Generic (for arbitrary commands, CI/CD tools, etc.)
{
  "command": "act -j build",
  "workingDir": "/path/to/project",
  "framework": "generic",
  "outputDir": "test_reports"
}

// Generic with security overrides
{
  "command": "sudo docker-compose -f docker-compose.test.yml up",
  "workingDir": "/path/to/project",
  "framework": "generic",
  "outputDir": "test_reports",
  "securityOptions": {
    "allowSudo": true
  }
}

Características de seguridad

El ejecutor de pruebas incluye funciones de seguridad integradas para evitar la ejecución de comandos potencialmente dañinos, en particular para el marco generic :

  1. Validación de comandos

    • Bloquea sudo y su por defecto

    • Previene comandos peligrosos como rm -rf /

    • Bloquea las operaciones de escritura del sistema de archivos fuera de ubicaciones seguras

  2. Saneamiento de variables ambientales

    • Filtra variables de entorno potencialmente peligrosas

    • Evita la anulación de variables críticas del sistema

    • Garantiza un manejo seguro de la ruta

  3. Seguridad configurable

    • Anular las restricciones de seguridad cuando sea necesario a través de securityOptions

    • Control detallado sobre las funciones de seguridad

    • Configuración segura predeterminada para el uso de pruebas estándar

Opciones de seguridad que puedes configurar:

{
  "securityOptions": {
    "allowSudo": false,        // Allow sudo commands
    "allowSu": false,          // Allow su commands
    "allowShellExpansion": true, // Allow shell expansion like $() or backticks
    "allowPipeToFile": false   // Allow pipe to file operations (> or >>)
  }
}

Soporte de pruebas de Flutter

El ejecutor de pruebas incluye soporte mejorado para pruebas de Flutter:

  1. Configuración del entorno

    • Configuración automática del entorno de Flutter

    • Configuración de PATH y PUB_CACHE

    • Verificación de la instalación de Flutter

  2. Manejo de errores

    • Recopilación de seguimientos de pila

    • Manejo de errores de aserción

    • Captura de excepciones

    • Detección de fallos en las pruebas

  3. Procesamiento de salida

    • Captura completa de la salida de prueba

    • Preservación del seguimiento de la pila

    • Informe detallado de errores

    • Conservación de la salida sin procesar

Soporte para pruebas de óxido

El ejecutor de pruebas proporciona soporte específico para cargo test de Rust:

  1. Configuración del entorno

    • Establece automáticamente RUST_BACKTRACE=1 para obtener mejores mensajes de error

  2. Análisis de salida

    • Analiza los resultados de pruebas individuales

    • Captura mensajes de error detallados para pruebas fallidas

    • Identifica las pruebas ignoradas

    • Extrae información resumida

Soporte de pruebas genéricas

Para las canalizaciones de CI/CD, las acciones de GitHub a través act o cualquier otra ejecución de comando, el marco genérico proporciona:

  1. Análisis automático de salida

    • Intenta segmentar la salida en bloques lógicos

    • Identifica los encabezados de sección

    • Detecta indicadores de aprobación/reprobación

    • Proporciona una estructura de salida razonable incluso para formatos desconocidos

  2. Integración flexible

    • Funciona con comandos de shell arbitrarios

    • No hay requisitos de formato específicos

    • Perfecto para la integración con herramientas como act , Docker y scripts personalizados.

  3. Características de seguridad

    • Validación de comandos para evitar operaciones dañinas

    • Se puede configurar para permitir permisos elevados específicos cuando sea necesario

Formato de salida

El ejecutor de pruebas produce una salida estructurada mientras preserva la salida de prueba completa:

interface TestResult {
  name: string;
  passed: boolean;
  output: string[];
  rawOutput?: string;  // Complete unprocessed output
}

interface TestSummary {
  total: number;
  passed: number;
  failed: number;
  duration?: number;
}

interface ParsedResults {
  framework: string;
  tests: TestResult[];
  summary: TestSummary;
  rawOutput: string;  // Complete command output
}

Los resultados se guardan en el directorio de salida especificado:

  • test_output.log : Salida de prueba sin procesar

  • test_errors.log : Mensajes de error, si los hay

  • test_results.json : Resultados de pruebas estructuradas

  • summary.txt : Resumen legible para humanos

Desarrollo

Configuración

  1. Clonar el repositorio

  2. Instalar dependencias:

    npm install
  3. Construir el proyecto:

    npm run build

Ejecución de pruebas

npm test

El conjunto de pruebas incluye pruebas para todos los marcos compatibles y verifica escenarios de pruebas exitosos y fallidos.

CI/CD

El proyecto utiliza GitHub Actions para la integración continua:

  • Pruebas automatizadas en Node.js 18.x y 20.x

  • Resultados de pruebas cargados como artefactos

  • Dependabot configurado para actualizaciones de dependencias automatizadas

Contribuyendo

  1. Bifurcar el repositorio

  2. Crea tu rama de funciones

  3. Confirme sus cambios

  4. Empujar hacia la rama

  5. Crear una solicitud de extracción

Licencia

Este proyecto está licenciado bajo la licencia MIT: consulte el archivo de LICENCIA para obtener más detalles.

Available Tools

1 tool
run_testsC

Run tests and capture output

ParametersJSON Schema
NameRequiredDescriptionDefault
commandYesTest command to execute (e.g., "bats tests/*.bats")
envNoEnvironment variables for test execution
frameworkYesTesting framework being used
outputDirNoDirectory to store test results
securityOptionsNoSecurity options for command execution
timeoutNoTest execution timeout in milliseconds (default: 300000)
workingDirYesWorking directory for test execution

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries full burden but offers minimal behavioral insight. It mentions 'capture output' but doesn't describe output format, error handling, side effects, or security implications. The description doesn't contradict annotations (none exist), but fails to disclose important behavioral traits.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely concise at just four words, front-loaded with the core action and outcome. Every word earns its place with zero redundancy or unnecessary elaboration.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex tool with 7 parameters, nested objects, no annotations, and no output schema, the description is insufficient. It doesn't explain return values, error conditions, security considerations, or typical usage patterns that would help an agent understand this execution tool's behavior.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, providing detailed parameter documentation. The description adds no additional parameter semantics beyond the schema's comprehensive coverage, so it meets the baseline of 3 for high schema coverage without compensating value.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Run tests and capture output' clearly states the action (run tests) and outcome (capture output), but lacks specificity about what types of tests or how they're executed. It doesn't distinguish from siblings (none exist), but remains somewhat vague about scope and implementation details.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives, prerequisites, or typical scenarios. With no sibling tools mentioned, differentiation isn't needed, but there's still no context about appropriate use cases or limitations.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 1 tool updatev1.0.0
    • First observedrun_tests

TDQS

C2.9/5.0

Scored across 1 tool

Disambiguation5/5

With only one tool, there is no possibility of ambiguity or overlap between tools. The tool's purpose is clearly defined as running tests and capturing output, making it distinct by default.

Naming Consistency5/5

The single tool name 'run_tests' follows a clear verb_noun pattern, which is consistent within this minimal set. There are no other tools to compare against, so no inconsistency can arise.

Tool Count2/5

A single tool is generally too few for most server purposes, as it limits functionality and flexibility. For a test runner, one tool might cover basic execution but lacks operations like listing tests, filtering, or managing test suites, making it feel thin and under-scoped.

Completeness2/5

The tool surface is severely incomplete for a test runner domain. It only provides execution without supporting operations such as listing available tests, retrieving results, configuring test runs, or handling test environments, leading to significant gaps that could cause agent failures.

Maintenance

ActivityInactive
ResponsivenessUnresponsive

Related MCP Connectors

  • Direct access to Cypress tests results and accessibility reports in your AI workflow.

  • Approved test intent, reviewed Playwright automation and run evidence, inside your editor.

  • Manage test suites, run tests, view results, and automate QA workflows via AI with testRigor.

  • Run, debug, and triage tests from your IDE using natural language, no dashboard switching, no manual data transfers. The TestMu AI (formerly LambdaTest) MCP Server is a single remote server exposing four tool suites: HyperExecute — analyze your project, generate YAML configs and test runner commands, then monitor jobs and sessions. Automation — pull a TestID's details plus command, network, and console logs into one chat for instant root-cause analysis. Includes mobile app upload. SmartUI — explain pixel, layout, DOM, and perceptual changes in a visual regression run, with context-aware React/HTML/CSS fixes. Accessibility — audit any public URL or a local React app against WCAG and get ready-to-apply remediation steps. Connects over https://mcp.lambdatest.com/mcp using OAuth 2.1 — no API keys in your config. One-click install in Cursor; works with Claude, GitHub Copilot, Cline, and any MCP client. Tests execute on the TestMu AI cloud: 3,000+ browsers and 10,000+ real devices.

Related MCP Servers

  • A
    license
    Not graded
    quality
    B
    maintenance
    Facilitates isolated code execution within Docker containers, enabling secure multi-language script execution and integration with language models like Claude via the Model Context Protocol.
    5
    MIT
  • A
    license
    Not graded
    quality
    D
    maintenance
    Provides a standardized interface for interacting with Rocketlane's tools and services through the Model Context Protocol, enabling unified API access.
    1
    MIT
  • A
    license
    Not graded
    quality
    D
    maintenance
    Provides a standardized interface for interacting with Neon's tools and services through a unified API via the Model Context Protocol.
    MIT