Test Runner MCP
Ejecutor de pruebas MCP
Un servidor de Protocolo de Contexto de Modelo (MCP) para ejecutar y analizar resultados de pruebas de múltiples entornos de prueba. Este servidor proporciona una interfaz unificada para ejecutar pruebas y procesar sus resultados, y admite:
Bats (Sistema de pruebas automatizadas Bash)
Pytest (marco de pruebas de Python)
Pruebas de Flutter
Jest (Marco de pruebas de JavaScript)
Pruebas Go
Pruebas de óxido (prueba de carga)
Genérico (para ejecución de comandos arbitrarios)
Instalación
npm install test-runner-mcpRelated MCP server: Tailscale MCP Server
Prerrequisitos
Es necesario instalar los siguientes marcos de prueba para sus respectivos tipos de prueba:
Murciélagos:
apt-get install batsobrew install batsPytest:
pip install pytestFlutter: siga la guía de instalación de Flutter
Broma:
npm install --save-dev jestGo: Siga la guía de instalación de Go
Rust: siga la guía de instalación de Rust
Uso
Configuración
Agregue el ejecutor de pruebas a su configuración de MCP (por ejemplo, en claude_desktop_config.json o cline_mcp_settings.json ):
{
"mcpServers": {
"test-runner": {
"command": "node",
"args": ["/path/to/test-runner-mcp/build/index.js"],
"env": {
"NODE_PATH": "/path/to/test-runner-mcp/node_modules",
// Flutter-specific environment (required for Flutter tests)
"FLUTTER_ROOT": "/opt/homebrew/Caskroom/flutter/3.27.2/flutter",
"PUB_CACHE": "/Users/username/.pub-cache",
"PATH": "/opt/homebrew/Caskroom/flutter/3.27.2/flutter/bin:/usr/local/bin:/usr/bin:/bin"
}
}
}
}Nota: Para las pruebas de Flutter, asegúrese de reemplazar:
/opt/homebrew/Caskroom/flutter/3.27.2/fluttercon su ruta de instalación real de Flutter/Users/username/.pub-cachecon su ruta de caché de publicación actualActualice PATH para incluir las rutas reales de su sistema
Puede encontrar estos valores ejecutando:
# Get Flutter root
flutter --version
# Get pub cache path
echo $PUB_CACHE # or default to $HOME/.pub-cache
# Get Flutter binary path
which flutterEjecución de pruebas
Utilice la herramienta run_tests con los siguientes parámetros:
{
"command": "test command to execute",
"workingDir": "working directory for test execution",
"framework": "bats|pytest|flutter|jest|go|rust|generic",
"outputDir": "directory for test results",
"timeout": "test execution timeout in milliseconds (default: 300000)",
"env": "optional environment variables",
"securityOptions": "optional security options for command execution"
}Ejemplo para cada marco:
// Bats
{
"command": "bats test/*.bats",
"workingDir": "/path/to/project",
"framework": "bats",
"outputDir": "test_reports"
}
// Pytest
{
"command": "pytest test_file.py -v",
"workingDir": "/path/to/project",
"framework": "pytest",
"outputDir": "test_reports"
}
// Flutter
{
"command": "flutter test test/widget_test.dart",
"workingDir": "/path/to/project",
"framework": "flutter",
"outputDir": "test_reports",
"FLUTTER_ROOT": "/opt/homebrew/Caskroom/flutter/3.27.2/flutter",
"PUB_CACHE": "/Users/username/.pub-cache",
"PATH": "/opt/homebrew/Caskroom/flutter/3.27.2/flutter/bin:/usr/local/bin:/usr/bin:/bin"
}
// Jest
{
"command": "jest test/*.test.js",
"workingDir": "/path/to/project",
"framework": "jest",
"outputDir": "test_reports"
}
// Go
{
"command": "go test ./...",
"workingDir": "/path/to/project",
"framework": "go",
"outputDir": "test_reports"
}
// Rust
{
"command": "cargo test",
"workingDir": "/path/to/project",
"framework": "rust",
"outputDir": "test_reports"
}
// Generic (for arbitrary commands, CI/CD tools, etc.)
{
"command": "act -j build",
"workingDir": "/path/to/project",
"framework": "generic",
"outputDir": "test_reports"
}
// Generic with security overrides
{
"command": "sudo docker-compose -f docker-compose.test.yml up",
"workingDir": "/path/to/project",
"framework": "generic",
"outputDir": "test_reports",
"securityOptions": {
"allowSudo": true
}
}Características de seguridad
El ejecutor de pruebas incluye funciones de seguridad integradas para evitar la ejecución de comandos potencialmente dañinos, en particular para el marco generic :
Validación de comandos
Bloquea
sudoysupor defectoPreviene comandos peligrosos como
rm -rf /Bloquea las operaciones de escritura del sistema de archivos fuera de ubicaciones seguras
Saneamiento de variables ambientales
Filtra variables de entorno potencialmente peligrosas
Evita la anulación de variables críticas del sistema
Garantiza un manejo seguro de la ruta
Seguridad configurable
Anular las restricciones de seguridad cuando sea necesario a través de
securityOptionsControl detallado sobre las funciones de seguridad
Configuración segura predeterminada para el uso de pruebas estándar
Opciones de seguridad que puedes configurar:
{
"securityOptions": {
"allowSudo": false, // Allow sudo commands
"allowSu": false, // Allow su commands
"allowShellExpansion": true, // Allow shell expansion like $() or backticks
"allowPipeToFile": false // Allow pipe to file operations (> or >>)
}
}Soporte de pruebas de Flutter
El ejecutor de pruebas incluye soporte mejorado para pruebas de Flutter:
Configuración del entorno
Configuración automática del entorno de Flutter
Configuración de PATH y PUB_CACHE
Verificación de la instalación de Flutter
Manejo de errores
Recopilación de seguimientos de pila
Manejo de errores de aserción
Captura de excepciones
Detección de fallos en las pruebas
Procesamiento de salida
Captura completa de la salida de prueba
Preservación del seguimiento de la pila
Informe detallado de errores
Conservación de la salida sin procesar
Soporte para pruebas de óxido
El ejecutor de pruebas proporciona soporte específico para cargo test de Rust:
Configuración del entorno
Establece automáticamente RUST_BACKTRACE=1 para obtener mejores mensajes de error
Análisis de salida
Analiza los resultados de pruebas individuales
Captura mensajes de error detallados para pruebas fallidas
Identifica las pruebas ignoradas
Extrae información resumida
Soporte de pruebas genéricas
Para las canalizaciones de CI/CD, las acciones de GitHub a través act o cualquier otra ejecución de comando, el marco genérico proporciona:
Análisis automático de salida
Intenta segmentar la salida en bloques lógicos
Identifica los encabezados de sección
Detecta indicadores de aprobación/reprobación
Proporciona una estructura de salida razonable incluso para formatos desconocidos
Integración flexible
Funciona con comandos de shell arbitrarios
No hay requisitos de formato específicos
Perfecto para la integración con herramientas como
act, Docker y scripts personalizados.
Características de seguridad
Validación de comandos para evitar operaciones dañinas
Se puede configurar para permitir permisos elevados específicos cuando sea necesario
Formato de salida
El ejecutor de pruebas produce una salida estructurada mientras preserva la salida de prueba completa:
interface TestResult {
name: string;
passed: boolean;
output: string[];
rawOutput?: string; // Complete unprocessed output
}
interface TestSummary {
total: number;
passed: number;
failed: number;
duration?: number;
}
interface ParsedResults {
framework: string;
tests: TestResult[];
summary: TestSummary;
rawOutput: string; // Complete command output
}Los resultados se guardan en el directorio de salida especificado:
test_output.log: Salida de prueba sin procesartest_errors.log: Mensajes de error, si los haytest_results.json: Resultados de pruebas estructuradassummary.txt: Resumen legible para humanos
Desarrollo
Configuración
Clonar el repositorio
Instalar dependencias:
npm installConstruir el proyecto:
npm run build
Ejecución de pruebas
npm testEl conjunto de pruebas incluye pruebas para todos los marcos compatibles y verifica escenarios de pruebas exitosos y fallidos.
CI/CD
El proyecto utiliza GitHub Actions para la integración continua:
Pruebas automatizadas en Node.js 18.x y 20.x
Resultados de pruebas cargados como artefactos
Dependabot configurado para actualizaciones de dependencias automatizadas
Contribuyendo
Bifurcar el repositorio
Crea tu rama de funciones
Confirme sus cambios
Empujar hacia la rama
Crear una solicitud de extracción
Licencia
Este proyecto está licenciado bajo la licencia MIT: consulte el archivo de LICENCIA para obtener más detalles.
Available Tools
1 toolrun_testsC
Run tests and capture output
| Name | Required | Description | Default |
|---|---|---|---|
| command | Yes | Test command to execute (e.g., "bats tests/*.bats") | |
| env | No | Environment variables for test execution | |
| framework | Yes | Testing framework being used | |
| outputDir | No | Directory to store test results | |
| securityOptions | No | Security options for command execution | |
| timeout | No | Test execution timeout in milliseconds (default: 300000) | |
| workingDir | Yes | Working directory for test execution |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries full burden but offers minimal behavioral insight. It mentions 'capture output' but doesn't describe output format, error handling, side effects, or security implications. The description doesn't contradict annotations (none exist), but fails to disclose important behavioral traits.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is extremely concise at just four words, front-loaded with the core action and outcome. Every word earns its place with zero redundancy or unnecessary elaboration.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex tool with 7 parameters, nested objects, no annotations, and no output schema, the description is insufficient. It doesn't explain return values, error conditions, security considerations, or typical usage patterns that would help an agent understand this execution tool's behavior.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, providing detailed parameter documentation. The description adds no additional parameter semantics beyond the schema's comprehensive coverage, so it meets the baseline of 3 for high schema coverage without compensating value.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description 'Run tests and capture output' clearly states the action (run tests) and outcome (capture output), but lacks specificity about what types of tests or how they're executed. It doesn't distinguish from siblings (none exist), but remains somewhat vague about scope and implementation details.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives, prerequisites, or typical scenarios. With no sibling tools mentioned, differentiation isn't needed, but there's still no context about appropriate use cases or limitations.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
1 tool update
v1.0.0- First observed
run_tests
TDQS
Scored across 1 tool
With only one tool, there is no possibility of ambiguity or overlap between tools. The tool's purpose is clearly defined as running tests and capturing output, making it distinct by default.
The single tool name 'run_tests' follows a clear verb_noun pattern, which is consistent within this minimal set. There are no other tools to compare against, so no inconsistency can arise.
A single tool is generally too few for most server purposes, as it limits functionality and flexibility. For a test runner, one tool might cover basic execution but lacks operations like listing tests, filtering, or managing test suites, making it feel thin and under-scoped.
The tool surface is severely incomplete for a test runner domain. It only provides execution without supporting operations such as listing available tests, retrieving results, configuring test runs, or handling test environments, leading to significant gaps that could cause agent failures.
Maintenance
Related MCP Connectors
Direct access to Cypress tests results and accessibility reports in your AI workflow.
Approved test intent, reviewed Playwright automation and run evidence, inside your editor.
Manage test suites, run tests, view results, and automate QA workflows via AI with testRigor.
Run, debug, and triage tests from your IDE using natural language, no dashboard switching, no manual data transfers. The TestMu AI (formerly LambdaTest) MCP Server is a single remote server exposing four tool suites: HyperExecute — analyze your project, generate YAML configs and test runner commands, then monitor jobs and sessions. Automation — pull a TestID's details plus command, network, and console logs into one chat for instant root-cause analysis. Includes mobile app upload. SmartUI — explain pixel, layout, DOM, and perceptual changes in a visual regression run, with context-aware React/HTML/CSS fixes. Accessibility — audit any public URL or a local React app against WCAG and get ready-to-apply remediation steps. Connects over https://mcp.lambdatest.com/mcp using OAuth 2.1 — no API keys in your config. One-click install in Cursor; works with Claude, GitHub Copilot, Cline, and any MCP client. Tests execute on the TestMu AI cloud: 3,000+ browsers and 10,000+ real devices.
Related MCP Servers
- AlicenseNot gradedqualityBmaintenanceFacilitates isolated code execution within Docker containers, enabling secure multi-language script execution and integration with language models like Claude via the Model Context Protocol.5MIT
- AlicenseBqualityAmaintenanceProvides seamless integration with Tailscale's CLI commands and REST API, enabling automated network management and monitoring through a standardized Model Context Protocol interface.18603 npm136MIT
- AlicenseNot gradedqualityDmaintenanceProvides a standardized interface for interacting with Rocketlane's tools and services through the Model Context Protocol, enabling unified API access.1MIT
- AlicenseNot gradedqualityDmaintenanceProvides a standardized interface for interacting with Neon's tools and services through a unified API via the Model Context Protocol.MIT