Skip to main content
Glama

测试运行器 MCP

模型上下文协议 (MCP) 服务器,用于运行和解析来自多个测试框架的测试结果。该服务器提供统一的接口来执行测试并处理其输出,支持:

  • Bats(Bash 自动测试系统)

  • Pytest(Python测试框架)

  • Flutter 测试

  • Jest(JavaScript 测试框架)

  • Go 测试

  • 防锈测试(货物测试)

  • 通用(用于任意命令执行)

安装

npm install test-runner-mcp

Related MCP server: Tailscale MCP Server

先决条件

需要针对各自的测试类型安装以下测试框架:

用法

配置

将测试运行器添加到您的 MCP 设置中(例如,在claude_desktop_config.json或cline_mcp_settings.json中):

{
  "mcpServers": {
    "test-runner": {
      "command": "node",
      "args": ["/path/to/test-runner-mcp/build/index.js"],
      "env": {
        "NODE_PATH": "/path/to/test-runner-mcp/node_modules",
        // Flutter-specific environment (required for Flutter tests)
        "FLUTTER_ROOT": "/opt/homebrew/Caskroom/flutter/3.27.2/flutter",
        "PUB_CACHE": "/Users/username/.pub-cache",
        "PATH": "/opt/homebrew/Caskroom/flutter/3.27.2/flutter/bin:/usr/local/bin:/usr/bin:/bin"
      }
    }
  }
}

注意:对于 Flutter 测试,请确保替换:

  • /opt/homebrew/Caskroom/flutter/3.27.2/flutter替换为你的实际 Flutter 安装路径

  • /Users/username/.pub-cache替换为你的实际 pub 缓存路径

  • 更新 PATH 以包含系统的实际路径

您可以通过运行以下命令找到这些值:

# Get Flutter root
flutter --version

# Get pub cache path
echo $PUB_CACHE   # or default to $HOME/.pub-cache

# Get Flutter binary path
which flutter

运行测试

使用具有以下参数的run_tests工具:

{
  "command": "test command to execute",
  "workingDir": "working directory for test execution",
  "framework": "bats|pytest|flutter|jest|go|rust|generic",
  "outputDir": "directory for test results",
  "timeout": "test execution timeout in milliseconds (default: 300000)",
  "env": "optional environment variables",
  "securityOptions": "optional security options for command execution"
}

每个框架的示例:

// Bats
{
  "command": "bats test/*.bats",
  "workingDir": "/path/to/project",
  "framework": "bats",
  "outputDir": "test_reports"
}

// Pytest
{
  "command": "pytest test_file.py -v",
  "workingDir": "/path/to/project",
  "framework": "pytest",
  "outputDir": "test_reports"
}

// Flutter
{
  "command": "flutter test test/widget_test.dart",
  "workingDir": "/path/to/project",
  "framework": "flutter",
  "outputDir": "test_reports",
  "FLUTTER_ROOT": "/opt/homebrew/Caskroom/flutter/3.27.2/flutter",
  "PUB_CACHE": "/Users/username/.pub-cache",
  "PATH": "/opt/homebrew/Caskroom/flutter/3.27.2/flutter/bin:/usr/local/bin:/usr/bin:/bin"
}

// Jest
{
  "command": "jest test/*.test.js",
  "workingDir": "/path/to/project",
  "framework": "jest",
  "outputDir": "test_reports"
}

// Go
{
  "command": "go test ./...",
  "workingDir": "/path/to/project",
  "framework": "go",
  "outputDir": "test_reports"
}

// Rust
{
  "command": "cargo test",
  "workingDir": "/path/to/project",
  "framework": "rust",
  "outputDir": "test_reports"
}

// Generic (for arbitrary commands, CI/CD tools, etc.)
{
  "command": "act -j build",
  "workingDir": "/path/to/project",
  "framework": "generic",
  "outputDir": "test_reports"
}

// Generic with security overrides
{
  "command": "sudo docker-compose -f docker-compose.test.yml up",
  "workingDir": "/path/to/project",
  "framework": "generic",
  "outputDir": "test_reports",
  "securityOptions": {
    "allowSudo": true
  }
}

安全功能

测试运行器包含内置的安全功能,以防止执行潜在的有害命令,特别是对于generic框架:

  1. 命令验证

    • 默认阻止sudo和su

    • 防止危险命令,如rm -rf /

    • 阻止在安全位置之外进行文件系统写入操作

  2. 环境变量清理

    • 过滤掉潜在的危险环境变量

    • 防止覆盖关键系统变量

    • 确保安全路径处理

  3. 可配置的安全性

    • 必要时通过securityOptions覆盖安全限制

    • 对安全功能的细粒度控制

    • 标准测试使用的默认安全设置

您可以配置的安全选项:

{
  "securityOptions": {
    "allowSudo": false,        // Allow sudo commands
    "allowSu": false,          // Allow su commands
    "allowShellExpansion": true, // Allow shell expansion like $() or backticks
    "allowPipeToFile": false   // Allow pipe to file operations (> or >>)
  }
}

Flutter 测试支持

测试运行器包括对 Flutter 测试的增强支持:

  1. 环境设置

    • 自动 Flutter 环境配置

    • PATH 和 PUB_CACHE 设置

    • Flutter 安装验证

  2. 错误处理

    • 堆栈跟踪收集

    • 断言错误处理

    • 异常捕获

    • 测试失败检测

  3. 输出处理

    • 完成测试输出捕获

    • 堆栈跟踪保存

    • 详细的错误报告

    • 原始输出保存

Rust 测试支持

测试运行器为 Rust 的cargo test提供了特定的支持:

  1. 环境设置

    • 自动设置 RUST_BACKTRACE=1 以获得更好的错误消息

  2. 输出解析

    • 解析单个测试结果

    • 捕获失败测试的详细错误消息

    • 识别被忽略的测试

    • 提取摘要信息

通用测试支持

对于 CI/CD 管道、通过act执行的 GitHub Actions 或任何其他命令执行,通用框架提供:

  1. 自动输出分析

    • 尝试将输出分割成逻辑块

    • 标识节标题

    • 检测通过/失败指标

    • 即使对于未知格式也能提供合理的输出结构

  2. 灵活集成

    • 可与任意 shell 命令配合使用

    • 无特定格式要求

    • 非常适合与act 、Docker 和自定义脚本等工具集成

  3. 安全功能

    • 命令验证以防止有害操作

    • 必要时可配置为允许特定的提升权限

输出格式

测试运行器生成结构化输出,同时保留完整的测试输出:

interface TestResult {
  name: string;
  passed: boolean;
  output: string[];
  rawOutput?: string;  // Complete unprocessed output
}

interface TestSummary {
  total: number;
  passed: number;
  failed: number;
  duration?: number;
}

interface ParsedResults {
  framework: string;
  tests: TestResult[];
  summary: TestSummary;
  rawOutput: string;  // Complete command output
}

结果保存在指定的输出目录中:

  • test_output.log :原始测试输出

  • test_errors.log :错误消息(如果有)

  • test_results.json :结构化测试结果

  • summary.txt :人类可读的摘要

发展

设置

  1. 克隆存储库

  2. 安装依赖项:

    npm install
  3. 构建项目:

    npm run build

运行测试

npm test

该测试套件包括所有受支持框架的测试,并验证成功和失败的测试场景。

持续集成/持续交付

该项目使用 GitHub Actions 进行持续集成:

  • 在 Node.js 18.x 和 20.x 上进行自动化测试

  • 测试结果已上传为工件

  • Dependabot 配置为自动依赖项更新

贡献

  1. 分叉存储库

  2. 创建你的功能分支

  3. 提交你的更改

  4. 推送到分支

  5. 创建拉取请求

执照

该项目根据 MIT 许可证获得许可 - 有关详细信息,请参阅LICENSE文件。

Available Tools

1 tool
run_testsC

Run tests and capture output

ParametersJSON Schema
NameRequiredDescriptionDefault
commandYesTest command to execute (e.g., "bats tests/*.bats")
envNoEnvironment variables for test execution
frameworkYesTesting framework being used
outputDirNoDirectory to store test results
securityOptionsNoSecurity options for command execution
timeoutNoTest execution timeout in milliseconds (default: 300000)
workingDirYesWorking directory for test execution

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries full burden but offers minimal behavioral insight. It mentions 'capture output' but doesn't describe output format, error handling, side effects, or security implications. The description doesn't contradict annotations (none exist), but fails to disclose important behavioral traits.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely concise at just four words, front-loaded with the core action and outcome. Every word earns its place with zero redundancy or unnecessary elaboration.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex tool with 7 parameters, nested objects, no annotations, and no output schema, the description is insufficient. It doesn't explain return values, error conditions, security considerations, or typical usage patterns that would help an agent understand this execution tool's behavior.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, providing detailed parameter documentation. The description adds no additional parameter semantics beyond the schema's comprehensive coverage, so it meets the baseline of 3 for high schema coverage without compensating value.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Run tests and capture output' clearly states the action (run tests) and outcome (capture output), but lacks specificity about what types of tests or how they're executed. It doesn't distinguish from siblings (none exist), but remains somewhat vague about scope and implementation details.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives, prerequisites, or typical scenarios. With no sibling tools mentioned, differentiation isn't needed, but there's still no context about appropriate use cases or limitations.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 1 tool updatev1.0.0
    • First observedrun_tests

TDQS

C2.9/5.0

Scored across 1 tool

Disambiguation5/5

With only one tool, there is no possibility of ambiguity or overlap between tools. The tool's purpose is clearly defined as running tests and capturing output, making it distinct by default.

Naming Consistency5/5

The single tool name 'run_tests' follows a clear verb_noun pattern, which is consistent within this minimal set. There are no other tools to compare against, so no inconsistency can arise.

Tool Count2/5

A single tool is generally too few for most server purposes, as it limits functionality and flexibility. For a test runner, one tool might cover basic execution but lacks operations like listing tests, filtering, or managing test suites, making it feel thin and under-scoped.

Completeness2/5

The tool surface is severely incomplete for a test runner domain. It only provides execution without supporting operations such as listing available tests, retrieving results, configuring test runs, or handling test environments, leading to significant gaps that could cause agent failures.

Maintenance

ActivityInactive
ResponsivenessUnresponsive

Related MCP Connectors

  • Direct access to Cypress tests results and accessibility reports in your AI workflow.

  • Approved test intent, reviewed Playwright automation and run evidence, inside your editor.

  • Manage test suites, run tests, view results, and automate QA workflows via AI with testRigor.

  • Run, debug, and triage tests from your IDE using natural language, no dashboard switching, no manual data transfers. The TestMu AI (formerly LambdaTest) MCP Server is a single remote server exposing four tool suites: HyperExecute — analyze your project, generate YAML configs and test runner commands, then monitor jobs and sessions. Automation — pull a TestID's details plus command, network, and console logs into one chat for instant root-cause analysis. Includes mobile app upload. SmartUI — explain pixel, layout, DOM, and perceptual changes in a visual regression run, with context-aware React/HTML/CSS fixes. Accessibility — audit any public URL or a local React app against WCAG and get ready-to-apply remediation steps. Connects over https://mcp.lambdatest.com/mcp using OAuth 2.1 — no API keys in your config. One-click install in Cursor; works with Claude, GitHub Copilot, Cline, and any MCP client. Tests execute on the TestMu AI cloud: 3,000+ browsers and 10,000+ real devices.

Related MCP Servers

  • A
    license
    Not graded
    quality
    B
    maintenance
    Facilitates isolated code execution within Docker containers, enabling secure multi-language script execution and integration with language models like Claude via the Model Context Protocol.
    5
    MIT
  • A
    license
    Not graded
    quality
    D
    maintenance
    Provides a standardized interface for interacting with Rocketlane's tools and services through the Model Context Protocol, enabling unified API access.
    1
    MIT
  • A
    license
    Not graded
    quality
    D
    maintenance
    Provides a standardized interface for interacting with Neon's tools and services through a unified API via the Model Context Protocol.
    MIT