Test Reporter MCP Server
Server Configuration
Describes the environment variables required to run the server.
| Name | Required | Description | Default |
|---|---|---|---|
| TEST_LEDGER_API_KEY | Yes | Your API key from the dashboard | |
| TEST_LEDGER_API_URL | No | Custom API URL (default: https://app-api.testledger.dev) | https://app-api.testledger.dev |
| TEST_LEDGER_PROJECT_ID | No | Default project ID to use for queries |
Instructions
Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.
This server publishes no instructions, or was last inspected before Glama recorded them.
Capabilities
Features and capabilities supported by this server
Protocol revision2025-11-25
| Capability | Details |
|---|---|
| tools | {} |
Tools
Functions exposed to the LLM to take actions
| Name | Description |
|---|---|
| get_test_historyB | Get historical pass/fail/flaky statistics for a specific test. Use this to understand how often a test fails and its overall reliability. Returns health_status (healthy/flaky/broken/disabled/insufficient_data) from the test_health view for AI pre-filtering decisions. |
| get_failure_patternsB | Analyze when and how tests fail to identify patterns. Returns failure rates by hour, day of week, version, browser/site, and duration analysis. |
| get_correlated_failuresA | Find tests that tend to fail together with a given test. High correlation suggests shared setup issues, test pollution, or dependencies. |
| get_flaky_testsA | Get a list of flaky tests (tests that fail then pass on retry) across the project, sorted by flakiness rate. Note: This scans all tests - use smaller 'days' values for faster results. |
| get_flaky_specsA | Get flaky specs from pre-computed materialized view. Faster than get_flaky_tests as it uses cached data refreshed hourly. Returns spec-level flakiness (not individual test level). |
| get_recent_failuresA | Get the most recent test failures for quick triage. Useful for seeing what's currently broken. For faster results, provide a spec_file filter. |
| get_test_trendB | Get trend data for a test over time, useful for seeing if a test is getting more or less reliable. |
| get_failure_screenshotsB | Get screenshots from recent test failures. Returns presigned S3 URLs that can be viewed with the Read tool to see exactly what the UI looked like when the test failed. |
| get_consecutive_failuresA | Get tests that are failing consecutively (broken tests, not flaky). Returns tests where the last 2+ runs have failed, with timing info (last_passed_date, first_failed_date) useful for identifying which merge broke them. |
Prompts
Interactive templates invoked by user choice
| Name | Description |
|---|---|
No prompts | |
Resources
Contextual data attached and managed by the client
| Name | Description |
|---|---|
No resources | |
TDQS
Scored across 9 tools
Most tools target distinct analytical questions (patterns, correlation, screenshots, consecutive failures). The main overlap is get_flaky_tests vs get_flaky_specs (test-level vs spec-level) and some tension between get_test_history and get_test_trend, but the descriptions explicitly clarify the differences. Boundaries are mostly clear.
Every tool follows the consistent get_<entity> snake_case pattern (get_failure_patterns, get_test_history, get_flaky_tests, etc.). Naming is fully predictable with no stylistic deviations.
Nine tools is well within a sensible range for a test analytics domain and each tool maps to a distinct analytical query. No bloat or redundancy that would suggest padding.
The surface covers failure patterns, flakiness, correlation, trends, screenshots, and consecutive failures well. Minor gaps exist—no tool to fetch individual run detail or list/search tests directly—but core triage and reliability-analysis workflows are covered.