Gaffer MCP Server
This server connects AI assistants to Gaffer test analytics, providing test health, history, coverage, and report/upload tools via MCP.
Test health & trends: project health score, pass rate, flaky test count, trend direction, slowest tests, and comparisons between commits/runs.
Test history & flaky detection: pass/fail history by test name/path, flaky test identification, root-cause clustering of failures, and search across past failures by error, stack trace, or test name.
Test runs & results: list recent runs with filters (commit, branch, status) and get detailed results with stack traces.
Coverage analysis: overall coverage summary, file/path coverage, untested files, and high-risk areas with both low coverage and failures.
Reports & uploads: raw report URLs (HTML/JSON/XML), signed browser-navigable links (30 min), upload status checks, and test result uploads (the only write operation; rate-limited and audit-logged).
Project management: list all accessible projects using a user API key and discover available functions.
Agentic workflows: chain up to 20 API calls in a single execute_code invocation (30s timeout) for complex tasks.
Click on "Install Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Gaffer MCP ServerWhich tests are flaky in my project?"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
@gaffer-sh/mcp
MCP (Model Context Protocol) server for Gaffer - give your AI assistant memory of your tests.
What is this?
This MCP server connects AI coding assistants like Claude Code and Cursor to your Gaffer test history and coverage data. It runs in code mode: three MCP tools over a namespace of 17 functions — 16 read-only analytics functions plus upload_test_results. It allows AI to:
Check your project's test health (pass rate, flaky tests, trends)
Look up the history of specific tests to understand stability
Get context about test failures when debugging
Analyze code coverage and identify untested areas
Browse all your projects (with user API Keys)
Access test report files (HTML reports, coverage, etc.)
Related MCP server: Koalr
Prerequisites
A Gaffer account with test results uploaded
An API Key from Account Settings > API Keys
Setup
Claude Code (CLI)
The easiest way to add the Gaffer MCP server is via the Claude Code CLI:
claude mcp add gaffer -e GAFFER_API_KEY=gaf_your_api_key_here -- npx -y @gaffer-sh/mcpClaude Code (Manual)
Alternatively, add to your Claude Code settings (~/.claude.json or project .claude/settings.json):
{
"mcpServers": {
"gaffer": {
"command": "npx",
"args": ["-y", "@gaffer-sh/mcp"],
"env": {
"GAFFER_API_KEY": "gaf_your_api_key_here"
}
}
}
}Cursor
Add to .cursor/mcp.json in your project:
{
"mcpServers": {
"gaffer": {
"command": "npx",
"args": ["-y", "@gaffer-sh/mcp"],
"env": {
"GAFFER_API_KEY": "gaf_your_api_key_here"
}
}
}
}How this server works
This server uses code mode. Instead of exposing one MCP tool per API call, it exposes three tools plus a codemode namespace you call from JavaScript. Fewer tool definitions occupy the context window, and a single execution can chain several calls.
MCP tool | What it does |
| Run JavaScript against |
| Find available functions by keyword. An empty query lists all of them. |
| List projects. Registered only when the token is a user API Key ( |
const health = await codemode.get_project_health({ projectId: "proj_abc" });
if (health.flakyTestCount > 0) {
const flaky = await codemode.get_flaky_tests({ projectId: "proj_abc" });
return { health, flaky };
}
return { health };Functions available via execute_code
Function | Category | Description |
| health | Health score, pass rate, flaky count, trend |
| testing | Pass/fail history for a specific test |
| testing | Tests with high flip rates (pass↔fail) |
| testing | Recent test runs, filterable by commit/branch/status |
| testing | Parsed individual results for one run |
| testing | Failed tests grouped by root cause |
| testing | Slowest tests by P95 duration |
| testing | Compare test performance between commits or runs |
| testing | Search failures by error or test-name pattern, or list all recent failures |
| coverage | Overall coverage metrics and trend |
| coverage | Coverage for specific files or paths |
| coverage | Files below a coverage threshold |
| coverage | Files with low coverage AND test failures |
| reports | Report file URLs for a test run |
| reports | Signed browser-navigable report URL (30 min) |
| uploads | Whether CI results are uploaded and processed |
| uploads | Upload test results (write) — rate-limited and audit-logged |
Every function except upload_test_results is read-only.
Function Reference
list_projects
List all projects you have access to.
Input:
organizationId(optional),limit(optional, default: 50)Returns: List of projects with IDs, names, and organization info
Example: "What projects do I have in Gaffer?"
get_project_health
Get the health metrics for a project.
Input:
projectId(required),days(optional, default: 30)Returns: Health score (0-100), pass rate, test run count, flaky test count, trend
Example: "What's the health of my test suite?"
get_test_history
Get the pass/fail history for a specific test.
Input:
projectId(required),testNameorfilePath(one required),limit(optional)Returns: History of runs with status, duration, branch, commit, errors
Example: "Is the login test flaky? Check its history"
get_flaky_tests
Get the list of flaky tests in a project.
Input:
projectId(required),threshold(optional, default: 0.1),days(optional),limit(optional)Returns: List of flaky tests with flip rates, transition counts, run counts
Example: "Which tests are flaky in my project?"
list_test_runs
List recent test runs with optional filtering.
Input:
projectId(required),commitSha(optional),branch(optional),status(optional),limit(optional)Returns: List of test runs with pass/fail/skip counts, commit and branch info
Example: "What tests failed in the last commit?"
get_test_run_details
Get parsed test results for a specific test run.
Input:
testRunId(required),projectId(required),status(optional filter),limit(optional)Returns: Individual test results with name, status, duration, file path, errors
Example: "Show me all failed tests from this test run"
get_report
Get URLs for report files uploaded with a test run.
Input:
testRunId(required)Returns: List of files with filename, size, content type, download URL
Example: "Get the Playwright report for the latest test run"
get_report_browser_url
Get a browser-navigable URL for viewing a test report.
Input:
projectId(required),testRunId(required),filename(optional)Returns: Signed URL valid for 30 minutes
Example: "Give me a link to view the test report"
get_slowest_tests
Get the slowest tests in a project, sorted by P95 duration.
Input:
projectId(required),days(optional),limit(optional),framework(optional),branch(optional)Returns: List of tests with average and P95 duration, run count
Example: "Which tests are slowing down my CI pipeline?"
compare_test_metrics
Compare test metrics between two commits or test runs.
Input:
projectId(required),testName(required),beforeCommit/afterCommitORbeforeRunId/afterRunIdReturns: Before/after metrics with duration change and percentage
Example: "Did my fix make this test faster?"
get_coverage_summary
Get the coverage metrics summary for a project.
Input:
projectId(required),days(optional, default: 30)Returns: Line/branch/function coverage percentages, trend, report count, lowest coverage files
Example: "What's our test coverage?"
get_coverage_for_file
Get coverage metrics for specific files or paths.
Input:
projectId(required),filePath(required - exact or partial match)Returns: List of matching files with line/branch/function coverage
Example: "What's the coverage for our API routes?"
get_untested_files
Get files with little or no test coverage.
Input:
projectId(required),maxCoverage(optional, default: 10%),limit(optional)Returns: List of files below threshold sorted by coverage (lowest first)
Example: "Which files have no tests?"
find_uncovered_failure_areas
Find code areas with both low coverage AND test failures (high risk).
Input:
projectId(required),days(optional),coverageThreshold(optional, default: 80%)Returns: Risk areas ranked by score, with file path, coverage %, failure count
Example: "Where should we focus our testing efforts?"
get_failure_clusters
Group failed tests by root cause using error message similarity.
Input:
projectId(required),testRunId(required)Returns: Clusters of failed tests grouped by similar error messages, with representative error and test count
Example: "Are these 15 failures from the same bug?"
search_failures
Search past failures by error message, stack trace, or test name — or list every failure in the window.
Input:
query(optional — omit to return all failures),projectId(required forgaf_keys),searchIn(optional:errors/names/all, defaultall),days(optional, default: 30),branch(optional),limit(optional, default: 20)Returns: Matching failures with test name, error message, run and commit context, plus
truncatedwhen scan caps cut the list shortExample: "Have we seen this connection-refused error before?" / "What failed in the last 7 days?"
get_upload_status
Check if CI results have been uploaded and processed.
Input:
projectId(required),sessionId(optional),commitSha(optional),branch(optional)Returns: Upload session(s) with processing status, linked test runs and coverage reports
Example: "Are my test results ready for commit abc123?"
upload_test_results
Upload structured test results. This is the only function that writes.
Use it when you have results in hand — parsed from CI output or a runner's JSON report — and no Gaffer CLI is available to upload them.
Input:
projectId(required forgaf_keys),framework(required),tests(required),branch,commitSha,ciProvider,startedAt,finishedAt,coverageReturns:
uploadSessionId, the generatedrunId, and the derived pass/fail/skip summaryExample: "Upload these 42 parsed pytest results so we can track them"
runId, the run timestamps and the summary are derived from tests — pass
startedAt/finishedAt only if you know the real wall-clock window.
Two constraints worth knowing:
Not idempotent. Each call creates a new run, so a retry after an uncertain failure produces a duplicate. Check
get_upload_statusinstead of retrying.Rate-limited per project, and every call is written to the project's audit log with the id of the credential that made it.
Processing is asynchronous: results take a few seconds to become visible to the read functions.
Agentic CI Workflows
These workflows show how an AI agent diagnoses CI failures, waits for results, and finds coverage gaps. Each step is a codemode function, so a whole chain runs inside one execute_code call rather than one round-trip per step.
Workflow: Diagnose CI Failures
list_test_runs(projectId, status="failed")
→ get_test_run_details(projectId, testRunId, status="failed")
→ get_failure_clusters(projectId, testRunId)
→ get_test_history(projectId, testName="...")
→ compare_test_metrics(projectId, testName, beforeCommit, afterCommit)Find the failed test run
Get individual failure details with stack traces
Group failures by root cause — often 15 failures are 2-3 bugs
Check if each failure is new (regression) or recurring
Verify fixes by comparing before/after
Workflow: Wait for Results
get_upload_status(projectId, commitSha="abc123")
→ poll until processingStatus="completed"
→ get_test_run_details(projectId, testRunId)Check if results for a commit have been uploaded
Wait for processing to complete
Use linked test run IDs to get results
Workflow: Find Coverage Gaps
find_uncovered_failure_areas(projectId)
→ get_untested_files(projectId)
→ get_coverage_for_file(projectId, filePath="src/critical/")Find files with both low coverage and test failures (highest risk)
Find files with no coverage at all
Drill into specific directories for targeted analysis
Function Quick Reference
Agent Question | Function |
"What failed?" |
|
"Same root cause?" |
|
"Seen this error before?" |
|
"Is it flaky?" |
|
"Is this new?" |
|
"Did my fix work?" |
|
"Are results ready?" |
|
"What's untested?" |
|
"What's slow?" |
|
Prioritizing Coverage Improvements
When using coverage tools to improve your test suite, combine coverage data with codebase exploration for best results:
1. Understand Code Utilization
Before targeting files purely by coverage percentage, explore which code is actually critical:
Find entry points: Look for route definitions, event handlers, exported functions - these reveal what code actually executes in production
Find heavily-imported files: Files imported by many others are high-value targets
Identify critical business logic: Look for files handling auth, payments, data mutations, or core domain logic
2. Prioritize by Impact
Low coverage alone doesn't indicate priority. Consider:
High utilization + low coverage = highest priority - Code that runs frequently but lacks tests
Large files with 0% coverage - More uncovered lines means bigger impact on overall coverage
Files with both failures and low coverage - Use
find_uncovered_failure_areasfor this
3. Use Path-Based Queries
The get_untested_files tool may return many frontend components. For backend or specific areas:
# Query specific paths with get_coverage_for_file
get_coverage_for_file(filePath="server/services")
get_coverage_for_file(filePath="src/api")
get_coverage_for_file(filePath="lib/core")4. Iterative Improvement
Get baseline with
get_coverage_summaryIdentify targets with
get_coverage_for_fileon critical pathsWrite tests for highest-impact files
Re-check coverage after CI uploads new results
Repeat
Authentication
User API Keys (Recommended)
User API Keys (gaf_ prefix) provide read-only access to all projects across your organizations. Get your API Key from: Account Settings > API Keys
Project Tokens
Project Tokens (gfr_ prefix) are designed for uploading test results and only provide access to a single project. When you use one, omit projectId — it resolves automatically. User API Keys are preferred for the MCP server because they enable list_projects and read across projects.
Environment Variables
Variable | Required | Description |
| Yes | Your Gaffer API Key (starts with |
| No | API base URL (default: |
Local Development
pnpm install
pnpm buildTest locally with Claude Code (use absolute path to built file):
{
"mcpServers": {
"gaffer": {
"command": "node",
"args": ["/absolute/path/to/dist/index.js"],
"env": {
"GAFFER_API_KEY": "gaf_..."
}
}
}
}License
MIT
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Servers
- AlicenseBqualityAmaintenanceIntegrates with Zebrunner Test Case Management to help QA teams manage test cases, test suites, and test execution data through AI assistants. Features intelligent validation, automated test code generation, and comprehensive reporting capabilities.49905AGPL 3.0
- Alicense-qualityCmaintenanceConnect engineering metrics, DORA performance, deploy risk scoring, and PR health to any AI assistant. Score PRs for deployment risk using a 36-signal model, query team health, incidents, coverage, and more.MIT

Tesults MCPofficial
Alicense-qualityDmaintenanceConnect AI agents to your test results, insights, and targets. Query test runs, failures, flaky tests, and regressions across frameworks including Playwright, Jest, Pytest, Cypress and more.37MIT- Alicense-qualityDmaintenanceEnables Claude Code to analyze flaky tests, find failure patterns, suggest fixes, and query test history, failure patterns, and correlated failures from Test Ledger.20MIT
Related MCP Connectors
BuildPulse CI test analytics for AI agents — flaky tests, coverage, and CI run history.
Generate answers & visualizations from your engineering data to track software development health.
Connect AI assistants to your GitHub-hosted Obsidian vault to seamlessly access, search, and analy…
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/gaffer-sh/mcp'
If you have feedback or need assistance with the MCP directory API, please join our Discord server