qa-toolkit-mcp
Server Configuration
Describes the environment variables required to run the server.
| Name | Required | Description | Default |
|---|---|---|---|
| QA_TOOLKIT_RUNS_DIR | No | Path to the folder containing test run reports. Defaults to './runs/'. | ./runs/ |
Instructions
Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.
This server publishes no instructions, or was last inspected before Glama recorded them.
Capabilities
Features and capabilities supported by this server
Protocol revision2025-11-25
| Capability | Details |
|---|---|
| tools | {
"listChanged": false
} |
| prompts | {
"listChanged": false
} |
| resources | {
"subscribe": false,
"listChanged": false
} |
| experimental | {} |
Tools
Functions exposed to the LLM to take actions
| Name | Description |
|---|---|
| qa_list_runsA | List available test runs from the configured runs directory. |
| qa_get_runA | Return a single test run by id. |
| qa_compare_runsA | Compare two test runs and categorize the differences. |
Prompts
Interactive templates invoked by user choice
| Name | Description |
|---|---|
| weekly_regression_review | Slash command: end-of-week regression review. The host injects this as a user message. The model is expected to use qa_list_runs / qa_get_run / qa_compare_runs to fulfill it. Args: suite: Optional suite filter (e.g., "api-regression"). Empty = all suites. days: How many days back to consider. Default 7. |
Resources
Contextual data attached and managed by the client
| Name | Description |
|---|---|
No resources | |
TDQS
Scored across 3 tools
The three tools have clearly distinct purposes: listing runs, getting a single run's details, and comparing two runs. There is no overlap or ambiguity.
All tool names follow the consistent pattern 'qa_<verb>_runs' (list_runs, get_run, compare_runs), using snake_case and a clear verb_noun structure.
Three tools is on the lower end but still reasonable for a focused test run analysis toolkit. The count matches the scope of listing, retrieving, and comparing runs.
The set covers the core operations of browsing and comparing test runs, but it lacks tools for creating, updating, or deleting runs, and flakiness detection is explicitly omitted. There are notable gaps for full lifecycle management.