agent-regress
by Gaurav7974
README.md
# agent-regress
agent-regress is a local MCP (Model Context Protocol) server that provides tools for AI coding agents to execute, evaluate, and diff regression tests for LLM agents. It integrates directly into your existing AI coding environment via the Model Context Protocol (MCP), allowing you to run regression tests against your agents whenever you change a prompt or model. To ensure fully trustworthy, verifiable regression detection, the MVP relies exclusively on deterministic code checks (exact match, regex, JSON schema, exit code) rather than delegating pass/fail judgment to an AI.
## Prerequisites
- Node.js (v18 or higher recommended)
- An MCP-compatible client (e.g., Claude Code, Cursor, or OpenCode)
## Installation
Currently, `agent-regress` runs from a local repository clone. (NPM publishing is not yet configured).
1. Clone this repository to your local machine.
2. Ensure Node.js (v18+) is installed.
3. Run `npm install` to install dependencies.
4. Run `npm run build` to compile the TypeScript source.
## Configuration
You must add this server to your MCP client's configuration file.
> **Note**: Only Kilo Code has been verified end-to-end in our testing. The configuration examples below for Claude Code, Cursor, OpenCode, and others are untested and provided based on general documentation. Contributions and corrections are welcome!
### For Claude Code
Edit `~/Library/Application Support/Claude/claude_desktop_config.json` (macOS) or the equivalent location on your OS. Add the following to the `mcpServers` object:
```json
{
"mcpServers": {
"agent-regress": {
"command": "node",
"args": ["/absolute/path/to/agent-regress/dist/index.js"]
}
}
}
```
### For Cursor
Edit `~/.cursor/mcp.json` and add the same configuration block.
## Writing `evals.yaml`
Create a file named `evals.yaml` in your project directory. This file defines the tests your agent must pass. The server reads this file from the current working directory where the MCP server is running.
```yaml
evals:
- name: "Check Error Extraction"
type: "CLI"
command: "python sentry_agent.py --issue 'TypeError: null reference'"
pass_criteria:
type: "deterministic"
check_type: "regex"
value: "Priority: (High|Medium)"
- name: "Check Health Endpoint"
type: "HTTP"
command: "http://localhost:3000/health"
method: "GET"
pass_criteria:
type: "deterministic"
check_type: "json_schema"
value: '{"type":"object","properties":{"status":{"const":"ok"}}}'
- name: "Check Apologetic Tone"
type: "CLI"
command: "python sentry_agent.py --issue 'API is slow'"
pass_criteria:
type: "semantic"
description: "The response must contain an apology for the inconvenience."
```
## Running Evals
*(Note: If you run `npx agent-regress` directly in a terminal, it will appear to hang. This is expected behavior! It is an MCP stdio server waiting for a client to connect and communicate via stdin/stdout, not a standalone CLI tool.)*
Once configured and your `evals.yaml` file is in place, prompt your AI client:
> "Run my agent evals using EVAL_PROMPT.md instructions."
The AI will call the agent-regress MCP tools to execute the tests and generate a report.
## Expected Output
The AI will output a markdown table summarizing the results. A regression (a test that previously passed but now fails) will be marked clearly.
| Eval Name | Status | Reason | Evaluated By |
| --- | --- | --- | --- |
| Check Error Extraction | 🔴 (Fail) | Output did not contain "Priority: High" | deterministic_code |
| Check Health Endpoint | 🟢 (Pass) | Output matched JSON schema | deterministic_code |
| Check Apologetic Tone | ⚪ (Skipped) | SKIPPED_SEMANTIC_MVP | host_ai |
**Regression Detected:** Check Error Extraction flipped from 🟢 to 🔴.
## Current Limitations
- **Semantic Evaluation Not Executed:** The MVP does not execute semantic checks. They will be marked as `SKIPPED_SEMANTIC_MVP`. This is intentional to ensure all reported results are 100% grounded in deterministic code.
- **No Severity Scoring:** The tool reports raw pass/fail and regressions. It does not assign severity scores or block gates based on severity.
This server cannot be deployed
Maintenance
ActivityStale
ResponsivenessNo issues