agent-eval-mcp
by ad-github1
README.md
# Enterprise AI Agent Evaluation & Deployment Platform
A dependency-light evaluation platform for RAG/wiki-quality AI agents. It scores agent outputs for faithfulness, retrieval relevance, hallucination risk, latency, and cost; produces CI-friendly quality gates; emits regression/canary reports; and exposes the workflow through a lightweight MCP-style stdio tool server.
## What Is Included
- JSONL evaluation case format for RAG/wiki workflows.
- Deterministic checks for:
- faithfulness to retrieved context and reference answer,
- retrieval relevance against question and expected keywords,
- hallucination risk from unsupported answer content,
- latency and cost thresholds.
- 100+ synthetic case generator.
- CI/CD-style suite-level and case-level quality gates.
- Markdown and JSON evaluation reports.
- Regression report comparing baseline and candidate runs.
- Canary promotion policy with traffic ramp decisions.
- OpenTelemetry-compatible JSONL traces/metrics.
- MCP-style stdio server exposing evaluation tools.
## Quick Start
```bash
git clone https://github.com/ad-github1/ENTERPRISE-AI-AGENT-EVALUATION-PLATFORM.git
cd ENTERPRISE-AI-AGENT-EVALUATION-PLATFORM
PYTHONPATH=src python3 -m agent_eval_platform generate-cases --count 120 --out examples/wiki_eval_cases.jsonl
PYTHONPATH=src python3 -m agent_eval_platform evaluate \
--cases examples/wiki_eval_cases.jsonl \
--gate examples/quality_gate.json \
--variant candidate \
--json-out reports/eval_result.json \
--markdown-out reports/eval_report.md \
--traces-out reports/traces.jsonl
PYTHONPATH=src python3 -m unittest discover -s tests
```
## Testing
Run the test suite:
```bash
PYTHONPATH=src python3 -m unittest discover -s tests
```
Current local result:
```text
Ran 5 tests in 0.015s
OK
```
After installation:
```bash
pip install -e .
agent-eval evaluate --cases examples/wiki_eval_cases.jsonl --gate examples/quality_gate.json
agent-eval-mcp
```
## Case Format
Each JSONL row contains one evaluated agent run:
```json
{
"case_id": "case-0001",
"question": "What contribution is Ada Lovelace known for in mathematics?",
"reference_answer": "Ada Lovelace is known for Analytical Engine notes.",
"expected_keywords": ["Ada Lovelace", "Analytical Engine"],
"retrieved_docs": [
{"doc_id": "wiki-1", "title": "Ada Lovelace", "text": "...", "score": 0.94}
],
"agent_answer": "Ada Lovelace is known for Analytical Engine notes.",
"latency_ms": 240.5,
"cost_usd": 0.0031,
"tags": ["wiki", "rag"]
}
```
## CI Quality Gate
The evaluator exits non-zero when `--fail-on-gate` is used and thresholds fail:
```bash
PYTHONPATH=src python3 -m agent_eval_platform evaluate \
--cases examples/wiki_eval_cases.jsonl \
--gate examples/quality_gate.json \
--fail-on-gate
```
See `.github/workflows/agent-eval.yml` for a GitHub Actions example.
## MCP-Style Tool Server
Run:
```bash
PYTHONPATH=src python3 -m agent_eval_platform.mcp_server
```
Supported JSON-RPC methods:
- `initialize`
- `tools/list`
- `tools/call` with:
- `run_evaluation_suite`
- `compare_regression`
- `decide_canary`
This is intentionally stdio and dependency-free. It follows the MCP tool shape closely enough for local agent integration demos without requiring the MCP Python SDK.
## Canary Workflow
```bash
PYTHONPATH=src python3 -m agent_eval_platform canary \
--result reports/eval_result.json \
--config examples/canary_config.json \
--json-out reports/canary_decision.json
```
The decision is `hold`, `increase_traffic`, or `promote` based on suite quality and minimum case coverage.
## Evaluation Results
### Evaluation Run
```bash
PYTHONPATH=src python3 -m agent_eval_platform evaluate \
--cases examples/wiki_eval_cases.jsonl \
--gate examples/quality_gate.json \
--variant candidate \
--json-out reports/eval_result.json \
--markdown-out reports/eval_report.md \
--traces-out reports/traces.jsonl
```
### Aggregate Metrics
| Metric | Value |
| --------------------------- | --------: |
| Evaluation cases | 120 |
| Pass rate | 82.5% |
| Average faithfulness | 0.825 |
| Average retrieval relevance | 0.838 |
| Average hallucination risk | 0.153 |
| p50 latency | 392.58 ms |
| p95 latency | 663.40 ms |
| p99 latency | 872.87 ms |
| Average cost | $0.00393 |
| Total cost | $0.47158 |
### Canary Decision
```bash
PYTHONPATH=src python3 -m agent_eval_platform canary \
--result reports/eval_result.json \
--config examples/canary_config.json \
--json-out reports/canary_decision.json
```
Result:
```json
{
"action": "hold",
"next_traffic_percent": 10.0,
"reasons": [
"pass_rate 0.825 < 0.900"
]
}
```
The canary policy correctly blocked promotion because the candidate run did not meet the configured 90% pass-rate threshold. This demonstrates how the platform can prevent low-quality agent versions from being promoted automatically.
### Observability
The evaluation emits OpenTelemetry-style JSONL traces to:
```text
reports/traces.jsonl
```
Each case generates spans for faithfulness, retrieval relevance, hallucination risk, and final case-level pass/fail status, enabling debugging of failed agent responses.
TDQS
B3/5.0
Scored across 3 tools
Disambiguation5/5
Each tool has a distinct purpose: running evaluations, comparing results, and applying deployment policy. There is no overlap.
Naming Consistency5/5
All tools follow a consistent verb_noun pattern (run_evaluation_suite, compare_regression, decide_canary), making them predictable.
Tool Count4/5
With 3 tools, the server covers the core evaluation workflow without being bloated. Slightly minimal but appropriate for a focused server.
Completeness4/5
The tools cover the main lifecycle: run suite, compare results, decide promotion. Minor gaps like suite management exist, but the core workflow is complete.
Maintenance
ActivityStale
ResponsivenessNo issues