eval
Evaluate retrieval quality against a golden query set, returning nDCG@10, recall@10, and MRR per query and overall for CI regression detection.
Instructions
Run the ranking-quality harness against a golden query set and return nDCG@10 / recall@10 / MRR per query and aggregated. Indexless in the sense that it never builds — consumes whatever index already lives at the project root (run index first if missing). Intended as a CI regression guard. MCP defaults to json: true so agents receive structured EvalReport JSON instead of the human-readable summary the CLI emits.
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| json | No | Emit the EvalReport as JSON to stdout. Default `true` in MCP context (agents want structured output) — note the CLI default is `false`. Set explicitly to `false` to fall back to the human-readable summary. | |
| bench | No | Path to the golden-set TOML. Defaults to the bundled `benches/ranking_golden/queries.toml` on the CLI side; pass this when running against a fixture. | |
| min_ndcg | No | Fail with non-zero exit if mean nDCG@10 drops below this floor. Default 0.0 (always succeed). CI pins a recorded floor. | |
| project_root | No | Absolute path to the project root (defaults to the MCP working directory) |