@cyanheads/evals-mcp-server
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@@cyanheads/evals-mcp-serverdraft an exact answer eval for 'What is 2+2?' with answer '4'"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Overview
Verifiable eval records, authored through a draft → review → surgical-revise → submit loop with server-enforced graders. Create a draft carrying its own executable grader, patch it surgically field by field, and submit through a committability gate that requires the gold to pass, a declared negative case to fail, and an independent verification to agree — then compile submitted records to JSONL, CSV, Inspect AI, or lm-evaluation-harness. Runs as a stdio process or a local Streamable HTTP server.
Tools
Tool | Description |
| Return the required and optional fields plus grader options for a task type. Call before drafting. |
| Create a draft eval record carrying its own grader; returns the parsed record, a review protocol, and a verification subagent prompt. |
| Read a draft or submitted record by id; the id is stable across submit. |
| Apply a surgical |
| Delete a draft record by id. Draft-only. |
| Run a grader spec against candidate answers and get PASS/REJECT per candidate, decoupled from any saved record. |
| Finalize a draft through the committability gate, then freeze it. |
| Browse and filter records by status, domain, task type, or tag. Returns a compact summary per record. |
| Compile submitted records to JSONL, CSV, Inspect AI, or lm-evaluation-harness and write the artifact under |
Resources
Resource | Description |
| A single draft or submitted record by id — the same payload |
All record data is also reachable through the tool surface — evals_get_record for a single record, evals_list_records to browse. The resource is a convenience mirror for clients that support resources, not the access path.
Related MCP server: agent-eval-mcp
Capability reference
evals_describe_schema tool
Static — derived from the record and grader Zod schemas, no disk or runtime state
task_typeis one ofnumeric,exact_answer,set_answer,mcq,regex_answer,json_answer,free_responseReturns the gold shape, applicable grader kind(s), required/optional fields, and per-type authoring notes (e.g.
mcqneedschoices,free_responseneeds anllm_rubricgrader)
evals_create_draft tool
Validates against the
task_typediscriminated union and persists the draft;mcqrequireschoices,free_responserequires anllm_rubricgraderRuns a self-consistency check — the grader must PASS against
goldand eachdiscrimination.positive, and REJECT eachdiscrimination.negativeReturns the normalized record, a per-field review protocol, a ready-to-paste verification-subagent prompt, and what's still required before submit
Accepts optional draft-time
verificationevidence andcaptures(EvalsIDs) when provenance is already in handTyped errors:
grader_unexecutable,task_type_constraint,mcq_choice_mismatchStays
draft— passing self-consistency proves the grader discriminates, not that the gold is correct
evals_get_record tool
Reads by
id, stable across submit — resolves whether the record is still a draft or already submittedReturns the full record, including its grader, discrimination cases, and verification evidence
not_foundwhen no record matches; recovery points toevals_list_records
evals_revise_draft tool
Explicit
set(dotted-path → value),append(dotted-path → array items), andunset(dotted paths) operations — never a full-record rewriteCannot target
task_typeor server-owned fields — start a new draft to change the discriminantRe-validates the full record shape and per-task-type constraints after the patch, and re-runs self-consistency since the grader may have moved
Returns the updated record and an itemized
changedlist (op, path, before, after)Draft-only —
record_frozenon a submitted idTyped errors:
not_found,record_frozen,invalid_patch_path,task_type_constraint,mcq_choice_mismatch
evals_discard_draft tool
Deletes a draft record by
draft_idDraft-only —
record_frozenwhen the id refers to a submitted recordA missing id reports
not_foundrather than a distinct "already discarded" error — effectively idempotent
evals_run_check tool
Runs a grader spec against one or more
candidates(strings, numbers, objects, or arrays) without touching a saved recordReturns PASS/REJECT and a
detailper candidate, plus theresolvedcomparison value (e.g. the math.js-evaluated numeric target)goldapplies only to gold-relative kinds (exact_match); it's a no-op for target-embedding kinds likenumericandmcqllm_rubriccannot run here — submission relies on recorded independent verification insteadTyped errors:
grader_unexecutable,mcq_choice_mismatch
evals_submit_draft tool
The committability gate: the gold must PASS its grader, ≥1 declared negative must be REJECTED, and a recorded, decorrelated independent verification must agree with the gold
Resolves and embeds any
capturesfromEVALS_CAPTURE_DIR, cross-checking the gold against the authoritative captured valueRejects duplicates by
content_hash;confirm(orEVALS_REQUIRE_CONFIRMATION) can require human confirmation through multi-round input before finalizingOn pass, flips the record to
submitted, stampssubmitted_atand achecksum, and freezes it; otherwise refuses and the record stays a draftfree_responseis admitted on recorded independent verification alone and flaggedserver_verified: falseTyped errors:
not_found,record_frozen,verification_incomplete,grader_failed_on_gold,verification_disagrees_with_gold,missing_negative_case,negative_case_passed,duplicate,decorrelation_violation,capture_unresolved,submit_declined
evals_list_records tool
Filters by
status(draft/submitted),domain,task_type, ortag; up to 500 per call (default 50)Returns a compact summary per record (id, status, task_type, domain, tags, timestamps), newest-first — not full records
Discloses truncation (
shown,cap, total count) when the limit is hit, so a partial set is never mistaken for the whole corpus
evals_export_records tool
Formats:
jsonl(lossless),csv(flattened, lossy summary),inspect(UK AISI Inspect AI),lm-eval(EleutherAI lm-evaluation-harness)Optional
domain/task_type/tagfilterOnly
submittedrecords are exported — drafts are skippedWrites the artifact under
exports/and returns its path, record count, byte size, and a short preview instead of dumping it inline
eval://record/{id} resource
Returns the same payload as
evals_get_record, asapplication/jsonidcomes fromevals_list_recordsor a draft/submit responsenot_foundwhen no record matches
Features
Built on @cyanheads/mcp-ts-core: stdio and Streamable HTTP transports, pluggable auth (none / jwt / oauth), swappable storage (in-memory, filesystem, Supabase, Cloudflare KV/R2/D1), structured logging with optional OpenTelemetry tracing.
Eval authoring:
A
draft → review → surgical-revise → submitloop, with the server acting as both scribe (normalize, persist, compile) and adversarial checker (runs the record's own grader, rejects what doesn't hold up)Records are a Zod
discriminatedUnionontask_type—numeric,exact_answer,set_answer,mcq,regex_answer,json_answer,free_responseA typed grader DSL serialized with each record — deterministic kinds (
numericvia math.js,exact_match,set_match,regex,mcq,json_match) run server-side;llm_rubricrelies on recorded independent verificationAn enforced committability gate at submit: the gold must pass its own grader, ≥1 negative must be rejected, and a recorded decorrelated verification must agree with the gold
Plain JSON files under
EVALS_DATA_DIR— inspectable, diffable, version-controllable records, with drafts, submitted records, and exports kept separate
Agent-friendly output:
Instructional responses —
evals_create_draftandevals_revise_draftreturn the parsed record parroted back, a per-field review protocol, and a ready-to-paste verification-subagent promptSelf-consistency verdicts — every draft/revise response reports per-positive and per-negative pass/reject results, not just a boolean
Truncation disclosure —
evals_list_recordsreportsshown/cap/ total count when the limit is hit, so a partial set is never mistaken for the whole corpusTyped refusal — the submit gate fails with a typed
reasonplus a recovery hint, so a rejected record tells the agent exactly what to fix
Getting started
Add the following to your MCP client configuration file. Set EVALS_DATA_DIR to a writable folder — the server manages drafts/, submitted/, and exports/ under it.
{
"mcpServers": {
"evals-mcp-server": {
"type": "stdio",
"command": "bunx",
"args": ["@cyanheads/evals-mcp-server@latest"],
"env": {
"MCP_TRANSPORT_TYPE": "stdio",
"MCP_LOG_LEVEL": "info",
"EVALS_DATA_DIR": "/absolute/path/to/evals-data"
}
}
}
}Or with npx (no Bun required):
{
"mcpServers": {
"evals-mcp-server": {
"type": "stdio",
"command": "npx",
"args": ["-y", "@cyanheads/evals-mcp-server@latest"],
"env": {
"MCP_TRANSPORT_TYPE": "stdio",
"MCP_LOG_LEVEL": "info",
"EVALS_DATA_DIR": "/absolute/path/to/evals-data"
}
}
}
}Or with Docker:
{
"mcpServers": {
"evals-mcp-server": {
"type": "stdio",
"command": "docker",
"args": [
"run", "-i", "--rm",
"-e", "MCP_TRANSPORT_TYPE=stdio",
"-e", "EVALS_DATA_DIR=/data",
"-v", "evals-data:/data",
"ghcr.io/cyanheads/evals-mcp-server:latest"
]
}
}
}For Streamable HTTP, set the transport and start the server:
MCP_TRANSPORT_TYPE=http MCP_HTTP_PORT=3010 EVALS_DATA_DIR=./evals-data bun run start:http
# Server listens at http://localhost:3010/mcpPrerequisites
Bun v1.4.0 or higher (or Node.js v24+).
A writable directory for
EVALS_DATA_DIR. No external API key is required.
Installation
Clone the repository:
git clone https://github.com/cyanheads/evals-mcp-server.gitNavigate into the directory:
cd evals-mcp-serverInstall dependencies:
bun installConfigure environment:
cp .env.example .env
# edit .env and set EVALS_DATA_DIRConfiguration
All server configuration is validated at startup via Zod schemas in src/config/server-config.ts.
Variable | Description | Default |
| Root folder for record JSON; the store manages |
|
| When |
|
| Default | — |
| Directory of framework-written tool-call captures; when set, | — |
| Transport: |
|
| Port for the HTTP server. |
|
| HTTP session mode: |
|
| Auth mode: |
|
| Log level (RFC 5424). |
|
| Enable OpenTelemetry instrumentation. |
|
See .env.example for the full list of optional overrides.
Running the server
Local development
Build and run:
# One-time build bun run rebuild # Run the built server bun run start:stdio # or bun run start:httpRun checks and tests:
bun run devcheck # Lint, format, typecheck, security bun run test # Vitest test suite bun run lint:mcp # Validate MCP definitions against spec
Docker
docker build -t evals-mcp-server .
docker run --rm -e MCP_TRANSPORT_TYPE=stdio -e EVALS_DATA_DIR=/data -v evals-data:/data evals-mcp-serverThe Dockerfile defaults to HTTP transport, stateful session mode, and logs to /var/log/evals-mcp-server. OpenTelemetry peer dependencies are installed by default — build with --build-arg OTEL_ENABLED=false to omit them.
Project structure
Directory | Purpose |
|
|
| Server-specific environment variable parsing and validation with Zod. |
| Tool definitions ( |
| Resource definitions ( |
| The record schema, draft builder, and submit gate. |
| Deterministic grader DSL execution and the committability check. |
| On-disk JSON record CRUD, the draft→submitted move, and export writes. |
| Compiling submitted records to JSONL/CSV/Inspect/lm-eval. |
| Unit and integration tests mirroring |
Development guide
See CLAUDE.md/AGENTS.md for development guidelines and architectural rules. The short version:
Handlers throw, framework catches — no
try/catchin tool logicUse
ctx.logfor request-scoped logging; records persist to disk via therecord-storeservice, notctx.stateRegister new tools and resources in the
createApp()arrays insrc/index.tsThe server is the source of truth — validate inputs, run the grader as a hard gate, and never admit a record on assertion alone
Contributing
Issues are welcome. Run checks and tests before submitting:
bun run devcheck
bun run testLicense
Apache-2.0 — see LICENSE for details.
This server cannot be deployed
Maintenance
Related MCP Connectors
MCP-native AI evaluation: rubric audits, eval suites, and proof reports for AI/LLM output.
Prompt evals over MCP: run a prompt on your dataset, score each output 1-5 with an LLM judge.
Remote MCP for Gemini upgrade evals, prompt regressions, output diffs, and eval receipts.
Pay-per-call AI evaluation MCP server. Score LLM outputs against benchmark rubrics via Workers AI.
Related MCP Servers
- FlicenseNot gradedqualityDmaintenanceMCP-based code evaluation harness — sandboxed execution + LLM quality scoring.-
- AlicenseBqualityDmaintenanceAn MCP-style stdio server for evaluating AI agent outputs, enabling CI-friendly quality gates, regression comparisons, and canary promotion decisions.3MIT

Agent-Townofficial
AlicenseNot gradedqualityDmaintenanceA neutral verification court for AI tools that ranks MCP servers by executing them against ground truth and recording results. Enables agents to consult execution records, contribute verdicts, and challenge claims.Apache 2.0- AlicenseBqualityAmaintenanceEnables building and running custom LLM benchmarks with multi-judge evaluation, supporting GUI, MCP client, and CLI usage for ranked, auditable results.3MIT