Technical Answer Validator
This server exposes one MCP tool for reviewing a single answer against a caller-provided rubric.
Call
evaluate_answerwithrubricandanswer.Provide required concepts (1–50), accepted synonyms keyed by concept, optional numeric requirements (
value,unit,tolerance), and optionalrequired_count.Get matched/missing concepts, numeric check details, score, verdict (
correct/partial/wrong), andreview_required: true.No question bank is bundled; it is an assistive practice tool, not an official exam grader.
Uses local stdio MCP with no API key; optional local usage counts and opt-in anonymous telemetry.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Technical Answer ValidatorValidate 'HTTP is stateless' against concepts: stateless; synonyms: connectionless; numeric: none"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Technical Answer Validator
A tiny, deterministic answer review tool for AI agents, available over MCP stdio and as a REST API. Each caller supplies the concepts, accepted synonyms, numeric requirements, and answer text for a single evaluation. It does not include a question bank or answer corpus.
This is an assistive practice tool, not an official certification exam grader. Keyword matching can miss semantically correct paraphrases and can accept misleading surface matches. Users should review the supplied rubric and every result.
Run locally
Python 3.10+; the REST server uses the standard library. The MCP adapter uses the official Python SDK 2.2.0.
$env:TAV_API_KEY = "replace-with-a-long-random-secret-at-least-24-characters"
python -m tav_apiThe service listens on 127.0.0.1:8080 by default. To change it, set TAV_HOST and TAV_PORT. Local single-key mode refuses to start without a TAV_API_KEY of at least 24 characters. For deployment, create a distinct key per caller with python scripts/create_api_key.py CLIENT_ID and configure TAV_API_KEYS as a JSON object mapping each client ID to the generated SHA-256 digest. Store the one-time raw key with that client; do not store or commit it in this repository. When TAV_API_KEYS is set, it takes precedence over TAV_API_KEY.
Related MCP server: mcp-llm-eval
MCP for AI agents
Install the pinned MCP SDK 2.2.0 in a virtual environment from the committed lockfile:
uv sync --locked
.\.venv\Scripts\Activate.ps1The stdio MCP server exposes one tool: evaluate_answer(rubric, answer). Configure an MCP host with the absolute path to the environment's Python and mcp_server.py. Example Claude Desktop configuration (replace paths):
{
"mcpServers": {
"technical-answer-validator": {
"command": "C:\\path\\to\\technical-answer-validator\\.venv\\Scripts\\python.exe",
"args": ["C:\\path\\to\\technical-answer-validator\\mcp_server.py"]
}
}
}After publishing to PyPI, run with uvx --from technical-answer-validator tav-mcp. For local development use the .venv Python plus mcp_server.py. For Codex CLI or another MCP host, use its stdio server configuration with that command and script path. Restart the host, then ask it to list tools and call evaluate_answer. The stdio transport is local to the user's agent host and needs no internet endpoint or API key.
The Codex TOML template is codex-mcp-config.example.toml; the Claude Desktop JSON template is claude-mcp-config.example.json. Replace both placeholder paths with absolute paths. Merge the block into the host configuration; do not overwrite other MCP servers or global settings.
Local MCP usage counts
The MCP server keeps local daily aggregate counts for tool calls, successful evaluations, invalid requests, and internal errors. It never stores rubric or answer text and sends nothing over the network. Counts are retained for 90 days in ~/.technical-answer-validator/usage.sqlite3 (Windows: the user's home directory). Run tav-mcp-stats to print a JSON report. Set TAV_ANALYTICS=off in the MCP server environment to disable counting; set TAV_ANALYTICS_DB to choose another local database path. These counts remain on the user's computer unless the user voluntarily shares the report.
To voluntarily share anonymous usage counts with the maintainers, enable remote telemetry in the MCP server environment only after reviewing this notice:
TAV_TELEMETRY=on
TAV_TELEMETRY_URL=https://technical-answer-validator-telemetry.chl1591204.workers.dev/v1/eventRemote telemetry is off by default. When enabled, each tool call sends only a random persistent installation ID, package version, and outcome (success, invalid_request, or error) to the configured Cloudflare Worker. The ID is random and is not derived from a user, device, or account identifier; it lets us estimate distinct installations, not distinct people. Answers, rubrics, tool outputs, IP addresses, host names, and request contents are not included in the event payload. Cloudflare necessarily receives network connection metadata such as the sender IP to deliver the request; our Worker does not log or store it. Analytics Engine retains event data for three months. Disable by setting TAV_TELEMETRY=off and optionally remove ~/.technical-answer-validator/telemetry-id to reset the random ID. Telemetry failures never interrupt evaluations.
The collector is deployed at the URL above and accepts events only at /v1/event. To read aggregate usage, create a Cloudflare API token restricted to Account Analytics Read for the account that owns the Worker, then run python scripts/query_telemetry.py from this repository. The script prompts for the token without echoing or saving it. The report groups call totals by UTC day, package version, and outcome, and estimates distinct installations; it does not identify people. Do not put the token in source control, an MCP configuration, or a command line.
PyPI 0.1.1 does not include this opt-in telemetry code. It will be available to package users in 0.1.2; collection starts only after that version is installed and a user explicitly enables both environment variables above. The local-only tav-mcp-stats command continues to work independently.
Request
POST /v1/evaluate
{
"rubric": {
"required_concepts": ["isolation", "lockout tag"],
"accepted_synonyms": {"isolation": ["energy isolation"]},
"numeric_requirements": [],
"required_count": 2
},
"answer": "Apply energy isolation and attach a lockout tag."
}accepted_synonyms keys must exactly match a concept. numeric_requirements is an optional array such as [{"value":"10","unit":"kN","tolerance":"0"}]. Numbers in an answer are only checked when explicit requirements are provided. Concept score is matched concepts / required_count (defaults to the number of concepts), capped at 1.0; numeric failures apply a 50% score penalty. The verdict thresholds are correct >= 0.8, partial >= 0.4, otherwise wrong.
Response and errors
Successful requests return api_version, status, score, verdict, matched/missing concepts, numeric check details, and review_required: true.
Errors use JSON { "error": { "code": "...", "message": "..." } }. Statuses include 400 (invalid request), 401 (missing/invalid key), 404, 405, 413 (body over 64 KiB), and 429 (over 60 requests/minute per client ID). The rate limit is 60 requests/minute per authenticated client ID, in memory, and resets when the process restarts. Daily request/success/client-error/rate-limited counters are persisted in SQLite without answer text and expire after 90 days. GET /v1/usage returns only the caller's current UTC-day counts. Retain and back up the usage volume as desired; it contains client IDs and aggregates only.
Privacy and deployment limits
The stdio MCP option runs locally and sends no usage information to a central service unless the user explicitly enables remote telemetry as described above. Local aggregate call counts can be disabled separately. The REST API has separate server-side aggregates per API client.
The server does not log request bodies or answers. It stores daily counts keyed by client ID and request timestamps in process memory for REST rate limiting. The stdio MCP option runs locally inside the agent host and sends no requests to this HTTP server. The REST API is containerized and keeps usage counters in a persistent volume. Before public service, terminate TLS at a reverse proxy, set proxy-level rate/concurrency limits, deploy from a secret manager, monitor the host, and publish a data-retention/contact policy. The app-level per-client rate limit resets on restart and is not a substitute for edge controls.
Verify
python -m unittest discover -s tests -vThe OpenAPI contract is in openapi.yaml; the draft official MCP Registry descriptor is server.json, with publication steps in PUBLISHING.md. Run locally with Docker Compose after copying .env.example to .env and adding a private key; Compose publishes the service only on loopback, so configure an HTTPS reverse proxy separately. compose.yaml persists aggregate usage in a named volume and applies a read-only root filesystem, dropped Linux capabilities, and resource limits.
These tests check API and grading behavior; they do not establish professional exam accuracy.
Available Tools
1 toolevaluate_answerA
Check an answer against caller-provided concepts, synonyms, and numeric requirements.
Rubric fields: required_concepts (1-50 strings), optional accepted_synonyms keyed by concept, optional numeric_requirements [{value, unit?, tolerance?}], and optional required_count. No question bank is bundled. Review the returned result; it is not an official exam grade.
| Name | Required | Description | Default |
|---|---|---|---|
| answer | Yes | ||
| rubric | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations exist, so the description carries the full burden. It usefully discloses two behavioral facts: no question bank is bundled (all rubric data must come from the caller) and the returned result is not an official exam grade (a reliability caveat). It does not explain scoring/partial-credit behavior or how missing concepts are treated, leaving meaningful gaps for an unannotated evaluation tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the action in the first sentence, then lists rubric fields compactly, then two short caveat sentences. No filler, though the field enumeration is dense and could be tidier.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained. Given the free-form rubric object, the field enumeration is exactly the missing structured information; the only residual gap is the meaning/format of the 'answer' string.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and rubric is an untyped object with additionalProperties:true, so the schema alone is nearly useless. The description compensates well by enumerating the rubric keys and their shapes (required_concepts 1-50 strings, accepted_synonyms keyed by concept, numeric_requirements {value, unit?, tolerance?}, required_count). The 'answer' parameter is never explained, which keeps it from a 5.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource: check an answer against a caller-supplied rubric of concepts, synonyms, and numeric requirements. The purpose is unambiguous even without siblings to distinguish from. It stops short of 5 only because no alternative/scope framing (e.g., what kinds of answers are supported) is given.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by 'caller-provided concepts, synonyms, and numeric requirements' and by the note that no question bank is bundled, which tells the agent it must supply the full rubric. There is no explicit when-to-use/when-not guidance or alternative tool named, so this is minimum-viable context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
1 tool update
v0.1.1- First observed
evaluate_answer
TDQS
Scored across 1 tool
With only a single tool, there is no possibility of overlap or misselection. The tool's purpose—validating an answer against caller-supplied concepts, synonyms, and numeric requirements—is unambiguous.
The lone tool uses a clear snake_case verb_noun convention (evaluate_answer) that is fully self-consistent. There are no competing names or styles to create inconsistency.
A single tool is at the thin end of the scale given the rubric explicitly treats 1-2 tools as borderline. The validation domain is narrow enough that one monolithic tool is defensible, but there is no granularity for separate concerns.
The tool covers the core validation surface well: required concepts, synonyms, numeric requirements with tolerance, and required_count, with a sensible disclaimer about grade authority. Some gaps exist (e.g., no batch evaluation or question-bank management), but agents can work around them.
Maintenance
Related MCP Connectors
MCP server providing access to the Scorecard API to evaluate and optimize LLM systems.
Pay-per-call AI evaluation MCP server. Score LLM outputs against benchmark rubrics via Workers AI.
Remote MCP server for deterministic educational practice-assessment score conversions.
Evidence-readiness MCP server: validate, audit, and score briefs, memos, and evidence packs.
Related MCP Servers
AlicenseAqualityAmaintenanceMCP server that gives AI coding agents direct access to evaluation tools.23Apache 2.0- AlicenseAqualityCmaintenanceA local MCP server that packages LLM evaluation gates as reusable CI/CD primitives, enabling AI agents to run datasets against models, score responses, and enforce quality thresholds.10MIT
- AlicenseBqualityDmaintenanceAn MCP-style stdio server for evaluating AI agent outputs, enabling CI-friendly quality gates, regression comparisons, and canary promotion decisions.3MIT
- AlicenseNot gradedqualityCmaintenanceMCP server for grounded agentic Q&A over customer feedback, exposing typed tools to query a feedback corpus and return answers with citations to specific record IDs or a refusal when unsupported.MIT