jev-mcp
Provides a baseline adapter in the evaluation harness for comparing JEV judgments against OpenAI models, supporting logprobs-based probability extraction and verbalized probabilities.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@jev-mcpEvaluate this ticket: is it urgent, and which team should handle it?"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
JEV MCP · Structured Judgments & LLM Evaluation
Connect JEV to MCP clients and compare its judgments against general-purpose LLMs using shared datasets and measurable accuracy.
Agents are good at producing text and bad at producing answers you can branch on. Ask one "is this
ticket urgent?" and you get back a sentence you then have to parse, with no number attached — no way
to tell a confident yes from a coin flip. This server exposes TypeSafe's Jev classifier as a single
MCP tool that returns typed answers with probabilities, so the agent gets 0.94 and moves on.
agent ──── evaluate ────▶ jev-mcp ──── POST /v1/systemone ────▶ TypeSafe
(stdio) (this) (Jev)
◀─── typed JSON ─── ◀─── probabilities ────────
jev-eval ── same question ─▶ JEV ─┐
─▶ LLM ─┴─▶ Brier · calibration · McNemarTwo halves: an MCP server that exposes the classifier to your agents, and an evaluation harness that tells you whether it is actually beating whatever you were using before.
Requirements
Node.js 20 or newer
A TypeSafe API key from console.typesafe.ai
Related MCP server: Jev MCP
Install
Not published to npm yet, so install from the repository:
npm install -g github:arunav25/jev-mcpOr clone it, which is what you want if you plan to run the evaluation harness:
git clone https://github.com/arunav25/jev-mcp.git
cd jev-mcp && npm install && npm linkThe bare name
jev-mcpon npm belongs to an unrelated project. This package publishes as@arunav25/jev-mcp; until it is published, use one of the commands above.
Then point your agents at it. The key has to be in your environment before you run this, because agents launch the server without your shell, so its value is written into each client's config:
export TYPESAFE_API_KEY=sk-...
jev-mcp installThat registers the server with Claude Code, Claude Desktop and Codex, skipping any that aren't installed. Restart Claude Desktop afterwards. To see what it would do first:
jev-mcp install --dry-run
jev-mcp install --client codex # or limit it to oneAny MCP client that speaks stdio will do. Run jev-mcp doctor to get the exact launch command, then:
{
"mcpServers": {
"jev": {
"command": "/usr/local/bin/node",
"args": ["/usr/local/lib/node_modules/@arunav25/jev-mcp/src/cli.js", "serve"],
"env": { "TYPESAFE_API_KEY": "sk-..." }
}
}
}Using it
One tool, evaluate. Give it the material to judge and one or more questions:
{
"state": {
"subject": "Payouts failing",
"body": "Help! My payouts have been failing for 3 days."
},
"questions": {
"is_urgent": {
"type": "noul",
"instructions": "Does the sender need a response today rather than this week?"
},
"department": {
"type": "choice",
"instructions": "Which team should own this ticket?",
"criteria": {
"billing": "Payments, payouts, refunds, invoices",
"technical": "Bugs, outages, API errors",
"sales": "Pricing and plan changes",
"unclear": "Not enough information to route"
}
}
}
}Back comes the answer under the same keys you used:
{
"model": "jev-latest",
"answers": {
"is_urgent": { "type": "noul", "noul": 0.94 },
"department": {
"type": "choice",
"choice": "billing",
"probabilities": { "billing": 0.88, "technical": 0.09, "sales": 0.02, "unclear": 0.01 },
"confidence": 0.81
}
}
}Question types
Type | Criteria | Returns |
| optional |
|
| required map of option → description or |
|
| required ordered array of 2+ level descriptions |
|
score answers are 0-indexed: N levels answer between 0 and N-1, so 3.87 over 5 levels sits
between the fourth and fifth level — not 3.87 out of 5. Read it through the legend the response
returns, which names each index.
Getting good answers
stateis the only thing a question can see. Put every fact the decision rests on there.Question keys aren't sent to the model. Calling one
urgentexplains nothing — the instructions have to define what urgent means here.One judgment per question. Two decisions in one set of instructions blur the distribution.
Questions in a call can't see each other's answers. They run together over the same state. If step two depends on step one, make two calls.
Give
choicean escape hatch when the input might match nothing.Put observations in
state, not conclusions. A verdict you already reached reads as evidence for itself, and the probability comes back as your own conclusion with a number attached. Word instructions as the condition to test, not the answer you expect.scorelevels have to describe real situations. A bare 1–5 scale gives the model nothing to anchor on.0.5 on a
noulmeans uncertain, not "medium amount of the thing you asked about". Andconfidencemeasures how concentrated the distribution is, not whether the answer is right.
Malformed questions are caught here, before the request goes out — a choice with no criteria or a
score with one level comes back as a message the agent can fix, not a 422 and a wasted round trip.
Measuring whether it's actually better
jev-eval is a harness for answering "which one is more accurate?" with a number that survives
scrutiny. It exists because the usual version of that comparison — run ten inputs through both, eyeball
the outputs — cannot detect anything. Here is the arithmetic:
Effect to detect | Disagreement rate | Labelled items needed |
90% → 92% (2 points) | 15% | 2,941 |
90% → 92% (2 points) | 25% | 4,904 |
85% → 90% (5 points) | 20% | 626 |
80% → 90% (10 points) | 25% | 194 |
70% → 85% (15 points) | 30% | 103 |
80% power, α = 0.05, paired McNemar. Run jev-eval power --from 0.9 --to 0.92 for your own numbers.
A two-point gap over ten calls is three orders of magnitude short of conclusive. Any ranking drawn from it is a coin flip wearing a number.
The loop
jev-eval init datasets/urgency --question "Does this message need a response today rather than this week?"
# put your real items in datasets/urgency/items.jsonl, one JSON object per line:
# {"id": "t-001", "state": {"subject": "...", "body": "..."}}
jev-eval label datasets/urgency --rater rater-1 # never shows you a model's answer
jev-eval label datasets/urgency --rater rater-2 # a second rater bounds what's resolvable
jev-eval agreement datasets/urgency # Cohen's kappa between the two
jev-eval run datasets/urgency --system jev --run jev
jev-eval run datasets/urgency --system openai:model=gpt-4o-mini,mode=logprobs --run llm
jev-eval compare datasets/urgency -a jev -b llmBoth systems are asked the identical question — it lives in the dataset's config.json, not in the
command — and compare scores them only on items both covered, so the per-item difficulty cancels.
The harness scores noul (binary) questions only. choice and score questions work through the
MCP server but have no scoring path here yet — comparing them needs different metrics (macro-F1 and
a confusion matrix; MAE and rank correlation respectively).
Latency, tokens and cost — the part that needs no labels
jev-eval run datasets/urgency --system jev --run jev
jev-eval run datasets/urgency --system openai:model=gpt-4o-mini,mode=logprobs --run llm
jev-eval perf datasets/urgency -a jev -b llm # no labels involvedEvery call is timed and its token counts recorded, so perf reports p50/p90/p95/p99 latency, tokens
per call, and a paired median latency difference with a confidence interval — paired because a long
ticket is a long prompt for both systems, and median because one retry in the tail would otherwise
decide it.
This matters more than it first looks. An accuracy gap of a point or two needs thousands of labelled items to establish; a system that is twice as slow or five times dearer is unmistakable across fifty unlabelled ones. So the operational comparison is available on day one, before any labelling starts, and it is often what the decision actually turns on.
Cost is reported only from rates you supply, because they change and are per-account. Add them to the
dataset's config.json, keyed by run name:
{
"pricing": {
"jev": { "inputPer1M": 0.00, "outputPer1M": 0.00, "currency": "USD" },
"llm": { "inputPer1M": 0.00, "outputPer1M": 0.00, "currency": "USD" }
}
}Without them the latency and token numbers still appear; the cost line says what is missing.
What it reports
Brier score and log loss — proper scoring rules, computed on the probability itself rather than on which side of 0.5 it fell. This is the headline, not accuracy.
Calibration (ECE + reliability table) — of the things it called 0.9, how many happened? A model that is 85% accurate and honest about it beats one that is 87% accurate and says 0.99 every time.
AUC — ranking quality, independent of any threshold.
Accuracy, precision, recall, F1 at 0.5 and at the best available threshold.
95% bootstrap intervals on every one, seeded so a rerun reproduces exactly.
McNemar's test on the paired decisions — exact binomial below 25 discordant pairs, where the chi-square approximation misleads.
When an interval spans zero, the report says so in those words and tells you how many more items you would need. "No difference detected" is a result; "System A won" from a 10-item sample is not.
Getting a probability out of a general LLM
The baseline adapter has two modes, and the choice matters more than the model does:
mode=logprobsconstrains the reply to one token and reads the distribution overYes/No. This is the fair comparison — a real probability, not a stated one.mode=verbalizedasks the model to say a number. Convenient, and reliably badly calibrated: models pile up on 0.8/0.9/0.95. Use it to reproduce what a hand-rolled comparison actually measures, and read its ECE knowing part of the gap is the interface, not the model.
The Anthropic adapter is verbalized-only, since the Messages API exposes no logprobs.
Before trusting any of it
Run jev-eval agreement first. If two people labelling the same items score κ below about 0.6, the
question is ambiguous and no amount of data will separate the systems — the ceiling on what an eval
can resolve is how consistently humans can answer it. Fix the question, then collect labels.
Commands
Command | What it does |
| Run the MCP server over stdio. This is what agents invoke. |
| Register with Claude Code, Claude Desktop and Codex. |
| Show the resolved key status, endpoint, launch command and config path. |
| Latency, tokens and cost for one or two runs. Needs no labels. |
| Evaluation harness — see above. |
Environment
Variable | Purpose |
| Required. |
| Override the API host. Useful for staging and tests. |
| Only for the eval harness' baseline adapter. |
| Only for the eval harness' baseline adapter. |
Every TYPESAFE_* variable in your shell is carried into the client configs by install.
Behaviour worth knowing
429,529and transport failures are retried four times with jittered exponential backoff, and aRetry-Afterheader is always honoured over the computed delay.401and422fail straight away — retrying a bad key or a bad request only wastes time.Responses are capped at 8 MB and rejected past it, not truncated. The read stops at the first chunk over the line, so an oversized reply is never fully buffered, and it is not retried — a body that couldn't be read whole is not one to decide from, and asking again returns the same body.
The API's response JSON is forwarded to the agent byte for byte rather than re-serialized, so nothing in it is rewritten through a double on the way out.
installnever touches an MCP server it didn't create. If something is already registered under the name and its launch path isn't this package, it is left alone and reported;--nameregisters under a different one. Codex is the exception — it exposes no config read path, so an existing entry there is replaced.The Claude Desktop config is written via a temp file and a rename, so a failed write can't truncate a file that also holds your own preferences. Every other key in it is preserved.
stdoutcarries the MCP protocol and nothing else; all diagnostics go tostderr.
One limit you have to work around
Numbers reaching the tool have already been parsed as IEEE-754 doubles by the JSON-RPC layer, so an
integer above 9007199254740991 arrives with its last digits gone, and two distinct ids can turn up
identical. This is upstream of anything the server can fix. Send long identifiers as strings —
the tool description tells the agent so, but it is worth knowing yourself.
Development
npm install
npm test # node:test, no test runner dependency
node src/cli.js doctor
node bin/eval.js --helpLicense
MIT — see LICENSE.
Available Tools
1 toolevaluateEvaluate with JevARead-onlyIdempotent
Judge some content against one or more typed questions and get back probabilities rather than prose — a yes/no likelihood (noul), a pick from a named set (choice), or a position on ordered levels (score). Use it wherever you would otherwise ask a model for an answer and then parse the reply.
| Name | Required | Description | Default |
|---|---|---|---|
| model | No | Model identifier. Defaults to "jev-latest". | |
| state | Yes | The material being judged, and only that. Every question in the call reads this same value and nothing else, so anything the decision depends on has to appear here. Prefer an object with named fields when the input has parts. | |
| questions | Yes | Questions keyed by an id of your choosing; answers come back under those same ids. Questions in one call run together over the same state and cannot see each other's answers — split dependent judgments across calls. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover readOnly, openWorld, and idempotent hints. The description adds critical behavioral details: questions in one call cannot see each other's answers, and every question reads only the same state. It also clarifies the output is probabilities, not prose. This goes beyond the annotations and is valuable.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences, front-loaded with the core purpose and output types. It includes the key usage guidance without any filler. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with nested objects and no output schema, the description adequately explains the output format and key constraints. It does not mention error handling or the 'model' parameter, but the schema covers the model default. Overall, the information an agent needs to call it correctly is present.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description does not add new parameter-level semantics beyond what the schema already provides for 'state' and 'questions'. It does explain the overall concept and the split-dependent-judgments rule, but that is more behavioral than parameter-specific. The schema carries the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: judging content against typed questions and returning structured probabilities. It distinguishes three output types (noul, choice, score) and frames the use case as an alternative to parsing raw model replies. This is specific and actionable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides explicit usage guidance: 'Use it wherever you would otherwise ask a model for an answer and then parse the reply.' Since there are no sibling tools, this is sufficient. It does not mention when not to use it, but the guidance is clear and context-rich.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
1 tool update
v0.1.0- First observed
evaluate
TDQS
Scored across 1 tool
With only one tool, there is no possibility of confusing its purpose with another tool. The tool's description clearly defines its unique role of evaluating content against typed questions and returning structured probabilities.
A single tool named 'evaluate' cannot be inconsistent with anything else, and the name is a clear, verb-based descriptor that matches its function. No pattern mixing or naming conflicts exist to penalize.
The server exposes only one tool for what appears to be a broad evaluation domain. The description suggests a wide range of uses, but one tool provides no supporting workflows, making the server feel too thin for its apparent scope.
The tool covers a single evaluation action, but the broader domain of evaluation likely includes question/template management, batch evaluation, or result history. As it stands, agents can perform isolated evaluations but have no way to manage or reuse evaluation setups, creating dead ends.
Maintenance
Related MCP Connectors
A paid remote MCP for HyperFrames, built to return verdicts, receipts, usage logs, and audit-ready J
A paid remote MCP for Equibles, built to return verdicts, receipts, usage logs, and audit-ready JSON
Evidence-readiness MCP server: validate, audit, and score briefs, memos, and evidence packs.
A paid remote MCP for Skybridge, built to return verdicts, receipts, usage logs, and audit-ready JSO
Related MCP Servers
- AlicenseNot gradedqualityBmaintenanceEnables MCP clients to consult TypeSafe's Jev through a judge tool, answering narrow typed questions with calibrated probabilities instead of prose.597 npm1MIT
- AlicenseNot gradedqualityAmaintenanceEnables MCP clients to submit bounded semantic-uncertainty judgments to the pinned TypeSafe Jev API, with tools for yes/no, choice, and score evaluations plus optional evidence or context selection.MIT
- AlicenseAqualityBmaintenanceEnables MCP clients to get calibrated probability judgments for choice, score, yes/no, and batch classification questions via the TypeSafe System One API.53MIT
- AlicenseNot gradedqualityCmaintenanceEnables MCP-capable agents to run TypeSafe's Jev judgment model as typed yes/no, choice, and score tools, with calibrated probabilities, confidence thresholds, escalation for uncertain or non-judgment tasks, and an optional action gate that fails open.1MIT