Skip to main content
Glama

JEV MCP · Structured Judgments & LLM Evaluation

Connect JEV to MCP clients and compare its judgments against general-purpose LLMs using shared datasets and measurable accuracy.

Agents are good at producing text and bad at producing answers you can branch on. Ask one "is this ticket urgent?" and you get back a sentence you then have to parse, with no number attached — no way to tell a confident yes from a coin flip. This server exposes TypeSafe's Jev classifier as a single MCP tool that returns typed answers with probabilities, so the agent gets 0.94 and moves on.

  agent  ──── evaluate ────▶  jev-mcp  ──── POST /v1/systemone ────▶  TypeSafe
 (stdio)                     (this)                                    (Jev)
         ◀─── typed JSON ───           ◀─── probabilities ────────

  jev-eval ── same question ─▶ JEV  ─┐
                             ─▶ LLM ─┴─▶ Brier · calibration · McNemar

Two halves: an MCP server that exposes the classifier to your agents, and an evaluation harness that tells you whether it is actually beating whatever you were using before.

Requirements

Related MCP server: Jev MCP

Install

Not published to npm yet, so install from the repository:

npm install -g github:arunav25/jev-mcp

Or clone it, which is what you want if you plan to run the evaluation harness:

git clone https://github.com/arunav25/jev-mcp.git
cd jev-mcp && npm install && npm link

The bare name jev-mcp on npm belongs to an unrelated project. This package publishes as @arunav25/jev-mcp; until it is published, use one of the commands above.

Then point your agents at it. The key has to be in your environment before you run this, because agents launch the server without your shell, so its value is written into each client's config:

export TYPESAFE_API_KEY=sk-...
jev-mcp install

That registers the server with Claude Code, Claude Desktop and Codex, skipping any that aren't installed. Restart Claude Desktop afterwards. To see what it would do first:

jev-mcp install --dry-run
jev-mcp install --client codex   # or limit it to one

Any MCP client that speaks stdio will do. Run jev-mcp doctor to get the exact launch command, then:

{
  "mcpServers": {
    "jev": {
      "command": "/usr/local/bin/node",
      "args": ["/usr/local/lib/node_modules/@arunav25/jev-mcp/src/cli.js", "serve"],
      "env": { "TYPESAFE_API_KEY": "sk-..." }
    }
  }
}

Using it

One tool, evaluate. Give it the material to judge and one or more questions:

{
  "state": {
    "subject": "Payouts failing",
    "body": "Help! My payouts have been failing for 3 days."
  },
  "questions": {
    "is_urgent": {
      "type": "noul",
      "instructions": "Does the sender need a response today rather than this week?"
    },
    "department": {
      "type": "choice",
      "instructions": "Which team should own this ticket?",
      "criteria": {
        "billing": "Payments, payouts, refunds, invoices",
        "technical": "Bugs, outages, API errors",
        "sales": "Pricing and plan changes",
        "unclear": "Not enough information to route"
      }
    }
  }
}

Back comes the answer under the same keys you used:

{
  "model": "jev-latest",
  "answers": {
    "is_urgent": { "type": "noul", "noul": 0.94 },
    "department": {
      "type": "choice",
      "choice": "billing",
      "probabilities": { "billing": 0.88, "technical": 0.09, "sales": 0.02, "unclear": 0.01 },
      "confidence": 0.81
    }
  }
}

Question types

Type

Criteria

Returns

noul

optional {"true": …, "false": …}

noul — probability the condition holds

choice

required map of option → description or null

choice, probabilities, confidence

score

required ordered array of 2+ level descriptions

score, legend, probabilities, confidence

score answers are 0-indexed: N levels answer between 0 and N-1, so 3.87 over 5 levels sits between the fourth and fifth level — not 3.87 out of 5. Read it through the legend the response returns, which names each index.

Getting good answers

  • state is the only thing a question can see. Put every fact the decision rests on there.

  • Question keys aren't sent to the model. Calling one urgent explains nothing — the instructions have to define what urgent means here.

  • One judgment per question. Two decisions in one set of instructions blur the distribution.

  • Questions in a call can't see each other's answers. They run together over the same state. If step two depends on step one, make two calls.

  • Give choice an escape hatch when the input might match nothing.

  • Put observations in state, not conclusions. A verdict you already reached reads as evidence for itself, and the probability comes back as your own conclusion with a number attached. Word instructions as the condition to test, not the answer you expect.

  • score levels have to describe real situations. A bare 1–5 scale gives the model nothing to anchor on.

  • 0.5 on a noul means uncertain, not "medium amount of the thing you asked about". And confidence measures how concentrated the distribution is, not whether the answer is right.

Malformed questions are caught here, before the request goes out — a choice with no criteria or a score with one level comes back as a message the agent can fix, not a 422 and a wasted round trip.

Measuring whether it's actually better

jev-eval is a harness for answering "which one is more accurate?" with a number that survives scrutiny. It exists because the usual version of that comparison — run ten inputs through both, eyeball the outputs — cannot detect anything. Here is the arithmetic:

Effect to detect

Disagreement rate

Labelled items needed

90% → 92% (2 points)

15%

2,941

90% → 92% (2 points)

25%

4,904

85% → 90% (5 points)

20%

626

80% → 90% (10 points)

25%

194

70% → 85% (15 points)

30%

103

80% power, α = 0.05, paired McNemar. Run jev-eval power --from 0.9 --to 0.92 for your own numbers.

A two-point gap over ten calls is three orders of magnitude short of conclusive. Any ranking drawn from it is a coin flip wearing a number.

The loop

jev-eval init datasets/urgency --question "Does this message need a response today rather than this week?"
# put your real items in datasets/urgency/items.jsonl, one JSON object per line:
#   {"id": "t-001", "state": {"subject": "...", "body": "..."}}

jev-eval label datasets/urgency --rater rater-1     # never shows you a model's answer
jev-eval label datasets/urgency --rater rater-2     # a second rater bounds what's resolvable
jev-eval agreement datasets/urgency                 # Cohen's kappa between the two

jev-eval run datasets/urgency --system jev --run jev
jev-eval run datasets/urgency --system openai:model=gpt-4o-mini,mode=logprobs --run llm

jev-eval compare datasets/urgency -a jev -b llm

Both systems are asked the identical question — it lives in the dataset's config.json, not in the command — and compare scores them only on items both covered, so the per-item difficulty cancels.

The harness scores noul (binary) questions only. choice and score questions work through the MCP server but have no scoring path here yet — comparing them needs different metrics (macro-F1 and a confusion matrix; MAE and rank correlation respectively).

Latency, tokens and cost — the part that needs no labels

jev-eval run datasets/urgency --system jev --run jev
jev-eval run datasets/urgency --system openai:model=gpt-4o-mini,mode=logprobs --run llm
jev-eval perf datasets/urgency -a jev -b llm     # no labels involved

Every call is timed and its token counts recorded, so perf reports p50/p90/p95/p99 latency, tokens per call, and a paired median latency difference with a confidence interval — paired because a long ticket is a long prompt for both systems, and median because one retry in the tail would otherwise decide it.

This matters more than it first looks. An accuracy gap of a point or two needs thousands of labelled items to establish; a system that is twice as slow or five times dearer is unmistakable across fifty unlabelled ones. So the operational comparison is available on day one, before any labelling starts, and it is often what the decision actually turns on.

Cost is reported only from rates you supply, because they change and are per-account. Add them to the dataset's config.json, keyed by run name:

{
  "pricing": {
    "jev": { "inputPer1M": 0.00, "outputPer1M": 0.00, "currency": "USD" },
    "llm": { "inputPer1M": 0.00, "outputPer1M": 0.00, "currency": "USD" }
  }
}

Without them the latency and token numbers still appear; the cost line says what is missing.

What it reports

  • Brier score and log loss — proper scoring rules, computed on the probability itself rather than on which side of 0.5 it fell. This is the headline, not accuracy.

  • Calibration (ECE + reliability table) — of the things it called 0.9, how many happened? A model that is 85% accurate and honest about it beats one that is 87% accurate and says 0.99 every time.

  • AUC — ranking quality, independent of any threshold.

  • Accuracy, precision, recall, F1 at 0.5 and at the best available threshold.

  • 95% bootstrap intervals on every one, seeded so a rerun reproduces exactly.

  • McNemar's test on the paired decisions — exact binomial below 25 discordant pairs, where the chi-square approximation misleads.

When an interval spans zero, the report says so in those words and tells you how many more items you would need. "No difference detected" is a result; "System A won" from a 10-item sample is not.

Getting a probability out of a general LLM

The baseline adapter has two modes, and the choice matters more than the model does:

  • mode=logprobs constrains the reply to one token and reads the distribution over Yes/No. This is the fair comparison — a real probability, not a stated one.

  • mode=verbalized asks the model to say a number. Convenient, and reliably badly calibrated: models pile up on 0.8/0.9/0.95. Use it to reproduce what a hand-rolled comparison actually measures, and read its ECE knowing part of the gap is the interface, not the model.

The Anthropic adapter is verbalized-only, since the Messages API exposes no logprobs.

Before trusting any of it

Run jev-eval agreement first. If two people labelling the same items score κ below about 0.6, the question is ambiguous and no amount of data will separate the systems — the ceiling on what an eval can resolve is how consistently humans can answer it. Fix the question, then collect labels.

Commands

Command

What it does

jev-mcp serve

Run the MCP server over stdio. This is what agents invoke.

jev-mcp install

Register with Claude Code, Claude Desktop and Codex. --name avoids a name collision.

jev-mcp doctor

Show the resolved key status, endpoint, launch command and config path.

jev-eval perf

Latency, tokens and cost for one or two runs. Needs no labels.

jev-eval …

Evaluation harness — see above. jev-eval --help lists its subcommands.

Environment

Variable

Purpose

TYPESAFE_API_KEY

Required.

TYPESAFE_BASE_URL

Override the API host. Useful for staging and tests.

OPENAI_API_KEY

Only for the eval harness' baseline adapter.

ANTHROPIC_API_KEY

Only for the eval harness' baseline adapter.

Every TYPESAFE_* variable in your shell is carried into the client configs by install.

Behaviour worth knowing

  • 429, 529 and transport failures are retried four times with jittered exponential backoff, and a Retry-After header is always honoured over the computed delay. 401 and 422 fail straight away — retrying a bad key or a bad request only wastes time.

  • Responses are capped at 8 MB and rejected past it, not truncated. The read stops at the first chunk over the line, so an oversized reply is never fully buffered, and it is not retried — a body that couldn't be read whole is not one to decide from, and asking again returns the same body.

  • The API's response JSON is forwarded to the agent byte for byte rather than re-serialized, so nothing in it is rewritten through a double on the way out.

  • install never touches an MCP server it didn't create. If something is already registered under the name and its launch path isn't this package, it is left alone and reported; --name registers under a different one. Codex is the exception — it exposes no config read path, so an existing entry there is replaced.

  • The Claude Desktop config is written via a temp file and a rename, so a failed write can't truncate a file that also holds your own preferences. Every other key in it is preserved.

  • stdout carries the MCP protocol and nothing else; all diagnostics go to stderr.

One limit you have to work around

Numbers reaching the tool have already been parsed as IEEE-754 doubles by the JSON-RPC layer, so an integer above 9007199254740991 arrives with its last digits gone, and two distinct ids can turn up identical. This is upstream of anything the server can fix. Send long identifiers as strings — the tool description tells the agent so, but it is worth knowing yourself.

Development

npm install
npm test          # node:test, no test runner dependency
node src/cli.js doctor
node bin/eval.js --help

License

MIT — see LICENSE.

Available Tools

1 tool
evaluateEvaluate with JevA
Read-onlyIdempotent

Judge some content against one or more typed questions and get back probabilities rather than prose — a yes/no likelihood (noul), a pick from a named set (choice), or a position on ordered levels (score). Use it wherever you would otherwise ask a model for an answer and then parse the reply.

ParametersJSON Schema
NameRequiredDescriptionDefault
modelNoModel identifier. Defaults to "jev-latest".
stateYesThe material being judged, and only that. Every question in the call reads this same value and nothing else, so anything the decision depends on has to appear here. Prefer an object with named fields when the input has parts.
questionsYesQuestions keyed by an id of your choosing; answers come back under those same ids. Questions in one call run together over the same state and cannot see each other's answers — split dependent judgments across calls.

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover readOnly, openWorld, and idempotent hints. The description adds critical behavioral details: questions in one call cannot see each other's answers, and every question reads only the same state. It also clarifies the output is probabilities, not prose. This goes beyond the annotations and is valuable.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences, front-loaded with the core purpose and output types. It includes the key usage guidance without any filler. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with nested objects and no output schema, the description adequately explains the output format and key constraints. It does not mention error handling or the 'model' parameter, but the schema covers the model default. Overall, the information an agent needs to call it correctly is present.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description does not add new parameter-level semantics beyond what the schema already provides for 'state' and 'questions'. It does explain the overall concept and the split-dependent-judgments rule, but that is more behavioral than parameter-specific. The schema carries the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: judging content against typed questions and returning structured probabilities. It distinguishes three output types (noul, choice, score) and frames the use case as an alternative to parsing raw model replies. This is specific and actionable.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides explicit usage guidance: 'Use it wherever you would otherwise ask a model for an answer and then parse the reply.' Since there are no sibling tools, this is sufficient. It does not mention when not to use it, but the guidance is clear and context-rich.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 1 tool updatev0.1.0
    • First observedevaluate

TDQS

A4/5.0

Scored across 1 tool

Disambiguation5/5

With only one tool, there is no possibility of confusing its purpose with another tool. The tool's description clearly defines its unique role of evaluating content against typed questions and returning structured probabilities.

Naming Consistency5/5

A single tool named 'evaluate' cannot be inconsistent with anything else, and the name is a clear, verb-based descriptor that matches its function. No pattern mixing or naming conflicts exist to penalize.

Tool Count2/5

The server exposes only one tool for what appears to be a broad evaluation domain. The description suggests a wide range of uses, but one tool provides no supporting workflows, making the server feel too thin for its apparent scope.

Completeness2/5

The tool covers a single evaluation action, but the broader domain of evaluation likely includes question/template management, batch evaluation, or result history. As it stands, agents can perform isolated evaluations but have no way to manage or reuse evaluation setups, creating dead ends.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    B
    maintenance
    Enables MCP clients to consult TypeSafe's Jev through a judge tool, answering narrow typed questions with calibrated probabilities instead of prose.
    597 npm
    1
    MIT
  • A
    license
    Not graded
    quality
    A
    maintenance
    Enables MCP clients to submit bounded semantic-uncertainty judgments to the pinned TypeSafe Jev API, with tools for yes/no, choice, and score evaluations plus optional evidence or context selection.
    MIT
  • A
    license
    Not graded
    quality
    C
    maintenance
    Enables MCP-capable agents to run TypeSafe's Jev judgment model as typed yes/no, choice, and score tools, with calibrated probabilities, confidence thresholds, escalation for uncertain or non-judgment tasks, and an optional action gate that fails open.
    1
    MIT