Skip to main content
Glama

Jev MCP

An MCP server that gives your coding agents a judge: twelve purpose-built tools that hand Jev a state and a set of typed questions and come back with a verdict — the chosen option or probability, the full distribution, confidence, and an action (auto, review or escalate).

One URL, one key, one model. The endpoint is configuration, not code: the same server talks to TypeSafe, OpenRouter, Vercel AI Gateway, or any other System One-compatible endpoint.

Jev is TypeSafe AI's System One decision model: state in, typed answers out, no generated text. See the TypeSafe docs for the model itself.

What it is not

It is not a chat model, not a browser tool, and not a wrapper around somebody's prompt pack. The question design lives in the tools, not in your prompt: it does not translate answers into its own vocabulary beyond spelling the yes/no question the way your endpoint expects, and it never substitutes a value for a malformed answer.

Related MCP server: jev-code

Install from scratch

A machine with nothing installed. Five steps; the details live in the sections below, so nothing is written twice.

1. Get the server running

Desktop app — no Node needed. Take the installer from Releases and drag it to Applications. The tray then owns the server: Start/Stop, Edit .env, Open logs. Harnesses attach to it, Claude Desktop through the .mcpb bridge — see Harnesses. Signature, notarization and the first-launch confirmation are covered under Desktop app.

From source — Node 22.18 or newer. The server runs the TypeScript sources directly.

git clone git@github.com:rtf6x/jev-mcp.git
cd jev-mcp
npm install
cp .env.example .env      # then set JEV_URL and JEV_API_KEY
npm start                 # MCP over HTTP on http://127.0.0.1:18791/mcp

Without the app, a host that spawns a process runs the server itself: npm run stdio, wired as in Harnesses.

2. Point it at a provider

JEV_URL, JEV_API_KEY, and the model and yes/no spelling that endpoint expects: the table and the key pages are in Configure. The tray app keeps its own .env in its config directory; a server started by hand reads the repository's.

3. Wire your harness

One command for Claude Code, a JSON block for omp, opencode and any other Streamable-HTTP client: Harnesses.

4. Install the skill

The server registers twelve tools; the skill is what makes an agent call them instead of answering from its own reading. skills/jev/ ships here as a copy — copy it into the directory your client reads:

cp -R skills/jev ~/.claude/skills/            # Claude Code (or .claude/skills in one project)
cp -R skills/jev ~/.config/opencode/skills/   # opencode
cp -R skills/jev ~/.agents/skills/            # omp, Codex, generic agents

The line that does the work is the skill's description: it is the only text a host keeps in context before the skill is opened, so it carries the mandate ("judgment goes through these tools, not your own reading") rather than a list of capabilities. The server sends the same mandate as instructions in the initialize result, but Claude Code does not pass that field to the model (open issues in anthropics/claude-code), so the skill is what a client actually reads.

Claude Desktop loads plugins rather than skill directories, so the skill does not reach it this way; there the server is what matters.

5. Check that it works

curl -sS http://127.0.0.1:18791/health        # the endpoint and the model actually in force
claude mcp list                               # jev ... Connected

Then call one tool — in a harness, or over HTTP with an initialize first. A first call that comes back with a verdict, probabilities and a usage block is the whole chain working.

When it does not work

  • 401 Authentication failed — a real environment variable beats .env, so a key exported in the shell the app was launched from wins over the app's own file. /health names the endpoint and the model in force, and the key's prefix says which provider it belongs to (vck_ Vercel AI Gateway, jev_ TypeSafe, sk-or-v1_ OpenRouter). Launch the tray app from a clean environment: env -u JEV_API_KEY open "/Applications/Jev MCP.app".

  • Nothing listens on the port — the sidecar is not running. Start it in the tray (or npm start in the checkout) and read Open logs.

  • "Jev MCP.app" is damaged and cannot be opened — an ad-hoc build rather than a broken download: xattr -dr com.apple.quarantine "/Applications/Jev MCP.app".

  • Launch at startup records a temporary path — a quarantined app runs from a randomized AppTranslocation copy. Launch it once from /Applications first.

Configure

Variable

Default

Meaning

JEV_URL

—

Endpoint that receives the request. Used verbatim.

JEV_API_KEY

—

Bearer key for that endpoint.

JEV_MODEL

jev-latest

Model id sent in the request body.

JEV_QUESTION_TYPE

noul

Spelling of a yes/no question: noul or boolean.

MCP_HOST

127.0.0.1

Interface the HTTP transport binds.

MCP_HTTP_PORT

18791

Port for the HTTP transport.

JEV_TIMEOUT_MS

60000

Whole-request deadline for the upstream call.

JEV_MCP_ENV

<cwd>/.env

Path of the env file to read.

JEV_MCP_EXIT_WITH_PARENT

0

1 makes the server exit when its stdin closes: the tray app pipes stdin, so the server leaves when the app does instead of holding the port as an orphan. Leave it off for a host that runs without a parent pipe.

Real environment variables win over .env entries; the file only fills what the environment does not carry. No keys are read from anywhere else.

Known endpoints:

Provider

JEV_URL

JEV_MODEL

JEV_QUESTION_TYPE

TypeSafe (official)

https://api.typesafe.ai/v1/systemone

jev-latest

noul

OpenRouter

https://openrouter.ai/api/alpha/decisions

typesafe/jev-1.13

noul

Vercel AI Gateway

https://ai-gateway.vercel.sh/v1/evaluate

typesafe-ai/jev

boolean

Key pages: console.typesafe.ai/keys, openrouter.ai/settings/keys, Vercel AI Gateway API keys.

Run

npm start          # MCP over Streamable HTTP: http://127.0.0.1:18791/mcp (+ /health)
npm run stdio      # MCP over stdio, for hosts that spawn a process

Harnesses

The app is the server. Every host below points at the same running process, so there is one .env, one endpoint and one place to read the logs.

omp, Claude Code, opencode and any other Streamable-HTTP client:

{
  "mcpServers": {
    "jev": { "type": "http", "url": "http://127.0.0.1:18791/mcp" }
  }
}

Claude Desktop speaks only local stdio, so it gets a bridge, not a server. Prefer the extension: installing jev-mcp-desktop-<version>.mcpb (Releases, or the packed file npm run pack:mcpb leaves in the repository root) lists the server under the manifest's own name — Jev MCP, with its description and a "Jev MCP server URL" field — and runs the bridge on Desktop's bundled Node. The hand-wired entry below is the CLI-managed alternative (claude mcp add, a config file you keep in git); there the server shows the raw key as its name:

{
  "mcpServers": {
    "jev": { "command": "node", "args": ["/absolute/path/to/jev-mcp/mcpb/dist/server/index.mjs"] }
  }
}

Both are the same bridge at the same port (JEV_MCP_URL, default http://127.0.0.1:18791/mcp). Pick one — two of them surface the twelve tools twice, so installing the extension means emptying that entry ("mcpServers": {}), not adding beside it. With the app stopped the bridge says so and exits; it never starts a server of its own.

Without the app, a host that can spawn a process runs the server itself:

{
  "mcpServers": {
    "jev": {
      "command": "node",
      "args": ["/absolute/path/to/jev-mcp/mcp-server/src/index.ts", "--stdio"],
      "env": { "JEV_MCP_ENV": "/absolute/path/to/jev-mcp/.env" }
    }
  }
}

That path owns its own .env (the one JEV_MCP_ENV names, or the repository's), lives only as long as the host, and exists for a machine where the app is not wanted — everything else is better served by the app.

A config that already carries a jev entry gets that entry replaced, never a second one: a JSON object keeps the last duplicate key, so the stale entry wins and the client keeps talking to the old endpoint while the file still looks right. The server answers tools/list and one instructions string in the initialize result — no prompts and no resources; the skill that tells an agent which tool to call ships here as skills/jev/, a copy of ai-skillset/skills/jev, which stays the source of truth.

Desktop app

desktop-app/ is a tray app that runs the server for you: it spawns the sidecar on launch, keeps it in the menu bar, and gives you the settings without a terminal.

npm run icons                       # once: prepares the icon set from desktop-app/icons-src/
npm run sidecar -- aarch64-apple-darwin   # once per target: bundles the server into the app
npm run desktop                     # builds the .app

The tray shows the status and the configured endpoint, and offers Start/Stop, Launch at startup, Edit .env and Open logs. Its .env is its own, in the app's config directory (~/Library/Application Support/cc.rootfox.jev-mcp-desktop/.env on macOS); it is created from .env.example on first launch and handed to the sidecar through JEV_MCP_ENV. The sidecar carries Node inside it, so the app needs no Node installation on the machine it runs on.

The app carries two icons, because macOS wants two different things: the colour app icon (a square tile the system shows in the Finder and the installer) and the menu-bar glyph, drawn from its alpha alone and repainted by the system for a light, dark or highlighted bar. desktop-app/icons-src/README.md holds the sources and how to rebuild them.

Install a release

Downloads are signed with the team's Developer ID Application certificate and notarized by Apple, so the .dmg opens with a plain double-click and the old «"Jev MCP.app" is damaged and cannot be opened» is gone. The release workflow takes the certificate and the notarization credentials from the repository secrets — APPLE_CERTIFICATE, APPLE_CERTIFICATE_PASSWORD, APPLE_TEAM_ID, APPLE_ID, APPLE_PASSWORD (an app-specific password; or APPLE_API_KEY, APPLE_API_ISSUER, APPLE_API_KEY_P8 instead of the last three). The certificate itself is the team's, reused from Tail MCP; the material and the recipe live in ~/tail-mcp-signing/.

macOS may still ask you to confirm the first launch of a downloaded copy («… is an app downloaded from the internet. Are you sure you want to open it?»). That is the standard quarantine confirmation, not a signature problem — click Open once and later launches are silent. While that flag is present macOS may also run the app from a randomized read-only copy instead of /Applications (ps xw | grep "Jev MCP" shows a /private/var/folders/…/AppTranslocation/… path). The app itself does not care, but Launch at startup does: the login item would record that temporary path. Launch the app once from /Applications before enabling it, or clear the flag:

xattr -dr com.apple.quarantine "/Applications/Jev MCP.app"

With no certificate the build falls back to an ad-hoc signature — and then macOS answers «"Jev MCP.app" is damaged and cannot be opened», with that line as the only way in. A certificate present without notarization credentials stops the release on purpose: a signed-but-unnotarized build still trips Gatekeeper, so shipping it would only move the problem. Releases up to v0.1.1 are ad-hoc and need the line above; v0.1.4 is the first notarized one.

The tools

The server exposes twelve community judgment tools (ported from jkudish/jev-mcp, MIT — see THIRD-PARTY-NOTICES.md). Each one turns a state into typed questions, posts one request to JEV_URL, and validates the answers fail-closed.

Tool

What it answers

Arguments (defaults in brackets)

jev_verify

How the evidence relates to each claim: verified / contradicted / unsupported

claims, evidence, auto_accept [0.8]

jev_screen

Whether fetched text is safe to read: injection, substance, relevance → pass / review / block / skip

text, purpose, block_at [0.75], review_at [0.25]

jev_noul

A calibrated probability per proposition (likely / unlikely / uncertain)

propositions, context, auto_accept [0.85]

jev_find

Which candidate best answers a query, and whether any candidate does

query, candidates, top_k [5]

jev_rerank

Every candidate's relevance, sorted

query, candidates, top_k [all]

jev_classify

One class per item from a shared catalog, with a margin-gated decision

items, classes, purpose, context, auto_accept [0.85], minimum_margin [0.5]

jev_decide

Which bounded alternative fits the evidence and priorities: escape hatches (ask_user / investigate / none) plus a per-requirement check

decision, evidence, priorities, candidates, requirements, escape_hatches [true], escalate_on_contradiction [false]

jev_compare

same_fact / contradicts / different_facts for two passages, optionally per named aspect

passage_a, passage_b, aspects, purpose, auto_accept [0.85], minimum_margin [0.5]

jev_extract

Which regex candidate is the field's true value, returned verbatim

document, fields (id, pattern, flags, description), auto_accept [0.85], minimum_margin [0.5]

jev_audit

Whether an extracted value is wrong: hallucinated, off-target, incomplete, wrong format, or wrongly omitted

source, records, wrong_at [0.7]

jev_review

Whether a proposed patch may be applied: four 0..2 rubric scores → auto / review / escalate

request, one of diff or files, tests, auto_accept [0.8], review_at [min(0.5, auto_accept)], composite_floor [0.7]

jev_gate

A patch review and completion claims against evidence, in one call

request, one of diff or files, claims, evidence, tests, the jev_review thresholds

Question types

Every question is one of the Jev primitives: choice (pick one of the criteria), score (position on an ordered rubric) and a yes/no question. The yes/no question is written in whichever spelling the configured endpoint expects — JEV_QUESTION_TYPE=noul or boolean — so the same server talks to TypeSafe, OpenRouter and Vercel AI Gateway with no provider branch. The answer readers accept noul, probability or bool for the yes/no probability, whichever the endpoint reports.

Results

Each tool returns one JSON text block: tool, model, provider (endpoint, or none when no call was needed), usage (input_tokens / output_tokens / cost, read from either the snake_case or the camelCase spelling, null when the endpoint reported none), plus the tool's own fields — usually a status, one or more actions, and per-item results:

{
  "tool": "jev_classify",
  "summary": { "items": 2, "auto": 1, "review": 1, "invalid_response": 0, "by_class": { "billing": 1 } },
  "thresholds": { "auto_accept": 0.85, "minimum_margin": 0.5 },
  "results": [
    { "id": "refund", "classification": "billing", "probabilities": { "billing": 0.95, "technical": 0.05 },
      "confidence": 0.9, "margin": 0.9, "top_probability": 0.95, "decision": "auto" }
  ],
  "usage": { "input_tokens": 120, "output_tokens": 24, "cost": null }
}

Fail-closed rules

An action comes from thresholds, never from hope, and every tool exposes its thresholds per call: auto requires the top probability to clear auto_accept at jev_verify (where it is the answer confidence), jev_noul and jev_find, the winner-to-runner-up margin to clear minimum_margin as well at jev_classify, jev_compare and jev_extract, and jev_screen routes on block_at / review_at instead. jev_review and jev_gate require safe_to_apply and the rubric confidence at auto_accept with the weighted composite at composite_floor; jev_audit escalates the whole audit when any value's max-gated P(wrong) reaches wrong_at. Everything else is review (or escalate in the review family), and truncated or incomplete input can only make an action stronger — never auto.

A missing or malformed answer is status: "invalid_response" with action review, never a guessed value: a distribution that does not cover the criteria or does not sum to one, a choice that is not the highest-probability option, a score outside the rubric, a confidence that is present but malformed, a regex that timed out, a candidate universe that was capped, or an answer that is simply absent. An endpoint failure — transport, or an HTTP error status — is a tool error carrying the upstream status and body. There is no fallback and no retry: a paid call is never silently repeated.

Tune the thresholds against your own data — confidence is not correctness.

Development

npm test           # unit and end-to-end tests against a stub endpoint, no key needed
npm run typecheck  # tsc --noEmit
npm run build      # esbuild bundle: mcp-server/dist/jev-mcp.cjs
npm run sidecar -- <rust-target-triple>   # SEA executable for the tray app
npm run desktop:dev                       # tray app in dev mode

The living plan is docs/plans/jev-mcp.md.

The tray app compiles with a Rust toolchain of 1.77 or newer (Tauri v2's minimum). On a machine where Homebrew's rust shadows a newer rustup toolchain, put rustup's first: PATH="$HOME/.cargo/bin:$PATH" npm run desktop.

Licence

MIT. The twelve tools are ported from jkudish/jev-mcp (MIT); see THIRD-PARTY-NOTICES.md for the upstream commit, the adapted files and the licence texts.

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    A
    maintenance
    Enables coding agents to send natural language and JSON application state to a judge tool that returns typed judgments (yes/no probability, choice, or score) with model and usage metadata, and fails gracefully without blocking workflows.
    428 npm
    4
    Apache 2.0