Skip to main content
Glama
README.md
# Jev MCP

An MCP server that gives your coding agents a **judge**: twelve purpose-built tools that hand Jev
a state and a set of typed questions and come back with a verdict — the chosen option or
probability, the full distribution, confidence, and an action (`auto`, `review` or `escalate`).

One URL, one key, one model. The endpoint is configuration, not code: the same server talks to
TypeSafe, OpenRouter, Vercel AI Gateway, or any other System One-compatible endpoint.

Jev is TypeSafe AI's System One decision model: state in, typed answers out, no generated text.
See the [TypeSafe docs](https://docs.typesafe.ai/) for the model itself.

## What it is not

It is not a chat model, not a browser tool, and not a wrapper around somebody's prompt pack. The
question design lives in the tools, not in your prompt: it does not translate answers into its own
vocabulary beyond spelling the yes/no question the way your endpoint expects, and it never
substitutes a value for a malformed answer.

## Install from scratch

A machine with nothing installed. Five steps; the details live in the sections below, so nothing is
written twice.

### 1. Get the server running

**Desktop app — no Node needed.** Take the installer from
[Releases](https://github.com/rtf6x/jev-mcp/releases) and drag it to Applications. The tray then owns
the server: Start/Stop, **Edit .env**, Open logs. Harnesses attach to it, Claude Desktop through the
`.mcpb` bridge — see [Harnesses](#harnesses). Signature, notarization and the first-launch
confirmation are covered under [Desktop app](#desktop-app).

**From source — Node 22.18 or newer.** The server runs the TypeScript sources directly.

```bash
git clone git@github.com:rtf6x/jev-mcp.git
cd jev-mcp
npm install
cp .env.example .env      # then set JEV_URL and JEV_API_KEY
npm start                 # MCP over HTTP on http://127.0.0.1:18791/mcp
```

**Without the app**, a host that spawns a process runs the server itself: `npm run stdio`, wired as
in [Harnesses](#harnesses).

### 2. Point it at a provider

`JEV_URL`, `JEV_API_KEY`, and the model and yes/no spelling that endpoint expects: the table and the
key pages are in [Configure](#configure). The tray app keeps its own `.env` in its config directory;
a server started by hand reads the repository's.

### 3. Wire your harness

One command for Claude Code, a JSON block for omp, opencode and any other Streamable-HTTP client:
[Harnesses](#harnesses).

### 4. Install the skill

The server registers twelve tools; the skill is what makes an agent call them instead of answering
from its own reading. `skills/jev/` ships here as a copy — copy it into the directory your client
reads:

```bash
cp -R skills/jev ~/.claude/skills/            # Claude Code (or .claude/skills in one project)
cp -R skills/jev ~/.config/opencode/skills/   # opencode
cp -R skills/jev ~/.agents/skills/            # omp, Codex, generic agents
```

The line that does the work is the skill's `description`: it is the only text a host keeps in context
**before** the skill is opened, so it carries the mandate ("judgment goes through these tools, not
your own reading") rather than a list of capabilities. The server sends the same mandate as
`instructions` in the `initialize` result, but Claude Code does not pass that field to the model
(open issues in `anthropics/claude-code`), so the skill is what a client actually reads.

Claude Desktop loads plugins rather than skill directories, so the skill does not reach it this way;
there the server is what matters.

### 5. Check that it works

```bash
curl -sS http://127.0.0.1:18791/health        # the endpoint and the model actually in force
claude mcp list                               # jev ... Connected
```

Then call one tool — in a harness, or over HTTP with an `initialize` first. A first call that comes
back with a verdict, probabilities and a usage block is the whole chain working.

### When it does not work

- **`401 Authentication failed`** — a real environment variable beats `.env`, so a key exported in
  the shell the app was launched from wins over the app's own file. `/health` names the endpoint and
  the model in force, and the key's prefix says which provider it belongs to (`vck_` Vercel AI
  Gateway, `jev_` TypeSafe, `sk-or-v1_` OpenRouter). Launch the tray app from a clean environment:
  `env -u JEV_API_KEY open "/Applications/Jev MCP.app"`.
- **Nothing listens on the port** — the sidecar is not running. Start it in the tray (or `npm start`
  in the checkout) and read Open logs.
- **`"Jev MCP.app" is damaged and cannot be opened`** — an ad-hoc build rather than a broken
  download: `xattr -dr com.apple.quarantine "/Applications/Jev MCP.app"`.
- **Launch at startup records a temporary path** — a quarantined app runs from a randomized
  `AppTranslocation` copy. Launch it once from `/Applications` first.

## Configure

| Variable | Default | Meaning |
| --- | --- | --- |
| `JEV_URL` | — | Endpoint that receives the request. Used verbatim. |
| `JEV_API_KEY` | — | Bearer key for that endpoint. |
| `JEV_MODEL` | `jev-latest` | Model id sent in the request body. |
| `JEV_QUESTION_TYPE` | `noul` | Spelling of a yes/no question: `noul` or `boolean`. |
| `MCP_HOST` | `127.0.0.1` | Interface the HTTP transport binds. |
| `MCP_HTTP_PORT` | `18791` | Port for the HTTP transport. |
| `JEV_TIMEOUT_MS` | `60000` | Whole-request deadline for the upstream call. |
| `JEV_MCP_ENV` | `<cwd>/.env` | Path of the env file to read. |
| `JEV_MCP_EXIT_WITH_PARENT` | `0` | `1` makes the server exit when its stdin closes: the tray app pipes stdin, so the server leaves when the app does instead of holding the port as an orphan. Leave it off for a host that runs without a parent pipe. |

Real environment variables win over `.env` entries; the file only fills what the environment
does not carry. No keys are read from anywhere else.

Known endpoints:

| Provider | `JEV_URL` | `JEV_MODEL` | `JEV_QUESTION_TYPE` |
| --- | --- | --- | --- |
| TypeSafe (official) | `https://api.typesafe.ai/v1/systemone` | `jev-latest` | `noul` |
| OpenRouter | `https://openrouter.ai/api/alpha/decisions` | `typesafe/jev-1.13` | `noul` |
| Vercel AI Gateway | `https://ai-gateway.vercel.sh/v1/evaluate` | `typesafe-ai/jev` | `boolean` |

Key pages: [console.typesafe.ai/keys](https://console.typesafe.ai/keys),
[openrouter.ai/settings/keys](https://openrouter.ai/settings/keys),
Vercel AI Gateway API keys.

## Run

```bash
npm start          # MCP over Streamable HTTP: http://127.0.0.1:18791/mcp (+ /health)
npm run stdio      # MCP over stdio, for hosts that spawn a process
```

### Harnesses

**The app is the server.** Every host below points at the same running process, so there is one
`.env`, one endpoint and one place to read the logs.

omp, Claude Code, opencode and any other Streamable-HTTP client:

```json
{
  "mcpServers": {
    "jev": { "type": "http", "url": "http://127.0.0.1:18791/mcp" }
  }
}
```

Claude Desktop speaks only local stdio, so it gets a **bridge**, not a server. Prefer the
**extension**: installing `jev-mcp-desktop-<version>.mcpb`
([Releases](https://github.com/rtf6x/jev-mcp/releases), or the packed file `npm run pack:mcpb`
leaves in the repository root) lists the server under the manifest's own name — *Jev MCP*, with its
description and a "Jev MCP server URL" field — and runs the bridge on Desktop's bundled Node. The
hand-wired entry below is the CLI-managed alternative (`claude mcp add`, a config file you keep in
git); there the server shows the raw key as its name:

```json
{
  "mcpServers": {
    "jev": { "command": "node", "args": ["/absolute/path/to/jev-mcp/mcpb/dist/server/index.mjs"] }
  }
}
```

Both are the same bridge at the same port (`JEV_MCP_URL`, default `http://127.0.0.1:18791/mcp`).
Pick one — two of them surface the twelve tools twice, so installing the extension means emptying
that entry (`"mcpServers": {}`), not adding beside it. With the app stopped the bridge says so and
exits; it never starts a server of its own.

**Without the app**, a host that can spawn a process runs the server itself:

```json
{
  "mcpServers": {
    "jev": {
      "command": "node",
      "args": ["/absolute/path/to/jev-mcp/mcp-server/src/index.ts", "--stdio"],
      "env": { "JEV_MCP_ENV": "/absolute/path/to/jev-mcp/.env" }
    }
  }
}
```

That path owns its own `.env` (the one `JEV_MCP_ENV` names, or the repository's), lives only as long
as the host, and exists for a machine where the app is not wanted — everything else is better served
by the app.

A config that already carries a `jev` entry gets that entry **replaced**, never a second one: a
JSON object keeps the last duplicate key, so the stale entry wins and the client keeps talking to
the old endpoint while the file still looks right. The server answers `tools/list` and one
`instructions` string in the `initialize` result — no prompts and no resources; the skill that tells
an agent which tool to call ships here as `skills/jev/`, a copy of `ai-skillset/skills/jev`, which
stays the source of truth.

## Desktop app

`desktop-app/` is a tray app that runs the server for you: it spawns the sidecar on launch,
keeps it in the menu bar, and gives you the settings without a terminal.

```bash
npm run icons                       # once: prepares the icon set from desktop-app/icons-src/
npm run sidecar -- aarch64-apple-darwin   # once per target: bundles the server into the app
npm run desktop                     # builds the .app
```

The tray shows the status and the configured endpoint, and offers Start/Stop, Launch at
startup, **Edit .env** and Open logs. Its `.env` is its own, in the app's config directory
(`~/Library/Application Support/cc.rootfox.jev-mcp-desktop/.env` on macOS); it is created from
`.env.example` on first launch and handed to the sidecar through `JEV_MCP_ENV`. The sidecar
carries Node inside it, so the app needs no Node installation on the machine it runs on.

The app carries two icons, because macOS wants two different things: the colour app icon (a
square tile the system shows in the Finder and the installer) and the menu-bar glyph, drawn
from its alpha alone and repainted by the system for a light, dark or highlighted bar.
`desktop-app/icons-src/README.md` holds the sources and how to rebuild them.

### Install a release

Downloads are signed with the team's **Developer ID Application** certificate and notarized by
Apple, so the `.dmg` opens with a plain double-click and the old «"Jev MCP.app" is damaged and
cannot be opened» is gone. The release workflow takes the certificate and the notarization
credentials from the repository secrets — `APPLE_CERTIFICATE`, `APPLE_CERTIFICATE_PASSWORD`,
`APPLE_TEAM_ID`, `APPLE_ID`, `APPLE_PASSWORD` (an app-specific password; or `APPLE_API_KEY`,
`APPLE_API_ISSUER`, `APPLE_API_KEY_P8` instead of the last three). The certificate itself is the
team's, reused from Tail MCP; the material and the recipe live in `~/tail-mcp-signing/`.

macOS may still ask you to confirm the **first** launch of a downloaded copy («… is an app
downloaded from the internet. Are you sure you want to open it?»). That is the standard
quarantine confirmation, not a signature problem — click **Open** once and later launches are
silent. While that flag is present macOS may also run the app from a randomized read-only copy
instead of `/Applications` (`ps xw | grep "Jev MCP"` shows a
`/private/var/folders/…/AppTranslocation/…` path). The app itself does not care, but **Launch at
startup** does: the login item would record that temporary path. Launch the app once from
`/Applications` before enabling it, or clear the flag:

```bash
xattr -dr com.apple.quarantine "/Applications/Jev MCP.app"
```

With **no certificate** the build falls back to an ad-hoc signature — and then macOS answers
«"Jev MCP.app" is damaged and cannot be opened», with that line as the only way in. A certificate
present **without** notarization credentials stops the release on purpose: a
signed-but-unnotarized build still trips Gatekeeper, so shipping it would only move the problem.
Releases up to `v0.1.1` are ad-hoc and need the line above; `v0.1.4` is the first notarized one.

## The tools

The server exposes twelve community judgment tools (ported from
[jkudish/jev-mcp](https://github.com/jkudish/jev-mcp), MIT — see `THIRD-PARTY-NOTICES.md`). Each
one turns a state into typed questions, posts **one** request to `JEV_URL`, and validates the
answers fail-closed.

| Tool | What it answers | Arguments (defaults in brackets) |
| --- | --- | --- |
| `jev_verify` | How the evidence relates to each claim: `verified` / `contradicted` / `unsupported` | `claims`, `evidence`, `auto_accept` [0.8] |
| `jev_screen` | Whether fetched text is safe to read: injection, substance, relevance → `pass` / `review` / `block` / `skip` | `text`, `purpose`, `block_at` [0.75], `review_at` [0.25] |
| `jev_noul` | A calibrated probability per proposition (`likely` / `unlikely` / `uncertain`) | `propositions`, `context`, `auto_accept` [0.85] |
| `jev_find` | Which candidate best answers a query, and whether any candidate does | `query`, `candidates`, `top_k` [5] |
| `jev_rerank` | Every candidate's relevance, sorted | `query`, `candidates`, `top_k` [all] |
| `jev_classify` | One class per item from a shared catalog, with a margin-gated decision | `items`, `classes`, `purpose`, `context`, `auto_accept` [0.85], `minimum_margin` [0.5] |
| `jev_decide` | Which bounded alternative fits the evidence and priorities: escape hatches (`ask_user` / `investigate` / `none`) plus a per-requirement check | `decision`, `evidence`, `priorities`, `candidates`, `requirements`, `escape_hatches` [true], `escalate_on_contradiction` [false] |
| `jev_compare` | `same_fact` / `contradicts` / `different_facts` for two passages, optionally per named aspect | `passage_a`, `passage_b`, `aspects`, `purpose`, `auto_accept` [0.85], `minimum_margin` [0.5] |
| `jev_extract` | Which regex candidate is the field's true value, returned verbatim | `document`, `fields` (`id`, `pattern`, `flags`, `description`), `auto_accept` [0.85], `minimum_margin` [0.5] |
| `jev_audit` | Whether an extracted value is wrong: hallucinated, off-target, incomplete, wrong format, or wrongly omitted | `source`, `records`, `wrong_at` [0.7] |
| `jev_review` | Whether a proposed patch may be applied: four 0..2 rubric scores → `auto` / `review` / `escalate` | `request`, one of `diff` or `files`, `tests`, `auto_accept` [0.8], `review_at` [min(0.5, auto_accept)], `composite_floor` [0.7] |
| `jev_gate` | A patch review **and** completion claims against evidence, in one call | `request`, one of `diff` or `files`, `claims`, `evidence`, `tests`, the `jev_review` thresholds |

### Question types

Every question is one of the Jev primitives: `choice` (pick one of the criteria), `score`
(position on an ordered rubric) and a yes/no question. The yes/no question is written in
whichever spelling the configured endpoint expects — `JEV_QUESTION_TYPE=noul` or `boolean` — so
the same server talks to TypeSafe, OpenRouter and Vercel AI Gateway with no provider branch. The
answer readers accept `noul`, `probability` or `bool` for the yes/no probability, whichever the
endpoint reports.

### Results

Each tool returns one JSON text block: `tool`, `model`, `provider` (`endpoint`, or `none` when no
call was needed), `usage` (`input_tokens` / `output_tokens` / `cost`, read from either the
snake_case or the camelCase spelling, `null` when the endpoint reported none), plus the tool's own
fields — usually a `status`, one or more actions, and per-item results:

```jsonc
{
  "tool": "jev_classify",
  "summary": { "items": 2, "auto": 1, "review": 1, "invalid_response": 0, "by_class": { "billing": 1 } },
  "thresholds": { "auto_accept": 0.85, "minimum_margin": 0.5 },
  "results": [
    { "id": "refund", "classification": "billing", "probabilities": { "billing": 0.95, "technical": 0.05 },
      "confidence": 0.9, "margin": 0.9, "top_probability": 0.95, "decision": "auto" }
  ],
  "usage": { "input_tokens": 120, "output_tokens": 24, "cost": null }
}
```

### Fail-closed rules

An action comes from thresholds, never from hope, and every tool exposes its thresholds per call:
`auto` requires the top probability to clear `auto_accept` at `jev_verify` (where it is the answer
confidence), `jev_noul` and `jev_find`, the winner-to-runner-up margin to clear `minimum_margin`
as well at `jev_classify`, `jev_compare` and `jev_extract`, and `jev_screen` routes on
`block_at` / `review_at` instead. `jev_review` and `jev_gate` require `safe_to_apply` and the
rubric confidence at `auto_accept` with the weighted composite at `composite_floor`; `jev_audit`
escalates the whole audit when any value's max-gated P(wrong) reaches `wrong_at`. Everything else
is `review` (or `escalate` in the review family), and truncated or incomplete input can only make
an action stronger — never `auto`.

A missing or malformed answer is `status: "invalid_response"` with action `review`, never a
guessed value: a distribution that does not cover the criteria or does not sum to one, a `choice`
that is not the highest-probability option, a score outside the rubric, a confidence that is
present but malformed, a regex that timed out, a candidate universe that was capped, or an answer
that is simply absent. An endpoint failure — transport, or an HTTP error status — is a tool error
carrying the upstream status and body. There is no fallback and no retry: a paid call is never
silently repeated.

Tune the thresholds against your own data —
[confidence is not correctness](https://docs.typesafe.ai/confidence).

## Development

```bash
npm test           # unit and end-to-end tests against a stub endpoint, no key needed
npm run typecheck  # tsc --noEmit
npm run build      # esbuild bundle: mcp-server/dist/jev-mcp.cjs
npm run sidecar -- <rust-target-triple>   # SEA executable for the tray app
npm run desktop:dev                       # tray app in dev mode
```

The living plan is `docs/plans/jev-mcp.md`.

The tray app compiles with a Rust toolchain of 1.77 or newer (Tauri v2's minimum). On a machine
where Homebrew's `rust` shadows a newer rustup toolchain, put rustup's first:
`PATH="$HOME/.cargo/bin:$PATH" npm run desktop`.

## Licence

MIT. The twelve tools are ported from [jkudish/jev-mcp](https://github.com/jkudish/jev-mcp) (MIT);
see `THIRD-PARTY-NOTICES.md` for the upstream commit, the adapted files and the licence texts.