jev-mcp
Provides integration with Cloudflare Clef decision models, allowing MCP clients to ask typed yes/no, choice, and score questions about structured state and receive calibrated probabilities. Supports image inputs for Clef models with count, format, and size limits.
Provides integration with OpenAI's Decisions API in public beta through the openai/gpt-6-luna-decisions model, enabling typed decision questions over structured state and images, with calibrated probabilities and refusal reporting.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@jev-mcpAsk Jev if this support ticket is urgent: yes/no with a calibrated probability."
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
jev-mcp
A general-purpose MCP server for System One decision models such as TypeSafe Jev and Cloudflare Clef. It lets MCP clients ask typed questions (yes/no, choice, score) about structured state and get back calibrated probabilities.
MCP client → stdio → MCP tools → JEV Core → JevProvider → TypeSafe API or OpenRouterSee AGENTS.md for architecture and development rules, docs/typesafe-api-notes.md and docs/openrouter-notes.md for the verified API contracts, and CHANGELOG.md for release history.
Requirements
Node.js 20.12 or later
A TypeSafe API key, or an OpenRouter API key
Related MCP server: jevx-mcp
Setup
npm install
cp .env.example .env # then set the key for your provider in .env
npm run buildVariable | Required | Default |
| no |
|
| no | unset: only |
| when | — |
| no |
|
| when | — |
| no |
|
| no |
|
| no | unset: no Authorization header |
| no |
|
| no |
|
| no | unset: image |
| no |
|
| no |
|
| no |
|
| no |
|
| for | — (at least 32 characters; see Streamable HTTP) |
| no |
|
| no |
|
Checks before sending
Requests the server can tell will fail, or will be billed for nothing useful, are stopped before they reach the provider:
Check | Why |
Text input over | Clef accepted and billed inputs far beyond its documented context (118k tokens). 256,000 characters is about 64k tokens of English; text in Korean and similar scripts uses more tokens per character, so lower the limit if most input is such text |
Images to a model that cannot read them | Jev answers images with confident wrong results and bills them as text |
Clef rules: question IDs (letters, digits, | Clef returns 422 or 413 otherwise; the local message names the problem |
Image count, format and size limits | See Images |
Requests the provider rejects were not billed in testing, so these checks mainly prevent wasted spending on accepted requests and give clearer errors. Details: docs/openrouter-notes.md.
The stdio server loads .env from the package root when present. Variables already
set by the MCP client or shell take precedence.
Providers
JEV_PROVIDER selects one provider. It never switches to another provider on its
own, and a missing key for the selected provider is a configuration error.
| Models ( | Cost in results |
|
| not reported |
|
|
|
| Whatever your local System One server serves, for example | not reported (no per-request cost) |
Several providers in one server
JEV_PROVIDERS lets clients choose the provider per request, for example to ask Jev
and Clef the same question:
JEV_PROVIDER=typesafe # default when a request names none
JEV_PROVIDERS=typesafe,openrouter,localEvery question tool then takes an optional provider, and results report which
provider answered. jev.models lists the models of every provider, each tagged with
its provider; a provider that cannot be reached (such as a stopped local server)
appears under errors instead of failing the list.
Each listed provider needs its own key; one provider's key is never used for another, and a missing key is a configuration error.
Without
JEV_PROVIDERS, onlyJEV_PROVIDERis used, even if other keys are set.JEV_MODELapplies to the default provider. A request that selects TypeSafe or a local server without amodelusesjev-latest; OpenRouter needs an explicitmodel, such ascloudflare/clef-flashortypesafe/jev-1.13.An unknown
provideris rejected before anything is sent.
JEV_PROVIDER=local uses a server on your own machine or network: no data leaves it.
Setup with llama.cpp and Clef Flash, and measured results:
docs/local-provider.md.
On OpenRouter, Clef needs its full ID (cloudflare/clef-flash, not clef-flash),
accepts at most 64 questions per request, and Jev's context is 32k tokens instead of
64k. jev.models lists the decision models OpenRouter exposes. Details:
docs/openrouter-notes.md.
openai/gpt-6-luna-decisions is OpenAI's Decisions API, in public beta since
2026-10-06 ("we expect to GA in the coming weeks"), so its behavior may change. It
reads images and accepts long inputs, but it can refuse a question on content-policy
grounds, which fails the whole request; the server reports that as refused and does
not retry. Do not rely on it for production decisions before general availability.
Models differ in how confident they are. In testing Clef reported lower confidence
than Jev, and GPT-6 Luna mostly returned probabilities of 0 or 1, so set thresholds
per model from labeled examples.
Images
Every question tool (jev.evaluate, jev.noul, jev.choice, jev.score) takes an
optional images list, judged together with state. Images are not tied to any use
case: inspection photos, site photos, screenshots and documents all go through the
same field.
{
"state": { "line": "B", "note": "Customer return" },
"instructions": "Does the part in the photo show a visible defect?",
"images": [
{ "path": "D:/inspections/2026-10-07/part-0412.jpg" },
{ "data": "data:image/png;base64,iVBORw0KGgo..." }
]
}Source | Use | Rule |
| Agents on the same machine (Claude Code, Codex, VS Code) | Absolute path inside a |
| Programs that already hold the bytes | Base64 data URL |
Before anything is sent, JEV Core checks that there are at most 4 images, that each
is PNG, JPEG or WebP (from the bytes, not the name), and that each is at most 4 MiB
with 8 MiB in total. Image parts placed inside state are rejected.
For Clef through OpenRouter the practical limit is much lower than Clef documents: requests with more than about 384 KB of images in total fail there (measured, not documented), so they are rejected before sending. Downscale or recompress photos first; a 1024-pixel JPEG is usually well under the limit. GPT-6 Luna accepted a 624 KB photo.
Only image-capable models receive images: today cloudflare/clef,
cloudflare/clef-flash and openai/gpt-6-luna-decisions with
JEV_PROVIDER=openrouter, and local models the server lists with an image input
(for example Clef Flash in llama.cpp with --mmproj, which cannot load WebP).
Requests with images to any
other model, including Jev, fail before they are sent, because Jev answers images
with meaningless probabilities instead of an error. Details:
docs/openrouter-notes.md.
JEV_IMAGE_DIRS exists because the tools are called by AI agents: content an agent
reads could ask it to send a private file. Only files inside the listed directories
(resolved through symbolic links) can be read.
Clients
Verified with Claude Code, Codex CLI and ChatGPT (through OpenAI's Secure MCP Tunnel). Setup for each client, plus VS Code and Claude Desktop: docs/clients.md.
To build another application on jev-mcp (choosing tools, writing questions,
providers, errors, cost and decision thresholds), read the integration guide:
docs/usage.md. The server also sends a short usage guide to every
client in its MCP instructions, so a connected AI knows how to use the tools.
Streamable HTTP
Besides stdio, npm run start:http serves the same tools over Streamable HTTP at
http://127.0.0.1:8098/mcp, so several clients on this machine can share one
server. It needs JEV_HTTP_TOKEN (at least 32 characters), which every request
must send as Authorization: Bearer <token>. It listens on 127.0.0.1 only and
rejects requests whose Host or Origin is not a loopback address, which blocks
DNS rebinding. It is stateless: no sessions and no event stream. Remote access,
TLS and OAuth are not supported. Client setup:
docs/clients.md.
Docker
The same HTTP server also runs in Docker, which brings it back by itself after a restart:
docker compose up -d --build # build and start in the background
docker compose ps # shows (healthy) once it answers
docker compose logs # server output
docker compose down # stop and remove the containercompose.yaml reads .env at run time, so keys and JEV_HTTP_TOKEN
never enter the image, and publishes the port on 127.0.0.1:8098 only. Clients
connect exactly as to npm run start:http. The container runs as the unprivileged
node user with a read-only filesystem and no capabilities, and checks its own
health with an authenticated request. It restarts unless you stop it. To have it
running after a reboot, enable Start Docker Desktop when you sign in in Docker
Desktop's settings.
Inside the container the server listens on 0.0.0.0 (JEV_HTTP_HOST), because
Docker forwards the published port from outside the container. The Host, Origin
and token checks still apply. Keep the published host port equal to PORT. Paths in
.env such as JEV_IMAGE_DIRS refer to the host and are cleared in the container;
mount a folder to use image paths (see the comment in compose.yaml).
Inside the container 127.0.0.1 is the container itself, so compose.yaml sets
JEV_LOCAL_BASE_URL to http://host.docker.internal:8097. Docker Desktop forwards
that name to the host's loopback address, so a local server such as llama-server
keeps listening on 127.0.0.1 only and provider: "local" works through Docker.
Change the value in compose.yaml if your local server uses another port.
Stop the non-Docker server first if it is running: both use port 8098.
Retries
The stdio server wraps the provider in RetryingJevProvider, which retries
transient failures: rate limiting (429), overload (529), timeouts (including 408 and
OpenRouter's 524), HTTP 500/502/503/504 and network errors such as DNS failures.
Authentication, authorization, insufficient credits (402), invalid requests and
invalid responses are never retried.
Waits use exponential backoff with jitter (up to 0.5 s, then 1 s, 2 s, ..., capped
at 8 s). A retry-after header is honored when it is 8 s or less; a longer
requested wait ends retrying. Cancellation stops any pending wait. Each retry is
logged to stderr. After the last attempt, the original error is returned unchanged.
POST /v1/systemone is retried too. If a timed-out request was actually processed,
the retry is billed again.
Each attempt waits at most JEV_TIMEOUT_MS (default 30 s) for the provider. Raise it
for slow local servers or large images; a timeout counts as a transient failure and
is retried.
Logging
Logs go to stderr; stdout carries only MCP traffic. JEV_LOG_LEVEL picks how much:
Level | Adds |
| Unexpected server errors and startup failures |
| Retries, and the provider's payload (up to 2,000 characters) when a response is invalid, so the cause can be diagnosed |
| One startup line with the provider, model, timeout, retries and concurrency |
| One line per provider request: model, number of questions and images, duration, outcome, tokens and reported cost |
Logs never contain API keys, Authorization headers or base URLs. debug lines do not
include state or question text.
Use with Claude Code
This repository includes a project-scoped .mcp.json that starts
dist/transport/stdio.js. After npm run build, start Claude Code in the repository
root and approve the jev server when prompted. Check it with /mcp.
Claude Code exposes the tools with dots replaced by underscores, for example
mcp__jev__jev_evaluate and mcp__jev__jev_noul.
To use the server from another project, register it with an absolute path:
claude mcp add jev -- node /absolute/path/to/jev-mcp/dist/transport/stdio.jsTools
jev.evaluate
Evaluates state against a map of questions in one provider request.
{
"state": "Help! My payouts have been failing for 3 days.",
"questions": {
"is_urgent": { "type": "noul", "instructions": "Does this convey urgency?" },
"department": {
"type": "choice",
"instructions": "Which team should handle this?",
"criteria": { "billing": "Payments, refunds", "technical": "Bugs, outages" }
},
"frustration": {
"type": "score",
"instructions": "How frustrated is the customer?",
"criteria": ["Calm", "Frustrated", "Very angry"]
}
}
}model is optional and defaults to JEV_MODEL; surrounding whitespace is removed
from it. No other input is changed.
The result has model, answers (keyed by question ID) and usage. Failures come
back as tool errors with a kind such as invalid_input, authentication,
rate_limited or invalid_response. No answer is ever fabricated.
jev.evaluate_batch
Asks the same questions about many items (up to 100) in one tool call: classify
a list of tickets with a choice question, or score every search result and sort by
the score. Each item has its own state, optional id and optional images;
questions and model work as in jev.evaluate.
{
"items": [
{ "id": "t1", "state": "My payout failed again." },
{ "id": "t2", "state": "The app crashes when I log in." }
],
"questions": {
"team": {
"type": "choice",
"instructions": "Which team should handle this ticket?",
"criteria": { "billing": "Payments, refunds", "technical": "Bugs, outages" }
}
}
}The System One APIs take one state per request, so each item is a separate
provider request and is billed separately. At most JEV_MAX_CONCURRENCY requests
run at the same time.
Every item is checked before anything is sent; one invalid item rejects the call.
resultscome back in input order, each withstatusok(withanswersandusage),error(the item's own error), orskipped.After
authentication,authorization,payment_requiredorconfigurationerrors, items not yet sent areskipped, because they would fail the same way.summarycounts each status;usageadds up the successful items, withcostUsdonly when every successful item reported one.
jev.models
Lists the model names and aliases the provider accepts in model.
jev.noul, jev.choice, jev.score
Convenience tools for a single question. Each takes state, instructions, an
optional model and the question's criteria (optional for jev.noul), and
returns { model, answer, usage }. They wrap the input in a one-question
jev.evaluate request and run it through the same JEV Core path, so validation,
answer checks and errors are identical.
{
"state": "Help! My payouts have been failing for 3 days.",
"instructions": "Which team should handle this?",
"criteria": { "billing": "Payments, refunds", "technical": "Bugs, outages" }
}Use jev.evaluate to ask several questions about the same state: it answers them
in one request.
Scripts
Command | Purpose |
| Type-check sources and tests |
| Unit tests (no API key needed) |
| Compile to |
| Provider tests against the real TypeSafe API (uses |
| Build, then drive the stdio server over MCP against the real API |
| Start the stdio server |
| Start the Streamable HTTP server (needs |
| Windows: double-click to start the HTTP server in a window, or stop it |
GitHub Actions (.github/workflows/ci.yml) runs
typecheck, test and build on Node 20 and 22 for every push to main and every
pull request. It needs no secrets. The contract and e2e suites call the real API and
are run locally with a key.
License
This server cannot be deployed
Maintenance
Related MCP Connectors
Typed-question decisions from Jev: classify, moderate, route models and guard tool calls.
Calibrated judgments for text: yes/no probabilities, picks from your options, or scores.
A paid remote MCP for Pydantic AI structured output, built to return verdicts, receipts, usage logs,
A paid remote MCP for Equibles, built to return verdicts, receipts, usage logs, and audit-ready JSON
Related MCP Servers
- AlicenseAqualityBmaintenanceEnables MCP clients to get calibrated probability judgments for choice, score, yes/no, and batch classification questions via the TypeSafe System One API.5120 PyPI3MIT
- AlicenseAqualityCmaintenanceEnables running Jev AI typed decisions from any MCP client, returning choices, scores, or yes/no with calibrated probabilities.316 npmMIT
- AlicenseAqualityCmaintenanceEnables agents to query the TypeSafe Jev decision model for typed answers—yes/no, choice, and score questions—with calibrated probabilities, batched in a single request under a local context-budget guard.6MIT
- AlicenseNot gradedqualityAmaintenanceEnables agents and services to ask small decision models typed questions about text, JSON, images, or video, returning labels, scores, or yes/no answers over MCP tools or a Jev-compatible HTTP API.Apache 2.0