jev-mcp
Server Configuration
Describes the environment variables required to run the server.
| Name | Required | Description | Default |
|---|---|---|---|
| JEV_MCP_MODEL | No | Pin a Jev version, e.g. jev-1.12. | jev-latest |
| TYPESAFE_API_KEY | Yes | Your TypeSafe API key, required to use the Jev model. | |
| TYPESAFE_BASE_URL | No | Custom API endpoint for TypeSafe. |
Instructions
Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.
This server publishes no instructions, or was last inspected before Glama recorded them.
Capabilities
Features and capabilities supported by this server
Protocol revision2025-11-25
| Capability | Details |
|---|---|
| tools | {
"listChanged": false
} |
Tools
Functions exposed to the LLM to take actions
| Name | Description |
|---|---|
| jev_verifyA | Check each claim against provided evidence text with TypeSafe Jev. Returns per claim: verdict (verified | contradicted | unsupported), full probability distribution, confidence, and whether the verdict stands on its own (auto) or needs human review. Pattern: docs.typesafe.ai/cookbooks/citation_check. Pass reports, PR descriptions, or agent briefs as claims and their cited sources, diffs, or documents as evidence. |
| jev_screenA | Judge fetched or external text with TypeSafe Jev before an agent reads it: probability it contains instructions aimed at an AI agent (prompt injection), whether it has substantive content, and (when a purpose is given) whether it is relevant to the task. Returns a recommendation: pass | review | block | skip. Pattern: docs.typesafe.ai/cookbooks/llm_guardrails. |
| jev_noulA | Return a calibrated probability for each stated proposition with TypeSafe Jev, in one batched request: high means likely, low means unlikely, middling means genuinely uncertain. Supplied context informs the judgment but is not a proof guarantee; to test claims strictly against evidence, including whether the evidence is merely silent, use jev_verify instead. |
| jev_findA | Rank candidates against a plain-language query with TypeSafe Jev — no embeddings needed. One Choice scores every candidate id by how well it answers the query, plus a Noul checks whether any candidate addresses the query at all (so a confident 'top hit' cannot masquerade as an answer). Pattern: docs.typesafe.ai/cookbooks/semantic_find. Use for 'which file/note/line covers X' across up to 250 candidates. |
| jev_classifyA | Assign each item to one class from a shared catalog with TypeSafe Jev, in one batched request: the class catalog is sent once and every item becomes an independent Choice question. Returns per item: the chosen class, the full distribution, confidence, winner-to-runner-up margin, and an auto-versus-review decision. Auto requires both a high top probability (default 0.85) and a clear margin (default 0.50); everything else is flagged for review. Include a manual_review class in the catalog if you want an explicit escape hatch; the tool never invents one. |
| jev_decideA | One unresolved, bounded decision where semantic judgment over supplied evidence could change your plan: implementation alternatives, product tradeoffs with known preferences, workflow selection. Supply 2-6 candidates, evidence, and explicit priorities. Jev returns a Choice distribution over the candidates plus escape hatches (ask_user / investigate / none), and a per-candidate per-requirement supported / contradicted / unknown judgment for each optional requirement, all in one request. One call per unchanged decision; do not repeat a call to obtain a more pleasing answer. Use source inspection, tests, the user, or a reasoning model for open-ended research, routine choices, correctness proofs, or predicting user consent. High probability is not proof. |
| jev_rerankA | Rerank candidates against a query with TypeSafe Jev: one independent relevance probability per candidate, all in a single request, then sorted by score. Unlike jev_find (which picks one best answer), rerank scores every candidate so the full ordering survives. TypeSafe's rerank cookbook reports that on the CLERC benchmark this pattern lifted top-1 from 5% to 18% and top-10 from 38% to 62% (docs.typesafe.ai/cookbooks). Use for retrieval ordering, dedup triage, or feed ranking across up to 250 candidates. |
| jev_compareA | Judge the relation between two passages with TypeSafe Jev: same_fact, contradicts, or different_facts, with the full probability distribution, confidence, and an auto-versus-review decision. Optionally supply aspects (price, date, method, …) and each gets an independent per-aspect judgment in the same single request. Use for source reconciliation, changelog-vs-code drift, or merge sanity checks. The request supplies no evidence beyond the two passages, so a same_fact verdict means they agree with each other, not that they are true. |
| jev_extractA | Extract structured fields from a document with TypeSafe Jev as the picker, not the generator: your regex finds candidate substrings in code, Jev chooses which candidate is the field's true value, and the result is returned verbatim — never model-generated text. Fields with zero regex matches never reach the model (not_found); if no field has matches, no API call is made. Ambiguous picks are flagged for review. Use for prices, dates, version numbers, IDs, and anything with a recognizable shape; keep documents bounded. |
| jev_auditA | Audit extracted values against the text they claim to come from, before the values are trusted: one request with a per-value failure-mode battery (hallucinated / off-target / incomplete / wrong format, each framed so true = something is wrong) plus a dedicated omission check for empty values. Any value's P(wrong) at wrong_at escalates the whole audit; max-gated, never averaged. For multimodal intake: run your vision or ASR model first to produce a dense transcript of the image, scan, or recording, screen that transcript with jev_screen, then audit the extracted values against it here — the tool never sees pixels or audio, it audits two text artifacts against each other. Schema validation catches structural errors; it can flag a schema-valid fabrication against the supplied source text — a limited cross-check, not verification of the original. |
| jev_reviewA | Score a proposed diff against the request with TypeSafe Jev before the task is called done. Returns 0..2 rubric scores for correctness, spec match, test gap, and blast radius (the last two lower the weighted composite), a safe_to_apply probability, and an auto | review | escalate action. Auto requires safe_to_apply and min score confidence at auto_accept and the composite at composite_floor; truncated or malformed input never returns auto. Does not apply the patch or run tests. For a multi-file change, pass files instead of diff: the rubric is asked once per file in the same request and the action composes in code (auto only when every file is auto). Use jev_gate to also verify completion claims against evidence in the same call. |
| jev_gateA | Review a proposed patch and verify completion claims against supplied evidence in one TypeSafe Jev call. Auto only when the patch review is accepted and every claim is verified at or above auto_accept. Unsupported claims require review; confident contradictions, unknown confidence, or low confidence escalate. The request and claims are assertions to check, never proof; put supporting diff excerpts and test logs in evidence. Evidence is capped at 16 items and 200,000 characters in aggregate. Does not run tests or apply changes. Use jev_review for a patch without claims, jev_verify for claims without a patch review. |
Prompts
Interactive templates invoked by user choice
| Name | Description |
|---|---|
No prompts | |
Resources
Contextual data attached and managed by the client
| Name | Description |
|---|---|
No resources | |
TDQS
Scored across 12 tools
Each tool has a clearly distinct input/output contract (e.g., verify vs. noul vs. audit vs. compare), and descriptions explicitly cross-reference when to use alternatives (gate vs. review, find vs. rerank). Overlaps are mitigated rather than ambiguous.
All tools use consistent snake_case with the jev_ prefix, but jev_noul deviates from the otherwise uniform verb-based operation naming (verify, gate, screen, find, etc.), a minor inconsistency.
12 tools sit well within the 3-15 range and each maps to a distinct reasoning primitive (verification, screening, retrieval, classification, decision, review, etc.), so none feels redundant.
The surface covers a full cognitive toolchain: input screening, extraction, auditing, verification, comparison, retrieval/find/rerank, classification, decision, patch review, and gated completion. No obvious lifecycle gap for the stated Jev reasoning domain.