jev-mcp
This server is a local stdio MCP server that gives Cursor, Codex, and other MCP clients typed TypeSafe Jev judgments (Choice, Score, Noul) to route coding work, review patches, verify claims, screen content, and rank candidates without editing files or running tools.
Route the next coding step and select prepared tool calls in one request (
jev_step,jev_coding_loop,jev_tool_route).Review proposed patches for correctness, spec-match, test gaps, blast radius, and safe-to-apply (
jev_review).Verify factual claims against supplied evidence as verified/contradicted/unsupported (
jev_verify).Combine patch review and claim verification in one call (
jev_gate).Screen untrusted fetched or pasted text for prompt injection, substance, and relevance (
jev_screen).Rank supplied candidates (files, symbols, errors, skills) against a natural-language query (
jev_rank).Ask custom atomic noul/choice/score questions as an escape hatch (
jev_evaluate).Return typed answers, probabilities, confidence, token usage, and an action:
auto,review, orescalate.Support CLI diagnostics, deterministic mock mode, configurable timeouts, and safe policy limits on context, candidates, and claims.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@jev-mcpreview my latest changes before I commit"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
jev-mcp
A local stdio MCP server that gives Cursor, Codex, and other MCP clients typed TypeSafe Jev judgments. Jev returns Choice, Score, and Noul answers; the host agent still edits files and runs commands.
Tools
Tool | Purpose |
| Route the next step and select a prepared call in one request |
| Route the next step and decide whether a partner model is needed |
| Select an exact host-prepared tool call without generating arguments |
| Assess a proposed patch |
| Check claims against supplied evidence |
| Combine patch review and claim verification in one upstream call |
| Screen untrusted content before the host reads it |
| Rank candidates supplied by the host |
| Ask custom, atomic typed questions |
Results include typed answers, token usage, and an action: auto, review, or escalate. Confidence measures model certainty, not factual truth. Incomplete context never permits auto; reduce the input and submit it again for a complete judgment.
Question packs are MCP resources at jev://packs/{coding-loop,tool-route,step,review,verify,screen,rank,gate}.
Related MCP server: universal-jev-mcp
Coding with fewer partner-model turns
jev_step answers the whole loop turn in one request: it routes the step and, when the host supplies prepared calls, selects among them. A host that would otherwise call jev_coding_loop and then jev_tool_route spends one MCP round-trip instead of two, so it spends one host-model turn instead of two. Both original tools remain available.
Let host code execute known steps and prepare exact tool calls from an existing plan. When a semantic choice is needed, pass those calls to jev_step or jev_tool_route; either returns an executable call only for a confident, suitable selection with complete context and validated host facts. The host executes that call and routes again using the new observation. Jev never invents arguments or executes tools.
When a new plan or code may be needed, use jev_coding_loop with trusted execution facts. Its handoff distinguishes tool use, context gathering, review, user input, stopping, and a partner model. Invoke a generative partner only when partner_model.required is true; the legacy model_tier answer alone does not request a model turn. Uncertainty and escalation do not automatically spend a partner turn.
The prepared-call router accepts at most 32 candidates. Empty or wholly ineligible lists return locally with zero Jev usage. Other routing calls use Jev; this reduces unnecessary generative handoffs by policy, but live quality and cost savings have not been measured. See the tool contracts and host workflow.
Quick start
Install Node 20+ and run from a checkout:
npm ci
npm run buildSet a TypeSafe API key in your shell, then run diagnostics:
export TYPESAFE_API_KEY=ts_...
node dist/index.js doctor
node dist/index.js doctor --jsonPowerShell:
$env:TYPESAFE_API_KEY = 'ts_...'
node dist/index.js doctor --jsonFor a deterministic local demo, set JEV_MCP_MOCK=1 instead. Mock mode is for tests and demos, not production decisions. Neither the CLI nor MCP automatically reads .env; see configuration for explicit environment-file use.
Cursor: copy the MCP example into .cursor/mcp.json, replace its argument with the absolute path to this checkout's dist/index.js, and set the key in env.
Codex: register the absolute path:
codex mcp add jev --env TYPESAFE_API_KEY=ts_... -- node /absolute/path/to/jev-mcp/dist/index.jsCopy the agent skill into the project so the host knows when to call these tools. Detailed setup and the Windows checkout helper are in installation.
CLI
node dist/index.js # stdio MCP
node dist/index.js doctor # human-readable diagnostics on stderr
node dist/index.js doctor --json # structured diagnostics on stdout
node dist/index.js eval --stdin < request.jsonAn evaluation request contains state and a questions map:
{
"state": "Production payouts are failing. Urgent.",
"questions": {
"urgent": { "type": "noul", "instructions": "Is this urgent?" }
}
}Diagnostics do not log request content or API keys. Calls have a 30-second total deadline by default, configurable with JEV_MCP_TIMEOUT_MS. API failures and invalid responses return typed errors rather than fabricated judgments.
Limits and policy
State plus questions must fit the estimated 64,000-token total budget and the 32,000-token state-plus-longest-question budget. State may be shortened; results expose incomplete coverage and cannot automatically accept a judgment based on omitted context. Questions alone that exceed the budget are rejected.
Rank accepts unique candidate IDs and at most 5,000 supplied candidates, with at most 250 options per upstream call. Larger lists use repeated reduction rounds; each candidate text is capped at 2,000 characters. Verify and gate accept at most 1,000 claims. It ranks supplied candidates and does not index your repository. Arithmetic and date calculations belong in host code.
All nine tools expose an MCP output schema and return the same successful payload through both structuredContent and the JSON text content. Tool-route and fused-step judgments receive sanitized candidate descriptions and argument shapes; raw host arguments are retained only for the selected, locally validated call.
Development
npm test
npm run typecheck
npm run build
npm run test:package
npm run benchmarkThe regular suite runs without a key; the live test is skipped unless a key is present and mock mode is disabled. Package smoke testing builds and packs the project, installs the tarball into an isolated directory with npm --offline, then runs its shipped CLI. npm run benchmark builds the compiled server and measures the real MCP stdio transport in deterministic mock mode: sequential and concurrent calls, payload-size scaling, and candidate-count scaling. Use npm run benchmark:ci to apply broad sanity budgets; set JEV_BENCH_ITERATIONS, JEV_BENCH_CONCURRENCY, JEV_BENCH_MAX_P95_MS, or JEV_BENCH_MIN_RPS to tune a run. Run npm ci first to populate the dependency cache. No test publishes the package.
npm pack and npm publish build automatically through prepack. CI checks Node 20 and 22 on Windows and Linux, including the offline packed-install smoke test.
Documentation
Document | Contents |
Unreleased changes, compatibility notes, and validation | |
Request path, policy, limits, and errors | |
Arguments and outputs for all nine tools | |
Host configuration and Windows checkout | |
Environment, thresholds, diagnostics, tests | |
Calling guidance for the host | |
Short project guidance |
Available Tools
6 toolsjev_coding_loopJev coding-loop routerARead-onlyIdempotent
Call before spending a frontier turn on retry/stop/model-tier. One Jev fan-out returns next (continue|retry|ask_user|stop), model_tier (cheap|standard|reasoning), risk, focus, and noul flags done_enough / needs_more_context / tests_likely_fail. Policy in code maps confidence to action auto|review|escalate. Does not edit files.
| Name | Required | Description | Default |
|---|---|---|---|
| task | Yes | What the coding agent is trying to do | |
| model | No | ||
| extras | No | Optional extra JSON fields included in Jev state | |
| review_at | No | ||
| auto_accept | No | ||
| observation | Yes | Current turn: last diff, command output, test results, or blocker |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, openWorldHint, and destructiveHint=false. The description adds 'Does not edit files,' which is consistent, and reveals additional behavior: the policy mapping from confidence to action (auto|review|escalate) and the exact output flags. This goes beyond the annotations, though it could be richer about side effects or state changes.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single dense paragraph, front-loaded with the usage trigger. It packs a lot of information about outputs and policy without excessive fluff. It is appropriately sized for a router tool with complex output.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description must explain the return values, which it does reasonably well by listing the fields and flags. However, it leaves the optional parameters unexplained and the policy mechanism vague ('Policy in code maps confidence to action'). For a tool that is meant to be called before a decision, an agent might need more context on how to interpret risk/focus or how the optional parameters affect behavior.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 50%; the description does not compensate for the undocumented parameters (model, review_at, auto_accept). It explains the output structure but not the meaning of these optional inputs. For a tool with 6 parameters, the description should clarify what these fields do, especially since they are not described in the schema either.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: a router to call before spending a frontier turn on retry/stop/model-tier. It lists the specific outputs (next, model_tier, risk, focus, flags) and the policy mapping, making the tool's role distinct from the sibling evaluation/review/verify/screen/rank tools, though it does not explicitly name them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It provides a clear trigger: 'Call before spending a frontier turn on retry/stop/model-tier.' This gives an explicit condition for when to use the tool. It does not mention when not to use it or contrast with alternatives, but the context is specific and actionable.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
jev_evaluateEvaluate with JevARead-onlyIdempotent
Escape hatch: send shared state plus named noul/choice/score questions to TypeSafe Jev. Use when no other jev_* recipe fits. Jev does not write code or prose. Questions in one call run in parallel. Returns typed answers, probabilities, confidence, usage, and action auto|review|escalate.
| Name | Required | Description | Default |
|---|---|---|---|
| model | No | Override, default jev-latest | |
| state | Yes | Shared state to judge: text or JSON | |
| questions | Yes | Named noul, choice, and score questions evaluated in parallel |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description adds behavioral context beyond the annotations: parallel execution of questions, the fact that Jev does not write code or prose, and the return fields (typed answers, probabilities, confidence, usage, action). Since annotations already declare readOnly, idempotent, and non-destructive, the description enriches the agent's understanding of side effects and execution model.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three sentences with zero filler. It front-loads the purpose ('Escape hatch'), then gives the usage condition, a behavioral limitation, parallelism, and return summary. Every sentence earns its place and the structure is highly scannable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with nested objects and no output schema, the description provides a solid overview of what the tool returns (typed answers, probabilities, confidence, usage, action) and how it executes (parallel). It doesn't detail error cases or rate limits, but given the annotations already cover safety and the description covers the key behavior, it's reasonably complete for an agent to invoke correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 100% description coverage for all parameters, including the nested question structure. The description adds minimal extra semantic value beyond the schema (e.g., 'named' questions, parallel execution), but it doesn't clarify the 'criteria' field or provide additional guidance on constructing questions. Baseline 3 is appropriate given the schema's thoroughness.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: send shared state plus noul/choice/score questions to TypeSafe Jev for evaluation. It explicitly differentiates from siblings by saying 'Use when no other jev_* recipe fits', giving an agent a clear discriminator without opening other schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It provides an explicit condition for when to use this tool ('when no other jev_* recipe fits') and states what Jev does not do ('does not write code or prose'), which sets expectations. This is a clear directive for selection among siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
jev_rankJev candidate rankerARead-onlyIdempotent
Rank files, symbols, errors, or skills against a plain-language query. No embeddings. One Choice over candidate ids plus a Noul that the top hit actually answers the query (so a forced winner cannot masquerade as a match). Max 250 candidates per Jev call; larger lists are chunked then re-ranked. Pass candidates in; this server does not index the repo.
| Name | Required | Description | Default |
|---|---|---|---|
| model | No | ||
| query | Yes | What you are looking for, in natural language | |
| top_k | No | How many ranked candidates to return. Default 5. | |
| candidates | Yes | Candidates to rank. More than 250 are chunked, then the winners are re-ranked. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already mark the operation read-only, open-world, idempotent, and non-destructive. The description adds meaningful behavior beyond that: no embeddings, a top-hit relevance check, a 250-candidate cap with chunking/reranking, and a stateless input-only design. The only blemish is the unclear 'Noul' wording, which hampers full comprehension.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but mostly efficient: every sentence contributes either an input requirement, an algorithmic trait, or a constraint. It is front-loaded with the main action. The awkward 'One Choice over candidate ids plus a Noul' phrase and unmarked technical jargon reduce readability slightly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a stateless ranker, the description covers the main operational concerns: candidate format, maximum count, chunking behavior, lack of repo indexing, and the query-scoring approach. It lacks an explicit description of the return value shape, and no output schema exists to fill that gap, but an agent can likely call the tool correctly with what is given.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is high (75%), so the schema carries most parameter meaning. The description does add useful context for 'query' ('plain-language') and 'candidates' (pass them in, max 250, chunked), but the 'model' parameter remains undocumented and no explanation of output fields is provided.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Rank files, symbols, errors, or skills against a plain-language query.' It further distinguishes this tool from siblings by noting 'No embeddings' and 'Pass candidates in; this server does not index the repo.' The core purpose is unmistakable even before looking at the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly implies when to use it: rank a supplied list of candidates against a query, and do not expect repo indexing. However, it never names alternative sibling tools or explicitly says when another tool would be better, leaving routing largely to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
jev_reviewJev patch reviewARead-onlyIdempotent
Score a proposed diff against the request: correctness, spec-match, test-gap, blast-radius, plus noul safe_to_apply. Composite weights live in code. Call before declaring a fix done. Does not apply the patch.
| Name | Required | Description | Default |
|---|---|---|---|
| diff | Yes | Proposed patch, file excerpt, or change summary | |
| model | No | ||
| tests | No | Test output if any | |
| request | Yes | What the user asked for | |
| review_at | No | ||
| auto_accept | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already mark the tool as read-only, idempotent, and non-destructive. The description adds valuable behavioral context beyond those hints: it explicitly states the patch is not applied, that scoring criteria are used, and that composite weights are defined in code. This helps an agent understand side-effect-free, black-box behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is brief and front-loaded with the core action and criteria. Each sentence earns its place: purpose, usage timing, and non-application. The typo 'noul safe_to_apply' and the vague 'composite weights live in code' slightly reduce clarity and polish.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 6-parameter tool with no output schema, the description does not state what the tool returns (e.g., score, verdict, safe_to_apply value) and does not clarify the optional input parameters. It is adequate for deciding when to call it, but an agent invoking it may still be uncertain about the response shape and how to interpret review_at or auto_accept.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 50%; request, diff, and tests have descriptions, and the description reinforces their roles ('diff against the request,' 'test-gap'). However, optional parameters like model, review_at, and auto_accept receive no meaningful semantic clarification in either the schema or the description, so the description only partially compensates.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific action ('Score a proposed diff against the request') and enumerates concrete review dimensions (correctness, spec-match, test-gap, blast-radius, safe_to_apply). The title and description clearly distinguish it as a review/safety gate rather than an apply or verify tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit trigger: 'Call before declaring a fix done.' It also excludes a key non-behavior by saying the tool does not apply the patch. It does not name sibling alternatives or formal when-not-to-use conditions, so it stops short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
jev_screenJev content screenARead-onlyIdempotent
Judge fetched or pasted text before the agent reads it: prompt-injection probability, substance, and optional relevance to purpose. Recommendation: pass|review|block|skip. Use on untrusted web pages, issues, and pastes. Not for first-party repo files.
| Name | Required | Description | Default |
|---|---|---|---|
| text | Yes | Fetched or pasted text before the agent reads it | |
| model | No | ||
| purpose | No | What the agent is trying to do; enables relevance and skip | |
| block_at | No | ||
| review_at | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, openWorldHint=true, idempotentHint=true, and destructiveHint=false. The description adds behavioral context by explaining what the tool evaluates (injection, substance, relevance) and the output recommendation. It does not contradict annotations and provides useful operational detail beyond the safety profile.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two concise sentences. The first sentence states the main function and output, the second gives usage scope. It is front-loaded with the core purpose and contains no fluff. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has 5 parameters and no output schema, so the description must explain the expected return and any thresholds. It mentions the recommendation output but not the full output structure (e.g., whether it includes probability or substance scores). It also does not explain block_at and review_at parameters, which are essential for controlling the screening behavior. Given these gaps, the description is not fully complete for an agent to call it correctly without additional context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 40% (text and purpose have descriptions, while model, block_at, and review_at do not). The description mentions 'fetched or pasted text' (text) and 'optional relevance to purpose' (purpose), adding some meaning. However, it does not explain block_at and review_at, which likely set thresholds for the recommendation. Since coverage is low, the description should compensate more, but it only partially does.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's function: to judge fetched or pasted text for prompt-injection probability, substance, and optional relevance to purpose, and to produce a recommendation (pass|review|block|skip). It explicitly scopes usage to untrusted web pages, issues, and pastes, and excludes first-party repo files, which distinguishes it from sibling tools. The purpose is unambiguous and specific.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides explicit usage context: 'Use on untrusted web pages, issues, and pastes. Not for first-party repo files.' This tells the agent when to use the tool and when not to. It does not name alternative siblings directly, but the exclusion for repo files implies a different tool is appropriate for that case, which is adequate guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
jev_verifyJev claim verifierARead-onlyIdempotent
Check each claim against provided evidence (PR description, agent brief, docs, diffs). Returns per claim: verified|contradicted|unsupported, probabilities, confidence, and auto vs review. Prefer this over asking a chat model to 'double-check'.
| Name | Required | Description | Default |
|---|---|---|---|
| model | No | ||
| claims | Yes | Factual claims to check | |
| evidence | Yes | Source text, or a list of {id, text} documents | |
| auto_accept | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly/openWorld/idempotent/destructive hints with no contradiction. The description adds meaningful behavioral context beyond the annotations: the tri-state verdict (verified|contradicted|unsupported) aligns with and operationalizes the openWorldHint, and the 'auto vs review' distinction discloses that some verdicts are automated while others may require human judgment. It doesn't explain what triggers review, but adds real value over bare annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences with zero filler: the first states the core action, the second specifies the return contract, and the third gives usage preference. The purpose is front-loaded and each sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Since there is no output schema, the description correctly shoulders the burden of explaining return values, and it does so well with the verdict/probability/confidence taxonomy. The core workflow (required claims and evidence) is fully covered. The gaps are the optional parameters — model and auto_accept have no semantics in either the schema or description — but these are non-critical for a correct basic invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 50% (claims and evidence are described, model and auto_accept are not), so the baseline is 3. The description adds practical semantics for evidence by listing concrete types (PR description, agent brief, docs, diffs) beyond the schema's generic 'Source text, or a list of {id, text} documents'. However, the meaning of model and auto_accept remains undocumented, so the description only partially compensates for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource ('Check each claim against provided evidence') and enumerates the output contract (verified|contradicted|unsupported, probabilities, confidence, auto vs review), making the tool's function unambiguous. However, it does not explicitly distinguish itself from sibling tools like jev_review or jev_evaluate, and the 'prefer this over a chat model' alternative is not a named sibling.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The final sentence offers a usage directive: prefer this tool over asking a chat model to 'double-check'. This gives an implicit when-to-use signal, but there are no explicit conditions, exclusions, or routing guidance to sibling tools (jev_review, jev_evaluate, jev_rank) that might overlap. The guidance is implied rather than systematic.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
6 tool updates
v0.1.0- First observed
jev_coding_loop - First observed
jev_evaluate - First observed
jev_rank - First observed
jev_review - First observed
jev_screen - First observed
jev_verify
TDQS
Scored across 6 tools
Each recipe targets a distinct decision point (loop control, diff review, claim verification, text screening, candidate ranking), and `jev_evaluate` is explicitly scoped as an escape hatch rather than a competing operation. There is only a mild risk that agents reach for the generic evaluate tool instead of the specialized recipes.
All tools share the lowercase `jev_` prefix and use underscores, which gives the set a clear visual pattern. However, the names mix imperative verbs (`review`, `verify`, `screen`, `rank`, `evaluate`) with one noun-style name (`coding_loop`), a minor consistency deviation.
Six tools is a well-scoped size for a decision-support server. Each tool fills a distinct role with no apparent bloat or redundant utility.
The tool surface covers the major agent workflow decision points: whether to continue, whether to apply a diff, whether claims hold up, whether text is safe to read, and which candidates best match a query. The generic `jev_evaluate` fallback plus usage/action outputs prevents obvious dead ends.
Maintenance
Related MCP Connectors
Paid deterministic data-quality and execution-verification tools for AI agents.
Lets coding agents check their own code for leaked secrets, risky dependencies and AI-code mistakes
11Pre-execution governance for AI agents. Deterministic PASS/FAIL/REVIEW verdicts, replayable proof.
Shared control plane for AI coding agents — tasks, memory, decisions, file locks. 12 tools.
Related MCP Servers
- AlicenseAqualityBmaintenanceEnables frontier coding agents to delegate routine probabilistic judgments to TypeSafe Jev, providing calibrated triage signals for failures, attempts, completion, context ranking, findings, risk, and generic evidence-grounded questions.7MIT
- AlicenseAqualityBmaintenanceEnables coding agents to compact conversation contexts verbatim, make fast decisions through choice, boolean, and rubric scoring, and enforce command safety guardrails.6MIT
- AlicenseNot gradedqualityAmaintenanceProvides coding agents with typed classification, yes/no checks, scoring, ranking, and question-answering tools that return calibrated probabilities for fast, reliable decisions.889 npm44MIT
- AlicenseAqualityBmaintenanceEnables coding agents to make offline, zero-cost decisions using schema-safe Choice/Score/Noul primitives, a confidence gatekeeper, planning, adversarial red-teaming, research, and RLVR-based self-improvement via 21 MCP tools.243MIT