jev-mcp
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@jev-mcpshould we retry, rollback, or change strategy?"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
jev-mcp
Unofficial, community-maintained MCP server for TypeSafe AI Jev.
简体中文 · Quick start · Examples · Security · FAQ
jev-mcp exposes TypeSafe AI's Jev decision model as four conservative,
read-only MCP tools for bounded probabilistic decisions.
It is intentionally designed as a second-opinion layer, not as a replacement for the active coding/reasoning model.
This project is not affiliated with, endorsed by, sponsored by, or an official product of TypeSafe AI or OpenAI.
Why this exists
Strong coding agents are good at open-ended reasoning, code generation, debugging, and repository-wide analysis. Jev is useful for a different class of problem: small, explicit decisions that benefit from a structured probability signal.
Typical examples:
retry vs rollback vs change strategy;
route to one of several known tools or subsystems;
estimate low / medium / high / critical change risk;
evaluate a yes/no gate against a threshold;
repeat the same bounded decision many times in an automated workflow.
The design rule is simple:
deterministic evidence
>
host-model repository-aware reasoning
>
Jev probabilistic adviceA Jev result is advisory. It is never proof, never ground truth, and never authorization for destructive or irreversible work.
Related MCP server: jev_mcp
Architecture
flowchart LR
U[User] --> H[Active host model / MCP client]
E[Tests · compiler · runtime · static analysis] -->|highest-priority evidence| H
H -->|stdio MCP| M[jev-mcp]
M -->|HTTPS + Bearer token| J[TypeSafe AI Jev API]
J -->|probabilistic advisory result| M
M -->|structured tool result| H
H --> O[Final decision / action]The API key stays in the local process environment or a gitignored .env file.
It is not placed in the MCP client configuration.
Privacy boundary: anything placed in a tool's state is sent to the
configured TypeSafe API endpoint. Send only the minimum necessary, preferably
redacted state. See Security and privacy model.
Tool surface
Tool | Use it for | Do not use it for |
| Choosing among 2–255 explicit alternatives | Open-ended design or coding |
| Routing among known tools/subsystems/workflows | Switching the user's selected model |
| Advisory low/medium/high/critical risk signal | Replacing tests or review |
| Yes/no probability against a threshold | Authorizing destructive actions |
All four tools are declared read-only. They do not edit files, execute shell commands, deploy infrastructure, or change the model selected by the user.
Successful responses include local metadata similar to:
{
"jev_mcp_advisory": {
"authority": "secondary_advisory",
"finalDecisionBy": "active_host_model",
"deterministicEvidenceOverrides": true,
"doNotTreatProbabilityAsFact": true
}
}That metadata is added by this MCP server; it is not Jev model output.
60-second install
Requirements:
Node.js 20+
a TypeSafe API key / applicable TypeSafe access and credits
an MCP host that can launch a local stdio server
Clone and set up:
git clone https://github.com/Afloat16/jev-mcp.git
cd jev-mcp
./setup.shWindows PowerShell:
git clone https://github.com/Afloat16/jev-mcp.git
cd jev-mcp
./setup.ps1The setup script reads the API key without echoing it, stores it only in the
local gitignored .env file when needed, installs dependencies, and runs local
checks.
For a more explicit walkthrough, see Quick start.
Codex configuration
Add this to ~/.codex/config.toml and replace the path:
[mcp_servers.jev]
command = "node"
args = ["/ABSOLUTE/PATH/TO/jev-mcp/dist/index.js"]Restart the MCP host or start a new session after configuration changes.
For users who explicitly want a conservative cross-project Codex policy, merge:
codex/AGENTS.jev-conservative.mdinto your own ~/.codex/AGENTS.md.
Do not copy unrelated existing instructions away.
Which integration should I use?
Situation | Recommended approach |
Everyday interactive coding | Strong host model alone, or TypeSafe's agent skill |
Learning Jev concepts / patterns | TypeSafe agent skill |
Stable callable decision tools |
|
CI / agent orchestration |
|
High-volume application logic | Direct SDK/API integration is often the cleanest |
Need Jev to replace tests/compiler | Do not use Jev for that |
The MCP is most useful when the tool boundary itself matters: repeatability, structured output, orchestration, explicit thresholds, or shared agent workflows.
Configuration
Variable | Required | Default | Description |
| yes | — | TypeSafe credential |
| no |
| Jev model override |
| no |
| API base URL |
| no |
| Request timeout, 250–120000 ms |
| no | project | Alternate env file path |
Existing process environment variables override values loaded from .env.
Secret-handling rules
Never commit
.env.Never put API keys in
README,AGENTS.md, MCP config, examples, issues, screenshots, or CI logs.If a key is ever pasted into a shared surface, rotate/revoke it.
Run
npm run secrets:checkbefore publishing changes.
The bundled scanner is a guardrail, not a complete DLP system.
Examples
See examples/README.md for redacted examples covering:
ambiguous CI failure routing;
change-risk assessment;
retry / rollback / escalate decisions;
yes/no gates with explicit thresholds.
All examples intentionally use synthetic data and placeholders.
Local verification
No TypeSafe API call:
npm run doctor
npm run checkInteractive MCP inspection:
npm run inspectA live tool invocation in MCP Inspector uses your own TypeSafe account and may consume provider credits.
Project status
Item | Status |
Interface | 4 read-only MCP tools |
Transport | local stdio |
Node.js | 20+ |
License | MIT |
npm publishing | intentionally disabled |
API dependency | TypeSafe-hosted System One API |
Stability | pre-1.0; behavior may evolve |
The project follows semantic versioning in spirit, but while it remains below
1.0.0, minor releases may refine tool schemas or behavior. Breaking changes
should be documented in CHANGELOG.md and migration notes.
Documentation
Contributing
Focused issues and pull requests are welcome. Please read CONTRIBUTING.md first and run:
npm run checkbefore opening a PR.
Security-sensitive reports should follow SECURITY.md, not public issue comments.
License
MIT. See LICENSE.
The license covers this repository's code only. Third-party services, names, APIs, trademarks, pricing, and availability remain subject to their respective owners. See NOTICE.
References
TypeSafe AI documentation: https://docs.typesafe.ai/
TypeSafe AI — Introducing System One Models & Jev: https://typesafe.ai/blog/introducing-system-one-models-and-jev
Model Context Protocol TypeScript SDK: https://github.com/modelcontextprotocol/typescript-sdk
Available Tools
4 toolsjev_decideJev DecideARead-only
Use as an independent second opinion only after inspecting enough context when a bounded decision has 2-255 explicit alternatives and genuine uncertainty remains. The active host model keeps final authority; do not mechanically follow Jev or use it for open-ended coding/prose.
| Name | Required | Description | Default |
|---|---|---|---|
| model | No | Optional Jev model override. | |
| state | Yes | Minimal relevant state/context for the decision. | |
| options | Yes | ||
| question | Yes | One atomic decision question. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=true, so the safety profile is covered. The description adds valuable behavioral context the annotations can't: the host model retains final authority, and output should not be mechanically followed. It omits latency/cost/determinism of the external call.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two dense sentences with zero waste: the usage condition and the authority caveat are front-loaded, and the exclusion ('open-ended coding/prose') closes the second sentence. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only advisory tool with no output schema, the description covers invocation conditions, boundaries, and authority semantics well. It does not describe the shape of the response or what the model override does, which leaves a small gap given no output schema exists to carry that.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 75%, so the schema already documents model, state, and question. The description's '2-255 explicit alternatives' maps meaning onto the options array and its bounds, but adds nothing for state or question beyond the schema. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific capability: producing an independent second opinion on a bounded decision with 2-255 explicit alternatives. It does not name its sibling tools (jev_gate, jev_route, jev_risk_score), so differentiation relies on the 'bounded decision' framing rather than explicit routing.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Strong when-to-use ('only after inspecting enough context when a bounded decision has 2-255 explicit alternatives and genuine uncertainty remains') and when-not-to-use ('do not... use it for open-ended coding/prose'). It doesn't explicitly reference the sibling tools as alternatives, which keeps it from a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
jev_gateJev GateARead-only
Use as an advisory second opinion for an atomic yes/no gate when a probability threshold would materially change the next action. A passed gate is not authorization for destructive, irreversible, production, security, or data-loss-sensitive actions.
| Name | Required | Description | Default |
|---|---|---|---|
| model | No | ||
| state | Yes | ||
| statement | Yes | ||
| threshold | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=true, so the safety profile is partly covered. The description adds substantive context beyond that: the tool is explicitly 'advisory' and its output is not authorization for irreversible or security-sensitive actions, which is important behavioral framing an agent needs before acting on a pass. It still does not say what a pass or fail returns, what the default 0.8 threshold implies, or what happens on a failed gate.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the operational trigger followed immediately by the critical caveat. Every clause carries weight and nothing is redundant.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 4-parameter tool with 0% schema description coverage and no output schema, the description leaves the central contract unstated: what 'state' and 'statement' are, what a gate pass/fail actually returns, and how the threshold maps to the outcome. The advisory/not-authorization framing is valuable but does not compensate for those gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across 4 parameters, so the description must carry the parameter burden and largely does not. It gestures at 'probability threshold' (mapping loosely to the threshold parameter) but never explains what state, statement, or model should contain, nor the 0-1 range and 0.8 default.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific capability: an atomic yes/no gate that acts as an advisory second opinion, conditioned on a probability threshold. An agent can grasp what the tool does. It does not, however, differentiate itself from the sibling tools (jev_decide, jev_route, jev_risk_score), which all sound like judgment/decision primitives.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives a real selection condition ('when a probability threshold would materially change the next action') and a genuine exclusion (passed gates are not authorization for destructive, irreversible, production, security, or data-loss-sensitive actions). It never names an alternative sibling or says when one of those should be preferred instead, so it stops short of full routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
jev_risk_scoreJev Risk ScoreARead-only
Use as an advisory second opinion for consequential or ambiguous-risk changes after inspecting relevant context. Returns low/medium/high/critical risk. It is never authorization to proceed and never overrides tests, deterministic evidence, primary-model reasoning, or review.
| Name | Required | Description | Default |
|---|---|---|---|
| model | No | ||
| state | Yes | Proposed action plus the minimal project context needed to assess its risk. | |
| question | No | How risky is this proposed change if executed as described? |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations mark it read-only and open-world; the description adds non-obvious behavioral limits: it is advisory only and explicitly subordinate to tests, evidence, reasoning, and review. It also states return categories, but does not describe model/state handling or output format beyond the category list.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tightly constructed sentences, front-loaded with the usage directive and return categories, then the critical limitation. No filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Purpose, usage boundaries, and return categories are covered, which matters given the absence of an output schema. However, with 33% schema coverage and optional model/question parameters, the description leaves parameter-level guidance to the schema and does not compensate for the gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 33%: state has a description, model and question do not. The description does not compensate by explaining these parameters, the anyOf state formats, or question semantics, so it adds little beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States the tool returns a risk score with four categories and frames it as an advisory second opinion. Its role is distinguishable from decision/gate siblings by the non-authorization and non-override framing, though no sibling is named explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Specifies when to use it: for consequential or ambiguous-risk changes after inspecting context. It also sets when-not boundaries: never authorization and never overrides tests, deterministic evidence, primary-model reasoning, or review. Alternatives are implied but not named.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
jev_routeJev RouteBRead-only
Use as a second opinion for uncertain routing among a fixed set of tools, subsystems, workstreams, or handling paths only when deterministic evidence does not already decide the route. Never changes the user-selected host model.
| Name | Required | Description | Default |
|---|---|---|---|
| model | No | ||
| state | Yes | Task state and only the context relevant to routing. | |
| objective | Yes | What the routing decision should optimize for. | |
| candidates | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=true, and the description is consistent with them. It adds one meaningful behavioral constraint — "Never changes the user-selected host model" — which clarifies the tool is advisory only. It says nothing about the shape of the result or latency, so the added context is modest.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences; the usage condition and the candidate scope are front-loaded. No filler or redundancy, though the second sentence about the host model is somewhat tangential to selection.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 4-parameter, 3-required tool with no output schema and half its parameters undocumented, the description covers the decision context but not the return value or what "second opinion" output looks like. Adequate but with clear gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 50%: state, objective, and candidates are documented in the schema, but the model parameter has no description in either place. The description text adds no parameter-level meaning at all, so it fails to compensate for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific function — producing a routing "second opinion" over a fixed candidate set (tools, subsystems, workstreams, handling paths) — which is concrete enough to distinguish it from a scoring or gating tool. It does not explicitly name or contrast with its siblings (jev_gate, jev_decide, jev_risk_score), so it falls short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives an explicit when-to-use condition: only for uncertain routing where deterministic evidence does not already decide the route. That is a real exclusion criterion. However, it never names the alternative tool to use when the route IS deterministic, so it stops at a 4.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
4 tool updates
v0.4.1- First observed
jev_decide - First observed
jev_gate - First observed
jev_risk_score - First observed
jev_route
TDQS
Scored across 4 tools
Each tool has a clearly distinct decision scope: jev_gate handles atomic yes/no gates, jev_decide handles bounded multi-alternative decisions, jev_route handles routing among fixed paths, and jev_risk_score produces risk levels. The descriptions reinforce these boundaries, so an agent can reliably select the right tool without overlap.
All tools share the jev_ prefix and use snake_case, which is predictable and readable. However, the suffixes mix verbs and nouns (decide, route vs. gate, risk_score), so the pattern is mostly rather than perfectly uniform.
Four tools is well-scoped for a focused advisory second-opinion server. Each tool covers a distinct decision category, and none appears redundant or trivial.
The surface covers the main advisory decision types: binary gate, bounded choice, routing, and risk scoring. Minor gaps exist around explaining or calibrating second opinions, but agents can work around them within the stated domain.
Maintenance
Related MCP Connectors
Four tools to check, watch, diagnose and verify AI agents and MCP servers. Free, read-only.
Free MCP window into a live autonomous machine-economy experiment: telemetry, hypothesis scoreboard.
Read-only MCP tools for AI agent discovery, structured resources, and NIULAI information.
Patterns for designing and reviewing AI skills, agents, and multi-agent workflows. Find guidance on context economy, delegation, verification, and tool design; inspect claims, worked examples, maturity labels, and source references. Five read-only tools let agents discover relevant patterns, compare concise cards, read specific sections, and explore relationships. Hosted Streamable HTTP at https://agentic-atlas.dev/mcp/ — no installation, account, or API key required. Browse the atlas at https://agentic-atlas.dev/.
Related MCP Servers
- AlicenseNot gradedqualityBmaintenanceEnables MCP hosts to query Jev's typed decision model—yes/no, choice, and score—with calibrated probabilities, while defaulting to an offline mock and disclosing all egress unless explicitly enabled.31Apache 2.0
- AlicenseNot gradedqualityBmaintenanceProvides MCP tools that wrap Jev's finite typed judgments to give agents a semantic decision-control layer for task routing, pre-execution risk gating, and completion verification, keeping final execution authority in code.MIT
- AlicenseNot gradedqualityBmaintenanceAn MCP server that exposes eleven typed decision tools—check, choose, score, judge, route, triage, guard, grep, rank, compact, and ask—so agents can make fast, branchable yes/no, option-pick, score, and filtering decisions on text via TypeSafe's Jev model.134 npm3MIT
- AlicenseNot gradedqualityCmaintenanceEnables MCP-capable agents to run TypeSafe's Jev judgment model as typed yes/no, choice, and score tools, with calibrated probabilities, confidence thresholds, escalation for uncertain or non-judgment tasks, and an optional action gate that fails open.1MIT