Skip to main content
Glama
Afloat16

jev-mcp

by Afloat16

jev-mcp

CI License: MIT Node.js MCP

Unofficial, community-maintained MCP server for TypeSafe AI Jev.

简体中文 · Quick start · Examples · Security · FAQ

jev-mcp exposes TypeSafe AI's Jev decision model as four conservative, read-only MCP tools for bounded probabilistic decisions.

It is intentionally designed as a second-opinion layer, not as a replacement for the active coding/reasoning model.

This project is not affiliated with, endorsed by, sponsored by, or an official product of TypeSafe AI or OpenAI.

Why this exists

Strong coding agents are good at open-ended reasoning, code generation, debugging, and repository-wide analysis. Jev is useful for a different class of problem: small, explicit decisions that benefit from a structured probability signal.

Typical examples:

  • retry vs rollback vs change strategy;

  • route to one of several known tools or subsystems;

  • estimate low / medium / high / critical change risk;

  • evaluate a yes/no gate against a threshold;

  • repeat the same bounded decision many times in an automated workflow.

The design rule is simple:

deterministic evidence
        >
host-model repository-aware reasoning
        >
Jev probabilistic advice

A Jev result is advisory. It is never proof, never ground truth, and never authorization for destructive or irreversible work.

Related MCP server: jev_mcp

Architecture

flowchart LR
    U[User] --> H[Active host model / MCP client]
    E[Tests · compiler · runtime · static analysis] -->|highest-priority evidence| H
    H -->|stdio MCP| M[jev-mcp]
    M -->|HTTPS + Bearer token| J[TypeSafe AI Jev API]
    J -->|probabilistic advisory result| M
    M -->|structured tool result| H
    H --> O[Final decision / action]

The API key stays in the local process environment or a gitignored .env file. It is not placed in the MCP client configuration.

Privacy boundary: anything placed in a tool's state is sent to the configured TypeSafe API endpoint. Send only the minimum necessary, preferably redacted state. See Security and privacy model.

Tool surface

Tool

Use it for

Do not use it for

jev_decide

Choosing among 2–255 explicit alternatives

Open-ended design or coding

jev_route

Routing among known tools/subsystems/workflows

Switching the user's selected model

jev_risk_score

Advisory low/medium/high/critical risk signal

Replacing tests or review

jev_gate

Yes/no probability against a threshold

Authorizing destructive actions

All four tools are declared read-only. They do not edit files, execute shell commands, deploy infrastructure, or change the model selected by the user.

Successful responses include local metadata similar to:

{
  "jev_mcp_advisory": {
    "authority": "secondary_advisory",
    "finalDecisionBy": "active_host_model",
    "deterministicEvidenceOverrides": true,
    "doNotTreatProbabilityAsFact": true
  }
}

That metadata is added by this MCP server; it is not Jev model output.

60-second install

Requirements:

  • Node.js 20+

  • a TypeSafe API key / applicable TypeSafe access and credits

  • an MCP host that can launch a local stdio server

Clone and set up:

git clone https://github.com/Afloat16/jev-mcp.git
cd jev-mcp
./setup.sh

Windows PowerShell:

git clone https://github.com/Afloat16/jev-mcp.git
cd jev-mcp
./setup.ps1

The setup script reads the API key without echoing it, stores it only in the local gitignored .env file when needed, installs dependencies, and runs local checks.

For a more explicit walkthrough, see Quick start.

Codex configuration

Add this to ~/.codex/config.toml and replace the path:

[mcp_servers.jev]
command = "node"
args = ["/ABSOLUTE/PATH/TO/jev-mcp/dist/index.js"]

Restart the MCP host or start a new session after configuration changes.

For users who explicitly want a conservative cross-project Codex policy, merge:

codex/AGENTS.jev-conservative.md

into your own ~/.codex/AGENTS.md.

Do not copy unrelated existing instructions away.

Which integration should I use?

Situation

Recommended approach

Everyday interactive coding

Strong host model alone, or TypeSafe's agent skill

Learning Jev concepts / patterns

TypeSafe agent skill

Stable callable decision tools

jev-mcp

CI / agent orchestration

jev-mcp or a direct SDK integration

High-volume application logic

Direct SDK/API integration is often the cleanest

Need Jev to replace tests/compiler

Do not use Jev for that

The MCP is most useful when the tool boundary itself matters: repeatability, structured output, orchestration, explicit thresholds, or shared agent workflows.

Configuration

Variable

Required

Default

Description

TYPESAFE_API_KEY

yes

—

TypeSafe credential

JEV_MODEL

no

jev-latest

Jev model override

TYPESAFE_BASE_URL

no

https://api.typesafe.ai

API base URL

TYPESAFE_TIMEOUT_MS

no

15000

Request timeout, 250–120000 ms

JEV_ENV_FILE

no

project .env

Alternate env file path

Existing process environment variables override values loaded from .env.

Secret-handling rules

  • Never commit .env.

  • Never put API keys in README, AGENTS.md, MCP config, examples, issues, screenshots, or CI logs.

  • If a key is ever pasted into a shared surface, rotate/revoke it.

  • Run npm run secrets:check before publishing changes.

The bundled scanner is a guardrail, not a complete DLP system.

Examples

See examples/README.md for redacted examples covering:

  • ambiguous CI failure routing;

  • change-risk assessment;

  • retry / rollback / escalate decisions;

  • yes/no gates with explicit thresholds.

All examples intentionally use synthetic data and placeholders.

Local verification

No TypeSafe API call:

npm run doctor
npm run check

Interactive MCP inspection:

npm run inspect

A live tool invocation in MCP Inspector uses your own TypeSafe account and may consume provider credits.

Project status

Item

Status

Interface

4 read-only MCP tools

Transport

local stdio

Node.js

20+

License

MIT

npm publishing

intentionally disabled

API dependency

TypeSafe-hosted System One API

Stability

pre-1.0; behavior may evolve

The project follows semantic versioning in spirit, but while it remains below 1.0.0, minor releases may refine tool schemas or behavior. Breaking changes should be documented in CHANGELOG.md and migration notes.

Documentation

Contributing

Focused issues and pull requests are welcome. Please read CONTRIBUTING.md first and run:

npm run check

before opening a PR.

Security-sensitive reports should follow SECURITY.md, not public issue comments.

License

MIT. See LICENSE.

The license covers this repository's code only. Third-party services, names, APIs, trademarks, pricing, and availability remain subject to their respective owners. See NOTICE.

References

Available Tools

4 tools
jev_decideJev DecideA
Read-only

Use as an independent second opinion only after inspecting enough context when a bounded decision has 2-255 explicit alternatives and genuine uncertainty remains. The active host model keeps final authority; do not mechanically follow Jev or use it for open-ended coding/prose.

ParametersJSON Schema
NameRequiredDescriptionDefault
modelNoOptional Jev model override.
stateYesMinimal relevant state/context for the decision.
optionsYes
questionYesOne atomic decision question.

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=true, so the safety profile is covered. The description adds valuable behavioral context the annotations can't: the host model retains final authority, and output should not be mechanically followed. It omits latency/cost/determinism of the external call.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two dense sentences with zero waste: the usage condition and the authority caveat are front-loaded, and the exclusion ('open-ended coding/prose') closes the second sentence. Every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only advisory tool with no output schema, the description covers invocation conditions, boundaries, and authority semantics well. It does not describe the shape of the response or what the model override does, which leaves a small gap given no output schema exists to carry that.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 75%, so the schema already documents model, state, and question. The description's '2-255 explicit alternatives' maps meaning onto the options array and its bounds, but adds nothing for state or question beyond the schema. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific capability: producing an independent second opinion on a bounded decision with 2-255 explicit alternatives. It does not name its sibling tools (jev_gate, jev_route, jev_risk_score), so differentiation relies on the 'bounded decision' framing rather than explicit routing.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Strong when-to-use ('only after inspecting enough context when a bounded decision has 2-255 explicit alternatives and genuine uncertainty remains') and when-not-to-use ('do not... use it for open-ended coding/prose'). It doesn't explicitly reference the sibling tools as alternatives, which keeps it from a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

jev_gateJev GateA
Read-only

Use as an advisory second opinion for an atomic yes/no gate when a probability threshold would materially change the next action. A passed gate is not authorization for destructive, irreversible, production, security, or data-loss-sensitive actions.

ParametersJSON Schema
NameRequiredDescriptionDefault
modelNo
stateYes
statementYes
thresholdNo

TDQS

A3.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=true, so the safety profile is partly covered. The description adds substantive context beyond that: the tool is explicitly 'advisory' and its output is not authorization for irreversible or security-sensitive actions, which is important behavioral framing an agent needs before acting on a pass. It still does not say what a pass or fail returns, what the default 0.8 threshold implies, or what happens on a failed gate.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with the operational trigger followed immediately by the critical caveat. Every clause carries weight and nothing is redundant.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 4-parameter tool with 0% schema description coverage and no output schema, the description leaves the central contract unstated: what 'state' and 'statement' are, what a gate pass/fail actually returns, and how the threshold maps to the outcome. The advisory/not-authorization framing is valuable but does not compensate for those gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across 4 parameters, so the description must carry the parameter burden and largely does not. It gestures at 'probability threshold' (mapping loosely to the threshold parameter) but never explains what state, statement, or model should contain, nor the 0-1 range and 0.8 default.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific capability: an atomic yes/no gate that acts as an advisory second opinion, conditioned on a probability threshold. An agent can grasp what the tool does. It does not, however, differentiate itself from the sibling tools (jev_decide, jev_route, jev_risk_score), which all sound like judgment/decision primitives.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives a real selection condition ('when a probability threshold would materially change the next action') and a genuine exclusion (passed gates are not authorization for destructive, irreversible, production, security, or data-loss-sensitive actions). It never names an alternative sibling or says when one of those should be preferred instead, so it stops short of full routing guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

jev_risk_scoreJev Risk ScoreA
Read-only

Use as an advisory second opinion for consequential or ambiguous-risk changes after inspecting relevant context. Returns low/medium/high/critical risk. It is never authorization to proceed and never overrides tests, deterministic evidence, primary-model reasoning, or review.

ParametersJSON Schema
NameRequiredDescriptionDefault
modelNo
stateYesProposed action plus the minimal project context needed to assess its risk.
questionNoHow risky is this proposed change if executed as described?

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations mark it read-only and open-world; the description adds non-obvious behavioral limits: it is advisory only and explicitly subordinate to tests, evidence, reasoning, and review. It also states return categories, but does not describe model/state handling or output format beyond the category list.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tightly constructed sentences, front-loaded with the usage directive and return categories, then the critical limitation. No filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Purpose, usage boundaries, and return categories are covered, which matters given the absence of an output schema. However, with 33% schema coverage and optional model/question parameters, the description leaves parameter-level guidance to the schema and does not compensate for the gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 33%: state has a description, model and question do not. The description does not compensate by explaining these parameters, the anyOf state formats, or question semantics, so it adds little beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States the tool returns a risk score with four categories and frames it as an advisory second opinion. Its role is distinguishable from decision/gate siblings by the non-authorization and non-override framing, though no sibling is named explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Specifies when to use it: for consequential or ambiguous-risk changes after inspecting context. It also sets when-not boundaries: never authorization and never overrides tests, deterministic evidence, primary-model reasoning, or review. Alternatives are implied but not named.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

jev_routeJev RouteB
Read-only

Use as a second opinion for uncertain routing among a fixed set of tools, subsystems, workstreams, or handling paths only when deterministic evidence does not already decide the route. Never changes the user-selected host model.

ParametersJSON Schema
NameRequiredDescriptionDefault
modelNo
stateYesTask state and only the context relevant to routing.
objectiveYesWhat the routing decision should optimize for.
candidatesYes

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=true, and the description is consistent with them. It adds one meaningful behavioral constraint — "Never changes the user-selected host model" — which clarifies the tool is advisory only. It says nothing about the shape of the result or latency, so the added context is modest.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences; the usage condition and the candidate scope are front-loaded. No filler or redundancy, though the second sentence about the host model is somewhat tangential to selection.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 4-parameter, 3-required tool with no output schema and half its parameters undocumented, the description covers the decision context but not the return value or what "second opinion" output looks like. Adequate but with clear gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 50%: state, objective, and candidates are documented in the schema, but the model parameter has no description in either place. The description text adds no parameter-level meaning at all, so it fails to compensate for the coverage gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific function — producing a routing "second opinion" over a fixed candidate set (tools, subsystems, workstreams, handling paths) — which is concrete enough to distinguish it from a scoring or gating tool. It does not explicitly name or contrast with its siblings (jev_gate, jev_decide, jev_risk_score), so it falls short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives an explicit when-to-use condition: only for uncertain routing where deterministic evidence does not already decide the route. That is a real exclusion criterion. However, it never names the alternative tool to use when the route IS deterministic, so it stops at a 4.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 4 tool updatesv0.4.1
    • First observedjev_decide
    • First observedjev_gate
    • First observedjev_risk_score
    • First observedjev_route

TDQS

A3.9/5.0

Scored across 4 tools

Disambiguation5/5

Each tool has a clearly distinct decision scope: jev_gate handles atomic yes/no gates, jev_decide handles bounded multi-alternative decisions, jev_route handles routing among fixed paths, and jev_risk_score produces risk levels. The descriptions reinforce these boundaries, so an agent can reliably select the right tool without overlap.

Naming Consistency4/5

All tools share the jev_ prefix and use snake_case, which is predictable and readable. However, the suffixes mix verbs and nouns (decide, route vs. gate, risk_score), so the pattern is mostly rather than perfectly uniform.

Tool Count5/5

Four tools is well-scoped for a focused advisory second-opinion server. Each tool covers a distinct decision category, and none appears redundant or trivial.

Completeness4/5

The surface covers the main advisory decision types: binary gate, bounded choice, routing, and risk scoring. Minor gaps exist around explaining or calibrating second opinions, but agents can work around them within the stated domain.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    B
    maintenance
    Enables MCP hosts to query Jev's typed decision model—yes/no, choice, and score—with calibrated probabilities, while defaulting to an offline mock and disclosing all egress unless explicitly enabled.
    31
    Apache 2.0
  • A
    license
    Not graded
    quality
    B
    maintenance
    Provides MCP tools that wrap Jev's finite typed judgments to give agents a semantic decision-control layer for task routing, pre-execution risk gating, and completion verification, keeping final execution authority in code.
    MIT
  • A
    license
    Not graded
    quality
    B
    maintenance
    An MCP server that exposes eleven typed decision tools—check, choose, score, judge, route, triage, guard, grep, rank, compact, and ask—so agents can make fast, branchable yes/no, option-pick, score, and filtering decisions on text via TypeSafe's Jev model.
    134 npm
    3
    MIT
  • A
    license
    Not graded
    quality
    C
    maintenance
    Enables MCP-capable agents to run TypeSafe's Jev judgment model as typed yes/no, choice, and score tools, with calibrated probabilities, confidence thresholds, escalation for uncertain or non-judgment tasks, and an optional action gate that fails open.
    1
    MIT