openjev-router
This MCP server exposes local OpenJev-backed probabilistic decision tools for categorical choices, binary judgments, tiered scoring, and backend status checks.
Check backend, model, and endpoint reachability with
decision_status.Make calibrated categorical decisions with
decide_choiceover candidate options, optionally using criteria and abstention.Judge whether a binary assertion is TRUE or FALSE for a given state with
decide_noul.Rate a state across ordered tiers (e.g. severity) with
decide_score, returning expected score and per-tier probabilities.Accept context as a JSON object (
state), plus optional criteria and abstention flags.Support agent routing, guardrails, intent detection, and severity rating via MCP.
Allows using Ollama as a backend for typed probabilistic decisions, including Choice, Noul, and Score, with calibrated probabilities and abstention via OpenJev-compatible clients.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@openjev-routerRoute this ticket to the correct department"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
mcp-agent-openjev
Local typed probabilistic decision service — Choice, Noul, Score with
calibrated probabilities and abstention — powered by
OpenJev (the tuned 27B Qwen3.5 decision
model) over Bionic's Open Responses endpoint in
LM Studio.
OpenJev is tuned so a single output position carries the decision: one score per
option letter, then a calibration step. LM Studio's /v1/completions and
/v1/chat/completions never return logprobs (and cap top_logprobs at 20), so
this project drives the one surface that does — /v1/responses — with the two
settings that make the readout exact: reasoning.effort: "none" (the model emits
exactly one token, the answer letter) and undistorted logits
(temperature: 1.0, top_p: 1.0, frequency_penalty: 0, presence_penalty: 0).
For whom: anyone with a local OpenJev GGUF in LM Studio who wants System-1 style typed decisions (routing, guardrails, intent detection, severity ratings) — and agents that want them as an MCP tool.
How it works
For a Choice over N options, the client:
labels the options
A,B,C, … and builds the OpenJev prompt (State:/Question:/Options:);sends it to
POST /v1/responseswithmax_tokens=1,reasoning.effort=none, and an uncappedtop_logprobs=64;reads each letter's log-probability at the answer position;
temperature-scales the logits (softmax,
T=0.85) for calibrated posterior probabilities;abstains (
value="UNKNOWN") when the winner's confidence falls belowJEV_ABSTAIN_THRESHOLD(autoscales it to1.25 / N), preserving the raw argmax intentative_value.
Related MCP server: llmstudio-mcp
Install
Requires uv and a running LM Studio with the OpenJev model loaded:
uv syncCopy .env.example to .env and set JEV_MODEL to the OpenJev model key LM
Studio serves (default openjev). Point JEV_BASE_URL at the LM Studio server
(default http://127.0.0.1:1234).
Usage
One-shot CLI
uv run mcp-agent-openjev doctor
uv run mcp-agent-openjev choice "unauthorized login from an unknown IP" `
-c billing -c tech_support -c security `
--criteria "pick the handling department"
uv run mcp-agent-openjev noul "the request is urgent" "this is a support request"
uv run mcp-agent-openjev score "PII exposed in a public bucket for 3 days" `
-t low -t medium -t high -t criticalScore tiers
A tier is either a label string or a dict with label (or value) and an
optional score:
// plain labels - weight is the position in the list
["low", "medium", "high", "critical"]
// explicit form - 'score' must equal the tier's position
[{"label": "low", "score": 0}, {"label": "critical", "score": 3}]score must match the tier's position because a tier's weight is its ordinal
position - that is what makes expected_score an expected tier index, which is
what callers compare against. A tier dict without label/value, a duplicate
label, or a score that disagrees with the position is rejected
(InvalidTierError, HTTP 400 on the service) rather than silently degraded to a
positional label - a degraded tier is rendered into the prompt as 0. 0, so the
model ends up ranking meaningless numbers. level_probabilities is always keyed
by tier label.
The response also echoes the scale it used, so the score never has to be read against an assumed convention:
{
"expected_score": 2.2056,
"level_probabilities": {"low": 0.00005, "high": 0.783, "critical": 0.211, "UNKNOWN": 0.0002},
"tier_weights": {"low": 0.0, "high": 2.0, "critical": 3.0, "UNKNOWN": 0.0},
"confidence": 0.783,
"abstained": false
}tier_weights is ordered by tier, and includes UNKNOWN (weight 0.0) when
abstention is enabled, so expected_score can be recomputed from
level_probabilities alone.
HTTP service
uv run mcp-agent-openjev http --host 127.0.0.1 --port 8377
curl http://localhost:8377/health
curl -X POST http://localhost:8377/v1/decide/choice -H "Content-Type: application/json" -d '{
"state": {"ticket": "unauthorized login from an unknown IP"},
"candidates": ["billing", "tech_support", "security"],
"criteria": {"billing": "payments", "security": "unauthorized access"}
}'Endpoints: POST /v1/decide/choice, POST /v1/decide/noul,
POST /v1/decide/score, GET /health. Choice and score accept optional
model / temperature / abstain_threshold overrides per request.
MCP server
uv run mcp-agent-openjev serve # stdio (default)
uv run mcp-agent-openjev serve --http # streamable-http on :8030Tools: decide_choice, decide_noul, decide_score, decision_status.
Wire it into opencode's MCP block:
"mcp-agent-openjev": {
"type": "local",
"command": ["uv", "--project", "<path-to-mcp_agent_openjev>", "run", "python", "-m", "mcp_agent_openjev", "serve"],
"environment": {}
}Examples
uv run python examples/ticket_router.py
uv run python examples/severity_score.pyDecisions
Type | Purpose | Returns |
| categorical decision over candidates |
|
| binary truth judgment |
|
| ordered tier evaluation |
|
Configuration
Variable | Default | Meaning |
|
| LM Studio server URL |
|
| Model key LM Studio serves |
| (empty) | Sent as |
|
| Softmax temperature scaling (OpenJev |
|
| Noul sigmoid scale (OpenJev |
|
| Noul sigmoid bias (OpenJev |
|
| Option orders to average over (accuracy mode; costs latency) |
|
| Abstention threshold; or |
|
| Backend request timeout (seconds) |
The released helper's READOUT_T, READOUT_NOUL_T, READOUT_NOUL_BIAS and
READOUT_PERMS variables are also honoured.
License
The adapter, service, CLI and MCP surface in this repository are MIT (see
LICENSE). Two upstream licenses are carried alongside because the calibration
mirrors their code and model:
LICENSE-APACHE-2.0.txt— Apache 2.0. Covers theopenjev-servercode this calibration mirrors (helper/,serve/) and OpenJev's Apache-2.0 open base model, as requested by the OpenJev model repo.CC BY-NC 4.0 — the OpenJev weights. Free for research and other non-commercial use with attribution; commercial use requires a separate licence (open a discussion on the model repo). The terms travel with the GGUF, so they apply however the model is served.
Available Tools
4 toolsdecide_choiceA
Select the single best candidate for the state, with calibrated probabilities.
state: dict of context to evaluate. candidates: the allowed options (strings). criteria: how to judge them (plain string or mapping). allow_abstain: when the best confidence is too low, value becomes UNKNOWN and tentative_value keeps the raw argmax. Returns ChoiceDecision JSON.
| Name | Required | Description | Default |
|---|---|---|---|
| state | Yes | ||
| criteria | No | ||
| candidates | Yes | ||
| allow_abstain | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the behavioral disclosure burden. It explains calibrated probabilities, the abstain mechanism (UNKNOWN plus tentative_value holding the raw argmax), and the JSON return type. This is strong transparency, though it does not clarify what 'calibrated probabilities' means in practice.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded, stating the core purpose first and then efficiently defining each parameter and the abstain behavior. Every sentence adds useful information without redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the output schema exists and the description covers parameters, behavioral nuances, and the return type, the tool is callable by an agent. The main missing piece is explicit guidance on how this tool relates to the sibling tools decide_score, decide_noul, and decision_status.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, and it does: state is a dict, candidates are strings, criteria can be a string or mapping, and allow_abstain controls low-confidence behavior. The only gap is that the semantics of a mapping-valued criteria object are not fully specified.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Select'), names the resource ('the single best candidate for the state'), and identifies the output type ('ChoiceDecision JSON'). It is clearly about candidate selection and distinguishes this from scoring/status siblings, though it does not explicitly name alternatives.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context for when to use the tool: given a state and allowed candidates, choose the best one according to criteria. It also explains the abstain option and output behavior. However, it does not explicitly state when not to use this tool or name sibling alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
decide_noulB
Judge whether a binary assertion is TRUE or FALSE for the given state.
| Name | Required | Description | Default |
|---|---|---|---|
| state | Yes | ||
| assertion | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden of disclosing behavior, but it only states the purpose. It does not mention whether the tool is deterministic, pure, read-only, or what side effects might occur, leaving an agent without critical safety or performance expectations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, well-formed sentence that communicates the core behavior directly without filler. The key qualifiers 'binary' and 'TRUE or FALSE' are front-loaded, making it easy to parse.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has a nested object parameter, no annotations, and no schema description coverage, the description is too sparse to fully inform an agent about acceptable assertion syntax or state structure. The existence of an output schema covers return values, but input guidance is lacking.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate by explaining parameter meaning. It adds only that the assertion is binary and judged against the state, but does not describe expected formats, constraints, or examples for 'state' or 'assertion'. This is minimal compensation for the absence of schema-provided details.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('judge') and resource ('binary assertion') and explicitly defines the possible outcomes (TRUE or FALSE). This semantically distinguishes it from siblings like decide_choice and decide_score, which imply non-binary outputs, even without naming them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies use when a binary true/false judgment on an assertion against a state is required, but it provides no explicit guidance on when not to use it or how it compares to sibling tools. The context is basic and not fully explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
decide_scoreC
Rate the state across ordered tiers (e.g. severity levels).
tiers: ordered list of strings, or dicts with 'label'/'value' and an optional 'score' weight. Returns ScoreDecision JSON with expected_score and per-tier probability mass.
| Name | Required | Description | Default |
|---|---|---|---|
| state | Yes | ||
| tiers | Yes | ||
| criteria | No | ||
| allow_abstain | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description must carry the full behavioral burden. It discloses that the tool returns a ScoreDecision JSON with expected_score and per-tier probability mass, which is useful, but it does not mention side effects (e.g., whether it is read-only), error behavior, determinism, or any constraints on the state object. For a decision tool, this is minimal.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise (two sentences) and front-loads the primary purpose. The second sentence adds needed detail about the tiers format and return type. It is efficient with no filler, though the tier format explanation is a bit dense.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given four parameters, nested objects, and an output schema that is not provided in the definition, the description is incomplete. It fails to describe the state object, criteria, and allow_abstain, and does not clarify the expected behavior or edge cases. The return format is mentioned but not elaborated. An agent would struggle to call this correctly without additional context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It explains the tiers parameter in some detail (list of strings or dicts with label/value and optional score weight), but provides no information about state, criteria, or allow_abstain. These parameters remain underdocumented, leaving the agent to guess their format and purpose.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a clear verb ('Rate') and a specific resource ('state across ordered tiers'), with an example ('severity levels') that grounds the purpose. It does not explicitly name sibling tools to differentiate, but the focus on tiers and the return type (ScoreDecision with probability mass) makes the function distinct within the decision family.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no guidance on when to use this tool versus the siblings (decision_status, decide_choice, decide_noul). It only hints at usage via 'ordered tiers,' but does not state prerequisites, exclusions, or alternative selection criteria. An agent has to infer the appropriate context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
decision_statusA
Report the configured backend, model, and endpoint reachability.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It communicates a read-only reporting behavior and indicates reachability checks, but it does not disclose potential side effects of probing endpoints, timeout behavior, or whether the tool only reads cached state. It is adequate but not deeply transparent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, tight sentence with no filler. It front-loads the action and clearly enumerates what is reported.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter status tool with an output schema present, the description covers the essential purpose and scope. The only notable gap is the lack of usage guidance relative to sibling tools, which slightly reduces completeness.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There are zero parameters and schema coverage is 100%, so the schema carries no load. According to the baseline for zero-parameter tools, a score of 4 is appropriate; there is no parameter ambiguity to resolve.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Report') and names the exact resource: configured backend, model, and endpoint reachability. This clearly differentiates it from the sibling decide_* tools, which are decision-oriented rather than status-reporting.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is given on when to use this tool versus the sibling decision tools. The distinction is only implicit through the tool name and description, but no explicit context or alternative is mentioned.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
4 tool updates
v0.1.0- First observed
decide_choice - First observed
decide_noul - First observed
decide_score - First observed
decision_status
TDQS
Scored across 4 tools
Each tool has a distinct purpose: decision_status reports system state, while the three decide_* tools handle different decision types (choice among options, binary assertion, and tiered scoring). No overlap or ambiguity between them.
Three tools follow the consistent 'decide_' prefix with action verbs, but decision_status deviates with a noun-based name. This is a minor inconsistency, but the pattern is still recognizable.
With 4 tools, the server is well-scoped for a decision-routing purpose. Each tool covers a necessary aspect without redundancy or bloat.
The surface covers status, choice selection, boolean judgment, and scoring, which aligns with the server's purpose. A possible gap is a tool for explaining or justifying decisions, but the provided functionality is solid.
Maintenance
Related MCP Connectors
MCP-first toolbox for agents: KV storage, auth, queue, and utility tools. Free in early access.
Free OpenAI-compatible inference with signed provenance receipts and 3 focused MCP tools.
LLM Orchestration Agent (Mcp)
OpenAI-compatible LLM MCP (7 tools); chat via balance key or x402 USDC on Base
Related MCP Servers
- AlicenseNot gradedqualityBmaintenanceEnables LLM agents to route responses as accept, verify, or ask-a-human based on token logprobs, and provides an MCP server for delegating generation to local models with confidence bands.MIT
- FlicenseAqualityCmaintenanceMCP server that connects LLM agents to a local LM Studio instance, enabling model management, OpenAI-compatible chat completions, text completions, and embeddings through a set of tools.91-
- AlicenseNot gradedqualityCmaintenanceExposes local LM Studio language models as MCP tools, enabling chat completions and model listing through a local OpenAI-compatible API without requiring API keys.MIT
- AlicenseNot gradedqualityBmaintenanceEnables coding agents to query a locally running Kev decision model through MCP tools, returning calibrated probabilities for typed questions such as yes/no, choice, and score.Apache 2.0