Skip to main content
Glama

mcp-agent-openjev

Local typed probabilistic decision service — Choice, Noul, Score with calibrated probabilities and abstention — powered by OpenJev (the tuned 27B Qwen3.5 decision model) over Bionic's Open Responses endpoint in LM Studio.

OpenJev is tuned so a single output position carries the decision: one score per option letter, then a calibration step. LM Studio's /v1/completions and /v1/chat/completions never return logprobs (and cap top_logprobs at 20), so this project drives the one surface that does — /v1/responses — with the two settings that make the readout exact: reasoning.effort: "none" (the model emits exactly one token, the answer letter) and undistorted logits (temperature: 1.0, top_p: 1.0, frequency_penalty: 0, presence_penalty: 0).

For whom: anyone with a local OpenJev GGUF in LM Studio who wants System-1 style typed decisions (routing, guardrails, intent detection, severity ratings) — and agents that want them as an MCP tool.

How it works

For a Choice over N options, the client:

  1. labels the options A, B, C, … and builds the OpenJev prompt (State: / Question: / Options:);

  2. sends it to POST /v1/responses with max_tokens=1, reasoning.effort=none, and an uncapped top_logprobs=64;

  3. reads each letter's log-probability at the answer position;

  4. temperature-scales the logits (softmax, T=0.85) for calibrated posterior probabilities;

  5. abstains (value="UNKNOWN") when the winner's confidence falls below JEV_ABSTAIN_THRESHOLD (auto scales it to 1.25 / N), preserving the raw argmax in tentative_value.

Related MCP server: llmstudio-mcp

Install

Requires uv and a running LM Studio with the OpenJev model loaded:

uv sync

Copy .env.example to .env and set JEV_MODEL to the OpenJev model key LM Studio serves (default openjev). Point JEV_BASE_URL at the LM Studio server (default http://127.0.0.1:1234).

Usage

One-shot CLI

uv run mcp-agent-openjev doctor

uv run mcp-agent-openjev choice "unauthorized login from an unknown IP" `
    -c billing -c tech_support -c security `
    --criteria "pick the handling department"

uv run mcp-agent-openjev noul "the request is urgent" "this is a support request"

uv run mcp-agent-openjev score "PII exposed in a public bucket for 3 days" `
    -t low -t medium -t high -t critical

Score tiers

A tier is either a label string or a dict with label (or value) and an optional score:

// plain labels - weight is the position in the list
["low", "medium", "high", "critical"]

// explicit form - 'score' must equal the tier's position
[{"label": "low", "score": 0}, {"label": "critical", "score": 3}]

score must match the tier's position because a tier's weight is its ordinal position - that is what makes expected_score an expected tier index, which is what callers compare against. A tier dict without label/value, a duplicate label, or a score that disagrees with the position is rejected (InvalidTierError, HTTP 400 on the service) rather than silently degraded to a positional label - a degraded tier is rendered into the prompt as 0. 0, so the model ends up ranking meaningless numbers. level_probabilities is always keyed by tier label.

The response also echoes the scale it used, so the score never has to be read against an assumed convention:

{
  "expected_score": 2.2056,
  "level_probabilities": {"low": 0.00005, "high": 0.783, "critical": 0.211, "UNKNOWN": 0.0002},
  "tier_weights": {"low": 0.0, "high": 2.0, "critical": 3.0, "UNKNOWN": 0.0},
  "confidence": 0.783,
  "abstained": false
}

tier_weights is ordered by tier, and includes UNKNOWN (weight 0.0) when abstention is enabled, so expected_score can be recomputed from level_probabilities alone.

HTTP service

uv run mcp-agent-openjev http --host 127.0.0.1 --port 8377
curl http://localhost:8377/health
curl -X POST http://localhost:8377/v1/decide/choice -H "Content-Type: application/json" -d '{
  "state": {"ticket": "unauthorized login from an unknown IP"},
  "candidates": ["billing", "tech_support", "security"],
  "criteria": {"billing": "payments", "security": "unauthorized access"}
}'

Endpoints: POST /v1/decide/choice, POST /v1/decide/noul, POST /v1/decide/score, GET /health. Choice and score accept optional model / temperature / abstain_threshold overrides per request.

MCP server

uv run mcp-agent-openjev serve              # stdio (default)
uv run mcp-agent-openjev serve --http       # streamable-http on :8030

Tools: decide_choice, decide_noul, decide_score, decision_status. Wire it into opencode's MCP block:

"mcp-agent-openjev": {
  "type": "local",
  "command": ["uv", "--project", "<path-to-mcp_agent_openjev>", "run", "python", "-m", "mcp_agent_openjev", "serve"],
  "environment": {}
}

Examples

uv run python examples/ticket_router.py
uv run python examples/severity_score.py

Decisions

Type

Purpose

Returns

Choice

categorical decision over candidates

value, probabilities, confidence, abstained, tentative_value

Noul

binary truth judgment

value, probability_true, confidence

Score

ordered tier evaluation

expected_score, level_probabilities, confidence

Configuration

Variable

Default

Meaning

JEV_BASE_URL

http://127.0.0.1:1234

LM Studio server URL

JEV_MODEL

openjev

Model key LM Studio serves

JEV_API_KEY

(empty)

Sent as Authorization: Bearer only when set; falls back to LM_STUDIO_API_KEY

JEV_TEMPERATURE

0.85

Softmax temperature scaling (OpenJev READOUT_T)

JEV_NOUL_T

1.829074

Noul sigmoid scale (OpenJev READOUT_NOUL_T)

JEV_NOUL_BIAS

0

Noul sigmoid bias (OpenJev READOUT_NOUL_BIAS)

JEV_PERMS

1

Option orders to average over (accuracy mode; costs latency)

JEV_ABSTAIN_THRESHOLD

auto

Abstention threshold; or auto (1.25 / N)

JEV_TIMEOUT

120

Backend request timeout (seconds)

The released helper's READOUT_T, READOUT_NOUL_T, READOUT_NOUL_BIAS and READOUT_PERMS variables are also honoured.

License

The adapter, service, CLI and MCP surface in this repository are MIT (see LICENSE). Two upstream licenses are carried alongside because the calibration mirrors their code and model:

  • LICENSE-APACHE-2.0.txt — Apache 2.0. Covers the openjev-server code this calibration mirrors (helper/, serve/) and OpenJev's Apache-2.0 open base model, as requested by the OpenJev model repo.

  • CC BY-NC 4.0 — the OpenJev weights. Free for research and other non-commercial use with attribution; commercial use requires a separate licence (open a discussion on the model repo). The terms travel with the GGUF, so they apply however the model is served.

Available Tools

4 tools
decide_choiceA

Select the single best candidate for the state, with calibrated probabilities.

state: dict of context to evaluate. candidates: the allowed options (strings). criteria: how to judge them (plain string or mapping). allow_abstain: when the best confidence is too low, value becomes UNKNOWN and tentative_value keeps the raw argmax. Returns ChoiceDecision JSON.

ParametersJSON Schema
NameRequiredDescriptionDefault
stateYes
criteriaNo
candidatesYes
allow_abstainNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the behavioral disclosure burden. It explains calibrated probabilities, the abstain mechanism (UNKNOWN plus tentative_value holding the raw argmax), and the JSON return type. This is strong transparency, though it does not clarify what 'calibrated probabilities' means in practice.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and front-loaded, stating the core purpose first and then efficiently defining each parameter and the abstain behavior. Every sentence adds useful information without redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the output schema exists and the description covers parameters, behavioral nuances, and the return type, the tool is callable by an agent. The main missing piece is explicit guidance on how this tool relates to the sibling tools decide_score, decide_noul, and decision_status.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, and it does: state is a dict, candidates are strings, criteria can be a string or mapping, and allow_abstain controls low-confidence behavior. The only gap is that the semantics of a mapping-valued criteria object are not fully specified.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Select'), names the resource ('the single best candidate for the state'), and identifies the output type ('ChoiceDecision JSON'). It is clearly about candidate selection and distinguishes this from scoring/status siblings, though it does not explicitly name alternatives.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear context for when to use the tool: given a state and allowed candidates, choose the best one according to criteria. It also explains the abstain option and output behavior. However, it does not explicitly state when not to use this tool or name sibling alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

decide_noulB

Judge whether a binary assertion is TRUE or FALSE for the given state.

ParametersJSON Schema
NameRequiredDescriptionDefault
stateYes
assertionYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden of disclosing behavior, but it only states the purpose. It does not mention whether the tool is deterministic, pure, read-only, or what side effects might occur, leaving an agent without critical safety or performance expectations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, well-formed sentence that communicates the core behavior directly without filler. The key qualifiers 'binary' and 'TRUE or FALSE' are front-loaded, making it easy to parse.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has a nested object parameter, no annotations, and no schema description coverage, the description is too sparse to fully inform an agent about acceptable assertion syntax or state structure. The existence of an output schema covers return values, but input guidance is lacking.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate by explaining parameter meaning. It adds only that the assertion is binary and judged against the state, but does not describe expected formats, constraints, or examples for 'state' or 'assertion'. This is minimal compensation for the absence of schema-provided details.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('judge') and resource ('binary assertion') and explicitly defines the possible outcomes (TRUE or FALSE). This semantically distinguishes it from siblings like decide_choice and decide_score, which imply non-binary outputs, even without naming them.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies use when a binary true/false judgment on an assertion against a state is required, but it provides no explicit guidance on when not to use it or how it compares to sibling tools. The context is basic and not fully explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

decide_scoreC

Rate the state across ordered tiers (e.g. severity levels).

tiers: ordered list of strings, or dicts with 'label'/'value' and an optional 'score' weight. Returns ScoreDecision JSON with expected_score and per-tier probability mass.

ParametersJSON Schema
NameRequiredDescriptionDefault
stateYes
tiersYes
criteriaNo
allow_abstainNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description must carry the full behavioral burden. It discloses that the tool returns a ScoreDecision JSON with expected_score and per-tier probability mass, which is useful, but it does not mention side effects (e.g., whether it is read-only), error behavior, determinism, or any constraints on the state object. For a decision tool, this is minimal.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise (two sentences) and front-loads the primary purpose. The second sentence adds needed detail about the tiers format and return type. It is efficient with no filler, though the tier format explanation is a bit dense.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given four parameters, nested objects, and an output schema that is not provided in the definition, the description is incomplete. It fails to describe the state object, criteria, and allow_abstain, and does not clarify the expected behavior or edge cases. The return format is mentioned but not elaborated. An agent would struggle to call this correctly without additional context.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It explains the tiers parameter in some detail (list of strings or dicts with label/value and optional score weight), but provides no information about state, criteria, or allow_abstain. These parameters remain underdocumented, leaving the agent to guess their format and purpose.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a clear verb ('Rate') and a specific resource ('state across ordered tiers'), with an example ('severity levels') that grounds the purpose. It does not explicitly name sibling tools to differentiate, but the focus on tiers and the return type (ScoreDecision with probability mass) makes the function distinct within the decision family.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives no guidance on when to use this tool versus the siblings (decision_status, decide_choice, decide_noul). It only hints at usage via 'ordered tiers,' but does not state prerequisites, exclusions, or alternative selection criteria. An agent has to infer the appropriate context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

decision_statusA

Report the configured backend, model, and endpoint reachability.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It communicates a read-only reporting behavior and indicates reachability checks, but it does not disclose potential side effects of probing endpoints, timeout behavior, or whether the tool only reads cached state. It is adequate but not deeply transparent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, tight sentence with no filler. It front-loads the action and clearly enumerates what is reported.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter status tool with an output schema present, the description covers the essential purpose and scope. The only notable gap is the lack of usage guidance relative to sibling tools, which slightly reduces completeness.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There are zero parameters and schema coverage is 100%, so the schema carries no load. According to the baseline for zero-parameter tools, a score of 4 is appropriate; there is no parameter ambiguity to resolve.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Report') and names the exact resource: configured backend, model, and endpoint reachability. This clearly differentiates it from the sibling decide_* tools, which are decision-oriented rather than status-reporting.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is given on when to use this tool versus the sibling decision tools. The distinction is only implicit through the tool name and description, but no explicit context or alternative is mentioned.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 4 tool updatesv0.1.0
    • First observeddecide_choice
    • First observeddecide_noul
    • First observeddecide_score
    • First observeddecision_status

TDQS

A3.6/5.0

Scored across 4 tools

Disambiguation5/5

Each tool has a distinct purpose: decision_status reports system state, while the three decide_* tools handle different decision types (choice among options, binary assertion, and tiered scoring). No overlap or ambiguity between them.

Naming Consistency4/5

Three tools follow the consistent 'decide_' prefix with action verbs, but decision_status deviates with a noun-based name. This is a minor inconsistency, but the pattern is still recognizable.

Tool Count5/5

With 4 tools, the server is well-scoped for a decision-routing purpose. Each tool covers a necessary aspect without redundancy or bloat.

Completeness4/5

The surface covers status, choice selection, boolean judgment, and scoring, which aligns with the server's purpose. A possible gap is a tool for explaining or justifying decisions, but the provided functionality is solid.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    B
    maintenance
    Enables LLM agents to route responses as accept, verify, or ask-a-human based on token logprobs, and provides an MCP server for delegating generation to local models with confidence bands.
    MIT
  • F
    license
    A
    quality
    C
    maintenance
    MCP server that connects LLM agents to a local LM Studio instance, enabling model management, OpenAI-compatible chat completions, text completions, and embeddings through a set of tools.
    9
    1
    -
  • A
    license
    Not graded
    quality
    B
    maintenance
    Enables coding agents to query a locally running Kev decision model through MCP tools, returning calibrated probabilities for typed questions such as yes/no, choice, and score.
    Apache 2.0