JevGuard MCP Server
JevGuard MCP Server is a zero-dependency MCP guardrail server that evaluates commands, code patches, and decisions with deterministic local caching and optional TypeSafe AI upstream evaluation.
Evaluate command safety: Classify shell commands as
ALLOW_AUTONOMOUS,REQUIRE_HUMAN_APPROVAL, orDENY_DESTRUCTIVE, with optional working directory and elevated-privilege context.Verify code patches: Audit diffs for security regressions, broken syntax, or critical impact and return
APPROVE,REQUEST_CHANGES, orREJECTwith risk levels under strict/balanced/permissive tolerances.Evaluate decisions: Choose among provided options, inject
UNRESOLVED_OR_OTHERfor out-of-domain ambiguity, and return confidence andCONFIDENT/AMBIGUOUS_STATEstatus.Run full deterministic pipeline: Prune state, inject neutral escapes, compute fingerprints, query local cache, call TypeSafe AI on misses, and calibrate certainty.
Calibrate answers: Detect low confidence (
top_prob < 0.40), flat distributions (dispersion_gap < 0.15), and boundary noul uncertainty near 0.50.Prune state: Clean nulls, empty collections, whitespace, and cyclic references from JSON payloads, with token estimates.
Cache fingerprints: Compute canonical SHA-256 hashes while masking volatile keys (timestamps, trace IDs, request IDs) for 0-token repeat evaluations.
Integrate as MCP server: Works over stdio with Claude Desktop, Cursor, LibreChat, and custom MCP clients using the MCP 2024-11-05 protocol.
Provide zero-dependency operation: Runs on Python stdlib only, with hardened SQLite caching, WAL mode, and graceful fallback to in-memory cache on errors.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@JevGuard MCP ServerAnalyze this AI response for false certainty and ambiguity"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
JevGuard MCP Server
Official Model Context Protocol (MCP) server for JevGuard. Provides a zero-dependency deterministic guardrail, local SQLite cache, and execution gate for AI coding agents and TypeSafe AI System One decision models. Available on PyPI as jevguard-mcp.
This package exposes JevGuard primitives through JSON-RPC 2.0 over standard input/output (stdio), adhering strictly to the MCP 2024-11-05 specification.
30-Second Quickstart
1. Installation
Install the package directly from PyPI:
pip install jevguard-mcpOr run directly without installation:
python -m jevguard_mcp.server2. Client Setup
Claude Desktop
Add to your claude_desktop_config.json:
{
"mcpServers": {
"jevguard": {
"command": "python",
"args": [
"-m",
"jevguard_mcp.server"
],
"env": {
"TYPESAFE_API_KEY": "your_typesafe_api_key_here"
}
}
}
}Configuration file paths:
Windows:
%APPDATA%\Claude\claude_desktop_config.jsonmacOS:
~/Library/Application Support/Claude/claude_desktop_config.jsonLinux:
~/.config/Claude/claude_desktop_config.json
Cursor IDE
Add in Cursor Settings under Features -> MCP Servers -> Add New MCP Server, or configure directly inside .cursor/mcp.json:
{
"mcpServers": {
"jevguard": {
"command": "python",
"args": [
"-m",
"jevguard_mcp.server"
],
"env": {
"TYPESAFE_API_KEY": "your_typesafe_api_key_here"
}
}
}
}LibreChat
Add to your librechat.yaml:
mcpServers:
jevguard:
type: stdio
command: python
args:
- "-m"
- "jevguard_mcp.server"
env:
TYPESAFE_API_KEY: "your_typesafe_api_key_here"Related MCP server: Omni-NLI
Available Tools
All tools use the canonical jevguard_* prefix to guarantee naming consistency across MCP registries. Legacy invocations without the prefix (evaluate_command_safety, verify_code_patch, evaluate_decision) remain fully supported aliases.
Tool Matrix
Tool | Primary Arguments | Target Function | Output Verdict |
|
| Gate shell and terminal commands |
|
|
| Audit diffs for security regressions |
|
|
| Resolve choices with neutral escape |
|
|
| Full RLCD decision pipeline | Typed answers and calibrated probabilities |
|
| Detect tie breaks and low margins |
|
|
| Strip dead keys, format space, break cycles | Sanitized mapping and token estimate |
|
| Mask volatile timestamps and hashes | Canonical SHA-256 fingerprint string |
Atomic Tools for Coding Agents
These high-level tools accept simple primitive arguments (str, bool, list[str]) to prevent LLMs from hallucinating complex nested question schemas.
1. jevguard_evaluate_command_safety (alias: evaluate_command_safety)
Evaluates whether a terminal command is destructive, requires human approval, or can execute autonomously:
Arguments:
command: str(required): Shell command to evaluate.working_dir: str = ""(optional): Target execution directory.elevated_privileges: bool = false(optional): Whether the command runs withsudoor administrator rights.
Pipeline: Evaluates boundary destruction probability (Noul), blast radius (Score), and policy recommendation (Choice) with certainty calibration.
Output: Returns an execution policy:
ALLOW_AUTONOMOUS,REQUIRE_HUMAN_APPROVAL, orDENY_DESTRUCTIVE.
2. jevguard_verify_code_patch (alias: verify_code_patch)
Verifies unified git diffs or code patches for regressions, broken syntax, or critical system impact:
Arguments:
patch_content: str(required): Unified diff or patch text.target_file: str(required): Target file path.risk_tolerance: str = "balanced"(optional): Risk threshold ("strict","balanced","permissive").
Pipeline: Calibrates regression probability and risk score against the configured risk tolerance threshold.
Output: Returns
approved(boolean),recommendation("APPROVE","REQUEST_CHANGES","REJECT"), andrisk_level("LOW","MEDIUM","HIGH","CRITICAL").
3. jevguard_evaluate_decision (alias: evaluate_decision)
Allows coding agents to resolve architectural or technical choices with a flat options list:
Arguments:
context: str(required): Background context and requirements.decision_question: str(required): Core decision question.options: list[str](required): Candidate options (for example,["PostgreSQL", "SQLite", "DuckDB"]).
Pipeline: Injects closed-world neutral escape (
UNRESOLVED_OR_OTHER) to catch out-of-distribution choices and calibrates probability dispersion.Domain options vs epistemic escape: If a caller provides an option named
Other(such as["PostgreSQL", "MySQL", "Other"]), selecting that option is treated as an intentional domain selection (is_escape_selected: False). In parallel, JevGuard MCP injects the canonicalUNRESOLVED_OR_OTHERalternative to catch genuine epistemic uncertainty, out-of-distribution prompts, and ambiguous choices without conflating them with caller options.Output: Returns
selected_option,confidence,is_escape_selected, andstatus(CONFIDENTorAMBIGUOUS_STATE).
Core JevGuard Primitives
4. jevguard_evaluate
Executes the full deterministic JevGuard evaluation pipeline:
Prunes incoming state data to eliminate empty keys and duplicate whitespace.
Normalizes question schemas and injects closed-world escape alternatives (
UNRESOLVED_OR_OTHER) to prevent false positives.Computes canonical SHA-256 fingerprints with volatile key masking.
Queries the zero-token cache on hit or dispatches upstream to TypeSafe AI when credentials are configured.
Calibrates response certainty and dispersion metrics.
5. jevguard_calibrate
Analyzes response probability distributions to prevent false certainty:
Flags low confidence when top probability falls below 0.40 (
top_prob < 0.40).Flags flat distributions when the gap between top and runner-up choices is below 0.15 (
dispersion_gap < 0.15).Evaluates boundary uncertainty for continuous noul probability ranges near 0.50 (
|prob - 0.50| < 0.12).Returns structured verdicts:
AMBIGUOUS_STATEorCONFIDENT.
6. jevguard_prune_state
Sanitizes structured input states:
Removes null values and empty strings or collections from mapping objects.
Normalizes and collapses repeated whitespace.
Detects circular references and replaces them with
<cyclic_ref>tokens.Calculates an input token count estimate.
7. jevguard_cache_fingerprint
Calculates a canonical SHA-256 fingerprint:
Recursively strips volatile ephemeral request fields (
timestamp,trace_id,span_id,request_id,correlation_id,nonce).Orders dictionary keys deterministically.
Produces identical hashes for semantically identical states regardless of key ordering or ephemeral trace variance.
Preserves domain date and time attributes (
created_at,updated_at) by default to prevent version collisions.
Live Verification Benchmark (5 Direct Calls vs 5 JevGuard MCP Calls)
A live comparison was conducted directly against the official TypeSafe AI endpoint (https://api.typesafe.ai/v1/systemone, model jev-latest) comparing 5 direct API calls against 5 JevGuard MCP tool calls from a development workstation.

Benchmark Summary
Scenario | Input Query Context | Direct API Latency | JevGuard Cache Latency | Decision / Guardrail Effect |
1. Incident Triage | Production latency spike | 741 ms | 0.099 ms (warm cache) | Ambiguity flagged on boundary severity |
2. Security Audit | Root command with path manipulation | 732 ms | 0.112 ms (warm cache) | Intercepted as |
3. Out-of-Domain Query | Corporate tax in Zurich | 749 ms | 0.098 ms (warm cache) | Escaped via |
4. Schema Modification | Malformed payload with timestamps | 728 ms | 0.105 ms (warm cache) | Ephemeral keys masked, cache matched |
5. Repeat Verification | Identical state with fresh trace ID | 735 ms | 0.095 ms (warm cache) | Local hit, 0 tokens billed upstream |
Empirical Findings
Local Cache Retrieval (0.099 ms): Repeated queries containing dynamic timestamps and trace IDs are intercepted locally. Volatile key masking matches the canonical SHA-256 fingerprint, avoiding WAN network roundtrips (~740 ms) and billing 0 tokens on cache hits.
Closed-World Trap Mitigation: In Scenario 3 (an off-topic inquiry about corporate tax offices in Zurich), the unguided model forced an incorrect classification (
credit_card_chargeback). JevGuard MCP injectedUNRESOLVED_OR_OTHER, routing the off-topic input to the neutral escape option.Ambiguity Calibration: In Scenario 1, boundary uncertainty on
is_outage(noul=0.49, distance 0.01 to threshold) and flat distribution onseverity(0.08 gap) were flagged asAMBIGUOUS_STATEusing default operational heuristics.Standard Library Overhead: Local middleware execution latency remained below 0.3 ms for cold requests and 0.099 ms for warm cache lookups.
Architectural Principles
Zero External Dependencies: Implemented strictly with the Python standard library (
sys,json,sqlite3,hashlib,urllib).Protocol Fidelity: Full compliance with the MCP 2024-11-05 standard, supporting initialize handshakes, ping, tool discovery, and tool execution.
Deterministic Local Layer: Canonical state sanitization, neutral escape injection, probability dispersion analysis, and SHA-256 fingerprint caching in SQLite.
Process Isolation: Runs as an independent stdio subprocess compatible with Claude Desktop, Cursor IDE, LibreChat, and custom MCP clients.
Robustness and Fault Tolerance
Hardened SQLite Concurrency:
Connections use
timeout=60.0andPRAGMA busy_timeout = 60000;to preventdatabase is lockedcontention under parallel agent execution.Operates with
PRAGMA journal_mode=WAL;andPRAGMA synchronous=NORMAL;for non-blocking concurrent reads and writes.Any unrecoverable lock, filesystem, or permission error transparently degrades to shared
:memory:without crashing or aborting execution.
Structured Exception Handling and Protocol Stability:
All tool executions are wrapped in defensive error handlers.
Failures (HTTP errors, timeouts, network interruptions, validation errors) return actionable JSON text payloads with
"fallback_action": "MANUAL_REVIEW_REQUIRED".Tool failures return actionable structured JSON error payloads with standard MCP
isError: true, while keeping the stdio transport cleanly connected so client environments (Cursor, Claude Desktop, Antigravity) never crash or drop sessions.
Third-Party Data Transmission Disclosure:
Live evaluations (cache misses or
bypass_cache=True) transmit the evaluatedcommand,patch_content, orstatepayload over encrypted HTTPS directly to the official TypeSafe AI endpoint (api.typesafe.ai).Ephemeral headers and keys are never forwarded across redirect chains (
NoRedirectHandlerblocks 301/302/303 redirect leakage).When deterministic cache hits occur, zero tokens are consumed and zero bytes leave the local host.
Running the Test Suite
Run the unit tests with Python standard unittest runner:
python -m unittest test_mcp_server.py -vAll 134 test cases execute in under 0.6 seconds with zero network dependencies.
Project Status and Validation Transparency
JevGuard MCP is an independent open-source runtime (v1.1.0) built solely with the Python standard library.
Key engineering notes:
The local server (protocol serialization, SQLite caching, state pruning, and calibration checks) is deterministic, while upstream evaluations from TypeSafe AI / Jev are probabilistic.
Default calibration thresholds (such as top probability below 0.40, margin below 0.15) represent operational heuristics for tie and uncertainty detection rather than parameters fitted on a specific domain corpus.
We welcome community peer review, external testing, and issue reports.
License
MIT License. Copyright (c) 2026 Seb4Ez.
Available Tools
7 toolsevaluate_command_safetyB
Evaluates terminal/shell command safety for autonomous agents. Determines whether a command is destructive, requires human approval, or can execute autonomously using Noul, Score, and Choice certainty calibration. Returns policy: ALLOW_AUTONOMOUS, REQUIRE_HUMAN_APPROVAL, or DENY_DESTRUCTIVE.
| Name | Required | Description | Default |
|---|---|---|---|
| command | Yes | Terminal command string to evaluate. | |
| timeout | No | HTTP request timeout in seconds (default: 30.0). | |
| working_dir | No | Working directory for execution context (optional). | |
| bypass_cache | No | Bypass deterministic cache lookup. | |
| elevated_privileges | No | Whether execution uses sudo or administrative privileges (optional, default: false). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. It states the tool 'evaluates' commands and returns a policy, but it does not disclose whether the tool itself executes commands, makes external network calls, or has side effects. It also does not mention authentication requirements, rate limits, or any destructive potential of the tool itself. This is a significant gap for a safety-related tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three sentences long and front-loaded with the core purpose. Each sentence adds value: purpose, evaluation method, and output policy. There is no redundant or filler content, and it is appropriately concise for the tool's complexity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description lists the three possible policy outcomes, which partially covers the return value, but there is no output schema to elaborate on additional fields like confidence scores or reasons. It does not explain the meaning of Noul, Score, and Choice, nor does it describe error behavior or edge cases (e.g., empty command). For a tool with five parameters and no output schema, the description is adequate but not fully complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% (all five parameters have descriptions in the input schema). The description adds no additional parameter-specific meaning beyond what the schema already provides. It mentions the calibration method (Noul, Score, Choice) but does not map these to parameters. Since the schema covers all parameters, a baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool evaluates terminal/shell command safety and returns a policy decision. The verb 'evaluates' with the resource 'terminal/shell command safety' is specific, and the three possible policy outcomes (ALLOW_AUTONOMOUS, REQUIRE_HUMAN_APPROVAL, DENY_DESTRUCTIVE) clarify its purpose. However, it does not explicitly differentiate itself from sibling tools like evaluate_decision or jevguard_evaluate, which may have overlapping evaluation purposes.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides context that this is for autonomous agents, but it does not state when to use this tool versus alternatives or when not to use it. There are no explicit exclusions or alternative recommendations. The usage is implied from the purpose rather than stated, which aligns with a score of 3.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
evaluate_decisionA
Evaluates architectural and implementation decisions with a simple list of options. Automatically injects closed-world neutral escape (UNRESOLVED_OR_OTHER) and calibrates probability dispersion.
| Name | Required | Description | Default |
|---|---|---|---|
| context | Yes | Context and constraints surrounding the decision. | |
| options | Yes | Candidate options list (e.g. ['A', 'B', 'C']). | |
| timeout | No | HTTP request timeout in seconds (default: 30.0). | |
| bypass_cache | No | Bypass deterministic cache lookup. | |
| decision_question | Yes | The specific question or decision to evaluate. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden and meaningfully discloses hidden behavior: automatic injection of UNRESOLVED_OR_OTHER and calibration of probability dispersion. This goes beyond what the schema reveals, though it does not describe result format or side effects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two concise sentences with no filler. The purpose is front-loaded and the behavioral details are presented efficiently, making every sentence earn its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers the core purpose and automatic behaviors, but with no output schema and no usage guidance, an agent still lacks clarity on what the tool returns or when to prefer it over related tools. It is adequate but not complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description does not add parameter-specific details beyond the schema; it only reinforces the 'simple list of options' aspect, which is already covered by the options parameter description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states a specific verb ('Evaluates') and resource ('architectural and implementation decisions'), and clarifies the input style ('simple list of options'). It does not explicitly differentiate from sibling tools like evaluate_command_safety or jevguard_evaluate, so it falls short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to choose this tool over its siblings, no exclusions, and no alternative tool mentions. 'Simple list of options' only implies a use condition; it does not establish when this tool is appropriate.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
jevguard_cache_fingerprintA
Calculates a canonical SHA-256 fingerprint from state and questions with volatile key masking (timestamp, trace_id, request_id) for 0-token deterministic caching.
| Name | Required | Description | Default |
|---|---|---|---|
| model | No | Target model identifier (default: jev-latest). | jev-latest |
| state | Yes | State payload dictionary to include in fingerprint computation. | |
| questions | No | Question definitions dictionary (optional). | |
| ignore_keys | No | List of volatile keys to mask in addition to standard defaults. | |
| auto_inject_escapes | No | Whether to consider escape injection logic when computing fingerprint (default: true). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
There are no annotations, so the description carries the full behavioral disclosure burden. It transparently explains the computation, masking of volatile keys, and deterministic caching intent, but it does not state the return format, whether the model parameter participates in the fingerprint, or any error/side-effect behavior. For a pure calculation tool this is acceptable but incomplete.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single dense sentence that front-loads the core action and purpose. Every phrase contributes meaning: the algorithm, the inputs, the masking behavior, and the intended caching benefit. No filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the moderate complexity of five parameters and nested objects, the description gives enough context to understand why the tool exists and roughly how it behaves. Since there is no output schema, a note about the exact return format would improve completeness, but 'Calculates a canonical SHA-256 fingerprint' sufficiently implies the output.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, giving a baseline of 3. The description adds meaning by explaining that ignore_keys masks volatile keys such as timestamp, trace_id, and request_id, and by framing the overall purpose of the parameters. It does not mention the model parameter's role in the fingerprint, but still adds value beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb ('Calculates') and resource ('canonical SHA-256 fingerprint from state and questions'), and adds distinguishing details like volatile key masking. It is clearly distinct from the evaluation/safety sibling tools, so an agent can tell what this tool is for without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'for 0-token deterministic caching' gives a clear intended use context, which helps an agent decide when to invoke this tool. However, it does not explicitly mention alternatives, when not to use it, or how it relates to sibling evaluation tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
jevguard_calibrateA
Evaluates probability distributions across answers to identify ambiguity, low confidence (top_prob < 0.40), and flat distributions (dispersion_gap < 0.15).
| Name | Required | Description | Default |
|---|---|---|---|
| answers | Yes | Dictionary mapping question names to answers with probabilities or confidence values. | |
| min_top_prob | No | Minimum confidence threshold for top choice (default: 0.40). | |
| min_dispersion_gap | No | Minimum probability gap between top choice and runner up (default: 0.15). | |
| noul_uncertainty_margin | No | Uncertainty margin around 0.50 boundary for noul probabilities (default: 0.12). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It reveals the internal decision logic (thresholds for top_prob and dispersion_gap) and implies a non-mutating evaluation, but it does not state whether the tool returns a report, modifies state, or requires specific permissions. This is a meaningful but not fatal gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, front-loaded sentence with zero filler. Every phrase contributes: the verb, the resource, and the two key detection criteria. It is compact and immediately scannable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has four parameters, a nested object, no annotations, and no output schema, so the description should compensate by explaining return values and side effects. It covers the evaluation logic but omits what the tool returns, whether it is read-only, and the role of noul_uncertainty_margin. An agent would not know what to expect from invoking it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3. The description adds value by explicitly linking min_top_prob to 'top_prob < 0.40' and min_dispersion_gap to flat distributions, clarifying the thresholds' roles. It does not mention noul_uncertainty_margin, but the schema already documents that parameter adequately.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly identifies a specific verb ('Evaluates') and resource ('probability distributions across answers'), and specifies the purpose (identify ambiguity, low confidence, flat distributions). However, it does not differentiate from sibling 'jevguard_evaluate', so it stops short of full sibling distinction.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The usage context is implied: use this tool when you need to assess answer distributions for ambiguity or low confidence. There is no explicit when/when-not guidance or mention of alternatives like jevguard_evaluate, so it lacks the explicit routing that a 5 would require.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
jevguard_evaluateA
Executes the deterministic JevGuard evaluation pipeline including state pruning, closed-world escape injection, certainty calibration, and 0-token caching.
| Name | Required | Description | Default |
|---|---|---|---|
| model | No | Target model identifier (default: jev-latest). | jev-latest |
| state | Yes | Input state payload dictionary or structure to evaluate against criteria. | |
| timeout | No | HTTP request timeout in seconds (default: 30.0). | |
| questions | Yes | Dictionary of question definitions mapping question keys to criteria (noul, score, choice). | |
| bypass_cache | No | Bypass deterministic cache lookup. | |
| auto_inject_escapes | No | Automatically inject neutral escape alternatives into categorical choices. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral disclosure burden and does add meaningful traits: deterministic execution, closed-world escape injection, certainty calibration, and 0-token caching. However, it does not disclose side effects, mutation risk, required permissions, or output behavior, leaving significant behavioral context unstated.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single dense sentence that fits a surprising amount of behavioral and scoping information: deterministic pipeline, four internal stages, and caching behavior. There is no filler or repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description gives a solid high-level picture and all parameters are schema-documented, but there is no output schema and no description of what the evaluation returns or how results should be interpreted. For a complex six-parameter tool, that is a meaningful gap, yet the core invocation details are present.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3 without needing extensive parameter explanation. The description adds conceptual context that maps pipeline stages to parameters such as bypass_cache and auto_inject_escapes, but it does not explain parameters directly or exceed what the schema already conveys.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('executes') and resource ('JevGuard evaluation pipeline'), then names four concrete pipeline stages. This clearly distinguishes it from sibling sub-tools such as jevguard_prune_state and jevguard_calibrate.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'pipeline including state pruning, ... certainty calibration, and 0-token caching' implies this is the aggregate evaluation tool, while siblings are individual stages. It gives clear context for selection, though it stops short of explicitly naming alternatives or stating when not to use it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
jevguard_prune_stateB
Sanitizes and prunes complex JSON state payloads by removing nulls, empty collections, collapsing whitespace, and protecting against cyclic references.
| Name | Required | Description | Default |
|---|---|---|---|
| state | Yes | The state payload dictionary or structure to sanitize and prune. | |
| prune_lists | No | Whether to strip empty values and nulls from lists (default: false). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It discloses several behaviors: removing nulls, empty collections, collapsing whitespace, and protecting against cyclic references. However, it doesn't disclose whether the operation mutates the input or returns a new object, whether it's destructive, or any side effects. The 'protecting against cyclic references' is a useful behavioral detail, but the mutation/return behavior is a significant gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence that packs in the core operations and the cyclic reference protection. It's concise and front-loaded with the main purpose. It could be slightly more structured (e.g., separating the pruning operations from the safety feature), but it's efficient and readable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with 2 parameters, 100% schema coverage, and no output schema, the description covers the main purpose and key behaviors. However, it doesn't explain the return value (sanitized state? success indicator?), which is important for an agent to know what to do with the result. It also doesn't clarify whether the input is mutated in place, which is a meaningful gap for a tool that 'prunes' state.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents both parameters. The description adds context about what 'sanitize' means (removing nulls, empty collections, collapsing whitespace) which helps understand the 'state' parameter's purpose. The 'prune_lists' parameter is not explicitly mentioned in the description, but the schema covers it. Baseline 3 is appropriate since the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's function: sanitizing and pruning JSON state payloads, with specific operations (removing nulls, empty collections, collapsing whitespace, protecting against cyclic references). It distinguishes itself from sibling tools like jevguard_cache_fingerprint and evaluate_* tools, which have different purposes. However, it doesn't explicitly name a sibling alternative for comparison, so it doesn't fully differentiate within a family of similar tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use this tool: when you need to sanitize/prune JSON state payloads. It doesn't explicitly state when not to use it or mention alternatives. The sibling tools are clearly different (evaluation, patching, fingerprinting), so the context is somewhat clear, but there's no explicit guidance on when to choose this over another tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
verify_code_patchA
Evaluates whether a code diff or patch introduces security regressions, broken syntax, or critical system impact under a configurable risk tolerance (strict, balanced, permissive).
| Name | Required | Description | Default |
|---|---|---|---|
| timeout | No | HTTP request timeout in seconds (default: 30.0). | |
| target_file | Yes | Path of the target file being modified. | |
| bypass_cache | No | Bypass deterministic cache lookup. | |
| patch_content | Yes | Diff or patch content to verify. | |
| risk_tolerance | No | Risk tolerance threshold for acceptance (strict, balanced, permissive; default: balanced). | balanced |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It does convey that the tool is evaluative and configurable, which implies a non-mutating verification rather than a system change. However, it does not disclose whether the check performs network calls, writes cache state, or returns a simple verdict versus a detailed report, leaving a meaningful transparency gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single sentence front-loads the operative verb, the object under evaluation, the three risk categories, and the configurable risk tolerance. There is no filler, redundancy, or burying of key information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description and full schema are enough for an agent to provide the required inputs and understand the tool's evaluation scope. However, there is no output schema and no statement of return behavior, so it is unclear whether the tool returns a boolean verdict, a list of findings, or a full report; operational details such as side effects are also absent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema already documents all five parameters with 100% coverage, so the description need not compensate for missing schema details. It adds contextual color about the risk categories being assessed and the risk-tolerance modes, but it does not provide per-parameter meaning beyond what the schema already offers.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific action ('Evaluates whether...') and a specific resource ('a code diff or patch'), and enumerates concrete risk categories: security regressions, broken syntax, and critical system impact. This clearly distinguishes it from sibling tools such as evaluate_command_safety, which target command safety rather than patch content.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The intended use is implied: call this when a code diff or patch needs pre-merge verification under a chosen risk tolerance. However, the description never explicitly contrasts this tool with evaluate_command_safety or the other siblings, nor does it state when not to use it, so an agent must infer the selection from the name and schema.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
6 tool updates
v1.0.1- Added
evaluate_command_safety - Added
evaluate_decision - Changed
jevguard_cache_fingerprint5 fields changed- changed
Input schema / properties / questions / descriptionPrevious value: -"Question definitions dictionary or list."New value: +"Question definitions dictionary (optional)." - removed
Input schema / properties / questions / oneOfRemoved value: -[ - { - "type": "object" - }, - { - "type": "array" - } -] - added
Input schema / properties / questions / typeAdded value: +"object" - changed
Input schema / properties / state / descriptionPrevious value: -"State payload to include in fingerprint computation."New value: +"State payload dictionary to include in fingerprint computation." - added
Input schema / properties / state / typeAdded value: +"object"
- Changed
jevguard_evaluate8 fields changed- removed
Input schema / properties / api_keyRemoved value: -{ - "description": "Optional TypeSafe AI API key (defaults to TYPESAFE_API_KEY environment variable).", - "type": "string" -} - removed
Input schema / properties / endpointRemoved value: -{ - "default": "https://api.typesafe.ai/v1/systemone", - "description": "Upstream API endpoint (default: https://api.typesafe.ai/v1/systemone).", - "type": "string" -} - removed
Input schema / properties / mock_answersRemoved value: -{ - "description": "Optional raw answers dictionary for testing or offline execution.", - "type": "object" -} - changed
Input schema / properties / questions / descriptionPrevious value: -"Dictionary or list of question definitions (noul, score, choice)."New value: +"Dictionary of question definitions mapping question keys to criteria (noul, score, choice)." - removed
Input schema / properties / questions / oneOfRemoved value: -[ - { - "type": "object" - }, - { - "type": "array" - } -] - added
Input schema / properties / questions / typeAdded value: +"object" - changed
Input schema / properties / state / descriptionPrevious value: -"Input state payload to evaluate against criteria."New value: +"Input state payload dictionary or structure to evaluate against criteria." - added
Input schema / properties / state / typeAdded value: +"object"
- Changed
jevguard_prune_state2 fields changed- changed
Input schema / properties / state / descriptionPrevious value: -"The state payload (dict, list, or primitive) to sanitize and prune."New value: +"The state payload dictionary or structure to sanitize and prune." - added
Input schema / properties / state / typeAdded value: +"object"
- Added
verify_code_patch
4 tool updates
v1.0.0- First observed
jevguard_cache_fingerprint - First observed
jevguard_calibrate - First observed
jevguard_evaluate - First observed
jevguard_prune_state
TDQS
Scored across 7 tools
Several tools occupy overlapping conceptual space: evaluate_decision, jevguard_calibrate, and jevguard_evaluate all involve evaluating, calibrating, or escaping closed-world options. Command safety and code patch verification are distinct, but an agent could easily misselect among the evaluation-related tools.
Naming is inconsistent: some tools use the jevguard_ prefix with snake_case (jevguard_prune_state, jevguard_calibrate), while others use bare action-style names (evaluate_command_safety, verify_code_patch). This mixed convention makes the tool set feel less predictable.
Seven tools is a reasonable size for a specialized evaluation server. The count is not excessive, though the overlapping evaluation/calibration tools could likely be consolidated without losing much functionality.
The tool set covers the core pipeline stages: state pruning, fingerprinting, safety evaluation, code patch verification, decision evaluation, calibration, and full pipeline execution. Minor gaps exist around direct inspection of cached fingerprints or explicit configuration of the full pipeline, but the core domain is covered.
Maintenance
Related MCP Connectors
Deterministic contextual decision arbitration and action routing for autonomous software. Takes current state, context, or intent plus caller-supplied candidate actions, state transitions, routes, refusals, escalations, tools, or models and returns a deterministic ordered candidate field. Also provides persistent machine representations for memory, retrieval, indexing, and downstream coherence measurement.
Context integrity for AI agents: evaluate freshness, confidence, provenance, and decision readiness.
Deterministic decision layer for autonomous agents: reproducible PROCEED/REVIEW/SKIP verdicts.
Deterministic reasoning stack for AI agents: simulate, decide & compute, plus cross-domain tools.
Related MCP Servers
- AlicenseNot gradedqualityFmaintenanceA Model Context Protocol server providing pre-curated canonical memory, prose/code provenance checking, and benchmark metrics to improve accuracy and reduce costs across AI tools.AGPL 3.0
- AlicenseNot gradedqualityCmaintenanceProvides natural language inference (NLI) capabilities via the Model Context Protocol, allowing AI agents to verify factual consistency and detect contradictions in text.3MIT
- AlicenseBqualityDmaintenanceProvides persistent memory, reasoning engine, agent-to-agent sharing, and immutable audit trail for AI agents via the Model Context Protocol.12MIT

LogicMem MCP Serverofficial
AlicenseBqualityDmaintenanceProvides persistent memory, reasoning, agent-to-agent sharing, and immutable audit trail for AI agents via the Model Context Protocol.121MIT