Skip to main content
Glama
seb4ez

JevGuard MCP Server

by seb4ez

JevGuard MCP Server

PyPI version Python 3.10+ Dependencies Protocol License: MIT

Official Model Context Protocol (MCP) server for JevGuard. Provides a zero-dependency deterministic guardrail, local SQLite cache, and execution gate for AI coding agents and TypeSafe AI System One decision models. Available on PyPI as jevguard-mcp.

This package exposes JevGuard primitives through JSON-RPC 2.0 over standard input/output (stdio), adhering strictly to the MCP 2024-11-05 specification.


30-Second Quickstart

1. Installation

Install the package directly from PyPI:

pip install jevguard-mcp

Or run directly without installation:

python -m jevguard_mcp.server

2. Client Setup

Claude Desktop

Add to your claude_desktop_config.json:

{
  "mcpServers": {
    "jevguard": {
      "command": "python",
      "args": [
        "-m",
        "jevguard_mcp.server"
      ],
      "env": {
        "TYPESAFE_API_KEY": "your_typesafe_api_key_here"
      }
    }
  }
}

Configuration file paths:

  • Windows: %APPDATA%\Claude\claude_desktop_config.json

  • macOS: ~/Library/Application Support/Claude/claude_desktop_config.json

  • Linux: ~/.config/Claude/claude_desktop_config.json

Cursor IDE

Add in Cursor Settings under Features -> MCP Servers -> Add New MCP Server, or configure directly inside .cursor/mcp.json:

{
  "mcpServers": {
    "jevguard": {
      "command": "python",
      "args": [
        "-m",
        "jevguard_mcp.server"
      ],
      "env": {
        "TYPESAFE_API_KEY": "your_typesafe_api_key_here"
      }
    }
  }
}

LibreChat

Add to your librechat.yaml:

mcpServers:
  jevguard:
    type: stdio
    command: python
    args:
      - "-m"
      - "jevguard_mcp.server"
    env:
      TYPESAFE_API_KEY: "your_typesafe_api_key_here"

Related MCP server: Omni-NLI

Available Tools

All tools use the canonical jevguard_* prefix to guarantee naming consistency across MCP registries. Legacy invocations without the prefix (evaluate_command_safety, verify_code_patch, evaluate_decision) remain fully supported aliases.

Tool Matrix

Tool

Primary Arguments

Target Function

Output Verdict

jevguard_evaluate_command_safety

command, working_dir, elevated_privileges

Gate shell and terminal commands

ALLOW_AUTONOMOUS, REQUIRE_HUMAN_APPROVAL, DENY_DESTRUCTIVE

jevguard_verify_code_patch

patch_content, target_file, risk_tolerance

Audit diffs for security regressions

APPROVE, REQUEST_CHANGES, REJECT

jevguard_evaluate_decision

context, decision_question, options

Resolve choices with neutral escape

CONFIDENT, AMBIGUOUS_STATE

jevguard_evaluate

state, questions, bypass_cache

Full RLCD decision pipeline

Typed answers and calibrated probabilities

jevguard_calibrate

answers, min_top_prob, min_dispersion_gap

Detect tie breaks and low margins

AMBIGUOUS_STATE, CONFIDENT

jevguard_prune_state

state

Strip dead keys, format space, break cycles

Sanitized mapping and token estimate

jevguard_cache_fingerprint

state, ignore_keys

Mask volatile timestamps and hashes

Canonical SHA-256 fingerprint string


Atomic Tools for Coding Agents

These high-level tools accept simple primitive arguments (str, bool, list[str]) to prevent LLMs from hallucinating complex nested question schemas.

1. jevguard_evaluate_command_safety (alias: evaluate_command_safety)

Evaluates whether a terminal command is destructive, requires human approval, or can execute autonomously:

  • Arguments:

    • command: str (required): Shell command to evaluate.

    • working_dir: str = "" (optional): Target execution directory.

    • elevated_privileges: bool = false (optional): Whether the command runs with sudo or administrator rights.

  • Pipeline: Evaluates boundary destruction probability (Noul), blast radius (Score), and policy recommendation (Choice) with certainty calibration.

  • Output: Returns an execution policy: ALLOW_AUTONOMOUS, REQUIRE_HUMAN_APPROVAL, or DENY_DESTRUCTIVE.

2. jevguard_verify_code_patch (alias: verify_code_patch)

Verifies unified git diffs or code patches for regressions, broken syntax, or critical system impact:

  • Arguments:

    • patch_content: str (required): Unified diff or patch text.

    • target_file: str (required): Target file path.

    • risk_tolerance: str = "balanced" (optional): Risk threshold ("strict", "balanced", "permissive").

  • Pipeline: Calibrates regression probability and risk score against the configured risk tolerance threshold.

  • Output: Returns approved (boolean), recommendation ("APPROVE", "REQUEST_CHANGES", "REJECT"), and risk_level ("LOW", "MEDIUM", "HIGH", "CRITICAL").

3. jevguard_evaluate_decision (alias: evaluate_decision)

Allows coding agents to resolve architectural or technical choices with a flat options list:

  • Arguments:

    • context: str (required): Background context and requirements.

    • decision_question: str (required): Core decision question.

    • options: list[str] (required): Candidate options (for example, ["PostgreSQL", "SQLite", "DuckDB"]).

  • Pipeline: Injects closed-world neutral escape (UNRESOLVED_OR_OTHER) to catch out-of-distribution choices and calibrates probability dispersion.

  • Domain options vs epistemic escape: If a caller provides an option named Other (such as ["PostgreSQL", "MySQL", "Other"]), selecting that option is treated as an intentional domain selection (is_escape_selected: False). In parallel, JevGuard MCP injects the canonical UNRESOLVED_OR_OTHER alternative to catch genuine epistemic uncertainty, out-of-distribution prompts, and ambiguous choices without conflating them with caller options.

  • Output: Returns selected_option, confidence, is_escape_selected, and status (CONFIDENT or AMBIGUOUS_STATE).


Core JevGuard Primitives

4. jevguard_evaluate

Executes the full deterministic JevGuard evaluation pipeline:

  • Prunes incoming state data to eliminate empty keys and duplicate whitespace.

  • Normalizes question schemas and injects closed-world escape alternatives (UNRESOLVED_OR_OTHER) to prevent false positives.

  • Computes canonical SHA-256 fingerprints with volatile key masking.

  • Queries the zero-token cache on hit or dispatches upstream to TypeSafe AI when credentials are configured.

  • Calibrates response certainty and dispersion metrics.

5. jevguard_calibrate

Analyzes response probability distributions to prevent false certainty:

  • Flags low confidence when top probability falls below 0.40 (top_prob < 0.40).

  • Flags flat distributions when the gap between top and runner-up choices is below 0.15 (dispersion_gap < 0.15).

  • Evaluates boundary uncertainty for continuous noul probability ranges near 0.50 (|prob - 0.50| < 0.12).

  • Returns structured verdicts: AMBIGUOUS_STATE or CONFIDENT.

6. jevguard_prune_state

Sanitizes structured input states:

  • Removes null values and empty strings or collections from mapping objects.

  • Normalizes and collapses repeated whitespace.

  • Detects circular references and replaces them with <cyclic_ref> tokens.

  • Calculates an input token count estimate.

7. jevguard_cache_fingerprint

Calculates a canonical SHA-256 fingerprint:

  • Recursively strips volatile ephemeral request fields (timestamp, trace_id, span_id, request_id, correlation_id, nonce).

  • Orders dictionary keys deterministically.

  • Produces identical hashes for semantically identical states regardless of key ordering or ephemeral trace variance.

  • Preserves domain date and time attributes (created_at, updated_at) by default to prevent version collisions.


Live Verification Benchmark (5 Direct Calls vs 5 JevGuard MCP Calls)

A live comparison was conducted directly against the official TypeSafe AI endpoint (https://api.typesafe.ai/v1/systemone, model jev-latest) comparing 5 direct API calls against 5 JevGuard MCP tool calls from a development workstation.

JevGuard MCP Benchmark

Benchmark Summary

Scenario

Input Query Context

Direct API Latency

JevGuard Cache Latency

Decision / Guardrail Effect

1. Incident Triage

Production latency spike

741 ms

0.099 ms (warm cache)

Ambiguity flagged on boundary severity

2. Security Audit

Root command with path manipulation

732 ms

0.112 ms (warm cache)

Intercepted as DENY_DESTRUCTIVE

3. Out-of-Domain Query

Corporate tax in Zurich

749 ms

0.098 ms (warm cache)

Escaped via UNRESOLVED_OR_OTHER

4. Schema Modification

Malformed payload with timestamps

728 ms

0.105 ms (warm cache)

Ephemeral keys masked, cache matched

5. Repeat Verification

Identical state with fresh trace ID

735 ms

0.095 ms (warm cache)

Local hit, 0 tokens billed upstream

Empirical Findings

  1. Local Cache Retrieval (0.099 ms): Repeated queries containing dynamic timestamps and trace IDs are intercepted locally. Volatile key masking matches the canonical SHA-256 fingerprint, avoiding WAN network roundtrips (~740 ms) and billing 0 tokens on cache hits.

  2. Closed-World Trap Mitigation: In Scenario 3 (an off-topic inquiry about corporate tax offices in Zurich), the unguided model forced an incorrect classification (credit_card_chargeback). JevGuard MCP injected UNRESOLVED_OR_OTHER, routing the off-topic input to the neutral escape option.

  3. Ambiguity Calibration: In Scenario 1, boundary uncertainty on is_outage (noul=0.49, distance 0.01 to threshold) and flat distribution on severity (0.08 gap) were flagged as AMBIGUOUS_STATE using default operational heuristics.

  4. Standard Library Overhead: Local middleware execution latency remained below 0.3 ms for cold requests and 0.099 ms for warm cache lookups.


Architectural Principles

  1. Zero External Dependencies: Implemented strictly with the Python standard library (sys, json, sqlite3, hashlib, urllib).

  2. Protocol Fidelity: Full compliance with the MCP 2024-11-05 standard, supporting initialize handshakes, ping, tool discovery, and tool execution.

  3. Deterministic Local Layer: Canonical state sanitization, neutral escape injection, probability dispersion analysis, and SHA-256 fingerprint caching in SQLite.

  4. Process Isolation: Runs as an independent stdio subprocess compatible with Claude Desktop, Cursor IDE, LibreChat, and custom MCP clients.


Robustness and Fault Tolerance

  1. Hardened SQLite Concurrency:

    • Connections use timeout=60.0 and PRAGMA busy_timeout = 60000; to prevent database is locked contention under parallel agent execution.

    • Operates with PRAGMA journal_mode=WAL; and PRAGMA synchronous=NORMAL; for non-blocking concurrent reads and writes.

    • Any unrecoverable lock, filesystem, or permission error transparently degrades to shared :memory: without crashing or aborting execution.

  2. Structured Exception Handling and Protocol Stability:

    • All tool executions are wrapped in defensive error handlers.

    • Failures (HTTP errors, timeouts, network interruptions, validation errors) return actionable JSON text payloads with "fallback_action": "MANUAL_REVIEW_REQUIRED".

    • Tool failures return actionable structured JSON error payloads with standard MCP isError: true, while keeping the stdio transport cleanly connected so client environments (Cursor, Claude Desktop, Antigravity) never crash or drop sessions.

  3. Third-Party Data Transmission Disclosure:

    • Live evaluations (cache misses or bypass_cache=True) transmit the evaluated command, patch_content, or state payload over encrypted HTTPS directly to the official TypeSafe AI endpoint (api.typesafe.ai).

    • Ephemeral headers and keys are never forwarded across redirect chains (NoRedirectHandler blocks 301/302/303 redirect leakage).

    • When deterministic cache hits occur, zero tokens are consumed and zero bytes leave the local host.


Running the Test Suite

Run the unit tests with Python standard unittest runner:

python -m unittest test_mcp_server.py -v

All 134 test cases execute in under 0.6 seconds with zero network dependencies.


Project Status and Validation Transparency

JevGuard MCP is an independent open-source runtime (v1.1.0) built solely with the Python standard library.

Key engineering notes:

  • The local server (protocol serialization, SQLite caching, state pruning, and calibration checks) is deterministic, while upstream evaluations from TypeSafe AI / Jev are probabilistic.

  • Default calibration thresholds (such as top probability below 0.40, margin below 0.15) represent operational heuristics for tie and uncertainty detection rather than parameters fitted on a specific domain corpus.

  • We welcome community peer review, external testing, and issue reports.


License

MIT License. Copyright (c) 2026 Seb4Ez.

Available Tools

7 tools
evaluate_command_safetyB

Evaluates terminal/shell command safety for autonomous agents. Determines whether a command is destructive, requires human approval, or can execute autonomously using Noul, Score, and Choice certainty calibration. Returns policy: ALLOW_AUTONOMOUS, REQUIRE_HUMAN_APPROVAL, or DENY_DESTRUCTIVE.

ParametersJSON Schema
NameRequiredDescriptionDefault
commandYesTerminal command string to evaluate.
timeoutNoHTTP request timeout in seconds (default: 30.0).
working_dirNoWorking directory for execution context (optional).
bypass_cacheNoBypass deterministic cache lookup.
elevated_privilegesNoWhether execution uses sudo or administrative privileges (optional, default: false).

TDQS

B3.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. It states the tool 'evaluates' commands and returns a policy, but it does not disclose whether the tool itself executes commands, makes external network calls, or has side effects. It also does not mention authentication requirements, rate limits, or any destructive potential of the tool itself. This is a significant gap for a safety-related tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is three sentences long and front-loaded with the core purpose. Each sentence adds value: purpose, evaluation method, and output policy. There is no redundant or filler content, and it is appropriately concise for the tool's complexity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description lists the three possible policy outcomes, which partially covers the return value, but there is no output schema to elaborate on additional fields like confidence scores or reasons. It does not explain the meaning of Noul, Score, and Choice, nor does it describe error behavior or edge cases (e.g., empty command). For a tool with five parameters and no output schema, the description is adequate but not fully complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% (all five parameters have descriptions in the input schema). The description adds no additional parameter-specific meaning beyond what the schema already provides. It mentions the calibration method (Noul, Score, Choice) but does not map these to parameters. Since the schema covers all parameters, a baseline of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool evaluates terminal/shell command safety and returns a policy decision. The verb 'evaluates' with the resource 'terminal/shell command safety' is specific, and the three possible policy outcomes (ALLOW_AUTONOMOUS, REQUIRE_HUMAN_APPROVAL, DENY_DESTRUCTIVE) clarify its purpose. However, it does not explicitly differentiate itself from sibling tools like evaluate_decision or jevguard_evaluate, which may have overlapping evaluation purposes.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides context that this is for autonomous agents, but it does not state when to use this tool versus alternatives or when not to use it. There are no explicit exclusions or alternative recommendations. The usage is implied from the purpose rather than stated, which aligns with a score of 3.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

evaluate_decisionA

Evaluates architectural and implementation decisions with a simple list of options. Automatically injects closed-world neutral escape (UNRESOLVED_OR_OTHER) and calibrates probability dispersion.

ParametersJSON Schema
NameRequiredDescriptionDefault
contextYesContext and constraints surrounding the decision.
optionsYesCandidate options list (e.g. ['A', 'B', 'C']).
timeoutNoHTTP request timeout in seconds (default: 30.0).
bypass_cacheNoBypass deterministic cache lookup.
decision_questionYesThe specific question or decision to evaluate.

TDQS

A3.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden and meaningfully discloses hidden behavior: automatic injection of UNRESOLVED_OR_OTHER and calibration of probability dispersion. This goes beyond what the schema reveals, though it does not describe result format or side effects.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two concise sentences with no filler. The purpose is front-loaded and the behavioral details are presented efficiently, making every sentence earn its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers the core purpose and automatic behaviors, but with no output schema and no usage guidance, an agent still lacks clarity on what the tool returns or when to prefer it over related tools. It is adequate but not complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description does not add parameter-specific details beyond the schema; it only reinforces the 'simple list of options' aspect, which is already covered by the options parameter description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states a specific verb ('Evaluates') and resource ('architectural and implementation decisions'), and clarifies the input style ('simple list of options'). It does not explicitly differentiate from sibling tools like evaluate_command_safety or jevguard_evaluate, so it falls short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to choose this tool over its siblings, no exclusions, and no alternative tool mentions. 'Simple list of options' only implies a use condition; it does not establish when this tool is appropriate.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

jevguard_cache_fingerprintA

Calculates a canonical SHA-256 fingerprint from state and questions with volatile key masking (timestamp, trace_id, request_id) for 0-token deterministic caching.

ParametersJSON Schema
NameRequiredDescriptionDefault
modelNoTarget model identifier (default: jev-latest).jev-latest
stateYesState payload dictionary to include in fingerprint computation.
questionsNoQuestion definitions dictionary (optional).
ignore_keysNoList of volatile keys to mask in addition to standard defaults.
auto_inject_escapesNoWhether to consider escape injection logic when computing fingerprint (default: true).

TDQS

A4.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

There are no annotations, so the description carries the full behavioral disclosure burden. It transparently explains the computation, masking of volatile keys, and deterministic caching intent, but it does not state the return format, whether the model parameter participates in the fingerprint, or any error/side-effect behavior. For a pure calculation tool this is acceptable but incomplete.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single dense sentence that front-loads the core action and purpose. Every phrase contributes meaning: the algorithm, the inputs, the masking behavior, and the intended caching benefit. No filler or redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the moderate complexity of five parameters and nested objects, the description gives enough context to understand why the tool exists and roughly how it behaves. Since there is no output schema, a note about the exact return format would improve completeness, but 'Calculates a canonical SHA-256 fingerprint' sufficiently implies the output.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, giving a baseline of 3. The description adds meaning by explaining that ignore_keys masks volatile keys such as timestamp, trace_id, and request_id, and by framing the overall purpose of the parameters. It does not mention the model parameter's role in the fingerprint, but still adds value beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb ('Calculates') and resource ('canonical SHA-256 fingerprint from state and questions'), and adds distinguishing details like volatile key masking. It is clearly distinct from the evaluation/safety sibling tools, so an agent can tell what this tool is for without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'for 0-token deterministic caching' gives a clear intended use context, which helps an agent decide when to invoke this tool. However, it does not explicitly mention alternatives, when not to use it, or how it relates to sibling evaluation tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

jevguard_calibrateA

Evaluates probability distributions across answers to identify ambiguity, low confidence (top_prob < 0.40), and flat distributions (dispersion_gap < 0.15).

ParametersJSON Schema
NameRequiredDescriptionDefault
answersYesDictionary mapping question names to answers with probabilities or confidence values.
min_top_probNoMinimum confidence threshold for top choice (default: 0.40).
min_dispersion_gapNoMinimum probability gap between top choice and runner up (default: 0.15).
noul_uncertainty_marginNoUncertainty margin around 0.50 boundary for noul probabilities (default: 0.12).

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It reveals the internal decision logic (thresholds for top_prob and dispersion_gap) and implies a non-mutating evaluation, but it does not state whether the tool returns a report, modifies state, or requires specific permissions. This is a meaningful but not fatal gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single, front-loaded sentence with zero filler. Every phrase contributes: the verb, the resource, and the two key detection criteria. It is compact and immediately scannable.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has four parameters, a nested object, no annotations, and no output schema, so the description should compensate by explaining return values and side effects. It covers the evaluation logic but omits what the tool returns, whether it is read-only, and the role of noul_uncertainty_margin. An agent would not know what to expect from invoking it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. The description adds value by explicitly linking min_top_prob to 'top_prob < 0.40' and min_dispersion_gap to flat distributions, clarifying the thresholds' roles. It does not mention noul_uncertainty_margin, but the schema already documents that parameter adequately.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly identifies a specific verb ('Evaluates') and resource ('probability distributions across answers'), and specifies the purpose (identify ambiguity, low confidence, flat distributions). However, it does not differentiate from sibling 'jevguard_evaluate', so it stops short of full sibling distinction.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The usage context is implied: use this tool when you need to assess answer distributions for ambiguity or low confidence. There is no explicit when/when-not guidance or mention of alternatives like jevguard_evaluate, so it lacks the explicit routing that a 5 would require.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

jevguard_evaluateA

Executes the deterministic JevGuard evaluation pipeline including state pruning, closed-world escape injection, certainty calibration, and 0-token caching.

ParametersJSON Schema
NameRequiredDescriptionDefault
modelNoTarget model identifier (default: jev-latest).jev-latest
stateYesInput state payload dictionary or structure to evaluate against criteria.
timeoutNoHTTP request timeout in seconds (default: 30.0).
questionsYesDictionary of question definitions mapping question keys to criteria (noul, score, choice).
bypass_cacheNoBypass deterministic cache lookup.
auto_inject_escapesNoAutomatically inject neutral escape alternatives into categorical choices.

TDQS

A3.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the behavioral disclosure burden and does add meaningful traits: deterministic execution, closed-world escape injection, certainty calibration, and 0-token caching. However, it does not disclose side effects, mutation risk, required permissions, or output behavior, leaving significant behavioral context unstated.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single dense sentence that fits a surprising amount of behavioral and scoping information: deterministic pipeline, four internal stages, and caching behavior. There is no filler or repetition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description gives a solid high-level picture and all parameters are schema-documented, but there is no output schema and no description of what the evaluation returns or how results should be interpreted. For a complex six-parameter tool, that is a meaningful gap, yet the core invocation details are present.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3 without needing extensive parameter explanation. The description adds conceptual context that maps pipeline stages to parameters such as bypass_cache and auto_inject_escapes, but it does not explain parameters directly or exceed what the schema already conveys.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('executes') and resource ('JevGuard evaluation pipeline'), then names four concrete pipeline stages. This clearly distinguishes it from sibling sub-tools such as jevguard_prune_state and jevguard_calibrate.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'pipeline including state pruning, ... certainty calibration, and 0-token caching' implies this is the aggregate evaluation tool, while siblings are individual stages. It gives clear context for selection, though it stops short of explicitly naming alternatives or stating when not to use it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

jevguard_prune_stateB

Sanitizes and prunes complex JSON state payloads by removing nulls, empty collections, collapsing whitespace, and protecting against cyclic references.

ParametersJSON Schema
NameRequiredDescriptionDefault
stateYesThe state payload dictionary or structure to sanitize and prune.
prune_listsNoWhether to strip empty values and nulls from lists (default: false).

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It discloses several behaviors: removing nulls, empty collections, collapsing whitespace, and protecting against cyclic references. However, it doesn't disclose whether the operation mutates the input or returns a new object, whether it's destructive, or any side effects. The 'protecting against cyclic references' is a useful behavioral detail, but the mutation/return behavior is a significant gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence that packs in the core operations and the cyclic reference protection. It's concise and front-loaded with the main purpose. It could be slightly more structured (e.g., separating the pruning operations from the safety feature), but it's efficient and readable.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with 2 parameters, 100% schema coverage, and no output schema, the description covers the main purpose and key behaviors. However, it doesn't explain the return value (sanitized state? success indicator?), which is important for an agent to know what to do with the result. It also doesn't clarify whether the input is mutated in place, which is a meaningful gap for a tool that 'prunes' state.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents both parameters. The description adds context about what 'sanitize' means (removing nulls, empty collections, collapsing whitespace) which helps understand the 'state' parameter's purpose. The 'prune_lists' parameter is not explicitly mentioned in the description, but the schema covers it. Baseline 3 is appropriate since the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's function: sanitizing and pruning JSON state payloads, with specific operations (removing nulls, empty collections, collapsing whitespace, protecting against cyclic references). It distinguishes itself from sibling tools like jevguard_cache_fingerprint and evaluate_* tools, which have different purposes. However, it doesn't explicitly name a sibling alternative for comparison, so it doesn't fully differentiate within a family of similar tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies when to use this tool: when you need to sanitize/prune JSON state payloads. It doesn't explicitly state when not to use it or mention alternatives. The sibling tools are clearly different (evaluation, patching, fingerprinting), so the context is somewhat clear, but there's no explicit guidance on when to choose this over another tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

verify_code_patchA

Evaluates whether a code diff or patch introduces security regressions, broken syntax, or critical system impact under a configurable risk tolerance (strict, balanced, permissive).

ParametersJSON Schema
NameRequiredDescriptionDefault
timeoutNoHTTP request timeout in seconds (default: 30.0).
target_fileYesPath of the target file being modified.
bypass_cacheNoBypass deterministic cache lookup.
patch_contentYesDiff or patch content to verify.
risk_toleranceNoRisk tolerance threshold for acceptance (strict, balanced, permissive; default: balanced).balanced

TDQS

A3.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden. It does convey that the tool is evaluative and configurable, which implies a non-mutating verification rather than a system change. However, it does not disclose whether the check performs network calls, writes cache state, or returns a simple verdict versus a detailed report, leaving a meaningful transparency gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single sentence front-loads the operative verb, the object under evaluation, the three risk categories, and the configurable risk tolerance. There is no filler, redundancy, or burying of key information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description and full schema are enough for an agent to provide the required inputs and understand the tool's evaluation scope. However, there is no output schema and no statement of return behavior, so it is unclear whether the tool returns a boolean verdict, a list of findings, or a full report; operational details such as side effects are also absent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema already documents all five parameters with 100% coverage, so the description need not compensate for missing schema details. It adds contextual color about the risk categories being assessed and the risk-tolerance modes, but it does not provide per-parameter meaning beyond what the schema already offers.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific action ('Evaluates whether...') and a specific resource ('a code diff or patch'), and enumerates concrete risk categories: security regressions, broken syntax, and critical system impact. This clearly distinguishes it from sibling tools such as evaluate_command_safety, which target command safety rather than patch content.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The intended use is implied: call this when a code diff or patch needs pre-merge verification under a chosen risk tolerance. However, the description never explicitly contrasts this tool with evaluate_command_safety or the other siblings, nor does it state when not to use it, so an agent must infer the selection from the name and schema.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 6 tool updatesv1.0.1
    • Addedevaluate_command_safety
    • Addedevaluate_decision
    • Changedjevguard_cache_fingerprint5 fields changed
      • changedInput schema / properties / questions / description
        Previous value: -"Question definitions dictionary or list."New value: +"Question definitions dictionary (optional)."
      • removedInput schema / properties / questions / oneOf
        Removed value: -[
        -  {
        -    "type": "object"
        -  },
        -  {
        -    "type": "array"
        -  }
        -]
      • addedInput schema / properties / questions / type
        Added value: +"object"
      • changedInput schema / properties / state / description
        Previous value: -"State payload to include in fingerprint computation."New value: +"State payload dictionary to include in fingerprint computation."
      • addedInput schema / properties / state / type
        Added value: +"object"
    • Changedjevguard_evaluate8 fields changed
      • removedInput schema / properties / api_key
        Removed value: -{
        -  "description": "Optional TypeSafe AI API key (defaults to TYPESAFE_API_KEY environment variable).",
        -  "type": "string"
        -}
      • removedInput schema / properties / endpoint
        Removed value: -{
        -  "default": "https://api.typesafe.ai/v1/systemone",
        -  "description": "Upstream API endpoint (default: https://api.typesafe.ai/v1/systemone).",
        -  "type": "string"
        -}
      • removedInput schema / properties / mock_answers
        Removed value: -{
        -  "description": "Optional raw answers dictionary for testing or offline execution.",
        -  "type": "object"
        -}
      • changedInput schema / properties / questions / description
        Previous value: -"Dictionary or list of question definitions (noul, score, choice)."New value: +"Dictionary of question definitions mapping question keys to criteria (noul, score, choice)."
      • removedInput schema / properties / questions / oneOf
        Removed value: -[
        -  {
        -    "type": "object"
        -  },
        -  {
        -    "type": "array"
        -  }
        -]
      • addedInput schema / properties / questions / type
        Added value: +"object"
      • changedInput schema / properties / state / description
        Previous value: -"Input state payload to evaluate against criteria."New value: +"Input state payload dictionary or structure to evaluate against criteria."
      • addedInput schema / properties / state / type
        Added value: +"object"
    • Changedjevguard_prune_state2 fields changed
      • changedInput schema / properties / state / description
        Previous value: -"The state payload (dict, list, or primitive) to sanitize and prune."New value: +"The state payload dictionary or structure to sanitize and prune."
      • addedInput schema / properties / state / type
        Added value: +"object"
    • Addedverify_code_patch
  2. 4 tool updatesv1.0.0
    • First observedjevguard_cache_fingerprint
    • First observedjevguard_calibrate
    • First observedjevguard_evaluate
    • First observedjevguard_prune_state

TDQS

B3.4/5.0

Scored across 7 tools

Disambiguation2/5

Several tools occupy overlapping conceptual space: evaluate_decision, jevguard_calibrate, and jevguard_evaluate all involve evaluating, calibrating, or escaping closed-world options. Command safety and code patch verification are distinct, but an agent could easily misselect among the evaluation-related tools.

Naming Consistency2/5

Naming is inconsistent: some tools use the jevguard_ prefix with snake_case (jevguard_prune_state, jevguard_calibrate), while others use bare action-style names (evaluate_command_safety, verify_code_patch). This mixed convention makes the tool set feel less predictable.

Tool Count4/5

Seven tools is a reasonable size for a specialized evaluation server. The count is not excessive, though the overlapping evaluation/calibration tools could likely be consolidated without losing much functionality.

Completeness4/5

The tool set covers the core pipeline stages: state pruning, fingerprinting, safety evaluation, code patch verification, decision evaluation, calibration, and full pipeline execution. Minor gaps exist around direct inspection of cached fingerprints or explicit configuration of the full pipeline, but the core domain is covered.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers