Skip to main content
Glama
Renwang-Huang

TypeSafe MCP

Arbitype

What is Arbitype?

Arbitype is an MCP-native typed decision layer for AI agents, powered by TypeSafe Jev. It turns probabilistic judgments into structured decision primitives that an agent or program can consume directly.

Arbitype is available on PyPI and the official MCP Registry.

NOTE

Arbitype is an independent open-source project. It is not an official TypeSafe AI product or an official integration for any particular agent host.

Related MCP server: daf-jev

Quick Start

The fastest way to connect an MCP host is a local STDIO server launched by uvx:

export TYPESAFE_API_KEY="your-key"
uvx arbitype

For a pinned, reproducible launch:

uvx --from 'arbitype==0.7.0' arbitype

Or install the package into the current environment:

python -m pip install arbitype
arbitype

The API key stays in the process environment. It is not an MCP argument and is never printed to standard output.

Host setup

Inspect detected hosts before writing anything:

arbitype setup --detect --dry-run

Then apply a reviewed plan:

arbitype setup codex
arbitype doctor --json

The setup flow is plan → unified diff → confirmation → backup → apply. Use claude, cursor, or vscode in place of codex. Detailed host formats, secret forwarding, --yes, --remove, ownership, and non-interactive behavior are documented in docs/HOST_SETUP.md.

Configuration

At minimum, set TYPESAFE_API_KEY. To select a model explicitly:

export TYPESAFE_API_KEY="your-key"
export TYPESAFE_MODEL="jev-latest"

See the full configuration reference for endpoints, timeouts, retries, and request/response limits.

Compatibility

arbitype is the canonical distribution, Python package, and CLI. Existing TypeSafe MCP users should migrate to Arbitype; see docs/REGISTRY_MIGRATION.md.

For the safe migration from the historical distribution:

python -m pip uninstall typesafe-mcp
python -m pip install arbitype

Surface

Canonical

Legacy compatibility

PyPI / CLI

arbitype

typesafe-mcp, typesafe-codex-mcp

Python imports

arbitype

typesafe_mcp, typesafe_codex_mcp

MCP tools

route, review

codex_route, codex_review

Why Arbitype?

Arbitype is designed around contracts and bounded decisions rather than free-form prose:

Principle

What it means

Contract-first

Requests and provider responses are validated before they reach the agent.

Safe by default

Credentials stay out of MCP tool arguments and generated host configuration.

Agent-oriented

Purpose-built classify, verify, gate, route, review, and score primitives.

Portable

Zero third-party runtime dependencies and standard MCP STDIO.

Measured

Public tool-selection and decision-stability evaluation assets.

The decision flow is:

free-form state
      ↓
  TypeSafe Jev
      ↓
probabilistic judgment
      ↓
    Arbitype
      ↓
typed decision + probability
      ↓
agent / code branch

Real Use Cases

Start with one of the small, runnable fixtures:

Scenario

Tool

What it demonstrates

Support routing

route

Choose one next action without executing it.

PR verification

verify

Check several claims independently.

Release review

review / gate

Separate holistic review from thresholded checks.

Agent next step

route

Keep the next workflow action bounded.

The example outputs are illustrative fixtures. They are not live model results, accuracy claims, or authorization decisions.

Tools

Arbitype advertises nine read-only, idempotent MCP tools:

Tool

Use when

Result

evaluate

You need a custom typed Jev question set.

Raw typed Jev response.

classify

You need one choice from unordered labels.

Choice and probability distribution.

score

You need one ordered rating or level.

Weighted score and distribution.

check

You need a probability for one bounded criterion.

Noul yes/no probability.

verify

You need several named claims checked independently.

Noul answer per claim.

gate

You need checks and thresholds transformed into a signal.

pass, review, or fail.

route

You need one suggested next action.

One action; no action execution.

review

You need a holistic quality or risk assessment.

Review decision and evidence.

health

You need local diagnostics or an explicit live check.

Configuration and optional provider health.

Tool Selection Guide

Need

Use

Do not substitute

One unordered label

classify

route, which selects an action

One ordered rating

score

classify, which has no order

One bounded proposition

check

verify, which handles multiple claims

Several named claims

verify

review, which assesses a whole object

Checks plus thresholds

gate

review, which is holistic

One next action

route

classify, which returns a category

Whole diff, plan, release, or report

review

verify, which answers claim by claim

Tool descriptions also state USE WHEN and DO NOT USE WHEN boundaries so hosts can select a primitive without hidden prompt conventions.

Probabilities and confidence are model signals, not proof. gate and review are advisory decision transformations, not authorization systems, security boundaries, or approval engines.

Architecture

flowchart LR
    host["AI Agent / MCP Host<br/>Codex · Claude · Cursor · VS Code"]
    arbitype["Arbitype<br/>Typed decision tools"]
    jev["TypeSafe Jev<br/>System One Model"]
    env["TYPESAFE_API_KEY<br/>process environment"]

    host -->|MCP| arbitype
    arbitype -->|validated HTTPS| jev
    jev -->|typed probabilistic judgment| arbitype
    arbitype -->|structured decision| host
    env -. credential .-> arbitype

The public product is Arbitype; TypeSafe Jev is the current provider. Provider configuration intentionally keeps the TYPESAFE_* names because the credential and endpoint belong to TypeSafe.

Benchmark Snapshot

Recorded on the public 120-case Jev-mediated tool-selection evaluation:

Metric

Result

Tool-selection accuracy

96.67%

Invalid-tool rate

0%

Schema-valid rate

100%

This is a Jev-mediated evaluation over Arbitype's advertised MCP tool catalog. It is not a Codex, Claude, or Cursor host benchmark and is not a general model-performance guarantee.

Details and reproducible assets:

Security Boundaries

Arbitype is a local MCP adapter and typed decision layer. It is not:

  • a sandbox for untrusted code;

  • an authorization or identity system;

  • a prompt-injection firewall;

  • a security approval boundary; or

  • an official TypeSafe AI product.

It does not execute actions suggested by route, edit files, run shell commands, or treat model probabilities as proof. Read SECURITY.md before using live credentials.

Documentation

Development

The test suite uses local fakes and does not need an API key:

python3 -m unittest discover -s tests -v
python3 -m compileall -q .
make check
make release-check

The official MCP Python SDK interoperability smoke test is in scripts/official_sdk_smoke.py. A live Jev check is opt-in and paid:

TYPESAFE_API_KEY="your-key" arbitype doctor --live

Normal CI does not call the provider or run paid benchmarks.

Available Tools

9 tools
checkA
Read-onlyIdempotent

USE WHEN you need one bounded yes/no proposition and its probability. DO NOT USE WHEN you need to verify multiple named claims (use verify), combine checks with thresholds (use gate), or assess an entire object or review package (use review).

ParametersJSON Schema
NameRequiredDescriptionDefault
modelNo
stateYes
instructionsYes
true_criteriaNo
false_criteriaNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
typeYes
modelYes
usageYes
answerYes
evaluationYes

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already provide readOnly, openWorld, idempotent, and non-destructive hints. The description adds the key behavioral contract of returning a bounded yes/no proposition with a probability, which is not covered by annotations. It does not explain how probability is derived, but the annotation coverage lowers the burden.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two short, front-loaded sentences with no filler. The use condition is stated first and exclusions follow, making it easy to scan.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has five loosely typed parameters with zero schema descriptions, and the description only covers when to use it, not how to construct inputs. Even with an output schema present, the lack of parameter guidance makes the description incomplete for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the description provides no explanation of state, instructions, true_criteria, false_criteria, or model. With five parameters and zero schema descriptions, the description must compensate but does not address any parameter meaning.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific purpose: returning one bounded yes/no proposition and its probability. It also explicitly distinguishes itself from verify, gate, and review, so an agent can tell it apart from siblings without inspecting schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives an explicit USE WHEN condition and a DO NOT USE WHEN section that names concrete alternatives (verify, gate, review). This leaves little ambiguity about when to select this tool over siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

classifyA
Read-onlyIdempotent

USE WHEN you need exactly one label from a closed, unordered set and its probability distribution. DO NOT USE WHEN labels are ordered levels (use score), options are next actions (use route), or inputs are claims (use verify).

ParametersJSON Schema
NameRequiredDescriptionDefault
modelNo
stateYes
labelsYes
instructionsYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
typeYes
modelYes
usageYes
answerYes
evaluationYes

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and non-destructive behavior, so the safety profile is covered. The description adds key behavioral context by specifying that the output is a probability distribution over a closed, unordered label set, which goes beyond the annotations and clarifies the tool's core semantics.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences, front-loaded with the core use case and followed by explicit exclusions. There is zero redundant text; every phrase earns its place. It is optimally structured for quick agent comprehension.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the presence of an output schema (not shown but indicated) and annotations covering safety, the description covers the essential behavioral contract. It states the purpose, the label-set nature, and the probability output. It does not elaborate on parameter semantics, but that gap is partially mitigated by the schema. The overall picture is sufficient for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description carries full responsibility for explaining parameters. It only implicitly references the 'labels' parameter as a closed set, but does not explain 'state', 'instructions', or the optional 'model'. The schema itself provides some clarity, but the description adds minimal parameter-level meaning.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description explicitly states the tool's purpose: to return exactly one label from a closed, unordered set along with its probability distribution. It also distinguishes itself from sibling tools by specifying what it is not for (ordered levels, next actions, claims). This provides a precise verb-resource pairing and clear differentiation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description includes both positive and negative usage guidance. It states when to use ('USE WHEN you need exactly one label...') and explicitly lists alternatives for cases not suited to this tool (score for ordered levels, route for next actions, verify for claims). This leaves no ambiguity about tool selection.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

evaluateA
Read-onlyIdempotent

USE WHEN you need raw typed Jev answers for one or more bounded noul, choice, or score questions. DO NOT USE WHEN you need prose, code generation, arithmetic, dates, or a specialized primitive such as route or gate.

ParametersJSON Schema
NameRequiredDescriptionDefault
modelNoOptional model id or alias.
stateYesText or structured evidence.
questionsYesMap of question id to a TypeSafe question.

Output Schema

ParametersJSON Schema
NameRequiredDescription
modelYes
usageYes
answersYes

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already establish read-only, idempotent, non-destructive behavior. The description adds that results are 'raw typed Jev answers' and that inputs are 'bounded' questions, giving useful behavioral context beyond the structured annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and front-loaded: the primary use case appears first, followed immediately by explicit exclusions. Every sentence earns its place, and there is no redundant restatement of the schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is complex with nested question schemas and an output schema, but the description covers the main behavioral contract and routing. The undefined term 'Jev' and lack of detail on how state feeds questions are minor gaps that the schema and output schema mostly absorb.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents model, state, and questions thoroughly. The description references question types but does not add parameter-level meaning beyond what the schema provides, which matches the baseline of 3.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states exactly what the tool does: returns raw typed Jev answers for bounded noul, choice, or score questions. It distinguishes itself from specialized siblings by explicitly excluding route/gate-style primitives and prose/code-generation tasks.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit 'USE WHEN' and 'DO NOT USE WHEN' guidance, naming both the intended task class and the excluded alternatives. This makes the selection decision clear without requiring the agent to infer context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

gateA
Read-onlyIdempotent

USE WHEN you need to combine multiple bounded checks and thresholds into pass/review/fail. DO NOT USE WHEN you need claim-by-claim evidence checking (use verify) or whole-object quality/risk assessment (use review). This is not authorization or a security boundary.

ParametersJSON Schema
NameRequiredDescriptionDefault
modelNo
stateYes
checksYes
pass_atNo
review_atNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
typeYes
checksYes
policyYes
decisionYes
evaluationYes

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already provide readOnly, idempotent, openWorld, and non-destructive hints, so the bar for additional disclosure is lower. The description adds meaningful context by defining the output categories (pass/review/fail) and explicitly disclaiming authorization/security semantics. No contradiction with annotations exists.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tightly worded sentences deliver the core purpose, usage boundaries, and alternatives with no filler. The most important guidance is front-loaded with 'USE WHEN,' making it scannable for an agent.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is moderately complex with nested objects, five parameters, and no schema descriptions, and the description covers selection and exclusion well. However, it does not explain how to construct state or checks, which are required parameters, leaving a significant gap for invocation. The output schema and annotations reduce some burden, but the missing parameter semantics prevent full completeness.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, but it only vaguely maps to parameters: 'bounded checks' suggests the checks property and 'thresholds' hints at pass_at/review_at. However, required parameters like state and the complex shape of checks are left unexplained, and the meaning of model is entirely absent. The description does not provide enough parameter-level detail for correct invocation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's function: combining 'multiple bounded checks and thresholds into pass/review/fail.' It also distinguishes itself from siblings by naming verify and review as alternatives for different use cases. The verb and resource are specific and unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly provides 'USE WHEN' and 'DO NOT USE WHEN' guidance, naming the sibling tools verify and review as the correct alternatives. It also adds a clear exclusion: 'This is not authorization or a security boundary,' which prevents misuse. This is exemplary routing guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

healthA
Read-onlyIdempotent

USE WHEN you need local Arbitype configuration or an explicitly requested live provider check. DO NOT USE WHEN you need a business decision; health does not evaluate state and does not make a network request unless live=true is explicit.

ParametersJSON Schema
NameRequiredDescriptionDefault
liveNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
liveYes
typeYes
modelNo
serverYes
statusNo
versionYes
base_urlNo
max_retriesNo
timeout_secondsNo
api_key_configuredYes

TDQS

A5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already provide readOnlyHint, openWorldHint, idempotentHint, and destructiveHint, but the description adds valuable behavioral context beyond them: health does not evaluate state and makes no network request unless live=true is explicit. This is non-obvious, safety-relevant behavior that the annotations alone would not convey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two tightly scoped sentences with zero filler. It front-loads the use case, then the exclusion, and ends with the key behavioral qualifier. Every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple one-parameter, read-only tool with an output schema and strong annotations, the description covers all essential ground: what it does, when to use it, when not to use it, and the parameter's behavioral impact. No critical decision or invocation detail is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description carries the burden of explaining the live parameter. It does this effectively by stating that live=true triggers an explicit network request, while default behavior is local only. This adds meaning beyond the raw boolean schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific purpose: retrieve local Arbitype configuration or perform an explicitly requested live provider check. It also distinguishes itself from evaluation-like sibling tools by asserting it does not evaluate state, so an agent can tell it apart from evaluate, classify, score, and verify.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit USE WHEN and DO NOT USE WHEN guidance, including a clear exclusion for business decisions and a precise condition for when a network request happens (live=true must be explicit). This fully routes the agent to the correct tool or away from it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

reviewA
Read-onlyIdempotent

USE WHEN you need an overall quality or risk review of a whole-object artifact such as a diff, plan, release, or test report against multiple checks. DO NOT USE WHEN you need claim-by-claim verification (use verify), a single proposition (use check), or next-action selection (use route). It returns signals and never edits files.

ParametersJSON Schema
NameRequiredDescriptionDefault
modelNo
stateYes
checksYes
pass_atNo
review_atNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
typeYes
checksYes
policyYes
decisionYes
evaluationYes

TDQS

A4.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, destructiveHint=false, and idempotentHint=true, covering the safety profile. The description adds 'returns signals and never edits files,' which reinforces non-mutation but does not disclose behavioral details like how pass_at/review_at affect the outcome.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences with zero filler; it front-loads the primary use case, then exclusions, then the key behavioral note. Every sentence earns its place and there is no repetition of schema or annotation data.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With nested schemas, an output schema, and safety annotations present, the description covers selection criteria, non-mutation, and the main parameters' roles. It is slightly incomplete on the meaning of threshold parameters, but the defaults and output schema make a valid invocation achievable.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description should compensate for the five parameters. It gives context for state ('whole-object artifact') and checks ('multiple checks'), but says nothing about model, pass_at, or review_at thresholds, leaving those to be inferred from names and defaults.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('review'), identifies the target resource as a 'whole-object artifact such as a diff, plan, release, or test report', and states the mode ('against multiple checks'). It explicitly contrasts with sibling tools verify, check, and route, so an agent can distinguish it without inspecting schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides explicit USE WHEN and DO NOT USE WHEN conditions. It names the alternatives (verify, check, route) for each exclusion case, leaving no ambiguity about when this tool should be selected over siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

routeA
Read-onlyIdempotent

USE WHEN you need exactly one next action from a closed set of actions. DO NOT USE WHEN you need a topic/category label (use classify), an ordered rating (use score), or action execution. This suggests an action; it never executes it.

ParametersJSON Schema
NameRequiredDescriptionDefault
modelNo
stateYes
actionsYes
instructionsNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
typeYes
modelYes
routeYes
usageYes
answerYes
evaluationYes

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false. The description adds the specific behavioral claim that the tool only suggests an action and never executes it, which is useful context beyond the generic read-only hint. It does not contradict annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact, front-loaded with the core use case, and contains no filler. Every sentence contributes: the first defines when to use it, the second defines when not to and names alternatives, and the third reinforces the no-execution behavior.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description fully covers intent and selection among siblings, and output schema exists to explain return values. However, with four input parameters, nested objects, and zero schema descriptions, the lack of guidance on how to construct state, actions, and instructions leaves a meaningful gap for an agent invoking the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate by explaining parameters. It only hints at 'a closed set of actions', which maps loosely to the actions parameter, but it does not explain the meaning or expected use of state, instructions, model, or how action values are interpreted.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific purpose: choosing exactly one next action from a closed set of actions. It also distinguishes itself from siblings by explicitly excluding classify (topic/category label) and score (ordered rating), and clarifies it never executes actions.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It provides explicit USE WHEN and DO NOT USE WHEN guidance, names concrete alternatives (classify, score, action execution), and tells the agent that action routing should use this tool while execution should not. This is strong usage direction.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

scoreA
Read-onlyIdempotent

USE WHEN you need a weighted rating over two or more ordered rubric levels. DO NOT USE WHEN choices are unordered categories (use classify), next actions (use route), or thresholded checks (use gate).

ParametersJSON Schema
NameRequiredDescriptionDefault
modelNo
stateYes
levelsYes
instructionsYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
typeYes
modelYes
usageYes
answerYes
evaluationYes

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, openWorldHint, idempotentHint, and destructiveHint=false, so the safety profile is covered. The description adds evaluation semantics but does not disclose additional behavioral traits such as how ratings are weighted, whether levels must be supplied in order, or what happens with malformed input.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two short sentences with no filler, and the primary usage condition is front-loaded while exclusions follow immediately. Every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description and annotations together explain when the tool is safe to use and how it differs from siblings, but they do not explain how the required parameters should be populated. Since the schema itself provides no descriptions and the output schema is present but not discussed, an agent still lacks key information needed for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, but it only hints at 'levels' being ordered rubric levels. It does not explain the roles of 'state', 'instructions', or 'model', leaving a significant semantic gap for an agent trying to construct a valid call.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description defines the tool as producing a weighted rating over two or more ordered rubric levels, which is a specific verb-plus-resource statement. It also explicitly contrasts it with classify, route, and gate, making the tool's identity clear even though its name is generic.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit USE WHEN and DO NOT USE WHEN guidance, naming concrete alternatives for unordered categories, next actions, and thresholded checks. This fully routes an agent to the correct sibling tool with no inference required.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

verifyA
Read-onlyIdempotent

USE WHEN you need to check multiple named claims against supplied evidence in one request. DO NOT USE WHEN you need one proposition (use check), an overall quality/risk review (use review), or a pass/review/fail threshold (use gate). Results are signals, not proof.

ParametersJSON Schema
NameRequiredDescriptionDefault
modelNo
stateYes
claimsYes
true_criteriaNo
false_criteriaNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
typeYes
answersYes
evaluationYes

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover read-only, idempotent, non-destructive, and open-world behavior. The description adds meaningful context beyond that: results are signals rather than proof, and the tool batches multiple claims into one request. This helps set expectations without contradicting annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two short sentences with no wasted words. The most important routing information is front-loaded, and the exclusion conditions are compactly listed.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description is sufficient for tool selection and gives a clear usage boundary, and the output schema exists to document return values. However, with five parameters, nested objects, and no schema-level descriptions, the lack of parameter-level guidance leaves a meaningful completeness gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate for the five parameters. It adds high-level meaning by implying that 'claims' are named propositions and 'state' is the supplied evidence, but it does not explain true_criteria, false_criteria, or model. Most parameter semantics remain undocumented.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's function: checking multiple named claims against supplied evidence in a single request. It also distinguishes itself from sibling tools by naming check, review, and gate and the conditions that select them, so an agent can disambiguate without opening schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It provides explicit USE WHEN and DO NOT USE WHEN guidance, naming alternatives (check, review, gate) and the exact scenarios for which they are appropriate. This leaves no ambiguity about when to select verify over its siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 6 tool updatesv0.6.0
    • Changedclassify2 fields changed
      • addedInput schema / properties / labels / maxProperties
        Added value: +255
      • addedInput schema / properties / labels / minProperties
        Added value: +1
    • Changedevaluate3 fields changed
      • changedInput schema / properties / questions / additionalProperties / oneOf
        Previous value: -[
        -  {
        -    "additionalProperties": false,
        -    "properties": {
        -      "criteria": {
        -        "additionalProperties": false,
        -        "properties": {
        -          "false": {
        -            "anyOf": [
        -              {
        -                "type": "string"
        -              },
        -              {
        -                "type": "object"
        -              },
        -              {
        -                "type": "array"
        -              },
        -              {
        -                "type": "null"
        -              }
        -            ]
        -          },
        -          "true": {
        -            "anyOf": [
        -              {
        -                "type": "string"
        -              },
        -              {
        -                "type": "object"
        -              },
        -              {
        -                "type": "array"
        -              },
        -              {
        -                "type": "null"
        -              }
        -            ]
        -          }
        -        },
        -        "type": "object"
        -      },
        -      "instructions": {
        -        "anyOf": [
        -          {
        -            "type": "string"
        -          },
        -          {
        -            "type": "object"
        -          },
        -          {
        -            "type": "array"
        -          },
        -          {
        -            "type": "null"
        -          }
        -        ]
        -      },
        -      "type": {
        -        "const": "noul"
        -      }
        -    },
        -    "required": [
        -      "type",
        -      "instructions"
        -    ],
        -    "type": "object"
        -  },
        -  {
        -    "additionalProperties": false,
        -    "properties": {
        -      "criteria": {
        -        "additionalProperties": {
        -          "anyOf": [
        -            {
        -              "type": "string"
        -            },
        -            {
        -              "type": "object"
        -            },
        -            {
        -              "type": "array"
        -            },
        -            {
        -              "type": "null"
        -            }
        -          ]
        -        },
        -        "maxProperties": 255,
        -        "minProperties": 1,
        -        "type": "object"
        -      },
        -      "instructions": {
        -        "anyOf": [
        -          {
        -            "type": "string"
        -          },
        -          {
        -            "type": "object"
        -          },
        -          {
        -            "type": "array"
        -          },
        -          {
        -            "type": "null"
        -          }
        -        ]
        -      },
        -      "type": {
        -        "const": "choice"
        -      }
        -    },
        -    "required": [
        -      "type",
        -      "instructions",
        -      "criteria"
        -    ],
        -    "type": "object"
        -  },
        -  {
        -    "additionalProperties": false,
        -    "properties": {
        -      "criteria": {
        -        "items": {
        -          "anyOf": [
        -            {
        -              "type": "string"
        -            },
        -            {
        -              "type": "object"
        -            },
        -            {
        -              "type": "array"
        -            },
        -            {
        -              "type": "null"
        -            }
        -          ]
        -        },
        -        "maxItems": 10,
        -        "minItems": 2,
        -        "type": "array"
        -      },
        -      "instructions": {
        -        "anyOf": [
        -          {
        -            "type": "string"
        -          },
        -          {
        -            "type": "object"
        -          },
        -          {
        -            "type": "array"
        -          },
        -          {
        -            "type": "null"
        -          }
        -        ]
        -      },
        -      "type": {
        -        "const": "score"
        -      }
        -    },
        -    "required": [
        -      "type",
        -      "instructions",
        -      "criteria"
        -    ],
        -    "type": "object"
        -  }
        -]New value: +[
        +  {
        +    "additionalProperties": false,
        +    "properties": {
        +      "criteria": {
        +        "additionalProperties": false,
        +        "properties": {
        +          "false": {
        +            "anyOf": [
        +              {
        +                "type": "string"
        +              },
        +              {
        +                "type": "object"
        +              },
        +              {
        +                "type": "array"
        +              },
        +              {
        +                "type": "null"
        +              }
        +            ]
        +          },
        +          "true": {
        +            "anyOf": [
        +              {
        +                "type": "string"
        +              },
        +              {
        +                "type": "object"
        +              },
        +              {
        +                "type": "array"
        +              },
        +              {
        +                "type": "null"
        +              }
        +            ]
        +          }
        +        },
        +        "type": "object"
        +      },
        +      "instructions": {
        +        "anyOf": [
        +          {
        +            "type": "string"
        +          },
        +          {
        +            "type": "object"
        +          },
        +          {
        +            "type": "array"
        +          },
        +          {
        +            "type": "null"
        +          }
        +        ]
        +      },
        +      "type": {
        +        "const": "noul"
        +      }
        +    },
        +    "required": [
        +      "type",
        +      "instructions"
        +    ],
        +    "type": "object"
        +  },
        +  {
        +    "additionalProperties": false,
        +    "properties": {
        +      "criteria": {
        +        "additionalProperties": {
        +          "anyOf": [
        +            {
        +              "type": "string"
        +            },
        +            {
        +              "type": "object"
        +            },
        +            {
        +              "type": "array"
        +            },
        +            {
        +              "type": "null"
        +            }
        +          ]
        +        },
        +        "maxProperties": 255,
        +        "minProperties": 1,
        +        "type": "object"
        +      },
        +      "instructions": {
        +        "anyOf": [
        +          {
        +            "type": "string"
        +          },
        +          {
        +            "type": "object"
        +          },
        +          {
        +            "type": "array"
        +          },
        +          {
        +            "type": "null"
        +          }
        +        ]
        +      },
        +      "type": {
        +        "const": "choice"
        +      }
        +    },
        +    "required": [
        +      "type",
        +      "instructions",
        +      "criteria"
        +    ],
        +    "type": "object"
        +  },
        +  {
        +    "additionalProperties": false,
        +    "properties": {
        +      "criteria": {
        +        "items": {
        +          "anyOf": [
        +            {
        +              "type": "string"
        +            },
        +            {
        +              "type": "object"
        +            },
        +            {
        +              "type": "array"
        +            }
        +          ]
        +        },
        +        "maxItems": 10,
        +        "minItems": 2,
        +        "type": "array"
        +      },
        +      "instructions": {
        +        "anyOf": [
        +          {
        +            "type": "string"
        +          },
        +          {
        +            "type": "object"
        +          },
        +          {
        +            "type": "array"
        +          },
        +          {
        +            "type": "null"
        +          }
        +        ]
        +      },
        +      "type": {
        +        "const": "score"
        +      }
        +    },
        +    "required": [
        +      "type",
        +      "instructions",
        +      "criteria"
        +    ],
        +    "type": "object"
        +  }
        +]
      • addedInput schema / properties / questions / maxProperties
        Added value: +64
      • addedInput schema / properties / questions / minProperties
        Added value: +1
    • Changedgate2 fields changed
      • addedInput schema / properties / checks / maxProperties
        Added value: +64
      • addedInput schema / properties / checks / minProperties
        Added value: +1
    • Changedreview1 field changed
      • addedInput schema / properties / checks / maxProperties
        Added value: +64
    • Changedscore1 field changed
      • changedInput schema / properties / levels / items / anyOf
        Previous value: -[
        -  {
        -    "type": "string"
        -  },
        -  {
        -    "type": "object"
        -  },
        -  {
        -    "type": "array"
        -  },
        -  {
        -    "type": "null"
        -  }
        -]New value: +[
        +  {
        +    "type": "string"
        +  },
        +  {
        +    "type": "object"
        +  },
        +  {
        +    "type": "array"
        +  }
        +]
    • Changedverify2 fields changed
      • addedInput schema / properties / claims / maxProperties
        Added value: +64
      • addedInput schema / properties / claims / minProperties
        Added value: +1
  2. 9 tool updatesv0.5.0
    • First observedcheck
    • First observedclassify
    • First observedevaluate
    • First observedgate
    • First observedhealth
    • First observedreview
    • First observedroute
    • First observedscore
    • First observedverify

TDQS

A4.3/5.0

Scored across 9 tools

Disambiguation5/5

Each tool has a clearly distinct purpose, reinforced by explicit DO NOT USE cross-references that map boundary cases to the correct tool. Even the broadest tool, evaluate, is positioned as the raw-answer fallback rather than a competitor to the specialized distribution, threshold, or routing tools.

Naming Consistency4/5

The tools all follow a consistent lowercase single-word imperative style, which is easy to predict and read. The only minor deviation is health, which is a noun rather than a verb, so the pattern is not perfectly uniform.

Tool Count5/5

Nine tools is a well-scoped count for a typed evaluation and decision primitive library. Each tool represents a distinct decision mode, and none feel redundant, superficial, or missing.

Completeness5/5

The surface covers raw typed evaluation, classification, scoring, whole-object review, single and multi-claim verification, threshold gating, routing, and health checks. For the apparent purpose of safe, typed decision primitives, there are no significant gaps or dead ends.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    A
    maintenance
    Enables AI tools to uniformly discover, inspect, and call tools, prompts, and resources from multiple upstream MCP servers through a small set of fixed MCP tools, over stdio or HTTP.
    2
    MIT
  • A
    license
    Not graded
    quality
    B
    maintenance
    Enables natural-language interaction with TypeSafe's Jev decision API, supporting mixed question calls, batch evaluation, model listing, and confidence or composite-score gates over stdio.
    6
    MIT
  • A
    license
    B
    quality
    C
    maintenance
    Enables AI agents to obtain typed judgments from TypeSafe's Jev System One models, including yes/no probabilities, multiple-choice selections with distributions, and rubric-based scores, directly usable in code.
    5
    2
    AGPL 3.0
  • A
    license
    Not graded
    quality
    C
    maintenance
    Enables MCP-capable agents to run TypeSafe's Jev judgment model as typed yes/no, choice, and score tools, with calibrated probabilities, confidence thresholds, escalation for uncertain or non-judgment tasks, and an optional action gate that fails open.
    1
    MIT