Skip to main content
Glama

RealityFirst MCP

A small MCP behavior layer for coding agents that should verify reality instead of trusting their own completion claims.

It does not replace your filesystem, terminal, GitHub, browser, deployment, or machine-control tools. It sits beside them and gives the agent a reusable discipline:

Use tools for facts. Use the model for decisions. Reality beats the report.

Why

Coding agents are often good at doing work and surprisingly bad at proving that they actually did it.

A typical failure looks like:

agent: "I wrote compact_summary.json"
disk:  "compact_summary placeholder"

RealityFirst turns completion claims into explicit evidence gates.

  • wrote/created -> read the exact target back

  • deleted -> prove exact absence

  • built -> terminal receipt + artifact check

  • tests passed -> test/terminal receipt

  • deployed/restarted -> fresh runtime probe

  • UNKNOWN_DO_NOT_REPLAY -> reconcile the same lineage before any replay

Related MCP server: Recommend Agentic Trust Layer

Core invariants

REALITY > CLAIM

READ WIDE, LOAD NARROW

FILE EXISTS != ACTIVE

HISTORY != CURRENT

fresh live Reality
> CURRENT law
> latest cutover
> active implementation/tests
> history
> model report

CREATED / WRITTEN / BUILT / DEPLOYED / PASSED / UPDATED / DELETED / SYNCED
=> independent readback / receipt / runtime proof

UNKNOWN SIDE EFFECT
=> NO BLIND REPLAY

REPLACE
=> PROVE -> CUTOVER -> RETIRE OLD

What it exposes

Tools

  • classify_claim — classify a claim and return its evidence gate

  • build_verification_plan — suggest the smallest independent verification plan

  • check_completion_evidence — PASS/FAIL/UNKNOWN for supplied evidence; labels are structurally validated

  • validate_evidence — reject fake hash/stat/readback labels that lack real payload structure

  • resolve_precedence — resolve current-vs-history conflicts

  • check_replay_safety — block blind replay on unresolved side effects

  • compact_evidence — keep decisive evidence without losing source refs

  • record_loss — create a structured anti-repeat/loss-ledger entry

Resources

  • realityfirst://policy/core

  • realityfirst://policy/completion

  • realityfirst://policy/precedence

Prompts

  • reality_first

  • read_wide

  • crosscheck_completion

Install and run

Requires Python 3.10+ and an MCP host.

Using uvx directly from GitHub:

uvx --from git+https://github.com/zone5101/realityfirst-mcp realityfirst-mcp

Or clone it:

git clone https://github.com/zone5101/realityfirst-mcp
cd realityfirst-mcp
uv sync
uv run realityfirst-mcp

The server uses stdio by default.

The current official MCP Python SDK 2.x is used (mcp>=2,<3).

MCP host config

Generic stdio configuration:

{
  "mcpServers": {
    "realityfirst": {
      "command": "uvx",
      "args": [
        "--from",
        "git+https://github.com/zone5101/realityfirst-mcp",
        "realityfirst-mcp"
      ]
    }
  }
}

For Cline, add that server in its MCP server configuration. Other MCP hosts use the same command/args shape even if the surrounding configuration file differs.

Suggested agent instruction

Adding the MCP server is useful; telling the agent when to call it makes it much more effective:

Use RealityFirst whenever you are about to trust or emit a completion/current-state claim.

Before saying created/written/built/deployed/passed/updated/deleted/synced:
1. classify the claim,
2. obtain independent evidence with the real filesystem/terminal/runtime tool,
3. call check_completion_evidence,
4. only claim completion on PASS.

For architecture/current-state questions, resolve conflicts with:
live Reality > CURRENT law > latest cutover > active implementation/tests > history.

Never blindly replay UNKNOWN_DO_NOT_REPLAY or unresolved side effects.

Example: catching a fake completion

# The agent says it wrote a file.
classify_claim("compact_summary.json was written")
# -> claim_type: file_written
# -> requires readback

# If all we have is the agent's own report:
check_completion_evidence(
    claim_type="file_written",
    evidence=[{"type": "model_report", "valid": True}]
)
# -> FAIL

# After the real filesystem tool reads the target:
check_completion_evidence(
    claim_type="file_written",
    evidence=[{"type": "readback", "valid": True, "ref": "file://compact_summary.json"}]
)
# -> PASS

Design boundary

RealityFirst is deliberately not another execution engine.

Agent
├── filesystem / terminal / GitHub / LivingOS / cloud tools
└── RealityFirst MCP
      ├── classify the claim
      ├── define required proof
      ├── resolve precedence
      └── reject unsupported completion

The authoritative side effect still belongs to the tool that owns it.

RealityFirst should never become a second mutation authority.

Development

uv sync --extra dev
uv run pytest -q

Test case zero is the failure that motivated the project: an agent reports that output files were written while disk reality still contains placeholders.

Status

0.1.1 — alpha.

The first release intentionally stays small: claim classification, evidence gates, current-state precedence, replay safety, compaction, and reusable MCP prompts/resources.

License

MIT

Available Tools

8 tools
build_verification_planC

Build a minimal Reality-first verification plan without executing external tools.

ParametersJSON Schema
NameRequiredDescriptionDefault
claimYes
claim_typeYes
available_toolsNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden. It discloses one important trait—no external tool execution—but remains silent on permissions, side effects, idempotency, and other operational characteristics.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The single sentence is front-loaded and contains no wasted words. However, for a tool with three parameters and no annotations, it is arguably too terse to be fully appropriate.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given three parameters, 0% schema coverage, and no annotations, the description is far from complete. It omits usage guidance, parameter semantics, and most behavioral details; only the output schema covers return values.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the description adds no meaning for claim, claim_type, or available_tools. There is no indication of expected format, syntax, or how these parameters affect the plan.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb 'Build' and resource 'verification plan', plus a distinguishing constraint 'without executing external tools'. It does not differentiate from sibling tools like validate_evidence or classify_claim, so it falls short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no when-to-use guidance, no conditions, and no alternatives. Sibling tools exist but are not referenced, leaving the agent to infer context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

check_completion_evidenceC

Check whether a completion/current-state claim has enough independent evidence.

ParametersJSON Schema
NameRequiredDescriptionDefault
evidenceYes
claim_typeYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full disclosure burden. It never explains what counts as 'enough' evidence, whether the check is read-only, what threshold or policy applies, or how the verdict is expressed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler. It is efficient, though the sparseness reflects missing content rather than tight editing.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be described, but with 0% parameter coverage, no annotations, and a non-trivial verification domain, the definition leaves the agent without the semantics needed to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate and it does not: neither the allowed claim_type values nor the shape/required fields of an evidence item are explained. Only the loose notion of 'independent evidence' is conveyed.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (check) and resource (completion/current-state claim) plus the criterion (enough independent evidence). It is clear what the tool evaluates, though it does not distinguish itself from the closely named sibling validate_evidence.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this versus validate_evidence, classify_claim, or build_verification_plan, and no stated prerequisites or exclusions. The agent must infer the routing from the name alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

check_replay_safetyC

Block blind replay when side effects or request lineage are unresolved.

ParametersJSON Schema
NameRequiredDescriptionDefault
stateYes
side_effect_stateNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. It reveals a single gating condition but does not state whether the tool actually performs a block, returns a decision, is read-only or mutating, requires special permissions, or what happens when the condition is not met. The behavioral picture is substantially incomplete.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single front-loaded sentence with no wasted words. However, given the tool's complexity and the completely undocumented parameters, it is too sparse to be appropriately sized for the task.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained. But with no annotations, 0% parameter description coverage, and two undocumented parameters, the one-sentence description leaves critical context missing for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, and the description does not mention either parameter. The meaning of 'state' and 'side_effect_state' is left entirely unexplained, so the description fails to compensate for the missing schema documentation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a conditional action (block blind replay) tied to a specific resource and conditions (side effects or request lineage unresolved). However, the tool name suggests a safety check rather than an enforcement action, and no sibling tool is mentioned or contrasted. The purpose is somewhat inferable but not sharply distinguished from siblings like validate_evidence or resolve_precedence.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies a trigger condition (when side effects or request lineage are unresolved), but it gives no explicit when-to-use guidance, no when-not-to-use guidance, and no alternatives among the sibling tools. An agent must infer the context entirely from the name and vague condition.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

classify_claimC

Classify an agent claim and return the evidence gate it must satisfy.

ParametersJSON Schema
NameRequiredDescriptionDefault
claimYes
claim_typeNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It does not disclose whether classification is deterministic, what categories exist, whether it mutates state, or what 'evidence gate' means operationally — a significant gap for a tool whose output apparently gates downstream behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no waste. It is efficient, though brevity here shades into under-specification rather than deliberate concision.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be described. But for a tool with zero parameter documentation, no annotations, and no usage routing, the description leaves an agent unable to construct a valid call with confidence or know when to make one.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and neither parameter has any schema-level description. The description does not explain what 'claim' should contain or what the optional 'claim_type' override does, so the meaning of both parameters is undocumented everywhere.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('classify') and resource ('agent claim') and hints at an output ('evidence gate'), which separates it somewhat from siblings like validate_evidence or resolve_precedence. However, 'evidence gate' and 'agent claim' are undefined domain terms, so an agent cannot tell precisely what classification taxonomy or output is produced.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to call this versus validate_evidence, check_completion_evidence, or the other six siblings. The condition under which classification precedes or replaces evidence validation is left entirely to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

compact_evidenceC

Compact evidence while retaining source refs and prioritizing direct Reality proof.

ParametersJSON Schema
NameRequiredDescriptionDefault
evidenceYes
max_itemsNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden. Compaction implies discarding material, yet it never states whether evidence is mutated, dropped, or merely summarized, nor whether the operation is reversible or permission-gated. The one behavioral claim made ('retaining source refs') is useful but far from complete.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no padding; the retained-refs constraint is stated before the prioritization rule. Only the unexplained 'Reality proof' jargon slightly undercuts clarity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return values need not be described, but for a mutation-flavored tool with zero annotations and 0% parameter coverage the definition is too thin to call safely. It omits the very behavioral facts an agent needs before compacting evidence.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description is expected to compensate. It never explains the shape of the `evidence` array items or what `max_items` truncation does to the retained refs, leaving both parameters undocumented in schema and prose alike.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

Names a specific verb+resource ('compact evidence') and adds two qualifiers (retain source refs, prioritize direct Reality proof). However 'compaction' semantics and the undefined jargon term 'Reality proof' leave the exact transform ambiguous, and it does not distinguish itself from siblings like validate_evidence or resolve_precedence.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no statement of when to use this tool, when not to, or which sibling to prefer. In a family of seven evidence-related tools, the agent gets no routing signal at all.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

record_lossB

Create a structured loss-ledger entry. This tool does not persist it.

ParametersJSON Schema
NameRequiredDescriptionDefault
claimYes
why_wrongYes
evidence_refsYes
corrected_formYes
prevention_ruleYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and it does deliver one critical trait: 'This tool does not persist it' clarifies this is a non-persisting/dry-run style operation, which matters for a tool named 'record_loss'. However, it omits permissions, idempotency, and side effects beyond the persistence caveat.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, front-loaded with the core action and immediately followed by the key caveat. No filler; every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with 5 required, undocumented parameters and no annotations, the description is too thin; it leaves parameter meaning and usage context entirely unaddressed. The output schema does relieve it of explaining return values, and the non-persistence caveat is a valuable addition, but the parameter gap dominates.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and there are 5 required parameters, so the description must compensate, but it adds no meaning for claim, why_wrong, corrected_form, evidence_refs, or prevention_rule. Only the phrase 'loss-ledger entry' gives vague thematic framing; an agent still lacks syntactic/format guidance for every required field.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Create a structured loss-ledger entry') that an agent can distinguish from siblings like classify_claim or validate_evidence. However, it offers no explicit differentiation from the other ledger/evidence-building tools in the sibling set.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool versus alternatives; the sibling set (validate_evidence, classify_claim, build_verification_plan) is never referenced. The only conditional information is the non-persistence note, which is behavioral rather than usage-directional.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

resolve_precedenceC

Resolve conflicting claims using Reality > CURRENT law > cutover > implementation/tests > history.

ParametersJSON Schema
NameRequiredDescriptionDefault
candidatesYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It discloses the ranking rule but not whether the tool mutates state, requires permissions, what happens to losing claims, or any side effects of 'resolving' a conflict.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no wasted words. Its terseness is efficient, though it borders on under-specification rather than genuine density.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return values need not be explained, but for a resolution tool the description omits what 'resolving' produces for the input candidates, how ties outside the ladder are handled, and any operational constraints.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% for the single 'candidates' parameter, an array of untyped objects. The description never mentions candidates or the shape each candidate must take, so it fails to compensate for the schema gap beyond the loose mapping to 'conflicting claims'.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a verb ('Resolve') and a resource ('conflicting claims'), and the precedence ladder hints at the decision logic. However, the terms 'Reality', 'CURRENT law', and 'cutover' are opaque without domain context, and nothing distinguishes this from siblings like classify_claim or validate_evidence.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The precedence order implies how conflicts are ranked but says nothing about when an agent should call this tool versus validate_evidence or check_replay_safety, and no prerequisites or exclusions are stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

validate_evidenceC

Validate evidence payload structure before it can satisfy a completion gate.

ParametersJSON Schema
NameRequiredDescriptionDefault
evidenceYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It does not state whether validation is read-only, what happens on failure (throw vs. structured result), whether it mutates or normalizes the payload, or any side effects — leaving the safety profile entirely implicit.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single tight sentence with the verb and resource front-loaded and no filler. It is efficient, though its brevity is partly a symptom of under-specification rather than disciplined concision.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be described. But for a validation tool with no annotations and an undocumented parameter, the description leaves out failure behavior and the relationship to the completion-gate workflow it references.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% for the single "evidence" parameter, so the description must compensate — but "payload structure" only restates the parameter name. It gives no information about the expected shape, required item fields, or valid values of the evidence array.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ("Validate") and resource ("evidence payload structure"), so the action is unambiguous. However, it does not distinguish itself from the sibling check_completion_evidence, which sounds like an overlapping concern, so the agent must guess which one to reach for.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The clause "before it can satisfy a completion gate" hints at sequencing but names no alternative tool or condition for choosing this over check_completion_evidence, compact_evidence, or build_verification_plan. No when-not guidance is given.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 8 tool updatesv0.1.1
    • First observedbuild_verification_plan
    • First observedcheck_completion_evidence
    • First observedcheck_replay_safety
    • First observedclassify_claim
    • First observedcompact_evidence
    • First observedrecord_loss
    • First observedresolve_precedence
    • First observedvalidate_evidence

TDQS

B3/5.0

Scored across 8 tools

Disambiguation4/5

Most tools have clearly distinct purposes (classify vs resolve vs compact). However, validate_evidence, check_completion_evidence, and check_replay_safety all perform gate-style checks on evidence, which could cause confusion without careful reading of descriptions.

Naming Consistency5/5

All tool names follow a consistent snake_case verb_noun pattern (validate_evidence, resolve_precedence, check_replay_safety, etc.). There are no mixed conventions or vague verbs.

Tool Count5/5

With 8 tools, the set is well-scoped for a verification/evidence-gate server. Each tool addresses a distinct step in the workflow, and there is no excessive proliferation.

Completeness3/5

The surface covers validation, classification, conflict resolution, planning, and loss recording, but record_loss explicitly does not persist entries, leaving no way to actually store or retrieve loss data. There is also no tool to execute verification or persist evidence, creating dead ends for some workflows.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    D
    maintenance
    Prevents autonomous agents from fabricating tool results by verifying HMAC-signed receipts against epistemic claim types, forcing re-grounding or escalation before actions commit.
    MIT
  • A
    license
    Not graded
    quality
    A
    maintenance
    Enables agents to classify every claim into evidence tiers (FACT/INFERENCE/SPECULATION/UNVERIFIED) with attached evidence, forcing deterministic, evidence-gated reporting without LLM or network calls.
    18 PyPI
    MIT
  • F
    license
    A
    quality
    C
    maintenance
    Enables agents to verify claims against a real falsification ledger, querying pre-registered refutation thresholds, recorded verdicts, and verifier drift metrics instead of guessing.
    5
    19 PyPI
    -