RealityFirst MCP
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@RealityFirst MCPI wrote compact_summary.json; verify the claim before I report success."
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
RealityFirst MCP
A small MCP behavior layer for coding agents that should verify reality instead of trusting their own completion claims.
It does not replace your filesystem, terminal, GitHub, browser, deployment, or machine-control tools. It sits beside them and gives the agent a reusable discipline:
Use tools for facts. Use the model for decisions. Reality beats the report.
Why
Coding agents are often good at doing work and surprisingly bad at proving that they actually did it.
A typical failure looks like:
agent: "I wrote compact_summary.json"
disk: "compact_summary placeholder"RealityFirst turns completion claims into explicit evidence gates.
wrote/created-> read the exact target backdeleted-> prove exact absencebuilt-> terminal receipt + artifact checktests passed-> test/terminal receiptdeployed/restarted-> fresh runtime probeUNKNOWN_DO_NOT_REPLAY-> reconcile the same lineage before any replay
Related MCP server: Recommend Agentic Trust Layer
Core invariants
REALITY > CLAIM
READ WIDE, LOAD NARROW
FILE EXISTS != ACTIVE
HISTORY != CURRENT
fresh live Reality
> CURRENT law
> latest cutover
> active implementation/tests
> history
> model report
CREATED / WRITTEN / BUILT / DEPLOYED / PASSED / UPDATED / DELETED / SYNCED
=> independent readback / receipt / runtime proof
UNKNOWN SIDE EFFECT
=> NO BLIND REPLAY
REPLACE
=> PROVE -> CUTOVER -> RETIRE OLDWhat it exposes
Tools
classify_claim— classify a claim and return its evidence gatebuild_verification_plan— suggest the smallest independent verification plancheck_completion_evidence— PASS/FAIL/UNKNOWN for supplied evidence; labels are structurally validatedvalidate_evidence— reject fake hash/stat/readback labels that lack real payload structureresolve_precedence— resolve current-vs-history conflictscheck_replay_safety— block blind replay on unresolved side effectscompact_evidence— keep decisive evidence without losing source refsrecord_loss— create a structured anti-repeat/loss-ledger entry
Resources
realityfirst://policy/corerealityfirst://policy/completionrealityfirst://policy/precedence
Prompts
reality_firstread_widecrosscheck_completion
Install and run
Requires Python 3.10+ and an MCP host.
Using uvx directly from GitHub:
uvx --from git+https://github.com/zone5101/realityfirst-mcp realityfirst-mcpOr clone it:
git clone https://github.com/zone5101/realityfirst-mcp
cd realityfirst-mcp
uv sync
uv run realityfirst-mcpThe server uses stdio by default.
The current official MCP Python SDK 2.x is used (mcp>=2,<3).
MCP host config
Generic stdio configuration:
{
"mcpServers": {
"realityfirst": {
"command": "uvx",
"args": [
"--from",
"git+https://github.com/zone5101/realityfirst-mcp",
"realityfirst-mcp"
]
}
}
}For Cline, add that server in its MCP server configuration. Other MCP hosts use the same command/args shape even if the surrounding configuration file differs.
Suggested agent instruction
Adding the MCP server is useful; telling the agent when to call it makes it much more effective:
Use RealityFirst whenever you are about to trust or emit a completion/current-state claim.
Before saying created/written/built/deployed/passed/updated/deleted/synced:
1. classify the claim,
2. obtain independent evidence with the real filesystem/terminal/runtime tool,
3. call check_completion_evidence,
4. only claim completion on PASS.
For architecture/current-state questions, resolve conflicts with:
live Reality > CURRENT law > latest cutover > active implementation/tests > history.
Never blindly replay UNKNOWN_DO_NOT_REPLAY or unresolved side effects.Example: catching a fake completion
# The agent says it wrote a file.
classify_claim("compact_summary.json was written")
# -> claim_type: file_written
# -> requires readback
# If all we have is the agent's own report:
check_completion_evidence(
claim_type="file_written",
evidence=[{"type": "model_report", "valid": True}]
)
# -> FAIL
# After the real filesystem tool reads the target:
check_completion_evidence(
claim_type="file_written",
evidence=[{"type": "readback", "valid": True, "ref": "file://compact_summary.json"}]
)
# -> PASSDesign boundary
RealityFirst is deliberately not another execution engine.
Agent
├── filesystem / terminal / GitHub / LivingOS / cloud tools
└── RealityFirst MCP
├── classify the claim
├── define required proof
├── resolve precedence
└── reject unsupported completionThe authoritative side effect still belongs to the tool that owns it.
RealityFirst should never become a second mutation authority.
Development
uv sync --extra dev
uv run pytest -qTest case zero is the failure that motivated the project: an agent reports that output files were written while disk reality still contains placeholders.
Status
0.1.1 — alpha.
The first release intentionally stays small: claim classification, evidence gates, current-state precedence, replay safety, compaction, and reusable MCP prompts/resources.
License
MIT
Available Tools
8 toolsbuild_verification_planC
Build a minimal Reality-first verification plan without executing external tools.
| Name | Required | Description | Default |
|---|---|---|---|
| claim | Yes | ||
| claim_type | Yes | ||
| available_tools | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It discloses one important trait—no external tool execution—but remains silent on permissions, side effects, idempotency, and other operational characteristics.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The single sentence is front-loaded and contains no wasted words. However, for a tool with three parameters and no annotations, it is arguably too terse to be fully appropriate.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given three parameters, 0% schema coverage, and no annotations, the description is far from complete. It omits usage guidance, parameter semantics, and most behavioral details; only the output schema covers return values.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the description adds no meaning for claim, claim_type, or available_tools. There is no indication of expected format, syntax, or how these parameters affect the plan.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb 'Build' and resource 'verification plan', plus a distinguishing constraint 'without executing external tools'. It does not differentiate from sibling tools like validate_evidence or classify_claim, so it falls short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no when-to-use guidance, no conditions, and no alternatives. Sibling tools exist but are not referenced, leaving the agent to infer context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
check_completion_evidenceC
Check whether a completion/current-state claim has enough independent evidence.
| Name | Required | Description | Default |
|---|---|---|---|
| evidence | Yes | ||
| claim_type | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full disclosure burden. It never explains what counts as 'enough' evidence, whether the check is read-only, what threshold or policy applies, or how the verdict is expressed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler. It is efficient, though the sparseness reflects missing content rather than tight editing.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be described, but with 0% parameter coverage, no annotations, and a non-trivial verification domain, the definition leaves the agent without the semantics needed to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate and it does not: neither the allowed claim_type values nor the shape/required fields of an evidence item are explained. Only the loose notion of 'independent evidence' is conveyed.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (check) and resource (completion/current-state claim) plus the criterion (enough independent evidence). It is clear what the tool evaluates, though it does not distinguish itself from the closely named sibling validate_evidence.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this versus validate_evidence, classify_claim, or build_verification_plan, and no stated prerequisites or exclusions. The agent must infer the routing from the name alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
check_replay_safetyC
Block blind replay when side effects or request lineage are unresolved.
| Name | Required | Description | Default |
|---|---|---|---|
| state | Yes | ||
| side_effect_state | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. It reveals a single gating condition but does not state whether the tool actually performs a block, returns a decision, is read-only or mutating, requires special permissions, or what happens when the condition is not met. The behavioral picture is substantially incomplete.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single front-loaded sentence with no wasted words. However, given the tool's complexity and the completely undocumented parameters, it is too sparse to be appropriately sized for the task.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained. But with no annotations, 0% parameter description coverage, and two undocumented parameters, the one-sentence description leaves critical context missing for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description does not mention either parameter. The meaning of 'state' and 'side_effect_state' is left entirely unexplained, so the description fails to compensate for the missing schema documentation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a conditional action (block blind replay) tied to a specific resource and conditions (side effects or request lineage unresolved). However, the tool name suggests a safety check rather than an enforcement action, and no sibling tool is mentioned or contrasted. The purpose is somewhat inferable but not sharply distinguished from siblings like validate_evidence or resolve_precedence.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies a trigger condition (when side effects or request lineage are unresolved), but it gives no explicit when-to-use guidance, no when-not-to-use guidance, and no alternatives among the sibling tools. An agent must infer the context entirely from the name and vague condition.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
classify_claimC
Classify an agent claim and return the evidence gate it must satisfy.
| Name | Required | Description | Default |
|---|---|---|---|
| claim | Yes | ||
| claim_type | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It does not disclose whether classification is deterministic, what categories exist, whether it mutates state, or what 'evidence gate' means operationally — a significant gap for a tool whose output apparently gates downstream behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no waste. It is efficient, though brevity here shades into under-specification rather than deliberate concision.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be described. But for a tool with zero parameter documentation, no annotations, and no usage routing, the description leaves an agent unable to construct a valid call with confidence or know when to make one.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and neither parameter has any schema-level description. The description does not explain what 'claim' should contain or what the optional 'claim_type' override does, so the meaning of both parameters is undocumented everywhere.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('classify') and resource ('agent claim') and hints at an output ('evidence gate'), which separates it somewhat from siblings like validate_evidence or resolve_precedence. However, 'evidence gate' and 'agent claim' are undefined domain terms, so an agent cannot tell precisely what classification taxonomy or output is produced.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to call this versus validate_evidence, check_completion_evidence, or the other six siblings. The condition under which classification precedes or replaces evidence validation is left entirely to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
compact_evidenceC
Compact evidence while retaining source refs and prioritizing direct Reality proof.
| Name | Required | Description | Default |
|---|---|---|---|
| evidence | Yes | ||
| max_items | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. Compaction implies discarding material, yet it never states whether evidence is mutated, dropped, or merely summarized, nor whether the operation is reversible or permission-gated. The one behavioral claim made ('retaining source refs') is useful but far from complete.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no padding; the retained-refs constraint is stated before the prioritization rule. Only the unexplained 'Reality proof' jargon slightly undercuts clarity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists so return values need not be described, but for a mutation-flavored tool with zero annotations and 0% parameter coverage the definition is too thin to call safely. It omits the very behavioral facts an agent needs before compacting evidence.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description is expected to compensate. It never explains the shape of the `evidence` array items or what `max_items` truncation does to the retained refs, leaving both parameters undocumented in schema and prose alike.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Names a specific verb+resource ('compact evidence') and adds two qualifiers (retain source refs, prioritize direct Reality proof). However 'compaction' semantics and the undefined jargon term 'Reality proof' leave the exact transform ambiguous, and it does not distinguish itself from siblings like validate_evidence or resolve_precedence.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no statement of when to use this tool, when not to, or which sibling to prefer. In a family of seven evidence-related tools, the agent gets no routing signal at all.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
record_lossB
Create a structured loss-ledger entry. This tool does not persist it.
| Name | Required | Description | Default |
|---|---|---|---|
| claim | Yes | ||
| why_wrong | Yes | ||
| evidence_refs | Yes | ||
| corrected_form | Yes | ||
| prevention_rule | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and it does deliver one critical trait: 'This tool does not persist it' clarifies this is a non-persisting/dry-run style operation, which matters for a tool named 'record_loss'. However, it omits permissions, idempotency, and side effects beyond the persistence caveat.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, front-loaded with the core action and immediately followed by the key caveat. No filler; every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with 5 required, undocumented parameters and no annotations, the description is too thin; it leaves parameter meaning and usage context entirely unaddressed. The output schema does relieve it of explaining return values, and the non-persistence caveat is a valuable addition, but the parameter gap dominates.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and there are 5 required parameters, so the description must compensate, but it adds no meaning for claim, why_wrong, corrected_form, evidence_refs, or prevention_rule. Only the phrase 'loss-ledger entry' gives vague thematic framing; an agent still lacks syntactic/format guidance for every required field.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Create a structured loss-ledger entry') that an agent can distinguish from siblings like classify_claim or validate_evidence. However, it offers no explicit differentiation from the other ledger/evidence-building tools in the sibling set.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this tool versus alternatives; the sibling set (validate_evidence, classify_claim, build_verification_plan) is never referenced. The only conditional information is the non-persistence note, which is behavioral rather than usage-directional.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
resolve_precedenceC
Resolve conflicting claims using Reality > CURRENT law > cutover > implementation/tests > history.
| Name | Required | Description | Default |
|---|---|---|---|
| candidates | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It discloses the ranking rule but not whether the tool mutates state, requires permissions, what happens to losing claims, or any side effects of 'resolving' a conflict.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no wasted words. Its terseness is efficient, though it borders on under-specification rather than genuine density.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists so return values need not be explained, but for a resolution tool the description omits what 'resolving' produces for the input candidates, how ties outside the ladder are handled, and any operational constraints.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for the single 'candidates' parameter, an array of untyped objects. The description never mentions candidates or the shape each candidate must take, so it fails to compensate for the schema gap beyond the loose mapping to 'conflicting claims'.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a verb ('Resolve') and a resource ('conflicting claims'), and the precedence ladder hints at the decision logic. However, the terms 'Reality', 'CURRENT law', and 'cutover' are opaque without domain context, and nothing distinguishes this from siblings like classify_claim or validate_evidence.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The precedence order implies how conflicts are ranked but says nothing about when an agent should call this tool versus validate_evidence or check_replay_safety, and no prerequisites or exclusions are stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
validate_evidenceC
Validate evidence payload structure before it can satisfy a completion gate.
| Name | Required | Description | Default |
|---|---|---|---|
| evidence | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It does not state whether validation is read-only, what happens on failure (throw vs. structured result), whether it mutates or normalizes the payload, or any side effects — leaving the safety profile entirely implicit.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single tight sentence with the verb and resource front-loaded and no filler. It is efficient, though its brevity is partly a symptom of under-specification rather than disciplined concision.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be described. But for a validation tool with no annotations and an undocumented parameter, the description leaves out failure behavior and the relationship to the completion-gate workflow it references.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for the single "evidence" parameter, so the description must compensate — but "payload structure" only restates the parameter name. It gives no information about the expected shape, required item fields, or valid values of the evidence array.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ("Validate") and resource ("evidence payload structure"), so the action is unambiguous. However, it does not distinguish itself from the sibling check_completion_evidence, which sounds like an overlapping concern, so the agent must guess which one to reach for.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The clause "before it can satisfy a completion gate" hints at sequencing but names no alternative tool or condition for choosing this over check_completion_evidence, compact_evidence, or build_verification_plan. No when-not guidance is given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
8 tool updates
v0.1.1- First observed
build_verification_plan - First observed
check_completion_evidence - First observed
check_replay_safety - First observed
classify_claim - First observed
compact_evidence - First observed
record_loss - First observed
resolve_precedence - First observed
validate_evidence
TDQS
Scored across 8 tools
Most tools have clearly distinct purposes (classify vs resolve vs compact). However, validate_evidence, check_completion_evidence, and check_replay_safety all perform gate-style checks on evidence, which could cause confusion without careful reading of descriptions.
All tool names follow a consistent snake_case verb_noun pattern (validate_evidence, resolve_precedence, check_replay_safety, etc.). There are no mixed conventions or vague verbs.
With 8 tools, the set is well-scoped for a verification/evidence-gate server. Each tool addresses a distinct step in the workflow, and there is no excessive proliferation.
The surface covers validation, classification, conflict resolution, planning, and loss recording, but record_loss explicitly does not persist entries, leaving no way to actually store or retrieve loss data. There is also no tool to execute verification or persist evidence, creating dead ends for some workflows.
Maintenance
Related MCP Connectors
Verified external-state intelligence for agents with corroborated evidence and ANSWER or WAIT.
Preflight, approve, and prove consequential agent actions with signed evidence and x402 tools.
Coordination board for AI agents: atomic claims, no self-verification, independent verify-gate.
Deterministic operations reconciliation for AI agents: COMPLETE, INCOMPLETE, or NEEDS_REVIEW.
Related MCP Servers
- AlicenseNot gradedqualityDmaintenancePrevents autonomous agents from fabricating tool results by verifying HMAC-signed receipts against epistemic claim types, forcing re-grounding or escalation before actions commit.MIT
- AlicenseNot gradedqualityBmaintenanceEnables agents to verify claims with evidence-based truth scores and confidence levels by running a deterministic pipeline of evidence lanes and adversarial checks.26MIT
- AlicenseNot gradedqualityAmaintenanceEnables agents to classify every claim into evidence tiers (FACT/INFERENCE/SPECULATION/UNVERIFIED) with attached evidence, forcing deterministic, evidence-gated reporting without LLM or network calls.18 PyPIMIT
- FlicenseAqualityCmaintenanceEnables agents to verify claims against a real falsification ledger, querying pre-registered refutation thresholds, recorded verdicts, and verifier drift metrics instead of guessing.519 PyPI-