Epistemic Warrant Protocol (EWP)
The server is an MCP façade for epistemic warrant evaluation over an evidence ledger.
Evaluate an EvidenceView or stored view at a given time with
ewp_warrant_now, returning the five normative warrant axes without persisting warrant.Persist a full EvidenceView with
ewp_evidence_view_put(requiresingest_attestationfor trusted origins; refuses WarrantView bodies).Load a stored EvidenceView by proposition_id with
ewp_evidence_view_get.Append a VerificationCheck to a stored view with
ewp_check_record.Append an EvidenceItem to a stored view with
ewp_evidence_record.Run a separate action gate with
ewp_may_act, requiring an explicit action object and not inferring permission from warrant alone.Request a bounded memory context packet via
ewp_memory_context, exposing axes, evidence ids, and warnings (OPEN/DEGRADED explicit), not a persona biography.Without the ingest role/token the server is evaluate-only; with ingest enabled it accepts writes from the trusted ingest pipeline.
Allows the MCP server to use a local SQLite database as an evidence store for the Epistemic Warrant Protocol, supporting persistence of evidence views and view completeness.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Epistemic Warrant Protocol (EWP)Run warrant_now on this evidence view with the reference-v1 policy."
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Epistemic Warrant Protocol (EWP)
EWP defines the deterministic boundary between what an AI agent’s memory contains and what the agent is epistemically justified in accepting.
Memory is evidence, not truth.
EWP-0.2.0. Policy reference-v2. To try it with Claude and your own data, see QUICKSTART.md. EWP is an epistemic protocol, not an action-authorization framework. Other projects use warrant to describe permission to act; EWP uses epistemic warrant to describe what an agent is justified in accepting. may_act() is deliberately a later gate.
Stores adapt to the protocol. The protocol does not inherit the store’s epistemology.
v0.2.0 binds verification to declared subjects[] as exact ids, scopes supersession to the proposition, treats future, missing, and unparsable times as unavailable at T, refuses invalid input (values outside the closed enums, views that mix propositions, duplicate or ambiguous ids, mistyped fields), stores immutable (proposition, view_id) snapshots over an append-only ledger, requires adapters to round-trip every field (Graphiti's unavoidable losses are listed), and ships an MCP façade where every write needs a server-side ingest role. Those judgments changed, so the policy is reference-v2. WarrantView carries protocol_version.
See NAME.md, docs/PROTOCOL.md, docs/DESIGN_NOTE.md, docs/implementer/, docs/PLATFORMS.md, docs/implementer/LIVE_ADAPTERS.md, docs/MCP_CONTRACT.md, CONFORMANCE.md, docs/OPEN_QUESTIONS.md. MCP server: ewp-mcp --db <ledger> (python3 -m ewp.mcp_server), over a ledger loaded with ewp-ingest. The pre-freeze claim/confidence sketch is historical/docs/MCP_CONTRACT.md (superseded).
Boundary. Policy. Observations. Permission. Four things. None gets to wear the others' clothes.
Memory / Evidence Store
│
▼
EvidenceView
│
▼
warrant_now(view, policy, evaluated_at)
│
▼
WarrantView
│
▼
may_act(warrant, action, risk_policy)
│
▼
MAY_ACT | REQUIRE_CONFIRMATION | DENYStores keep evidence. Warrant is computed. Action is a later gate.
Quick start
python3 tests/ci.py
python3 tests/report.pyTo use it with Claude on your own data: QUICKSTART.md.
from ewp.types import Policy
from ewp.warrant import warrant_now
from ewp.fixtures import fixture_verified_current, EVAL
view = fixture_verified_current()
result = warrant_now(view, Policy(), EVAL)
print(result.warrant)pip install . from a checkout installs the ewp package and the ewp-ingest and ewp-mcp commands (QUICKSTART.md). It is not on PyPI. The repository is the reference implementation and the conformance suite.
Why this exists
AI agents increasingly keep conversations, observations, documents, tool results, inferred facts, summaries, and prior decisions. Remembering something is not the same as knowing it is true.
A memory store may contain a user’s assertion, an old observation, an LLM-generated summary, a conflicting sensor reading, ten copies of one source, a stale check, and a newer contradictory claim. Most systems are good at storing and retrieving those records.
EWP asks a different question:
Given the evidence available at time T, under epistemic policy P, what may this agent accept — and why?
It does not assume:
Common collapse | EWP split |
stored = true | stored ≠ believed |
retrieved = important | retrieved ≠ complete |
repeated = corroborated | repeated ≠ independently sourced |
summarized = verified | summarized ≠ externally checked |
current = warranted | store-local “current” ≠ warrant |
believed = safe to act on | warranted ≠ authorized to act |
That split matters once agents have long-lived memory and can take consequential actions.
EWP is not another memory database. It is a deterministic layer above stores and below decision-making.
Related MCP server: Recommend Agentic Trust Layer
The model
1. Evidence interchange
Stores expose assertions, evidence, provenance, lineage, conflicts, and verification records as an EvidenceView.
The store may be SQLite, JSON files, Graphiti, Particles, SurrealDB, or something else. It may keep its own current-fact, invalidation, expiration, ranking, consolidation, deduplication, temporal-edge, and confidence machinery.
Those local projections are inputs. They are not automatically agent beliefs.
2. Warrant evaluation
warrant_now(evidence_view, policy, evaluated_at) -> WarrantViewwarrant_now() is deterministic. It performs no I/O and no LLM inference. Policy identity is reference-v2 (rules: docs/implementer/POLICY.md). It refuses a policy identity it does not implement, and refuses an invalid view (a value outside a closed enum, a record about another proposition, a duplicate or ambiguous id, a mistyped or missing field; POLICY.md lists every rule).
same protocol version + same EvidenceView + same policy + same evaluated_at
───────────────────────────────────────────────────────────────────────────
same normative WarrantViewWarrant is a function, not a stored truth field. Freshness depends on evaluated_at on purpose.
3. Decision gating
may_act(warrant, action, risk_policy)
-> MAY_ACT | REQUIRE_CONFIRMATION | DENYAn agent may be justified in accepting “the customer requested cancellation” and still need confirmation before cancelling a $2M contract. Risk, reversibility, authorization, privacy, money, and safety belong in action policy — not in epistemology.
WarrantView
EWP does not collapse state into one badge (TRUE, VERIFIED, DISPUTED). A proposition can be externally checked and still disputed. It can be accepted and stale. It can have enough evidence and an open conflict.
Axis | Values |
acceptance |
|
conflict |
|
verification |
|
currency |
|
sufficiency |
|
Example that a single enum cannot hold:
acceptance: TENTATIVE
conflict: OPEN
verification: EXTERNAL
currency: CURRENT
sufficiency: SUFFICIENTVerification is evidence, not a badge
EWP does not persist verified = true. A check has method, source, scope, observation time, result, and a freshness policy.
Proposition: server01 runs Windows Server 2022
Check: method tool_observation, result supports,
observed_at 2026-09-21T18:31Z, scope server01, origin toolresult is part of the check. opposes opens conflict and blocks ACCEPTED. inconclusive cannot raise EXTERNAL or HUMAN. Method name alone is never enough: an endogenous origin (extract, turn, summary, …) caps the class at INDIRECT, including human_attestation.
A check is about something. When a proposition declares subjects: [server01], only a check that also names server01 can raise EXTERNAL. A check on server02, on customer-42, or one that names no subject at all stays INDIRECT.
That check can become stale without ever having been false. Currency is computed at evaluation time from the checks that confer the chosen verification class. A later summary cannot refresh an old tool observation. History is not rewritten.
Source lineage
Ten transformations of one source are not ten independent confirmations.
Original article
├── summary → agent paraphrase
└── another summaryAll share one lineage_id. The evaluator counts unique lineage_id values on assertions, evidence, and checks — not repetitions.
SourceRef {
source_id, lineage_id, origin_type, origin_locator,
snapshot_id, content_hash, observed_at,
extractor_id?, parent_source_id?
}Three confidences that must not be mixed
Kind | Meaning |
Assertion confidence | How strongly the source or extractor stated the claim |
Retrieval score | How relevant a record looks to this query |
Warrant strength | How strongly policy permits acceptance |
Implementations MUST NOT use retrieval relevance, repetition count, memory strength, or assertion confidence as substitutes for warrant.
Core invariants
Assertions are not beliefs.
Beliefs are not truth.
Warrant is computed, not persisted as truth.
Verification is evidence with method, scope, source, and time — not a proposition badge.
Endogenous processing cannot manufacture external verification.
Derivation cannot manufacture provenance.
Contradiction must survive storage and retrieval.
Incomplete retrieval must be visible as
sufficiency=DEGRADED. Underreference-v2,DEGRADEDalso blocksACCEPTED.Epistemic policy and action policy are separate.
Semantically equivalent evidence must yield equivalent warrant independent of storage substrate.
Endogenous operations — retrieval, summarization, reflection, dreaming, consolidation, reranking, repetition, graph propagation, LLM critique, multi-agent agreement — may reorganize. They may not turn verification=NONE into verification=EXTERNAL without new external evidence. Ten agents repeating one hallucinated source are still one lineage.
If P is tentative and P → Q, inference does not make Q externally verified. Epistemic status taints forward.
If the store has A→P and B→¬P but search returns only A, sufficiency is DEGRADED. Losing evidence must not raise warrant.
What this repo is
The v0.2.0 reference kernel:
deterministic
warrant_now()— no network, no LLM, no hidden writesshared
ewp/classify.pyused by both reference evaluatorsseparate
may_act()14 canonical + 12 pathological + 26 hardening fixtures (laundering, subject binding, time at T, supersession scope, zero freshness, variant record propositions), plus 35 invalid views that must be refused
26 pinned golden
WarrantViews; all 52 fixtures' expected axes pinned indocs/implementer/an independent third evaluator (
docs/implementer/third_eval.py) written from the policy text aloneSQLite and JSON reference adapters (immutable snapshots; SQLite over an append-only ledger)
ewp-ingestto load your own evidence from JSON, andewp-mcpto serve it read-only to Claude (QUICKSTART.md)Graphiti-shaped semantic adapter (fake records)
live Mem0 mapping for the mem0ai 2.x clients, tested against a real OSS
mem0.Memory(mem0ai 2.2.0), and an experimental live Graphiti mapping (ewp/mem0_adapter.py,ewp/graphiti_client_adapter.py); every fixture round-trips through Mem0 with every field intact and through fake Graphiti with only its listed losses; livegraphiti-core 0.30.2is not validateddifferential fuzz tests: every field of every fixture replaced by malformed values (about 88,000 views) must be refused by all three evaluators or evaluated identically, every accepted malformed view must round-trip unchanged through every store, and every malformed MCP argument or JSON-RPC/HTTP request must get a specific error
MCP façade (
ewp/mcp_server.py,tests/test_mcp.py): MCP stdio transport or plain JSON-RPCPOST /mcp; contract indocs/MCP_CONTRACT.mdfour-stage runners and field-level diffs
SQLite and JSON stores safe for concurrent writers and readers (SQLite transactions; an OS file lock and atomic replaces for the JSON store)
Warrant evaluation does not require MCP. Persona files (MEMORY.md) are a generated checkout, not the system of record.
The pre-freeze claim/confidence sketch is historical/docs/MCP_CONTRACT.md (superseded). Earlier warrantmem sketches live in historical/. They are not the kernel.
Platforms
OpenClaw
OpenClaw consumes outbound MCP servers. Run the shipped façade and add it:
EWP_INGEST_TOKEN=... python3 -m ewp.mcp_server --http 127.0.0.1:8765 --db ./ewp.sqlite
openclaw mcp add ewp --url http://127.0.0.1:8765/mcpContract: docs/MCP_CONTRACT.md. Stdio is the standard MCP transport (newline-delimited JSON-RPC). HTTP is plain JSON-RPC, not MCP Streamable HTTP. Every write needs the server-side ingest role (--allow-ingest on stdio, the token on HTTP); without it the server is evaluate-only. Give the token to the ingest pipeline, not the agent. ewp_may_act gates only the latest stored evidence, at server time.
OpenClaw object | Role under EWP |
Daily notes, transcripts | Raw lineage. Append. Do not rewrite. |
| Working checkout of preferences. Regenerable. |
| Persona briefing from |
Dreaming / promotion | Candidate generator. Must not raise verification class. |
Local SQLite machine state | Cache. Not a second ledger. |
Dreaming may rewrite MEMORY.md. EWP treats that rewrite as a new assertion, not as verification. Compression must not turn “X is disputed” into “X.”
Claude, Codex, other MCP clients
Same contract over stdio: ewp-mcp --db ./ewp.sqlite, over a ledger created with ewp-ingest. Without --allow-ingest the server opens the ledger read-only. Point several clients at the same --db to share one ledger. Prompt rules: fluency is not recollection; conflict=OPEN is said out loud; DEGRADED means the view is incomplete; store-native write tools stay disconnected.
Graphiti / Mem0 / Particles / SQLite
Graphiti can be used as a temporal/entity evidence substrate or mirror. Inject lineage_id in episode metadata. invalid_at is store-local. valid_at is not a verification check. Search collapse marks the view DEGRADED. Mem0 is an extract-and-retrieve store: default origin is extract; retrieval score is not warrant. Particles is a good immutable substrate; writes go only through EWP's ingest role, never from the agent. SQLite and JSON prove store neutrality.
Full notes: docs/PLATFORMS.md. Live client mappings: docs/implementer/LIVE_ADAPTERS.md.
Compose EWP with action gates (OpenClaw allowlists, Tenuo, human approval). Do not merge those layers because they share the word warrant.
Why EWP sits above the store
System | Primary concern | EWP adds |
Graphiti / Zep | Temporal knowledge graph and retrieval | Store-independent warrant evaluation |
Mem0 | Agent memory storage and retrieval | Evidence lineage and deterministic warrant policy |
OpenClaw native memory | Inspectable agent context and memory | Separation of persona, evidence, and warrant |
EWP | Epistemic evaluation | Not a general-purpose memory store |
Graphiti adapts to the protocol. The protocol does not adapt to Graphiti.
Conformance
Schema serialization is not conformance. Behavior is.
python3 tests/ci.py
python3 tests/report.py
python3 tests/runner.py
python3 tests/runner_pathological.pyCI enforces the fixture, evaluator, policy, golden, and implementer-pack lock hashes; all 26 goldens; all 52 fixtures through the codec, SQLite, JSON, Mem0, and fake Graphiti with identical axes and fields; 35 invalid views refused by all three evaluators; fake-Graphiti isolation; the hardening pack; live-adapter mappings; the Mem0 adapter against a real mem0ai 2.2.0 client; the MCP façade, including the official MCP SDK client; the differential fuzz tests (evaluator parity, store neutrality, MCP arguments, JSON-RPC and HTTP transport); and the third evaluator. Changing a golden, the pack, the policy text, or the evaluator requires an explicit version change, then python3 tests/ci.py --write-lock. tests/report.py only says CONFORMANT for a CI stamp taken on the exact CI surface it is run on (kernel, adapters, MCP server, tests, implementer pack, examples, workflow).
Failure classes: INGEST_LOSS, ADAPTER_MAP_LOSS, WARRANT_MISMATCH, RETRIEVAL_LOSS, EXPECTED_DIVERGENCE.
EXPECTED_DIVERGENCE is an observation: the store’s “current fact” may disagree with warrant_now. That is the boundary working.
Pathological pack (among others): false supersession, open contradiction, real temporal upgrade, same text / two lineages, three wordings / one lineage, store-current vs verified, weak invalidating strong, maintenance expiry that must not erase a check, retrieval-induced false consensus, extractor polarity flip on one lineage, canonical-without-verification, invalidation cycles.
v0.2.0 lock
Epistemic Warrant Protocol EWP-0.2.0
Policy: reference-v2
Canonical: 14/14
Pathological: 12/12
Hardening: 26/26
Invalid refused: 35/35
SQLite PASS
JSON PASS
Fake Graphiti PASS
Mem0 (fake client) PASS
graphiti-core 0.30.2 — NOT VALIDATED
fixture_set_sha256:
4358b59fdda94f522abbbb8fc810c1b344da7aebc1ba116bc3fa614e3bb6c419
evaluator_set_sha256:
fd446401b46ca19626cbdf287ecd6c67f7bf82c231e5135b92393774ae464e57
policy_set_sha256:
8a9f94fffc7fa95473232da5fb1de8120c26a70c7624eb126c67021412b6405b
golden_set_sha256:
1db7cab34639a87ce35c36635d4b6cffa46ba08855107fc9a04a4b89348a81c8
implementer_pack_sha256:
700153b13a020ce7f6abcb876a1b6fab86ec79b3b8931656fd1b051c0276fa0eA store that produces a different answer has an adapter or conformance problem, not a license to move the goldens.
Layout
ewp/ kernel (classify, warrant, adapters, fixtures, MCP server)
plus live Graphiti/Mem0 mappings
tests/ conformance, goldens, adapter round trips, runners, CI,
live-adapter, real Mem0 client, MCP, and fuzz tests
docs/ PROTOCOL.md, design note, platforms, implementer pack, PDFs
docs/MCP_CONTRACT.md — shipped MCP façade
historical/ pre-freeze warrantmem ledger/MCP/Postgres sketches
including historical/docs/MCP_CONTRACT.md (superseded)
RELEASE.lock.json fixture + evaluator + policy + golden + implementer-pack hashes
pyproject.toml package metadata; ewp-ingest and ewp-mcp commands (pip install .)What EWP is not
Not a vector database, knowledge graph, memory engine, truth oracle, LLM fact-checker, or authorization framework.
It does not ask what is ultimately true. That may be inaccessible.
It asks a narrower, computable question: given this bounded evidence view, at this time, under this versioned policy, what may the agent accept — and why?
That answer can be reproduced, tested, inspected, and challenged.
Status
EWP-0.2.0, policy reference-v2. The 26 golden axes are unchanged from 0.1.0; their identity fields now read reference-v2 and EWP-0.2.0. The hardening pack (26) and the invalid pack (35) are locked. Known limits: conflict rows and lineage edges carry no timestamp, so they are not filtered by availability at T; assertions carry no polarity. See CHANGELOG.md.
New stores may reveal adapter bugs, retrieval loss, missing tests, or a genuine hole. They do not redefine warrant.
Stores keep evidence. Warrant is computed. Action is a later gate.
Copyright © 2026 Rob Koliha. MIT License.
Available Tools
8 toolsewp_check_recordA
Append a VerificationCheck: creates a new snapshot from the latest one. Requires the ingest role. Trusted origin requires ingest_attestation. Returns the new view_id.
| Name | Required | Description | Default |
|---|---|---|---|
| check | Yes | ||
| new_view_id | No | optional; content-derived when omitted | |
| proposition_id | Yes | ||
| ingest_attestation | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Since annotations are absent, the description carries the burden of behavioral disclosure. It covers the side effect (appends a snapshot), the precondition (ingest role), and the conditional requirement for trusted origin (ingest_attestation). It does not mention whether the operation is destructive, reversible, or how failures are handled. For a mutating append operation, this is a moderate disclosure, but it omits details like whether the original snapshot is preserved.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is short and front-loads the primary action. It packs essential information (returns new view_id, role requirement, trusted origin detail) into two sentences. It is efficient but slightly dense; it could be a bit more structured with bullet points, but overall every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (4 params, nested object, no output schema), the description covers core elements but leaves gaps: no explanation of what a VerificationCheck is, what the 'check' object should contain, or what the return format looks like beyond the new_view_id. The description explains some parameter semantics but not enough to fully use the tool without schema. The presence of sibling tools suggests the description could be more complete for an agent to select and invoke correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 25%, so the description must compensate, and it does: it explains that 'new_view_id' is optional and content-derived when omitted (adding meaning beyond the schema's simple 'optional' description), and it clarifies the role of 'ingest_attestation' as a requirement for trusted origin. It does not detail the structure of 'check', but the schema already uses a generic 'object' type, so the description could add value by specifying expected fields, but it doesn't.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Append a VerificationCheck') and the resource (a snapshot), and adds the key behavior of creating a new snapshot from the latest one. It mentions the return value (new view_id), which helps distinguish it from sibling tools like ewp_evidence_view_get. However, it does not explicitly contrast with siblings such as ewp_evidence_record, so it is clear but not fully differentiated.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides some usage context: it mentions the ingest role requirement and trusted origin requiring ingest_attestation, which hints at when this tool is appropriate (when caller has ingest role). It does not explicitly state when not to use it or point to alternatives, but the role requirement is a clear gate. Given the sibling set includes similar tools, more explicit when-to-use guidance would be helpful.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ewp_evidence_recordA
Append an EvidenceItem (polarity required): creates a new snapshot from the latest one. Requires the ingest role. Returns the new view_id.
| Name | Required | Description | Default |
|---|---|---|---|
| text | No | ||
| evidence | Yes | ||
| new_view_id | No | optional; content-derived when omitted | |
| proposition_id | Yes | ||
| ingest_attestation | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It discloses that this is a write operation (append and create), states the authorization requirement (ingest role), and names the return value (new view_id). The 'snapshot from the latest one' phrasing implies non-destructive append behavior, though side effects and failure modes are not detailed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is one compact sentence plus two short clauses, front-loading the core action and proceeding to role and return value. Every phrase earns its place and there is no redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description gives the essential behavioral, authorization, and return-value context, which is helpful given no annotations. But the nested evidence object and optional parameters remain under-documented, so an agent cannot confidently construct a complete call without inspecting the schema or external domain knowledge.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 20%, and the description adds little parameter meaning beyond saying polarity is required. It does not explain proposition_id, the evidence object fields (content, observed_at, source), text, ingest_attestation, or the relationship between the optional new_view_id input and the returned new view_id.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action ('Append an EvidenceItem') and a concrete outcome ('creates a new snapshot from the latest one' and 'Returns the new view_id'). It is clear but does not contrast itself with siblings such as ewp_evidence_view_put, so it is not fully differentiated.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The 'Append' verb plus 'creates a new snapshot' implies when to use it, and 'Requires the ingest role' gives an access prerequisite. However, there is no explicit when-to-use vs. alternatives and no mention of ewp_evidence_view_put/get or other sibling tools that might serve related view operations.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ewp_evidence_view_getA
Load a stored EvidenceView snapshot by proposition_id (latest snapshot unless view_id is given).
| Name | Required | Description | Default |
|---|---|---|---|
| view_id | No | ||
| proposition_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the burden of behavioral disclosure. It does communicate that this is a read-style operation ('Load') and clarifies snapshot selection behavior ('latest snapshot unless view_id is given'). However, it does not disclose error behavior, permissions, or whether any side effects occur.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence that is front-loaded with the core action and resource, then efficiently adds the parameter behavior. Every part earns its place with no repetition or filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple two-parameter getter, the description is mostly complete: it identifies the resource, the required parameter, and the optional parameter's behavior. However, with no output schema and no annotation coverage, it does not describe the shape of the returned snapshot or any error/not-found behavior, leaving a notable gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, but the description compensates by explaining both parameters: proposition_id is the basis for selecting the snapshot, and view_id optionally selects a specific snapshot instead of the latest. This gives the agent meaningful parameter semantics beyond the bare schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Load') and a clear resource ('stored EvidenceView snapshot'), and specifies the lookup key (proposition_id) with an optional variant (view_id). This clearly distinguishes it from the sibling ewp_evidence_view_put, which is the write counterpart.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context for when to use the tool: to retrieve a stored EvidenceView snapshot, with the default behavior of getting the latest snapshot unless a specific view_id is provided. It does not explicitly name alternatives or exclusions, but the usage intent is unambiguous.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ewp_evidence_view_putA
Persist an EvidenceView. Requires the ingest role. Refuses WarrantView bodies. Trusted origins require ingest_attestation=true.
| Name | Required | Description | Default |
|---|---|---|---|
| view | Yes | ||
| ingest_attestation | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full disclosure burden and does meaningful work: it states an auth requirement, a body-type restriction, and a conditional attestation flag. It does not disclose persistence semantics such as overwrite-versus-create behavior or failure modes, but the core behavioral constraints are present.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences front-load the action and follow with the two most important constraints. There is no filler or repetition of the schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with two parameters, no annotations, and no output schema, the description covers auth and the key conditional, but it omits what a trusted origin is, what happens on refusal, and any success/error behavior. An agent can invoke it correctly only if it already understands EvidenceView and trusted-origin context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It adds meaning to the 'view' parameter (must be an EvidenceView, not WarrantView) and clarifies that 'ingest_attestation' should be true for trusted origins, but it leaves the nested structure of the view object undocumented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource ('Persist an EvidenceView'), making the core action unmistakable. It also distinguishes from warrant-related siblings by explicitly refusing WarrantView bodies, so an agent can tell this is for evidence views, not warrant operations.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It provides clear invocation conditions: the caller must have the ingest role, WarrantView bodies must not be sent, and trusted origins must set ingest_attestation=true. It does not name alternative sibling tools or state when to prefer a different tool, so it stops short of fully explicit routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ewp_list_propositionsA
Find stored propositions by id or assertion text (case-insensitive substring). Returns ids, latest snapshot ids, and assertion texts. Discovery only: it does not say what is warranted; call ewp_memory_context next.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| query | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It discloses case-insensitive substring matching, the return fields (ids, latest snapshot ids, assertion texts), and the notable limitation that it does not determine warrant. The 'Find' verb and 'Discovery only' strongly imply a read-only operation. It stops short of stating side effects, ordering, or pagination, but these are less critical for a search tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two tight sentences with no filler. The main purpose and return values are front-loaded, and the caveat plus follow-up tool are given efficiently in the second sentence. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple search tool with two optional parameters and no output schema, the description covers the purpose, query semantics, return fields, and boundary of responsibility, and it names the next step. It does not disclose the default limit behavior or what happens when `query` is omitted, which are minor but relevant gaps given there is no output schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It clearly explains the `query` parameter via 'by id or assertion text (case-insensitive substring)'. However, the `limit` parameter is not mentioned at all; its behavior and default value are left to inference. This is partial compensation, not complete.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'Find stored propositions by id or assertion text', and lists what is returned. It also distinguishes itself from sibling ewp_memory_context by explicitly saying it does not say what is warranted and pointing the agent to call ewp_memory_context next. This clearly differentiates its role among the sibling tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'Discovery only' tells the agent when this tool is appropriate, and the explicit instruction 'call ewp_memory_context next' routes to the correct alternative when warrant information is needed. This provides clear usage context and an exclusion condition without leaving it to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ewp_may_actC
Separate action gate over the latest stored view at server time. Requires an action with risk (low|medium|high) and reversible (boolean). Refuses inline views, caller-built WarrantViews, and a view_id that is not the latest. risk_policy is operator-only (ingest role).
| Name | Required | Description | Default |
|---|---|---|---|
| action | Yes | ||
| view_id | No | optional guard; must be the latest snapshot | |
| risk_policy | No | ||
| evaluated_at | No | optional sanity check; the decision always uses server time | |
| proposition_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Since no annotations are provided, the description carries the full behavioral burden. It discloses that the tool requires a risk/reversible action, refuses certain view inputs, and that risk_policy is operator-only, which are useful constraints. However, it does not state whether the tool is read-only or has side effects, nor what the response format is, leaving the agent uncertain about the operation's nature.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three sentences and fairly compact, leading with the core purpose and then adding constraints. It is not overly verbose, but the phrasing could be tightened (e.g., 'Separate action gate' is unclear). Overall it is efficient.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (nested objects, 5 parameters, no output schema), the description is incomplete. It does not explain what the tool returns (e.g., a decision, denial reason), how to construct the action object correctly, or the relationship between proposition_id and view_id. The agent lacks enough information to invoke this tool properly without guessing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is low (40%), and the description partially compensates by explaining that view_id must be the latest snapshot, risk_policy is operator-only, and evaluated_at is a sanity check. However, it does not elaborate on the meaning of action, proposition_id, or the internal structure (kind, action_id), so several parameters remain under-documented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states the tool is a 'Separate action gate over the latest stored view at server time' and lists required inputs (risk, reversible), which gives a general sense of its function. However, the phrasing 'Separate action gate' is vague, and it does not clearly distinguish this from sibling tools like ewp_warrant_now or ewp_check_record, leaving the exact role ambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description mentions what it refuses (inline views, caller-built WarrantViews, non-latest view_id) but provides no guidance on when to use this tool versus alternatives. It does not name any sibling tool or specify the conditions that would make this the preferred choice, so the agent is left without routing context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ewp_memory_contextC
Bounded packet: axes, evidence ids, warnings. Not a persona biography. OPEN and DEGRADED are explicit. evaluated_at defaults to server time.
| Name | Required | Description | Default |
|---|---|---|---|
| query | No | ||
| view_id | No | ||
| evaluated_at | No | ||
| proposition_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description must carry the full burden, but it only adds a default for evaluated_at and vague statuses ('OPEN and DEGRADED are explicit'). It does not disclose whether the operation is read-only, whether it has side effects, or what the caller should expect regarding errors or state changes.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The text is compact and front-loads the core noun ('Bounded packet'), with no filler. Each sentence adds a distinct constraint or default, though the extreme terseness hurts overall clarity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has no output schema and no annotations, and the description does not explain the return structure, the meaning of axes/evidence ids/warnings, or how statuses relate to the requested packet. An agent would need outside knowledge to invoke this correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must explain parameters, but it only clarifies evaluated_at's server-time default. The required proposition_id and the optional query and view_id are left undefined, leaving an agent to guess how to populate the call.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description identifies an artifact ('Bounded packet: axes, evidence ids, warnings') and an exclusion ('Not a persona biography'), but it never states a verb or operation. An agent cannot tell whether ewp_memory_context retrieves, creates, or validates this packet, so the purpose remains vague rather than actionable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no statement of when to use this tool or when to prefer a sibling like ewp_evidence_view_get or ewp_list_propositions. 'Not a persona biography' only says what the result is not, not which task should route here.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ewp_warrant_nowA
Evaluate a stored EvidenceView (or a hypothetical inline one) at evaluated_at. Returns the five normative axes. Inline views have trusted origins demoted unless the caller holds the ingest role and sets ingest_attestation. Does not persist warrant.
| Name | Required | Description | Default |
|---|---|---|---|
| view | No | ||
| view_id | No | ||
| policy_id | No | ||
| evaluated_at | Yes | ||
| policy_version | No | ||
| proposition_id | No | ||
| ingest_attestation | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burdenency. It discloses key behavior: it returns five normative axes, does not persist a warrant, and demotes trusted origins for inline views unless the caller holds the ingest role and sets ingest_attestation. This is meaningful transparency, though it omits some details like required permissions for stored views or error behaviors.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three sentences, front-loaded with the core action and return value, and every sentence adds distinct information. It is compact without being vague.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has 7 parameters, nested objects, no annotations, and no output schema, so the description needs to carry more weight. It does not explain what the five normative axes are, which parameters apply to stored vs inline views, or the roles/requirements for policy_id, policy_version, and proposition_id. This is incomplete for an agent to reliably invoke the tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It clarifies evaluated_at, view/view_id, and ingest_attestation, but leaves policy_id, policy_version, proposition_id, and the view object semantics unexplained. With 7 parameters and no schema descriptions, this is a significant gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: evaluating an EvidenceView at a given time and returning five normative axes. It is not a tautology and gives a clear sense of the operation, though it does not name or explicitly distinguish sibling tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context for when to use the tool: to evaluate a stored or inline EvidenceView at evaluated_at, with the note that it does not persist the warrant. It also gives a conditional usage caveat for inline views and ingest_attestation, but it does not explicitly describe when to prefer this over sibling tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
6 tool updates
v0.2.1- Changed
ewp_check_record2 fields changed- added
Input schema / properties / new_view_idAdded value: +{ + "description": "optional; content-derived when omitted", + "type": "string" +} - removed
Input schema / properties / view_idRemoved value: -{ - "type": "string" -}
- Changed
ewp_evidence_record3 fields changed- added
Input schema / properties / evidence / propertiesAdded value: +{ + "polarity": { + "enum": [ + "supports", + "opposes" + ], + "type": "string" + } +} - added
Input schema / properties / evidence / requiredAdded value: +[ + "evidence_id", + "polarity", + "content", + "observed_at", + "source" +] - added
Input schema / properties / new_view_idAdded value: +{ + "description": "optional; content-derived when omitted", + "type": "string" +}
- Added
ewp_list_propositions - Changed
ewp_may_act6 fields changed- added
Input schema / properties / action / propertiesAdded value: +{ + "action_id": { + "type": "string" + }, + "kind": { + "type": "string" + }, + "reversible": { + "type": "boolean" + }, + "risk": { + "enum": [ + "low", + "medium", + "high" + ], + "type": "string" + } +} - added
Input schema / properties / action / requiredAdded value: +[ + "risk", + "reversible" +] - added
Input schema / properties / evaluated_at / descriptionAdded value: +"optional sanity check; the decision always uses server time" - added
Input schema / properties / view_idAdded value: +{ + "description": "optional guard; must be the latest snapshot", + "type": "string" +} - removed
Input schema / properties / warrantRemoved value: -{ - "type": "object" -} - changed
Input schema / requiredPrevious value: -[ - "action" -]New value: +[ + "action", + "proposition_id" +]
- Changed
ewp_memory_context1 field changed- changed
Input schema / requiredPrevious value: -[ - "proposition_id", - "evaluated_at" -]New value: +[ + "proposition_id" +]
- Changed
ewp_warrant_now1 field changed- added
Input schema / properties / ingest_attestationAdded value: +{ + "type": "boolean" +}
7 tool updates
v0.2.0- First observed
ewp_check_record - First observed
ewp_evidence_record - First observed
ewp_evidence_view_get - First observed
ewp_evidence_view_put - First observed
ewp_may_act - First observed
ewp_memory_context - First observed
ewp_warrant_now
TDQS
Scored across 8 tools
Most tools have distinct purposes: memory context, warrant evaluation, evidence persistence/retrieval, check/evidence appending, action gating, and discovery. However, ewp_evidence_view_put and ewp_check_record/ewp_evidence_record both create new snapshots and require ingest roles, which could cause some confusion about when to use each.
All tools share the ewp_ prefix and use snake_case, with a consistent verb_noun pattern (put, get, check, list, may_act). Minor deviation: ewp_may_act uses a modal verb rather than an imperative, and ewp_memory_context is a noun phrase rather than verb_noun.
Eight tools is well-scoped for a specialized epistemic warrant protocol. Each tool covers a distinct operation in the lifecycle: memory context, warrant evaluation, evidence view persistence/retrieval, record appending, action gating, and proposition discovery.
The core lifecycle is covered: put/get evidence views, append evidence/checks, evaluate warrant, gate actions, and list propositions. Minor gaps: no explicit delete/update for evidence views or propositions, and no tool to manage roles or ingest attestation beyond the inline parameter.
Maintenance
Related MCP Connectors
Evidence-bound second-opinion audit of an agent conclusion against caller-supplied evidence.
Evidence-readiness MCP server: validate, audit, and score briefs, memos, and evidence packs.
Deterministic authorization for one proposed AI agent action, returned with a signed receipt.
Deterministic allow/require_approval/deny verdicts for agent actions, before they happen.
Related MCP Servers
- AlicenseNot gradedqualityCmaintenanceEvidence-first delivery audit MCP server that evaluates task requirements against delivery evidence and returns a reproducible pass/needs_review/fail decision with a deterministic receipt.MIT
- AlicenseNot gradedqualityBmaintenanceEnables agents to verify claims with evidence-based truth scores and confidence levels by running a deterministic pipeline of evidence lanes and adversarial checks.26MIT
- AlicenseAqualityCmaintenanceDeterministically evaluates whether a proposed agent spend action matches a supplied policy, returning ELIGIBLE, DENY, or STEP_UP with stable reason codes. Provides local policy evidence only, not payment authorization.146 npmMIT
- AlicenseCqualityBmaintenanceEnables AI agents to validate candidate data pipelines against a trusted reference, running deterministic checks and returning evidence reports through MCP.1MIT