Skip to main content
Glama
HenricTBoehm

PADC Evidence Lab

by HenricTBoehm

PADC Evidence Lab v0.1.1.1

Source-Completeness and Provenance-Binding Hardened Prototype

Author: Henric T. Böhm — ORCID: 0009-0005-0349-6341

Prototype implementation of an audit-oriented evidence contract for agentic research workflows.

The core separation is:

Proposal != Evidence != Admissibility != Adjudication

The MCP agent may propose an epistemic status, but it cannot directly create source facts, page locators, extraction-completeness facts, provenance records, missingness records, or guard outcomes.

Research question and patch status

v0.1.0 asked whether the separation could be implemented. v0.1.1 asks a narrower adversarial question:

Does the implemented separation survive provenance counterexamples involving source completeness, historical artifact binding, and derived-state authority?

v0.1.1.1 is a narrowly scoped identity/immutability patch. It preserves the v0.1.1 architecture and adds the invariant that a stable server-derived artifact ID denotes immutable artifact content. Repeated creation of the same Missingness situation returns the already registered record. New claims are representation-normalized (Unicode NFC plus whitespace collapse) before claim hashing and new artifact storage; this normalizes representation, not meaning.

The target status is HARDENED REFERENCE PROTOTYPE — FREEZE CANDIDATE, not a validated PADC implementation, semantic evidence engine, or production audit system.

Related MCP server: aafp-commons

Epistemic statuses

  • S — Supported

  • P — Partially Supported

  • NS — Not Supported

  • C — Contradicted

  • NE — Non-Evaluable

S, P, and C require current server-generated EvidenceObject IDs. P additionally requires explicit conditions.

NS requires a current, claim-specific, protocol-complete DocumentReview over an explicitly declared complete PAGE_INDEXED_TEXT surface. The assessment binds to the exact review_id. A later review cannot retroactively revalidate an older NS assessment.

NE requires a server-generated MissingnessRecord. The assessment binds to the exact missingness_id. A later MissingnessRecord cannot retroactively revalidate an older NE assessment.

Passing a guard means only:

STRUCTURALLY ADMISSIBLE PROPOSAL FOR <status>

It does not mean that the status has been semantically adjudicated.

Declared extraction surface

v0.1.1.1 does not infer document completeness from the set of pages that happened to be ingested. Trusted ingestion must declare the expected PAGE_INDEXED_TEXT surface using either --expected-page-count or --expected-pages.

The resulting ExtractionManifest records, among other fields:

  • surface_type = PAGE_INDEXED_TEXT

  • expected and extracted page numbers

  • missing and unexpected page numbers

  • extraction_complete

  • declared_text_surface_complete

  • excluded modalities

  • extraction method and version

  • extraction_manifest_hash

A complete page-indexed text surface does not imply semantic exhaustion of the complete document. By default, the declared text surface explicitly excludes IMAGES, TABLE_STRUCTURE, and EMBEDDED_FILES.

Therefore:

protocol_complete_text != document_semantically_exhausted

Trust boundary

Trusted source ingestion is intentionally not exposed as an MCP tool.

A source is ingested out-of-band through padc-evidence-ingest. The MCP server can search and reference trusted state but no MCP tool exposes trusted ingestion.

This is an MCP capability boundary, not a complete operating-system security boundary. Protection against direct modification of the JSON state file additionally depends on host-level authority separation. If an agent has a separate shell or filesystem capability with write access to that state file, the MCP boundary alone does not prevent such modification.

The persistent audit log is audit-oriented. It is not append-only by cryptographic enforcement, externally attested, or tamper-resistant.

Install

python -m venv .venv
source .venv/bin/activate        # Windows: .venv\Scripts\activate
pip install -c constraints.txt -e ".[dev]"

The reference constraints pin MCP 2.3.0 and the dependency versions used for the v0.1.1/v0.1.1.1 core regression runs. See VALIDATION.md. The trusted page-extraction subprotocol remains version 0.1.1 because v0.1.1.1 does not change extraction semantics; it patches claim representation normalization and Missingness artifact identity.

Trusted ingestion

Prepare page-indexed text:

{
  "1": "First page text...",
  "2": "Second page text..."
}

The bundled example intentionally uses the page labels 1 and 12. If those are the trusted expected page labels for the declared text surface, ingest it as:

padc-evidence-ingest \
  --state ./data/evidence_state.json \
  --item-id DOC_001 \
  --title "Example Paper" \
  --pages-json ./examples/doc_001_pages.json \
  --collection COLL_MAIN \
  --expected-pages 1,12

For an ordinary contiguous twelve-page document, use --expected-page-count 12. If only pages 1 and 12 were then ingested, pages 2--11 would be recorded as missing and the declared text surface would remain incomplete.

If neither expected-page option is supplied, the ingested pages remain searchable, but declared_text_surface_complete is false and a complete NS review cannot be established.

Optionally add --artifact path/to/original.pdf to record the SHA-256 of the original artifact. Re-ingesting the same item_id creates a new immutable source revision. Historical assessment identity remains unchanged; current provenance may become invalid.

MCP tools

The MCP tool surface remains deliberately small:

  • list_sources()

  • find_evidence(item_id, query, limit=5)

  • run_document_review(item_id, claim, search_terms=[], semantic_queries=[])

  • validate_evidence_assessment(proposal)

  • build_evidence_matrix(collection_id)

There is deliberately no MCP tool for registering source text, assigning source revisions, creating MissingnessRecords directly, or declaring a review complete.

Semantic queries

The current review API retains semantic_queries as a declaration field for forward-compatible experiments, but no semantic-search adapter is executed in v0.1.1.1. Reviews record this as:

declared_semantic_queries_not_executed = [...]
semantic_search_executed = false
semantic_search_adapter = null

Two hashes separate contract from execution:

review_contract_hash        = hash(declared review contract)
executed_review_trace_hash  = hash(actually executed review trace)

Unexecuted semantic queries may affect the contract hash; they never masquerade as executed search in the execution trace.

Artifact identity and Missingness idempotence

For new server-derived Missingness artifacts, identity is content-stable:

MIS_1 = MIS_2  =>  dump(MIS_1) = dump(MIS_2)

After the deterministic missingness_id is computed, an existing record under that ID is returned unchanged rather than overwritten with a new created_at. New claim identity uses Unicode NFC normalization and collapses whitespace runs before hashing/storage. This prevents representational variants such as repeated spaces or canonically equivalent Unicode forms from creating or overwriting semantically intended duplicate Missingness artifacts. No semantic equivalence beyond this representation normalization is inferred.

Audit semantics

structurally_admissible is historical: it records whether the proposal passed the guards at submission time.

provenance_valid is revision-relative: it is re-computed against the current trusted source state using the exact supporting artifact IDs originally bound to the assessment.

Thus:

Assessment identity is historical.
Provenance validity is revision-relative.

and a valid matrix state may be:

structurally_admissible = true
provenance_valid        = false

A later DocumentReview cannot replace the review_id bound to an older NS assessment. A later MissingnessRecord cannot replace the missingness_id bound to an older NE assessment.

Collection membership is derived from the current SourceRecord rather than treated as an independent authority. Revisions therefore remove stale membership.

All assessment attempts, including rejected ones, are persisted. The log is not claimed to be a tamper-resistant ledger.

Tests

pytest

The regression suite covers, among other cases:

  • fabricated/unknown evidence IDs are rejected

  • page locators are server-derived

  • stale evidence is rejected after source revision

  • P requires and preserves explicit conditions

  • partial extraction cannot establish a complete text-surface review

  • expected document surface is not inferred from ingested dictionary keys

  • NS binds to the exact historical review_id

  • a new review cannot revalidate an old NS assessment

  • NE binds to an exact server-generated MissingnessRecord

  • a new MissingnessRecord cannot revalidate an old NE assessment

  • repeated creation of the same Missingness situation is idempotent and does not mutate created_at

  • Unicode/whitespace representation variants resolve to the same canonical Missingness identity

  • closing a later evidence gap changes current provenance, not historical admissibility

  • collection revision removes stale membership

  • declared semantic queries remain visibly unexecuted

  • model-reported confidence has no effect on guards

  • rejected attempts remain visible in the audit log

Scope

v0.1.1.1 intentionally does not add semantic adjudication, native PDF/Zotero ingestion, empirical PADC/DEA validation, multi-user authorization, a general governance-effectiveness claim, or a cryptographically tamper-resistant audit ledger.

The governing implementation principle remains:

An agent may propose an evidence status, but it may not create the trusted facts that make that status structurally admissible.

License

MIT License. See LICENSE.

Related MCP Connectors

Related MCP Servers

  • A
    license
    B
    quality
    B
    maintenance
    Enables software agents to maintain a signed, evidence-aware local ledger of claims, with policy-gated proposals, conflict queries, and constitution-pinned access through an MCP stdio server.
    7
    3
    Apache 2.0
  • F
    license
    Not graded
    quality
    B
    maintenance
    Enables agents to query a single-source ontology of types, entities, statements, evidence, and controlled vocabulary, including bitemporal "as of" and "as recorded" fact lookups. Supports writing by letting agents propose candidate statements while keeping approval, review transitions, and authoritative reasoning restricted to approved facts.
    -