Context Layer Lab
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Context Layer Labare all scheduled automations actually healthy right now?"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Context Layer Lab
A small, inspectable reference implementation for stopping AI agents from acting on stale operational context.
Open the local-first diagnostic: no install, loads a synthetic snapshot, fully interactive.
The failure it prevents
07:30 status dashboard: GREEN
08:02 nightly export: SUCCESS
08:40 docs build: FAILED
09:05 weekly report: PRESERVED_LOCAL
Naive answer: Yes, the dashboard is green.
Governed answer: No, two lanes need attention.The scenario is synthetic and public-safe, but the failure class is real: aggregate status commonly outlives the evidence it summarizes. This lab reconciles freshness, schedule state, and evidence specificity before it will answer "are all scheduled automations healthy?" It does not simply repeat whatever the last dashboard said.
Related MCP server: artagon-vault-mcp
What this proves
BM25F freshness-aware retrieval: 10/10 cases correct. Current records rank above stale and invalid ones; BM25F score orders records within each state. Whole-token, quality-flagged, bounded search, no embeddings. Full numbers, including a three-way naive vs. recency-only vs. governed contrast, in
docs/eval-report.md.Newer terminal receipts override an expired aggregate, and a not-yet-due lane is never mislabeled as failed (see the operational invariant matrix in
docs/architecture.md).Every decision-bearing summary, receipt, and lane schedule is bound to a source-linked typed assertion. A scenario outcome that contradicts its own evidence is rejected, not silently trusted.
17 deterministic cases pass against the real reconciliation and retrieval code, no model call:
npm run eval(7 record and operational, 10 retrieval).
This is not a RAG benchmark, vector database, or production authorization layer, and it does not claim schema validation makes information true. It is the smaller evidence-reconciliation layer that should exist before an agent is trusted to summarize operational state.
Quick start
Requires Node.js 22+.
npm install
npm run check
npm run demonpm run demo prints the scenario above from the real reconciliation code:
Question: Are all scheduled automations healthy?
As of: 2026-07-28T09:10:00Z
Naive answer: healthy. It repeats the dashboard observed at 2026-07-28T07:30:00Z.
Governed answer: attention. 2 of 3 due lanes need attention.
Why:
- The dashboard record is stale. The record expired at 2026-07-28T07:55:00Z.
- 3 newer run receipts supersede it.
- Docs Build: failed (docs-build-receipt).
- Weekly Report: preserved_local (weekly-report-receipt).
- Data Sync: not due yet, so it is not counted as failed.npm start starts the read-only MCP server on stdio. It prints nothing and
waits for a client; see MCP tools for the
client config.
To apply the pattern to a company knowledge base, see
docs/how-to-adopt.md.
Generate a portable diagnostic snapshot the console can open directly. Parsing and rendering happen in the browser; nothing is uploaded:
npm run diagnose -- --output snapshot.jsonTwo labs, one boundary
Context Layer Lab -> diagnose what current evidence supports
Governed Action Lab -> prepare/approve what may execute, under whose authority
-> execute/verify with what receiptThis lab establishes what current evidence supports. Its companion,
Governed Action Lab
(live console), starts at
that boundary and decides what may execute, under whose authority,
and with what receipt. They are one system in two repos, not two unrelated
projects. The real command sequence between them, with real output, is in
governed-action-lab's docs/pair-walkthrough.md.
Scope and limits
Status: reference implementation, not a production system. This
repository demonstrates governed context records, provenance, validity
windows, freshness-aware retrieval, and a naive-versus-governed failure
contrast. It does not compete with production context or memory platforms on
retrieval scale, storage architecture, access control, or poisoning defense.
Authorization and record-level access control are out of scope; recency comes
from an explicit validUntil, not inferred content. See
docs/architecture.md for the
full list.
Learn more
docs/architecture.md: MCP tools, retrieval, data contract, evaluations, adapters, and full development checks.docs/eval-report.md: auto-generated numbers and the naive/recency-only/governed contrast.docs/adr/: six architecture decision records.docs/how-to-adopt.md: how to put this pattern in front of a real agent, and what production still needs.CONTRIBUTING.md: development checks and how to add an eval case.
License
MIT
Available Tools
4 toolsexplain_sourceBRead-onlyIdempotent
Show a declared source and the exact claims it supports on one context record.
| Name | Required | Description | Default |
|---|---|---|---|
| recordId | Yes | ||
| sourceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false, so the safety profile is covered. The description adds that output is a source plus the claims it supports on one record, which is useful framing, but says nothing about result format, emptiness cases, or whether the source must already be declared.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, front-loaded sentence with no filler. It is efficient, though its brevity is part of why parameter and usage detail is missing.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only lookup tool with no output schema, the description covers the conceptual output but omits parameter meaning and usage routing. Adequate at a minimum-viable level, but leaves real gaps an agent would face when invoking it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for both required parameters. The description only obliquely implies the mapping ('one context record' -> recordId, 'declared source' -> sourceId) without explaining ID format, scope, or whether the source must pre-exist. It fails to compensate for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Show') and resource ('a declared source and the exact claims it supports'), so an agent can tell it produces a source-to-claims view. However, it does not distinguish itself from siblings like search_context or validate_record, leaving the boundary to inference.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is given on when to use this tool versus search_context, validate_record, or inspect_ingestion. The agent must guess the triggering scenario from the description alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
inspect_ingestionCRead-onlyIdempotent
Show the deterministic receipt linking one context record to its synthetic Markdown source.
| Name | Required | Description | Default |
|---|---|---|---|
| recordId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is fully covered elsewhere. The description adds only that the output is a 'deterministic receipt' tied to a source, but says nothing about what the receipt contains, what happens if the record has no source, or its size/shape — a gap since no output schema exists.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, front-loaded sentence with no filler; the artifact being returned is named immediately. It is efficiently worded, though the terseness contributes to the ambiguity noted elsewhere.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter read tool with annotations covering safety, the description is minimally viable but under-specified: the recordId input is undefined, the receipt's contents are unexplained, and no output schema exists to fall back on. An agent can guess the call signature but not what it will get back.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
One parameter (recordId) with 0% schema description coverage and no enum, so the schema contributes nothing beyond type and minLength. The description references 'one context record' but never clarifies whether recordId is that context record's ID, a source ID, or something else, leaving the parameter's meaning underdetermined.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Show') and a specific artifact (a 'deterministic receipt linking one context record to its synthetic Markdown source'), which is more than a tautology. However, 'deterministic receipt' is unexplained jargon, and the description gives no signal distinguishing it from the sibling explain_source, which plausibly covers the same record-to-source relationship.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance, no preconditions, and no mention of alternatives such as explain_source, search_context, or validate_record. The agent must infer from the name alone whether this is a lookup, an audit, or a debugging tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_contextBRead-onlyIdempotent
Search synthetic organizational context. Returns a bounded, ranked set: current records first, then stale, then invalid, with source IDs and visible freshness or evidence issues on every match.
| Name | Required | Description | Default |
|---|---|---|---|
| asOf | No | ||
| limit | No | ||
| query | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnly, idempotent, non-destructive), so the bar is lower. The description still earns credit by disclosing result behavior beyond annotations: bounded ranked output, ordering current > stale > invalid, source IDs, and visible freshness/evidence issues on every match. It omits edge behavior such as how limit truncates or how asOf affects filtering.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences, front-loaded with the action and then the return contract. No filler, no restatement of the name, every clause adds information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description usefully characterizes the return set, ordering, and freshness/evidence fields. However, for a 3-parameter tool with 0% schema coverage it leaves asOf and limit semantics completely unexplained, so an agent lacks what it needs to set those inputs correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must carry parameter meaning, and it largely does not. It never explains asOf (date-time semantics), limit (bound behavior), or query beyond the implicit notion of a search string. It says the result is 'bounded' but not how the limit parameter controls that.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Search synthetic organizational context') and even characterizes the result shape (ranked, bounded). It is distinguishable from siblings like validate_record and explain_source by the resource and operation, though it never names or explicitly contrasts them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use or when-not-to-use guidance. The agent gets no signal about when search_context is preferred over validate_record, inspect_ingestion, or explain_source, nor any prerequisite/context conditions. Usage is only implied by the verb 'Search'.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
validate_recordBRead-onlyIdempotent
Validate a context record's schema, provenance coverage, claim support, and freshness boundary.
| Name | Required | Description | Default |
|---|---|---|---|
| asOf | No | ||
| record | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint and destructiveHint=false, so safety and repeatability are covered. The description adds real substance by naming the four validation dimensions, but says nothing about what a validation result looks like, whether failure is an error or a payload, or how strict each check is.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence that states the verb, the target, and the four checks in one pass. No filler, no restatement of the name.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description carries the burden of explaining what a caller gets back from a validation call — a boolean, a list of violations, a severity scale — and it does not. Combined with an undocumented 'record' parameter, an agent cannot fully predict the call or its result.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the tool has two parameters. 'record' is an untyped freeform object with no documented shape, and the description only vaguely gestures at 'freshness boundary' to explain the optional asOf parameter. With the schema contributing nothing, the description fails to compensate for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb ('Validate') plus explicit resource ('a context record') and an enumeration of the four facets checked: schema, provenance coverage, claim support, and freshness boundary. It does not differentiate itself from the siblings (explain_source, inspect_ingestion, search_context), but those occupy clearly different domains, so an agent can still route correctly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description never states when to reach for this tool versus the siblings, nor any prerequisites (e.g., whether the record must already be ingested or fetched via search_context first). Usage is only implied by the tool's name and the enumerated checks.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
4 tool updates
v1.2.0- First observed
explain_source - First observed
inspect_ingestion - First observed
search_context - First observed
validate_record
TDQS
Scored across 4 tools
The four tools target distinct actions: search, validate, explain source claims, and inspect ingestion receipts. explain_source and inspect_ingestion both touch provenance, so minor overlap exists, but their descriptions make the different focuses clear enough.
All tool names follow a consistent snake_case verb_noun pattern: explain_source, inspect_ingestion, search_context, validate_record. There are no deviations or mixed conventions.
Four tools is well-scoped for a focused context-inspection and validation layer, and each tool has a clear role with no obvious redundancy. The count sits comfortably in the typical 3-15 range.
The surface covers search, validation, source-claim explanation, and ingestion receipt inspection, but lacks mutating lifecycle operations such as create, update, delete, or a direct get-by-ID. For a context layer, these are notable gaps unless the intended scope is strictly read-only inspection.
Maintenance
Related MCP Connectors
Guarded MCP server for agent-readable business truth, provenance, readiness, and discovery.
Read-only Remote MCP for externally grounded AI agent trust receipts.
Read-only MCP server for The Quiet Protocol's engines, benchmarks, proof, and business data.
Read-only MCP server for AIStatusDashboard status, incidents, metrics, and fallback recommendations.
Related MCP Servers
- AlicenseNot gradedqualityAmaintenanceRead-only MCP server that answers questions about turva.dev from its published data. Five tools return JSON: the service catalog with prices, contact and operator details, engagement principles and dated agent-readiness and security evidence with verification links. Connect over Streamable HTTP. You need no API key. The server does not scan other websites or run audits.MIT
- FlicenseNot gradedqualityCmaintenanceRead-only MCP server for querying an evidence-aware knowledge vault with temporal and provenance-aware data, supporting agent memory and semantic graph projections.-
- AlicenseNot gradedqualityAmaintenanceAn MCP server exposing scoped, read-only enterprise operations tools with fail-closed credential handling. It returns opaque approval IDs for mutations and requires a separate operator approval command to release one-time capabilities.MIT
- AlicenseAqualityCmaintenanceRead-only MCP server exposing a W3C PROV knowledge graph of verified facts with provenance, enabling AI agents to list, search, and check facts while enforcing that writes remain CLI-only.5MIT