context-lattice
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@context-latticeWhy did we change the retry policy in last week's sessions?"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Context Lattice
Context Lattice is a standalone, source-verifiable memory index for AI coding-agent conversation exports. Point it at JSON or JSONL produced by tools such as Codex, Cursor, Kiro, or your own agent, then retrieve a compact evidence bundle instead of replaying an entire history into the next context window.
Raw records remain immutable. Hierarchical summaries are lossy, disposable navigation indexes. Every result points back to the source file, JSON record, and SHA-256 hash.
Quick start
No API key or external database is required. Python 3.11+ is supported.
cd /path/to/context-lattice
python3 -m pip install -e .
context-lattice init --root ~/exports/coding-sessions/
context-lattice ingest ~/exports/coding-sessions/
context-lattice query "Why did we change the retry policy?"
context-lattice query "What is the latest payment reference?" --json --debug
context-lattice inspect
context-lattice verify
context-lattice doctorBy default the database is created at ~/.context-lattice/memory.db. Override it with
--db /path/to/memory.db. Ingestion accepts a single .json, a .jsonl, or a directory
tree containing both. init prints ready-to-paste MCP configuration using the same database
and source roots, avoiding CLI/server configuration drift.
Related MCP server: ThreadShelf MCP Server
Install as an MCP server
Context Lattice exposes memory_search, memory_get, memory_explain, memory_sync, and
memory_status over local stdio MCP. Configure the database and an explicit source
allowlist, then point any MCP-capable client at the installed executable:
{
"mcpServers": {
"context-lattice": {
"command": "context-lattice-mcp",
"env": {
"CONTEXT_LATTICE_DB": "/absolute/path/to/memory.db",
"CONTEXT_LATTICE_ALLOWED_ROOTS": "/absolute/path/to/agent/sessions"
}
}
}
}Multiple allowed roots use the platform path separator (: on macOS/Linux, ; on
Windows). memory_sync is the only mutating MCP operation. Search evidence is explicitly
marked untrusted, bounded by a token budget, and resolvable to source hashes. See
docs/MCP.md.
Input adapters
--adapter auto examines record envelopes and currently recognizes:
codex: event JSONL containingsession_metaandresponse_itemrecords;cursor: exported conversations containing messages, bubbles or turns;kiro: role/content JSONL with a metadata header;claude: Claude Code project JSONL with nested message blocks;canonical: the stable Context Lattice v1 format;generic: common role/content, speaker/text and nested-message shapes.
Provider formats can change. The Codex, Cursor and Kiro adapters are intentionally tolerant
and fixture-tested, but the canonical schema is the guaranteed integration boundary:
schemas/conversation-v1.schema.json. See
examples/canonical.json for the smallest complete example.
{
"schema": "context-lattice/v1",
"conversation_id": "release-planning",
"messages": [
{
"role": "user",
"content": "The deployment region is ap-south-1.",
"timestamp": "2026-08-20T10:00:00Z",
"fact_key": "deployment-region",
"entities": ["ap-south-1"]
}
]
}Records without recognizable conversational content are counted as skipped. Malformed or failed records are reported in the import result and stored in the import audit log. Re-importing the same file is idempotent.
Retrieval model
immutable, append-only raw events in SQLite;
original raw JSON plus file/record provenance and hashes;
FTS5/BM25 for exact identifiers;
an inverted sparse-postings index with an offline feature-hashing baseline;
fixed-fanout chronological summary trees;
reciprocal-rank fusion across lexical, semantic and hierarchical candidates;
correction chains through optional
fact_keyvalues;disagreement-triggered search expansion and confidence-based abstention;
explicit evidence-token budgets and inspectable retrieval traces.
The indexer and retriever are deterministic and make no LLM calls. The bundled semantic model is feature hashing, so it is portable and exact-repeatable but weaker than a learned embedding model. An external LLM may consume the evidence; it is not trusted to maintain the memory index. The embedder boundary can be replaced without changing the evidence contract.
The chronological hierarchy is a segment-tree-like navigation index. It cannot replace
semantic or lexical lookup: trees prune time ranges, while FTS5 and sparse postings locate
terms and concepts. Query-time dot products are aggregated inside SQLite, and only a
bounded root set and beam descend the tree. See docs/ARCHITECTURE.md.
Test and evaluate
python3 -m unittest discover -s tests -v
python3 -m context_lattice.cli eval --output benchmark-results.json
python3 -m context_lattice.cli golden-eval --output golden-results.json
python3 -m context_lattice.cli demo \
"What is the current meeting room for team-3?"The deterministic 50-question evaluation compares a recent-token window, rolling summary,
flat vector search and hierarchical hybrid retrieval. Its answer_accuracy is an evidence
sufficiency metric—not an LLM-judge score. See benchmark-results.json
and DESIGN.md.
The manually reviewed golden-v1 suite is the release gate for source recall, precision,
ranking, stale facts, unsupported results, citation integrity, and hierarchy branch recall.
It intentionally fails the command when thresholds regress. See
docs/GOLDEN_EVAL.md.
Production posture
The local-first core has atomic per-conversation index updates, WAL concurrency,
cross-process maintenance locking, immutable events, source verification, bounded inputs and
retrieval, nested-symlink-safe MCP allowlisting, schema compatibility checks, CI, and
deterministic release gates. context-lattice doctor checks database integrity, index
freshness, FTS5, permissions, SQLite, Python, and MCP. Provider formats remain unofficial and
can change; the canonical v1 schema is the stable integration boundary. Review
SECURITY.md before exposing anything beyond local stdio.
The concrete release checklist and current non-goals are in
docs/RELEASE_GATES.md.
Available Tools
5 toolsmemory_explainC
Search memory and include deterministic routing, fusion, and abstention diagnostics.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| query | Yes | ||
| token_budget | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description ''carries the full burden of disclose behavioral traits. 'Search memory' implies a read-only action, but the description does not explicitly state whether it unmutated or requires permissions, nor does it explain the nature of the diagnostics beyond jargon. The behavioral disclosure is insufficient for a tool with no annotatons.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is one concise, front-loaded sentence with no wasted wording or filler. It could use more depth, but that is a completeness issue rather than a conciseness or structure failure.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Even though there is an output schema, the description does not establish the key context: why anyone would choose this over memory_search, what "deterministic routing, fusion, and abstention" means in practitioner terms, or what behavioral tradeoffs come. For a specialized diagnostics tool with no annotations or clear parameter docs, this definition is incomplete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description needs to compensate for the three parameters explain. 'Query', 'limit', and 'token_budget' are not mentione, and 'token_budget' especially is not self-explaning. The description does not become more contribution for the parameter semantics.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action ('Search memory') and a scope ('include deterministic routing, fusion, and abstention diagnostics'), which makes it distinguishable from sibling memory_search. It is clear enough but does not explicitly name or contrast the sibling tools, so it falls short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The diagnostics wording implies the tool is for explaining or debugging memory retrieval decisions, but the description never explicitly says when to use this over memory_search, memory_get, or memory_sync, nor does it provide any exclusions or alternatives. Usage guidance is only implied, not stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
memory_getB
Resolve one Context Lattice citation and verify it against its original source file.
| Name | Required | Description | Default |
|---|---|---|---|
| citation | Yes | ||
| include_raw | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description anues "verify" against its original sourcefile, a non-obvious behavior that says the tool checks correctness and isn't just a raw accessor. It does not state outverification failures, permissions, or side effects, so the behavior is only partially disclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence with no wasted or repeated information. It fronts the main action and adds only the verification detail in the second clause.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is simple, but missing include_raw semantics and the format/form of a citation means the agent cannot safely choose correct arguments. The output schema may explain return but no the input field, and no annotation compensates.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description only clarifies that the citation targets a Context Lattice item. It does not describe include_raw, its default, or its effect, so an agent is left to guess how to request the raw variant. This is a significant gap for a two-parameter tool.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific action—resolving a single Context Lattice citation—and an additional verification behavior, which separates this from a plain search or sync operation. It does not fully demes multiple beats "Context Lattice" and "original source file" due to jargon, so no full 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies it is used when you already have a citation to resolve, rather than when searching or syncing. It never explicitly contrasts with siblings or explains when not to use it; the intended spot in a generate pipeline is inferred.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
memory_searchA
Search durable coding-agent memory and return bounded, untrusted, cited evidence.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| query | Yes | ||
| token_budget | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description bears the full burden and meaningfully warns that results are 'bounded and untrusted' and that the tool returns citations. That adds real behavioral context beyond the name and schema, though it does not explicitly state whether search mutates memory or what boundedness means in operation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One compact sentence with no filler. Every phrase earns its place: the resource, the search operation, and the caveats on returned evidence.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The output schema exists and the description gives the most important search-relevant caveat: evitable. However, with zero parameter coverage in the schema and no usage routing among siblings, the description is adequate but not complete enough for robust agent understanding.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the description never names or explains query, limit, or token_budget. 'Bounded' slightly hints at limits/token caps but does not map to specific parameters, leaving the required query and both optional controls semantically under-explained.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses the specific verb 'Search', identifies the resource as 'durable coding-agent memory', and states the return shape as 'bounded, untrusted, cited evidence'. This distinguishes memory_search from siblings such as memory_get or memory_explain, whose names imply other operations.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The task context is implied: use it when you need to search memory and retrieve cited evidence. However, it never explicitly contrasts this with memory_get, memory_explain, or other siblings, so the agent is left to infer when one should be preferred over another.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
memory_statusA
Report corpus, schema, index model, and allowed-source health without message content.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, this description carries the burden of transparency. It explicitly discloses that the tool does not report message content, which is an important behavioral boundary. The verb 'Report' also implies a non-mutating, observability-focused behavior despite no readOnlyHint annotation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, tight sentence with no filler. The key scope ('corpus, schema, index model, and allowed-source health') appears before the clarifying exclusion, making it easy to scan.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given that the tool has no parameters and has an output schema, the description provides all needed selection context: what health areas are reported and what is explicitly excluded. There is no missing information that would prevent an agent from invoking it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so no semantic contribution is required from the description. The schema fully covers the empty parameter set, meeting the baseline for parameterless tools.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a clear verb ('Report') with a specific resource: corpus, schema, index model, and allowed-source health. The added 'without message content' clause disambiguates it from memory_search and memory_get, which would return content.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The context is clear: this is a health/status tool, and the 'without message content' phrasing tells an agent not to use it when message content is needed. It does not explicitly name alternatives or contain a formal when-not-to-use rule, so it is slightly short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
memory_syncB
Incrementally ingest allowed agent-session roots and atomically update affected indexes.
| Name | Required | Description | Default |
|---|---|---|---|
| paths | No | ||
| adapter | No | auto |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden of behavioral detail. It does disclose 'incrementally ingest' and 'atomically update affected indexes', which is useful, but it doesn't explain side effects, failure behavior, permissions, idempotency, or what 'allowed' means in practice.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One sentence with no fluff. The core action 'incrementally ingest and atomically update' is clear and front-loaded, and every word earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with no annotations, undocumented parameters, and multiple siblings, it is under-specified. It omits usage direction, parameter semantics, and behavior side effects; the output schema helps but doesn't make up for those gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for the paths and adapter parameters. It doesn't name either parameter or explain what paths and adapter expect, nor how they affect synchronization.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Clearly specifies a distinct operation: incrementally ingest allowed agent-session roots and atomically update indexes. This distinguishes it from the read-oriented sibling tools memory_search, memory_get, memory_explain, and memory_status.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is given about when to call this tool, when not to call it, or how it relates to memory_search and memory_get. The only contexts is inferred from the phrase 'sync', but the description never says when synchronization should be triggered.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
5 tool updates
v0.3.0- First observed
memory_explain - First observed
memory_get - First observed
memory_search - First observed
memory_status - First observed
memory_sync
TDQS
Scored across 5 tools
The tools are mostly distinct in purpose: search returns evidence, get resolves a specific citation, explain adds diagnostics, sync ingests data, and status reports health. The only meaningful overlap is between memory_search and memory_explain, since both perform searches, but their different outputs make the distinction usable.
All tools share the consistent memory_ prefix and lower_snake_case convention. The minor deviation is memory_status, which names a state rather than an action, while the other four tools use memory_ followed by a verb-like operation.
Five tools is a well-scoped number for a memory/context server. Each tool covers a distinct part of the workflow: search, resolution, explanation, sync, and status.
The surface covers the core memory lifecycle: retrieval, verification, diagnostics, ingestion, and health monitoring. There is no explicit delete/forget or curation operation, but this is a minor gap because memory appears to be source-backed and managed through the sync tool.
Maintenance
Related MCP Connectors
MCP-native web evidence and claim verification: cited, source-grounded evidence for AI agents.
21Tamper-evident proof creation and verification for AI agents via MCP, A2A, and REST.
Bounded KVP, RAG search, and wipe receipts for agent jobs over remote MCP
Verified doc corpora for agents: grep-first retrieval, hashed pages, Merkle+RFC-3161 receipts
Related MCP Servers
- AlicenseNot gradedqualityAmaintenanceProvides verifiable, append-only context handoffs for Codex and MCP clients, enabling memory persistence and resume across chat sessions.1MIT
- AlicenseNot gradedqualityAmaintenanceEnables semantic search and retrieval of archived AI conversations from multiple providers via the MCP protocol, allowing thread continuation and integration with MCP-capable tools.10MIT
- AlicenseNot gradedqualityAmaintenanceEnables searching local OpenAI Codex conversation history by keywords, project, date, or role, listing sessions, and retrieving exact supporting messages through MCP tools.1MIT
- FlicenseNot gradedqualityBmaintenanceEnables AI coding agents to persistently store and retrieve conversation history with hybrid semantic and keyword search, cross-encoder reranking, session filtering, and archiving through MCP tools.-