Skip to main content
Glama

RecallLattice is a permanent local memory layer for AI agents. It turns durable facts into a searchable, typed knowledge lattice and returns only the context that fits the task and token budget.

No vector database. No embedding API. No hosted account. One portable SQLite file.

What makes it different

Capability

What RecallLattice does

Budgeted Context Forge

Packs ranked memories into a hard token ceiling instead of flooding the model context.

Hybrid Recall

Combines exact phrases, terms, stems, prefixes, typo similarity, tags, importance, strength, and pins.

Graph Neighborhoods

Traverses one to three hops around any memory and preserves typed relationships.

Batch Capture

Saves up to 100 facts in one MCP call while applying normal categorization, tags, deduplication, and auto-linking.

Suggested Connections

Finds useful missing graph edges without silently mutating the lattice.

Smart Deduplication

Merges exact duplicate facts, unions tags, and keeps the highest importance value.

Progressive Disclosure

Compact previews are the default; full records are retrieved only when needed.

Memory Timeline

Joins activity with memory previews for a chronological, inspectable history.

Spaced Review

Importance-aware decay, review queues, strength tracking, pins, streaks, and XP.

3D Observatory

Draggable graph, fullscreen neighborhoods, context lab, command palette, timeline, heatmap, and live health.

Related MCP server: Cortex

Start in sixty seconds

git clone https://github.com/Adam-ZS/RecallLattice.git
cd RecallLattice
./scripts/install.sh

Open the private dashboard:

http://127.0.0.1:8799

The installer creates ~/.recall-lattice, builds a persistent virtual environment, preserves an existing database, installs the MCP/dashboard dependencies, and enables the localhost-only dashboard service.

The smart memory pipeline

Save more signal

A batch passes through one consistent path:

  1. Normalize each fact.

  2. Infer a category when none is supplied.

  3. Extract useful technical tags.

  4. Merge exact duplicates instead of creating noise.

  5. Find related memories and create typed similarity edges.

  6. Store the canonical record in SQLite and synchronize FTS5.

Spend fewer tokens

A context request follows a separate retrieval path:

  1. Score exact, lexical, stemmed, prefix, fuzzy, and tag matches.

  2. Blend relevance with pin status, importance, and memory strength.

  3. Optionally add one-hop graph context.

  4. Deduplicate records.

  5. Fit the strongest content into an exact token budget.

import recall_lattice as brain

pack = brain.context_pack(
    "How does the deployment pipeline work?",
    token_budget=768,
    include_related=True,
)

print(pack["estimated_tokens"], pack["memories"])

Architecture

The standard-library core owns SQLite, FTS5, ranking, graph traversal, decay, and deduplication. MCP and FastAPI are optional interfaces around that same core, so dashboard writes and agent writes behave identically.

MCP configuration

Install with the all extra, or use ./scripts/install.sh:

python -m pip install 'recall-lattice-memory[all]'

Add the server to an MCP-compatible client:

{
  "mcpServers": {
    "recall_lattice": {
      "command": "/home/YOU/.recall-lattice/.venv/bin/python",
      "args": ["/home/YOU/.recall-lattice/recall_lattice_mcp.py"]
    }
  }
}

Hermes Agent YAML:

mcp_servers:
  recall_lattice:
    command: /home/YOU/.recall-lattice/.venv/bin/python
    args:
      - /home/YOU/.recall-lattice/recall_lattice_mcp.py
    enabled: true

Restart the client after changing its MCP configuration.

Nineteen MCP tools

Capture and retrieval

Tool

Purpose

recall_lattice_store

Store one durable fact with auto-category, tags, deduplication, and links.

recall_lattice_store_batch

Store up to 100 facts in one call.

recall_lattice_recall

Hybrid ranked recall with compact/full modes and optional graph context.

recall_lattice_context_pack

Produce a ranked context bundle inside a strict token budget.

recall_lattice_get

Retrieve one complete memory and its direct links.

Graph intelligence

Tool

Purpose

recall_lattice_neighborhood

Traverse one to three graph hops around a memory.

recall_lattice_suggest_links

Preview useful missing associations without mutation.

recall_lattice_link

Create a typed association.

recall_lattice_graph

Export all graph nodes and edges.

Memory lifecycle

Tool

Purpose

recall_lattice_pin

Protect and prioritize a critical memory.

recall_lattice_unpin

Remove pin priority.

recall_lattice_review

Refresh strength through spaced review.

recall_lattice_review_due

Find memories that need reinforcement.

recall_lattice_forget

Delete one memory by ID.

recall_lattice_prune

Preview or remove stale low-value records.

Observability and portability

Tool

Purpose

recall_lattice_stats

Return compact health, graph, category, and XP metrics.

recall_lattice_timeline

Return activity joined with memory previews.

recall_lattice_heatmap

Aggregate activity over time.

recall_lattice_export

Export portable memory records.

MCP examples

Capture a whole session efficiently

{
  "items": [
    {"content": "Project Aurora uses FastAPI", "category": "project", "importance": 8},
    {"content": "Production deploys require a signed tag", "category": "procedure", "importance": 9},
    "The dashboard binds to localhost"
  ],
  "source": "session-summary"
}

Forge task-specific context

{
  "query": "Aurora production deployment",
  "token_budget": 640,
  "include_related": true
}

Inspect a memory neighborhood

{
  "memory_id": 42,
  "depth": 2,
  "limit": 80
}

The 3D memory observatory

The dashboard is an operational surface, not a decorative landing page.

  • Context Forge — search, set a token ceiling, preview the exact packed memories, then copy.

  • Interactive brain — drag to rotate; double-click or press EXPAND for fullscreen.

  • Neighborhood focus — open a memory and jump directly into its two-hop subgraph.

  • Connection suggestions — inspect likely missing edges from the memory modal.

  • Batch capture — paste one fact per line and store the entire set in one operation.

  • Neural timeline — watch stores, recalls, reviews, links, and merges chronologically.

  • Command palette — press Ctrl/⌘ + K to capture, forge, explore, export, review, or open a random memory.

  • Keyboard navigation — N new memory, F Context Forge, G fullscreen graph, Esc close.

  • Pointer depth — memory cards respond spatially to cursor position.

  • Reduced motion — respects the operating-system preference.

The server binds to 127.0.0.1 by default because the dashboard has no authentication.

Python API

import recall_lattice as brain

saved = brain.store(
    "RecallLattice keeps its canonical store in SQLite",
    category="system",
    tags=["SQLite", "local-first"],
    importance=9,
)

matches = brain.recall("local canonical memory", limit=5)
neighbors = brain.neighborhood(saved["id"], depth=2)
suggestions = brain.suggest_links(saved["id"], limit=5)
timeline = brain.timeline(days=30, limit=100)

Use another database without editing source:

export RECALL_LATTICE_DB="$HOME/my-brain/memory.db"

REST API

Method

Endpoint

Purpose

GET

/api/search?q=...

Hybrid ranked search.

GET

/api/context-pack?q=...&token_budget=800

Budgeted context bundle.

GET

/api/neighborhood/{id}?depth=2

Typed graph neighborhood.

GET

/api/suggest-links/{id}

Missing-link candidates.

GET

/api/timeline?days=30

Activity with memory previews.

POST

/api/store

Store one form-encoded memory.

POST

/api/store-batch

Store a JSON array supplied in the items form field.

GET

/api/graph

Complete graph.

GET

/api/stats

Health and usage metrics.

GET

/api/export

Portable JSON export.

Interactive endpoint documentation is available at http://127.0.0.1:8799/docs.

Token efficiency

RecallLattice uses three layers to control context cost:

  1. Compact MCP serialization omits duplicated bookkeeping.

  2. Progressive disclosure keeps full content behind explicit retrieval.

  3. Context Forge enforces a caller-selected token ceiling.

A measured three-result compact recall reduced serialized output from 4,566 characters to 1,229 characters: 73.1% less context before applying a Context Forge budget.

Durability and privacy

  • SQLite runs in WAL mode.

  • FTS5 synchronization is maintained by database triggers.

  • Schema upgrades rebuild indexes from the canonical memories table.

  • The database is a single portable file.

  • Dashboard and API default to localhost.

  • No telemetry, hosted database, embedding provider, or API key is required.

  • Database files, exports, backups, and environment files are ignored by Git.

Backup manually:

cp ~/.recall-lattice/recall-lattice.db \
   ~/.recall-lattice/recall-lattice.db.backup-$(date +%Y%m%d)

Test and package

python -m unittest discover -s tests -v
python -m py_compile recall_lattice.py recall_lattice_protocol.py recall_lattice_mcp.py server.py
python -m pip wheel . --no-deps -w dist

The suite covers storage, exact deduplication, automatic links, FTS synchronization, hybrid ranking, typo tolerance, tag filters, access metadata, compact serialization, token-budget enforcement, batch capture, graph traversal, timeline previews, and all MCP schemas.

AI-agent setup

llms.txt gives coding agents a compact machine-readable installation guide, architecture summary, tool inventory, and operational constraints.

Contributing

Read CONTRIBUTING.md, open a focused issue, and submit changes through a pull request. main is protected from deletion and force pushes and requires linear history and resolved review conversations.

License

MIT — see LICENSE.

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    A
    maintenance
    Provides persistent, searchable memory for MCP-compatible agents, enabling recall by meaning, automatic decay, trust scoring, and cross-agent handoffs.
    5
    MIT
  • A
    license
    Not graded
    quality
    D
    maintenance
    Local-first AI memory layer with hybrid retrieval and brain-inspired namespaces. Enables agents to save, search, and manage memories directly via MCP tools.
    3 npm
    MIT
  • A
    license
    A
    quality
    C
    maintenance
    Enables AI agents to share a portable, user-level memory layer through MCP, allowing them to store, search, update, link, and consolidate facts with optional full-text and vector retrieval.
    7
    14 npm
    MIT
  • A
    license
    Not graded
    quality
    B
    maintenance
    Enables AI agents to store and semantically retrieve durable memories across sessions via MCP or REST, with tools for remembering, recalling, asking, updating, and forgetting memories.
    18 npm
    MIT