Skip to main content
Glama
README.md
<p align="center">
  <img src="docs/hero.svg" alt="RecallLattice — local-first memory infrastructure for AI agents" width="100%">
</p>

<p align="center">
  <a href="https://www.python.org/"><img src="https://img.shields.io/badge/Python-3.10--3.13-3776AB?logo=python&logoColor=white" alt="Python 3.10 through 3.13"></a>
  <a href="https://modelcontextprotocol.io/"><img src="https://img.shields.io/badge/MCP-19_tools-c1ff2e" alt="19 MCP tools"></a>
  <img src="https://img.shields.io/badge/storage-SQLite_FTS5-5ee7ff" alt="SQLite FTS5">
  <img src="https://img.shields.io/badge/embeddings-none-ff8c42" alt="No embeddings">
  <img src="https://img.shields.io/badge/data-100%25_local-34d399" alt="100 percent local">
  <a href="LICENSE"><img src="https://img.shields.io/badge/license-MIT-a970ff" alt="MIT License"></a>
  <a href="https://github.com/Adam-ZS/RecallLattice/releases"><img src="https://img.shields.io/badge/version-3.3.0-c1ff2e" alt="Version 3.3.0"></a>
</p>

RecallLattice is a permanent local memory layer for AI agents. It turns durable facts into a searchable, typed knowledge lattice and returns only the context that fits the task and token budget.

No vector database. No embedding API. No hosted account. One portable SQLite file.

<p align="center">
  <img src="docs/dashboard.png" alt="RecallLattice 3D memory observatory" width="100%">
</p>

## What makes it different

| Capability | What RecallLattice does |
|---|---|
| **Budgeted Context Forge** | Packs ranked memories into a hard token ceiling instead of flooding the model context. |
| **Hybrid Recall** | Combines exact phrases, terms, stems, prefixes, typo similarity, tags, importance, strength, and pins. |
| **Graph Neighborhoods** | Traverses one to three hops around any memory and preserves typed relationships. |
| **Batch Capture** | Saves up to 100 facts in one MCP call while applying normal categorization, tags, deduplication, and auto-linking. |
| **Suggested Connections** | Finds useful missing graph edges without silently mutating the lattice. |
| **Smart Deduplication** | Merges exact duplicate facts, unions tags, and keeps the highest importance value. |
| **Progressive Disclosure** | Compact previews are the default; full records are retrieved only when needed. |
| **Memory Timeline** | Joins activity with memory previews for a chronological, inspectable history. |
| **Spaced Review** | Importance-aware decay, review queues, strength tracking, pins, streaks, and XP. |
| **3D Observatory** | Draggable graph, fullscreen neighborhoods, context lab, command palette, timeline, heatmap, and live health. |

## Start in sixty seconds

```bash
git clone https://github.com/Adam-ZS/RecallLattice.git
cd RecallLattice
./scripts/install.sh
```

Open the private dashboard:

```text
http://127.0.0.1:8799
```

The installer creates `~/.recall-lattice`, builds a persistent virtual environment, preserves an existing database, installs the MCP/dashboard dependencies, and enables the localhost-only dashboard service.

## The smart memory pipeline

<p align="center">
  <img src="docs/context-pipeline.svg" alt="RecallLattice capture and context pipeline" width="100%">
</p>

### Save more signal

A batch passes through one consistent path:

1. Normalize each fact.
2. Infer a category when none is supplied.
3. Extract useful technical tags.
4. Merge exact duplicates instead of creating noise.
5. Find related memories and create typed similarity edges.
6. Store the canonical record in SQLite and synchronize FTS5.

### Spend fewer tokens

A context request follows a separate retrieval path:

1. Score exact, lexical, stemmed, prefix, fuzzy, and tag matches.
2. Blend relevance with pin status, importance, and memory strength.
3. Optionally add one-hop graph context.
4. Deduplicate records.
5. Fit the strongest content into an exact token budget.

```python
import recall_lattice as brain

pack = brain.context_pack(
    "How does the deployment pipeline work?",
    token_budget=768,
    include_related=True,
)

print(pack["estimated_tokens"], pack["memories"])
```

## Architecture

<p align="center">
  <img src="docs/architecture.svg" alt="RecallLattice system architecture" width="100%">
</p>

The standard-library core owns SQLite, FTS5, ranking, graph traversal, decay, and deduplication. MCP and FastAPI are optional interfaces around that same core, so dashboard writes and agent writes behave identically.

## MCP configuration

Install with the `all` extra, or use `./scripts/install.sh`:

```bash
python -m pip install 'recall-lattice-memory[all]'
```

Add the server to an MCP-compatible client:

```json
{
  "mcpServers": {
    "recall_lattice": {
      "command": "/home/YOU/.recall-lattice/.venv/bin/python",
      "args": ["/home/YOU/.recall-lattice/recall_lattice_mcp.py"]
    }
  }
}
```

Hermes Agent YAML:

```yaml
mcp_servers:
  recall_lattice:
    command: /home/YOU/.recall-lattice/.venv/bin/python
    args:
      - /home/YOU/.recall-lattice/recall_lattice_mcp.py
    enabled: true
```

Restart the client after changing its MCP configuration.

## Nineteen MCP tools

### Capture and retrieval

| Tool | Purpose |
|---|---|
| `recall_lattice_store` | Store one durable fact with auto-category, tags, deduplication, and links. |
| `recall_lattice_store_batch` | Store up to 100 facts in one call. |
| `recall_lattice_recall` | Hybrid ranked recall with compact/full modes and optional graph context. |
| `recall_lattice_context_pack` | Produce a ranked context bundle inside a strict token budget. |
| `recall_lattice_get` | Retrieve one complete memory and its direct links. |

### Graph intelligence

| Tool | Purpose |
|---|---|
| `recall_lattice_neighborhood` | Traverse one to three graph hops around a memory. |
| `recall_lattice_suggest_links` | Preview useful missing associations without mutation. |
| `recall_lattice_link` | Create a typed association. |
| `recall_lattice_graph` | Export all graph nodes and edges. |

### Memory lifecycle

| Tool | Purpose |
|---|---|
| `recall_lattice_pin` | Protect and prioritize a critical memory. |
| `recall_lattice_unpin` | Remove pin priority. |
| `recall_lattice_review` | Refresh strength through spaced review. |
| `recall_lattice_review_due` | Find memories that need reinforcement. |
| `recall_lattice_forget` | Delete one memory by ID. |
| `recall_lattice_prune` | Preview or remove stale low-value records. |

### Observability and portability

| Tool | Purpose |
|---|---|
| `recall_lattice_stats` | Return compact health, graph, category, and XP metrics. |
| `recall_lattice_timeline` | Return activity joined with memory previews. |
| `recall_lattice_heatmap` | Aggregate activity over time. |
| `recall_lattice_export` | Export portable memory records. |

## MCP examples

### Capture a whole session efficiently

```json
{
  "items": [
    {"content": "Project Aurora uses FastAPI", "category": "project", "importance": 8},
    {"content": "Production deploys require a signed tag", "category": "procedure", "importance": 9},
    "The dashboard binds to localhost"
  ],
  "source": "session-summary"
}
```

### Forge task-specific context

```json
{
  "query": "Aurora production deployment",
  "token_budget": 640,
  "include_related": true
}
```

### Inspect a memory neighborhood

```json
{
  "memory_id": 42,
  "depth": 2,
  "limit": 80
}
```

## The 3D memory observatory

The dashboard is an operational surface, not a decorative landing page.

- **Context Forge** — search, set a token ceiling, preview the exact packed memories, then copy.
- **Interactive brain** — drag to rotate; double-click or press `EXPAND` for fullscreen.
- **Neighborhood focus** — open a memory and jump directly into its two-hop subgraph.
- **Connection suggestions** — inspect likely missing edges from the memory modal.
- **Batch capture** — paste one fact per line and store the entire set in one operation.
- **Neural timeline** — watch stores, recalls, reviews, links, and merges chronologically.
- **Command palette** — press `Ctrl/⌘ + K` to capture, forge, explore, export, review, or open a random memory.
- **Keyboard navigation** — `N` new memory, `F` Context Forge, `G` fullscreen graph, `Esc` close.
- **Pointer depth** — memory cards respond spatially to cursor position.
- **Reduced motion** — respects the operating-system preference.

The server binds to `127.0.0.1` by default because the dashboard has no authentication.

## Python API

```python
import recall_lattice as brain

saved = brain.store(
    "RecallLattice keeps its canonical store in SQLite",
    category="system",
    tags=["SQLite", "local-first"],
    importance=9,
)

matches = brain.recall("local canonical memory", limit=5)
neighbors = brain.neighborhood(saved["id"], depth=2)
suggestions = brain.suggest_links(saved["id"], limit=5)
timeline = brain.timeline(days=30, limit=100)
```

Use another database without editing source:

```bash
export RECALL_LATTICE_DB="$HOME/my-brain/memory.db"
```

## REST API

| Method | Endpoint | Purpose |
|---|---|---|
| `GET` | `/api/search?q=...` | Hybrid ranked search. |
| `GET` | `/api/context-pack?q=...&token_budget=800` | Budgeted context bundle. |
| `GET` | `/api/neighborhood/{id}?depth=2` | Typed graph neighborhood. |
| `GET` | `/api/suggest-links/{id}` | Missing-link candidates. |
| `GET` | `/api/timeline?days=30` | Activity with memory previews. |
| `POST` | `/api/store` | Store one form-encoded memory. |
| `POST` | `/api/store-batch` | Store a JSON array supplied in the `items` form field. |
| `GET` | `/api/graph` | Complete graph. |
| `GET` | `/api/stats` | Health and usage metrics. |
| `GET` | `/api/export` | Portable JSON export. |

Interactive endpoint documentation is available at `http://127.0.0.1:8799/docs`.

## Token efficiency

RecallLattice uses three layers to control context cost:

1. Compact MCP serialization omits duplicated bookkeeping.
2. Progressive disclosure keeps full content behind explicit retrieval.
3. Context Forge enforces a caller-selected token ceiling.

A measured three-result compact recall reduced serialized output from 4,566 characters to 1,229 characters: **73.1% less context** before applying a Context Forge budget.

## Durability and privacy

- SQLite runs in WAL mode.
- FTS5 synchronization is maintained by database triggers.
- Schema upgrades rebuild indexes from the canonical memories table.
- The database is a single portable file.
- Dashboard and API default to localhost.
- No telemetry, hosted database, embedding provider, or API key is required.
- Database files, exports, backups, and environment files are ignored by Git.

Backup manually:

```bash
cp ~/.recall-lattice/recall-lattice.db \
   ~/.recall-lattice/recall-lattice.db.backup-$(date +%Y%m%d)
```

## Test and package

```bash
python -m unittest discover -s tests -v
python -m py_compile recall_lattice.py recall_lattice_protocol.py recall_lattice_mcp.py server.py
python -m pip wheel . --no-deps -w dist
```

The suite covers storage, exact deduplication, automatic links, FTS synchronization, hybrid ranking, typo tolerance, tag filters, access metadata, compact serialization, token-budget enforcement, batch capture, graph traversal, timeline previews, and all MCP schemas.

## AI-agent setup

[`llms.txt`](llms.txt) gives coding agents a compact machine-readable installation guide, architecture summary, tool inventory, and operational constraints.

## Contributing

Read [CONTRIBUTING.md](CONTRIBUTING.md), open a focused issue, and submit changes through a pull request. `main` is protected from deletion and force pushes and requires linear history and resolved review conversations.

## License

MIT — see [LICENSE](LICENSE).