Skip to main content
Glama
sequico
by sequico
README.md
# datalore-mcp

[![CI](https://github.com/sequico/datalore-mcp/actions/workflows/ci.yml/badge.svg)](https://github.com/sequico/datalore-mcp/actions/workflows/ci.yml)
[![CodeQL](https://github.com/sequico/datalore-mcp/actions/workflows/codeql.yml/badge.svg)](https://github.com/sequico/datalore-mcp/actions/workflows/codeql.yml)

**Shared, serverless, conflict-free memory for AI agents.**

`datalore-mcp` is a drop-in replacement for the official
[`@modelcontextprotocol/server-memory`](https://github.com/modelcontextprotocol/servers/tree/main/src/memory)
knowledge graph — same entities, observations and relations — but **multi-machine by design**:
several hosts keep their own copy and converge automatically, with **no server and no merge
conflicts**.

- **Conflict-free by construction.** State is an **LWW-Element-Set**: every mutation is an add or a
  remove (tombstone) in an append-only operation log, and an element is present exactly when its
  operation with the greatest `(timestamp, node, sequence)` is an add. Merges are idempotent,
  commutative and associative, so any two nodes that have seen the same operations end up with the
  same graph, in any order.
- **Per-node shards.** Each node writes only to its own log (`<node>.jsonl`) — it appends, and rewrites
  it only when compacting; a shard has a single writer by design, so the transport itself cannot
  produce conflicts.
- **Pluggable sync.** Shards live in a folder synced by Syncthing / Dropbox / git / a network mount,
  or as objects in any S3-compatible store (AWS S3, Cloudflare R2, Contabo, MinIO). The backend is a
  choice, not a lock-in.
- **Local-first.** With the `file` backend every node reads and writes its own shard on disk and sync
  happens out of band, so it keeps working offline; the `s3` backend reads and writes the bucket
  directly. A read folds all shards into the graph.

**A shared area is required.** datalore-mcp merges the shards; it does not move them between machines
by itself. Every node must point at the *same* shared area — a folder replicated by Syncthing /
Dropbox / git / a shared mount, or a shared object store — so each node's shard reaches the others
and the graph converges. Without a shared area each machine keeps only its own local graph.

## The problem

Agent memory today is local. The official memory server is a single JSONL file that every call
reads and rewrites in full, so:

- it lives on **one machine** — switch host and your agent forgets;
- sharing it over a synced folder (Dropbox, Syncthing, a network mount) causes **lost writes and
  corruption**, because whole-file read-modify-write is not safe across writers;
- the "shared" alternatives either need a **server or a database** you have to host, or ship your
  memory to a **cloud** — neither fits a private, serverless setup.

`datalore-mcp` fills that gap.

## Install

The package is on npm as **`datalore-mcp`**:

```sh
npx -y datalore-mcp
```

or install it globally with `npm install -g datalore-mcp`. You can also run it from a clone:

```sh
npm install
npm run build
node dist/index.js          # the MCP server on stdio
```

or with Docker (below). Point any MCP client that runs a stdio server at the server command
(Node.js ≥ 22). In OpenCode (V2) a local server goes under `mcp.servers`:

```jsonc
{
  "$schema": "https://opencode.ai/config.json",
  "mcp": {
    "servers": {
      "memory": {
        "type": "local",
        "command": ["npx", "-y", "datalore-mcp"],
        "environment": { "DATALORE_DIR": "~/.datalore" }
      }
    }
  }
}
```

Other MCP clients wrap the same command and environment in their own envelope; the clone alternative
is `"command": ["node", "/path/to/datalore-mcp/dist/index.js"]`. Because the tool surface is
identical, you can replace the `memory` server entry with `datalore-mcp` and change nothing else.

### Docker

```sh
docker build -t datalore-mcp .
docker run --rm -i -v "$HOME/.datalore:/data" -e DATALORE_DIR=/data datalore-mcp
```

The image runs the stdio server; mount a directory for the shards.

## Configuration

| Variable | Default | Meaning |
| --- | --- | --- |
| `DATALORE_BACKEND` | `file` | `file`, `memory` or `s3`. |
| `DATALORE_NODE_ID` | hostname | This node's shard name. Must be **unique** per machine and written by a **single server at a time**. |
| `DATALORE_DIR` | `~/.datalore` | Directory holding the shards (file backend). A leading `~` is expanded. |
| `DATALORE_S3_BUCKET` | — | Bucket (required for the s3 backend). |
| `DATALORE_S3_PREFIX` | `""` | Key prefix all shard objects live under. |
| `DATALORE_S3_REGION` | `us-east-1` | Region passed to the S3 client. |
| `DATALORE_S3_ENDPOINT` | — | Custom endpoint for MinIO / R2 / Contabo and similar. |
| `DATALORE_S3_FORCE_PATH_STYLE` | on when an endpoint is set | Path-style requests. |

S3 credentials come from the standard AWS chain (`AWS_ACCESS_KEY_ID`, `AWS_SECRET_ACCESS_KEY`,
`AWS_REGION`, profiles, instance roles, …).

A node id is a **single-writer shard**: exactly one server process should append to it at a time.
The normal setup is one memory server per machine; if you run several servers against the same store
concurrently, give each its own `DATALORE_NODE_ID`. Two writers that share a node id produce
conflicting operations with the same id, and a read then **fails loudly** instead of silently
dropping one of them.

`datalore-mcp import` and `datalore-mcp compact` write the shard too, so run them while the node's
server is stopped. `datalore-mcp compact` checks up front that the shard still holds exactly what it
read before rewriting it, so a forgotten running writer makes it fail loudly instead of losing that
writer's operations.

Within a node the timestamp is **monotonic**: if the wall clock steps back (NTP, a restored
snapshot, a manual `date`), a new operation keeps the previous timestamp rather than an older one,
so an earlier operation can never win over a later one of the same node.

## Sync backends

The backend is how a node reaches the **shared area**; that shared area — not the backend — is what
makes the nodes converge. The `file` and `s3` backends assume the area already exists and is
replicated outside datalore.

- **`file`** — shards are `*.jsonl` files in `DATALORE_DIR`. Point it at a folder replicated by
  Syncthing, Dropbox, a git working tree or a shared mount, and the copies converge.
- **`s3`** — one object per shard. Appends and rewrites are conditional writes on the object's
  version, retried on contention, so even a shared node id cannot lose an append silently (the store
  must support conditional writes). Uses `@aws-sdk/client-s3`, installed as an optional dependency
  and loaded only when the s3 backend is selected.
- **`memory`** — volatile, in-process only; useful for tests and throwaway sessions.

## CLI

The package exposes a **single command**. With no subcommand it runs the MCP server — the entry every
MCP client uses — so `datalore-mcp` and `datalore-mcp serve` are the same. The subcommands are
maintenance:

```sh
datalore-mcp                   # run the MCP memory server over stdio (same as `serve`)
datalore-mcp serve             # run the MCP memory server over stdio
datalore-mcp import <memory.jsonl> # migrate a file from the official server
datalore-mcp export            # print the folded knowledge graph as JSON
datalore-mcp merge             # fold every shard and report the merged state
datalore-mcp compact           # rewrite this node's shard, dropping shadowed operations
datalore-mcp query <text>      # search entities and print the matching subgraph
datalore-mcp help              # list the commands (--version prints the version)
```

## How it works

Every mutation is recorded as an **operation** in an append-only log, split into per-node shards:

- `create_entities` / `add_observations` / `create_relations` add elements;
- `delete_entities` / `delete_observations` / `delete_relations` write tombstones (a deleted entity
  also tombstones its observations and incident relations).

A read **folds** all shards: each element (entity, observation, relation) is an LWW register decided
by `(timestamp, node, sequence)`. Observations are shown while their entity is present, relations
while both endpoints are present. The fold depends only on the *set* of operations, never on the
order shards or lines were read.

Consequences, all covered by property tests:

- **Idempotent** — applying the same operations twice changes nothing.
- **Commutative** — order of operations across shards does not matter.
- **Associative** — folding shards in any grouping yields the same graph.
- **Convergent** — two nodes with the same operations produce the identical graph.

Ordering is deterministic: entities by name, relations by `(from, to, relationType)`, observations in
operation order (`timestamp`, `node`, `sequence`).

### Compaction

The log grows with every mutation. `datalore-mcp compact` rewrites this node's shard keeping, per
element, only the latest operation of each node. A dropped operation is always shadowed by a later
one of the same node for the same element — which wins in the total order — so the merged graph is
unchanged, tombstones included. The rewrite is atomic (a temporary file renamed over the shard on the
`file` backend, a conditional write on `s3`), so a crash midway can never truncate it. Only the local
shard is rewritten; every other shard is untouched, so compaction is conflict-free and other nodes
are unaffected.

## Drop-in compatibility

The same nine tools, so it slots under the `memory` server name with no agent changes:

`create_entities`, `create_relations`, `add_observations`, `delete_entities`, `delete_observations`,
`delete_relations`, `read_graph`, `search_nodes`, `open_nodes`.

The full graph is also exposed at the `memory://knowledge-graph` resource. The resource is
read-only; clients can subscribe to it (`resources/subscribe`) and receive
`notifications/resources/updated` whenever a mutation changes the graph.

Entities are keyed by name, and observations and relations are **sets**: adding the same observation
or relation twice has no effect. `delete_entities` also removes the entity's observations and its
incident relations. The tombstones cover what the deleting node could see: an observation added
concurrently on another node survives as an element, so if the entity is later created again that
observation reappears with it.

## Development

```sh
npm run check       # Biome: lint + format + import order
npm run typecheck   # tsc --noEmit
npm run test        # Vitest, including the CRDT property tests
npm run build       # tsc -> dist/
npm run gate        # all of the above
```

## Why `datalore`?

The name is an homage to *Star Trek: The Next Generation*. **"Datalore"** (Season 1, Episode 13,
first aired 18 January 1988) is the episode that introduces **Lore** — the twin brother of
**Data**, two Soong-type androids built by Dr. Noonien Soong. *Data* is the good one; *Lore* is the
flawed, emotional, malicious prototype. The episode title is itself a **portmanteau of Data and
Lore**, and both are played by the same actor, Brent Spiner.

The names fit a knowledge graph almost too well:

- **Data** — a *datum*: a single fact.
- **Lore** — the body of *knowledge and tradition* shared by a community.

So *datalore* = the facts **plus** the shared lore: exactly what a shared agent memory is. And the
twin motif maps onto this project's core problem — many copies of the same graph, on different
machines, that must stay consistent. Data's good twin and Lore's evil twin are the two failure
modes; in `datalore-mcp` the copies cannot drift apart, because the merge is conflict-free by
construction. The graph has **no evil twin**.

A small production gem: the episode was first pitched as a romance for Data with a female android.
It was **Brent Spiner himself who suggested the evil-twin plot** instead — so the twin who gave the
project its name also gave the episode its twist.

References:
- Memory Alpha — *Datalore (episode)*: <https://memory-alpha.fandom.com/wiki/Datalore_(episode)>
- Memory Alpha — *Lore*: <https://memory-alpha.fandom.com/wiki/Lore>
- Wikipedia — *Datalore*: <https://en.wikipedia.org/wiki/Datalore>

*This project is an independent fan homage. "Star Trek", "Data" and "Lore" are the property of
their respective rights holders; this project is not affiliated with or endorsed by them.*

## License

[MIT](LICENSE).