datalore-mcp
by sequico
README.md
# datalore-mcp
[](https://github.com/sequico/datalore-mcp/actions/workflows/ci.yml)
[](https://github.com/sequico/datalore-mcp/actions/workflows/codeql.yml)
**Shared, serverless, conflict-free memory for AI agents.**
`datalore-mcp` is a drop-in replacement for the official
[`@modelcontextprotocol/server-memory`](https://github.com/modelcontextprotocol/servers/tree/main/src/memory)
knowledge graph — same entities, observations and relations — but **multi-machine by design**:
several hosts keep their own copy and converge automatically, with **no server and no merge
conflicts**.
- **Conflict-free by construction.** State is an **LWW-Element-Set**: every mutation is an add or a
remove (tombstone) in an append-only operation log, and an element is present exactly when its
operation with the greatest `(timestamp, node, sequence)` is an add. Merges are idempotent,
commutative and associative, so any two nodes that have seen the same operations end up with the
same graph, in any order.
- **Per-node shards.** Each node writes only to its own log (`<node>.jsonl`) — it appends, and rewrites
it only when compacting; a shard has a single writer by design, so the transport itself cannot
produce conflicts.
- **Pluggable sync.** Shards live in a folder synced by Syncthing / Dropbox / git / a network mount,
or as objects in any S3-compatible store (AWS S3, Cloudflare R2, Contabo, MinIO). The backend is a
choice, not a lock-in.
- **Local-first.** With the `file` backend every node reads and writes its own shard on disk and sync
happens out of band, so it keeps working offline; the `s3` backend reads and writes the bucket
directly. A read folds all shards into the graph.
**A shared area is required.** datalore-mcp merges the shards; it does not move them between machines
by itself. Every node must point at the *same* shared area — a folder replicated by Syncthing /
Dropbox / git / a shared mount, or a shared object store — so each node's shard reaches the others
and the graph converges. Without a shared area each machine keeps only its own local graph.
## The problem
Agent memory today is local. The official memory server is a single JSONL file that every call
reads and rewrites in full, so:
- it lives on **one machine** — switch host and your agent forgets;
- sharing it over a synced folder (Dropbox, Syncthing, a network mount) causes **lost writes and
corruption**, because whole-file read-modify-write is not safe across writers;
- the "shared" alternatives either need a **server or a database** you have to host, or ship your
memory to a **cloud** — neither fits a private, serverless setup.
`datalore-mcp` fills that gap.
## Install
The package is on npm as **`datalore-mcp`**:
```sh
npx -y datalore-mcp
```
or install it globally with `npm install -g datalore-mcp`. You can also run it from a clone:
```sh
npm install
npm run build
node dist/index.js # the MCP server on stdio
```
or with Docker (below). Point any MCP client that runs a stdio server at the server command
(Node.js ≥ 22). In OpenCode (V2) a local server goes under `mcp.servers`:
```jsonc
{
"$schema": "https://opencode.ai/config.json",
"mcp": {
"servers": {
"memory": {
"type": "local",
"command": ["npx", "-y", "datalore-mcp"],
"environment": { "DATALORE_DIR": "~/.datalore" }
}
}
}
}
```
Other MCP clients wrap the same command and environment in their own envelope; the clone alternative
is `"command": ["node", "/path/to/datalore-mcp/dist/index.js"]`. Because the tool surface is
identical, you can replace the `memory` server entry with `datalore-mcp` and change nothing else.
### Docker
```sh
docker build -t datalore-mcp .
docker run --rm -i -v "$HOME/.datalore:/data" -e DATALORE_DIR=/data datalore-mcp
```
The image runs the stdio server; mount a directory for the shards.
## Configuration
| Variable | Default | Meaning |
| --- | --- | --- |
| `DATALORE_BACKEND` | `file` | `file`, `memory` or `s3`. |
| `DATALORE_NODE_ID` | hostname | This node's shard name. Must be **unique** per machine and written by a **single server at a time**. |
| `DATALORE_DIR` | `~/.datalore` | Directory holding the shards (file backend). A leading `~` is expanded. |
| `DATALORE_S3_BUCKET` | — | Bucket (required for the s3 backend). |
| `DATALORE_S3_PREFIX` | `""` | Key prefix all shard objects live under. |
| `DATALORE_S3_REGION` | `us-east-1` | Region passed to the S3 client. |
| `DATALORE_S3_ENDPOINT` | — | Custom endpoint for MinIO / R2 / Contabo and similar. |
| `DATALORE_S3_FORCE_PATH_STYLE` | on when an endpoint is set | Path-style requests. |
S3 credentials come from the standard AWS chain (`AWS_ACCESS_KEY_ID`, `AWS_SECRET_ACCESS_KEY`,
`AWS_REGION`, profiles, instance roles, …).
A node id is a **single-writer shard**: exactly one server process should append to it at a time.
The normal setup is one memory server per machine; if you run several servers against the same store
concurrently, give each its own `DATALORE_NODE_ID`. Two writers that share a node id produce
conflicting operations with the same id, and a read then **fails loudly** instead of silently
dropping one of them.
`datalore-mcp import` and `datalore-mcp compact` write the shard too, so run them while the node's
server is stopped. `datalore-mcp compact` checks up front that the shard still holds exactly what it
read before rewriting it, so a forgotten running writer makes it fail loudly instead of losing that
writer's operations.
Within a node the timestamp is **monotonic**: if the wall clock steps back (NTP, a restored
snapshot, a manual `date`), a new operation keeps the previous timestamp rather than an older one,
so an earlier operation can never win over a later one of the same node.
## Sync backends
The backend is how a node reaches the **shared area**; that shared area — not the backend — is what
makes the nodes converge. The `file` and `s3` backends assume the area already exists and is
replicated outside datalore.
- **`file`** — shards are `*.jsonl` files in `DATALORE_DIR`. Point it at a folder replicated by
Syncthing, Dropbox, a git working tree or a shared mount, and the copies converge.
- **`s3`** — one object per shard. Appends and rewrites are conditional writes on the object's
version, retried on contention, so even a shared node id cannot lose an append silently (the store
must support conditional writes). Uses `@aws-sdk/client-s3`, installed as an optional dependency
and loaded only when the s3 backend is selected.
- **`memory`** — volatile, in-process only; useful for tests and throwaway sessions.
## CLI
The package exposes a **single command**. With no subcommand it runs the MCP server — the entry every
MCP client uses — so `datalore-mcp` and `datalore-mcp serve` are the same. The subcommands are
maintenance:
```sh
datalore-mcp # run the MCP memory server over stdio (same as `serve`)
datalore-mcp serve # run the MCP memory server over stdio
datalore-mcp import <memory.jsonl> # migrate a file from the official server
datalore-mcp export # print the folded knowledge graph as JSON
datalore-mcp merge # fold every shard and report the merged state
datalore-mcp compact # rewrite this node's shard, dropping shadowed operations
datalore-mcp query <text> # search entities and print the matching subgraph
datalore-mcp help # list the commands (--version prints the version)
```
## How it works
Every mutation is recorded as an **operation** in an append-only log, split into per-node shards:
- `create_entities` / `add_observations` / `create_relations` add elements;
- `delete_entities` / `delete_observations` / `delete_relations` write tombstones (a deleted entity
also tombstones its observations and incident relations).
A read **folds** all shards: each element (entity, observation, relation) is an LWW register decided
by `(timestamp, node, sequence)`. Observations are shown while their entity is present, relations
while both endpoints are present. The fold depends only on the *set* of operations, never on the
order shards or lines were read.
Consequences, all covered by property tests:
- **Idempotent** — applying the same operations twice changes nothing.
- **Commutative** — order of operations across shards does not matter.
- **Associative** — folding shards in any grouping yields the same graph.
- **Convergent** — two nodes with the same operations produce the identical graph.
Ordering is deterministic: entities by name, relations by `(from, to, relationType)`, observations in
operation order (`timestamp`, `node`, `sequence`).
### Compaction
The log grows with every mutation. `datalore-mcp compact` rewrites this node's shard keeping, per
element, only the latest operation of each node. A dropped operation is always shadowed by a later
one of the same node for the same element — which wins in the total order — so the merged graph is
unchanged, tombstones included. The rewrite is atomic (a temporary file renamed over the shard on the
`file` backend, a conditional write on `s3`), so a crash midway can never truncate it. Only the local
shard is rewritten; every other shard is untouched, so compaction is conflict-free and other nodes
are unaffected.
## Drop-in compatibility
The same nine tools, so it slots under the `memory` server name with no agent changes:
`create_entities`, `create_relations`, `add_observations`, `delete_entities`, `delete_observations`,
`delete_relations`, `read_graph`, `search_nodes`, `open_nodes`.
The full graph is also exposed at the `memory://knowledge-graph` resource. The resource is
read-only; clients can subscribe to it (`resources/subscribe`) and receive
`notifications/resources/updated` whenever a mutation changes the graph.
Entities are keyed by name, and observations and relations are **sets**: adding the same observation
or relation twice has no effect. `delete_entities` also removes the entity's observations and its
incident relations. The tombstones cover what the deleting node could see: an observation added
concurrently on another node survives as an element, so if the entity is later created again that
observation reappears with it.
## Development
```sh
npm run check # Biome: lint + format + import order
npm run typecheck # tsc --noEmit
npm run test # Vitest, including the CRDT property tests
npm run build # tsc -> dist/
npm run gate # all of the above
```
## Why `datalore`?
The name is an homage to *Star Trek: The Next Generation*. **"Datalore"** (Season 1, Episode 13,
first aired 18 January 1988) is the episode that introduces **Lore** — the twin brother of
**Data**, two Soong-type androids built by Dr. Noonien Soong. *Data* is the good one; *Lore* is the
flawed, emotional, malicious prototype. The episode title is itself a **portmanteau of Data and
Lore**, and both are played by the same actor, Brent Spiner.
The names fit a knowledge graph almost too well:
- **Data** — a *datum*: a single fact.
- **Lore** — the body of *knowledge and tradition* shared by a community.
So *datalore* = the facts **plus** the shared lore: exactly what a shared agent memory is. And the
twin motif maps onto this project's core problem — many copies of the same graph, on different
machines, that must stay consistent. Data's good twin and Lore's evil twin are the two failure
modes; in `datalore-mcp` the copies cannot drift apart, because the merge is conflict-free by
construction. The graph has **no evil twin**.
A small production gem: the episode was first pitched as a romance for Data with a female android.
It was **Brent Spiner himself who suggested the evil-twin plot** instead — so the twin who gave the
project its name also gave the episode its twist.
References:
- Memory Alpha — *Datalore (episode)*: <https://memory-alpha.fandom.com/wiki/Datalore_(episode)>
- Memory Alpha — *Lore*: <https://memory-alpha.fandom.com/wiki/Lore>
- Wikipedia — *Datalore*: <https://en.wikipedia.org/wiki/Datalore>
*This project is an independent fan homage. "Star Trek", "Data" and "Lore" are the property of
their respective rights holders; this project is not affiliated with or endorsed by them.*
## License
[MIT](LICENSE).