Skip to main content
Glama
sequico
by sequico

datalore-mcp

CI CodeQL

Shared, serverless, conflict-free memory for AI agents.

datalore-mcp is a drop-in replacement for the official @modelcontextprotocol/server-memory knowledge graph — same entities, observations and relations — but multi-machine by design: several hosts keep their own copy and converge automatically, with no server and no merge conflicts.

  • Conflict-free by construction. State is an LWW-Element-Set: every mutation is an add or a remove (tombstone) in an append-only operation log, and an element is present exactly when its operation with the greatest (timestamp, node, sequence) is an add. Merges are idempotent, commutative and associative, so any two nodes that have seen the same operations end up with the same graph, in any order.

  • Per-node shards. Each node writes only to its own log (<node>.jsonl) — it appends, and rewrites it only when compacting; a shard has a single writer by design, so the transport itself cannot produce conflicts.

  • Pluggable sync. Shards live in a folder synced by Syncthing / Dropbox / git / a network mount, or as objects in any S3-compatible store (AWS S3, Cloudflare R2, Contabo, MinIO). The backend is a choice, not a lock-in.

  • Local-first. With the file backend every node reads and writes its own shard on disk and sync happens out of band, so it keeps working offline; the s3 backend reads and writes the bucket directly. A read folds all shards into the graph.

A shared area is required. datalore-mcp merges the shards; it does not move them between machines by itself. Every node must point at the same shared area — a folder replicated by Syncthing / Dropbox / git / a shared mount, or a shared object store — so each node's shard reaches the others and the graph converges. Without a shared area each machine keeps only its own local graph.

The problem

Agent memory today is local. The official memory server is a single JSONL file that every call reads and rewrites in full, so:

  • it lives on one machine — switch host and your agent forgets;

  • sharing it over a synced folder (Dropbox, Syncthing, a network mount) causes lost writes and corruption, because whole-file read-modify-write is not safe across writers;

  • the "shared" alternatives either need a server or a database you have to host, or ship your memory to a cloud — neither fits a private, serverless setup.

datalore-mcp fills that gap.

Related MCP server: mnemosis-mcp

Install

The package is on npm as datalore-mcp:

npx -y datalore-mcp

or install it globally with npm install -g datalore-mcp. You can also run it from a clone:

npm install
npm run build
node dist/index.js          # the MCP server on stdio

or with Docker (below). Point any MCP client that runs a stdio server at the server command (Node.js ≥ 22). In OpenCode (V2) a local server goes under mcp.servers:

{
  "$schema": "https://opencode.ai/config.json",
  "mcp": {
    "servers": {
      "memory": {
        "type": "local",
        "command": ["npx", "-y", "datalore-mcp"],
        "environment": { "DATALORE_DIR": "~/.datalore" }
      }
    }
  }
}

Other MCP clients wrap the same command and environment in their own envelope; the clone alternative is "command": ["node", "/path/to/datalore-mcp/dist/index.js"]. Because the tool surface is identical, you can replace the memory server entry with datalore-mcp and change nothing else.

Docker

docker build -t datalore-mcp .
docker run --rm -i -v "$HOME/.datalore:/data" -e DATALORE_DIR=/data datalore-mcp

The image runs the stdio server; mount a directory for the shards.

Configuration

Variable

Default

Meaning

DATALORE_BACKEND

file

file, memory or s3.

DATALORE_NODE_ID

hostname

This node's shard name. Must be unique per machine and written by a single server at a time.

DATALORE_DIR

~/.datalore

Directory holding the shards (file backend). A leading ~ is expanded.

DATALORE_S3_BUCKET

—

Bucket (required for the s3 backend).

DATALORE_S3_PREFIX

""

Key prefix all shard objects live under.

DATALORE_S3_REGION

us-east-1

Region passed to the S3 client.

DATALORE_S3_ENDPOINT

—

Custom endpoint for MinIO / R2 / Contabo and similar.

DATALORE_S3_FORCE_PATH_STYLE

on when an endpoint is set

Path-style requests.

S3 credentials come from the standard AWS chain (AWS_ACCESS_KEY_ID, AWS_SECRET_ACCESS_KEY, AWS_REGION, profiles, instance roles, …).

A node id is a single-writer shard: exactly one server process should append to it at a time. The normal setup is one memory server per machine; if you run several servers against the same store concurrently, give each its own DATALORE_NODE_ID. Two writers that share a node id produce conflicting operations with the same id, and a read then fails loudly instead of silently dropping one of them.

datalore-mcp import and datalore-mcp compact write the shard too, so run them while the node's server is stopped. datalore-mcp compact checks up front that the shard still holds exactly what it read before rewriting it, so a forgotten running writer makes it fail loudly instead of losing that writer's operations.

Within a node the timestamp is monotonic: if the wall clock steps back (NTP, a restored snapshot, a manual date), a new operation keeps the previous timestamp rather than an older one, so an earlier operation can never win over a later one of the same node.

Sync backends

The backend is how a node reaches the shared area; that shared area — not the backend — is what makes the nodes converge. The file and s3 backends assume the area already exists and is replicated outside datalore.

  • file — shards are *.jsonl files in DATALORE_DIR. Point it at a folder replicated by Syncthing, Dropbox, a git working tree or a shared mount, and the copies converge.

  • s3 — one object per shard. Appends and rewrites are conditional writes on the object's version, retried on contention, so even a shared node id cannot lose an append silently (the store must support conditional writes). Uses @aws-sdk/client-s3, installed as an optional dependency and loaded only when the s3 backend is selected.

  • memory — volatile, in-process only; useful for tests and throwaway sessions.

CLI

The package exposes a single command. With no subcommand it runs the MCP server — the entry every MCP client uses — so datalore-mcp and datalore-mcp serve are the same. The subcommands are maintenance:

datalore-mcp                   # run the MCP memory server over stdio (same as `serve`)
datalore-mcp serve             # run the MCP memory server over stdio
datalore-mcp import <memory.jsonl> # migrate a file from the official server
datalore-mcp export            # print the folded knowledge graph as JSON
datalore-mcp merge             # fold every shard and report the merged state
datalore-mcp compact           # rewrite this node's shard, dropping shadowed operations
datalore-mcp query <text>      # search entities and print the matching subgraph
datalore-mcp help              # list the commands (--version prints the version)

How it works

Every mutation is recorded as an operation in an append-only log, split into per-node shards:

  • create_entities / add_observations / create_relations add elements;

  • delete_entities / delete_observations / delete_relations write tombstones (a deleted entity also tombstones its observations and incident relations).

A read folds all shards: each element (entity, observation, relation) is an LWW register decided by (timestamp, node, sequence). Observations are shown while their entity is present, relations while both endpoints are present. The fold depends only on the set of operations, never on the order shards or lines were read.

Consequences, all covered by property tests:

  • Idempotent — applying the same operations twice changes nothing.

  • Commutative — order of operations across shards does not matter.

  • Associative — folding shards in any grouping yields the same graph.

  • Convergent — two nodes with the same operations produce the identical graph.

Ordering is deterministic: entities by name, relations by (from, to, relationType), observations in operation order (timestamp, node, sequence).

Compaction

The log grows with every mutation. datalore-mcp compact rewrites this node's shard keeping, per element, only the latest operation of each node. A dropped operation is always shadowed by a later one of the same node for the same element — which wins in the total order — so the merged graph is unchanged, tombstones included. The rewrite is atomic (a temporary file renamed over the shard on the file backend, a conditional write on s3), so a crash midway can never truncate it. Only the local shard is rewritten; every other shard is untouched, so compaction is conflict-free and other nodes are unaffected.

Drop-in compatibility

The same nine tools, so it slots under the memory server name with no agent changes:

create_entities, create_relations, add_observations, delete_entities, delete_observations, delete_relations, read_graph, search_nodes, open_nodes.

The full graph is also exposed at the memory://knowledge-graph resource. The resource is read-only; clients can subscribe to it (resources/subscribe) and receive notifications/resources/updated whenever a mutation changes the graph.

Entities are keyed by name, and observations and relations are sets: adding the same observation or relation twice has no effect. delete_entities also removes the entity's observations and its incident relations. The tombstones cover what the deleting node could see: an observation added concurrently on another node survives as an element, so if the entity is later created again that observation reappears with it.

Development

npm run check       # Biome: lint + format + import order
npm run typecheck   # tsc --noEmit
npm run test        # Vitest, including the CRDT property tests
npm run build       # tsc -> dist/
npm run gate        # all of the above

Why datalore?

The name is an homage to Star Trek: The Next Generation. "Datalore" (Season 1, Episode 13, first aired 18 January 1988) is the episode that introduces Lore — the twin brother of Data, two Soong-type androids built by Dr. Noonien Soong. Data is the good one; Lore is the flawed, emotional, malicious prototype. The episode title is itself a portmanteau of Data and Lore, and both are played by the same actor, Brent Spiner.

The names fit a knowledge graph almost too well:

  • Data — a datum: a single fact.

  • Lore — the body of knowledge and tradition shared by a community.

So datalore = the facts plus the shared lore: exactly what a shared agent memory is. And the twin motif maps onto this project's core problem — many copies of the same graph, on different machines, that must stay consistent. Data's good twin and Lore's evil twin are the two failure modes; in datalore-mcp the copies cannot drift apart, because the merge is conflict-free by construction. The graph has no evil twin.

A small production gem: the episode was first pitched as a romance for Data with a female android. It was Brent Spiner himself who suggested the evil-twin plot instead — so the twin who gave the project its name also gave the episode its twist.

References:

This project is an independent fan homage. "Star Trek", "Data" and "Lore" are the property of their respective rights holders; this project is not affiliated with or endorsed by them.

License

MIT.

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    A
    maintenance
    Self-hosted, local-first knowledge graph and memory server for AI agents. Enables agents to persist, recall, and organize knowledge through MCP with automatic distillation, deduplication, and cross-linking.
    MIT
  • A
    license
    C
    quality
    A
    maintenance
    Provides AI agents with a human-inspired memory layer via MCP, enabling episodic and semantic memory recall, forgetting curves, consolidation, and contradiction detection. It integrates with MCP clients to offer local-first, dependency-free memory management.
    98
    1
    MIT
  • F
    license
    Not graded
    quality
    B
    maintenance
    Provides a self-hosted shared memory service that lets AI agents capture and recall durable facts, decisions, and context across multiple tools and MCP-capable clients.
    3
    -