kura
This server is a long-term memory service for agents: store, retrieve by meaning, inspect, and switch between separate memories.
kura_recall: ask a natural question; returns relevant memories and walks [[links]].
kura_read: read a full memory by slug.
kura_map: show the whole index (one-line recognition triggers) to see what exists.
kura_doctor: health check — memory counts, dead links, islands, index drift.
kura_list: list the stores/modes this server holds.
kura_use: switch the session to another kura (store or mode).
HTTP surface: also offers /recall, /remember, /annotate, /index, /prefill, /memory, /doctor, /stores, /profile, and /health.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@kurarecall what we decided about the archive disk"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
蒸留蔵 — distill-kura
A long-term memory for agents that is distilled, not accumulated. Recall works by meaning, writing is gated by evidence, and one server can hold several separate memories — one per agent mode — so switching mode switches what the agent remembers.
Ships as a DeepSeek Harness plugin, an MCP server for any other host, an HTTP service, and a Python library. Standard library only; no vector database, no embeddings, no framework.
┌── recall ──────────────────────────────────────────────┐
│ question → names a memory? → deterministic hit ~2 ms│
│ else → whole index in one prompt → picked slugs │
│ → walk [[links]] → the neighbourhood ~0.4 s│
└────────────────────────────────────────────────────────┘
┌── distil ──────────────────────────────────────────────┐
│ journal → classed evidence → candidates → GATE │
│ → new? → composed → draft → judged → poured │
└────────────────────────────────────────────────────────┘Why this exists
Two failures kill an agent's long-term memory, and they kill it from opposite sides.
Retrieval by keyword misses the thing you needed. A question about "SSD inference chips" shares no word with a memory titled "running the 2.6T model off an SSD tier" — yet they are the same subject. Word search returns nothing; the agent answers from nowhere. The fix here is not embeddings but recognition: the entire index (one line per memory, written as a recognition trigger) goes into one prompt, and a small model names what bears on the question. An index of ~500 memories is around 6k tokens — a few percent of a modern context window, and it sits in the prefix cache.
Writing everything poisons the store. An agent asserts something; a naive distiller records the assertion as a fact; the next agent reads it back as ground truth and repeats it with more confidence. That loop is self-reinforcing, and prompt instructions do not stop it — measured, not assumed. So the write path is gated by deterministic Python: every candidate memory must carry quotes that exist character-for-character in the raw material, tagged with where they came from.
class | what it is | what it licenses |
| the human's own words | "they decided", "they asked" |
| machine output | numbers — the only source |
| a tool that was invoked | "this was done" |
| the agent's own prose | a judgement, in the first person, never a bare fact |
A quote that is not found verbatim is discarded. A candidate with no surviving quote is
thrown away. A number with no [TOOL] behind it is stripped. Text crediting the human
with a decision, when no [USER] quote survived, is refused at the last gate. Ideas are
welcome — they go to a seed file, never to the store, and graduate only when later
evidence confirms them.
Related MCP server: Synapto
Tier zero: recognition before intelligence
The recall above is the right tool for a question that shares no words with its memory. It is the wrong tool for a question that names what it wants — and in a working session, most questions do. A blind 40-question benchmark on a live 317-memory store split exactly along that line: the direct questions needed no intelligence at all, and everything else needed all of it.
So before the thinker runs, a deterministic recognizer gets one look. The design is transposed from the n-gram embedding table inside Qwen3.8-Flash-Next — many hash heads voting over one table, behind a gate — onto the index:
Five heads, each an independent recognition channel: exact name; IDF-weighted word tokens (identifiers, ports, katakana runs); character 3-grams with stop-grams (a gram present in over a fifth of the store drowns in its own collisions, so it is dropped); character 2-grams; and the head of each body.
Coverage scoring — each head's vote is normalized by what it could possibly have reached for this question, so one lucky rare gram cannot fake confidence.
An honesty gate — a hit is returned only when the top score clears an absolute bar and beats the runner-up by margin. Anything less and the fast path says nothing; the question falls through to the thinker, unchanged.
Blind-tested — the examiner wrote the 40 questions from the index alone, never seeing the implementation — against the live store over HTTP:
question type | tier zero | thinker tier |
direct (14) | 14/14, median 2.3 ms | 14/14, ~900 ms |
paraphrase (10) | silent → falls through | 10/10 |
semantic bridge (10) | silent → falls through | 10/10 |
not in the store (6) | 6/6 refused | 0/6 refused |
wrong answers, whole set | 0 | — |
What that buys:
The everyday case stops paying the intelligent price for a lookup. Direct recall drops from ~900 ms to ~2 ms, and the fall-through tax on every other question is about 2 ms.
It knows what it does not know. Zero wrong answers across the set is the gate working, not the heads being clever — everything uncertain goes to the model. On the six questions whose answers were not in the store, tier zero refused all six; the thinker tier answered something every time. Refusal is a feature this project keeps having to buy back.
A direct question now survives the thinker being down. Recall used to degrade straight to word overlap; the named memory comes back regardless.
Every reply says which tier answered —
how: "fastpath",fastpath_verdict,fastpath_ms— so a slow answer is never a mystery.
Configured under [fastpath] (enabled, on by default; gate), per-store
overridable like everything else. The row it will never win: a question that
shares no surface with its memory. That is the thinker's job, and the gate exists
to hand it over rather than guess.
Quick start
git clone https://github.com/lna-lab/distill-kura && cd distill-kura
pip install -e . # or just run: python3 -m distill_kura.cli
cp kura.example.toml kura.toml # edit: one model endpoint is enough to start
kura init main --path ~/kura/main # create an empty store
kura serve # http://127.0.0.1:8085curl -s -X POST localhost:8085/recall -H 'content-type: application/json' \
-d '{"question":"what did we decide about the archive disk?","hops":1}'Wear the index, so the agent always knows what is known:
kura weave # build the three-layer cloth
kura prefill # the block to put in the system promptFeed it your agent transcripts:
kura distill run # drink a batch → candidates → gate → drafts
kura distill drafts # look at what it wants to write
kura distill drain # the scribe re-reads each draft cold: pour / fix / toss
kura distill night # stay resident and do it whenever things go quietNothing enters the store until drain (or a hand-run pour). Drafts carry their
evidence in an HTML comment, so you can always see why a memory exists.
The resident map
Recall-by-tool answers "what do you know about X?" — but only once the agent has decided to ask. It never answers the question the agent does not think to ask: is there anything here at all? An agent that cannot see the map does not know what it is missing, so it guesses, and a confident guess about your household is precisely the failure this project exists to prevent.
So the index is also worn: a standing block in the system prompt, on every turn.
kura weave # re-weave the index into the three-layer cloth
kura prefill # print the block a host should injectThree layers, because detail only pays for recent things
A blind A/B test — 20 questions, fat index vs slimmed index, scored without knowing which was which — settled the shape:
band | fat | slim |
overall | 9 | 11 |
recent events | 4 | 1 |
doctrine | 1 | 4 |
cross-domain leaps | 1 | 4 |
The doctrine lines were byte-identical in both indexes, and the slim index still won that band: a lighter surround makes the standing lines work better. Detail is not the source of insight. It earns its place only where things are still moving.
layer | rule | line |
pinned | frontmatter | kept in full |
fresh | changed within | kept in full |
trigger | everything else | compressed to ~ |
Trigger lines are written by the scribe model and cached in a ledger keyed on the
description and the budget, so a re-weave in the steady state costs nothing. With no
model reachable the loom trims mechanically instead — a memory system must not go blank
because a GPU is down.
Age is not mtime. cp -r, a restore or a checkout resets every timestamp, the whole
index turns "fresh", nothing is trimmed, and the mechanism has silently switched itself
off. So the loom prefers a date written inside the memory, and distrusts any mtime
that a fifth of the store shares with one calendar day.
Where it goes, and why that is a cache decision
- id: kura
name: distill-kura
config: { store: eq, promptOrder: -50 } # before the personaA prefix cache is lost from the first changed byte onward — measured on one local server: an identical 4,029-token preamble reprices from 0.68 s to 0.14 s, appending at the end stays 0.14 s, and one word added at the front costs the whole cache (0.66 s). The persona commonly carries a clock, so it changes every minute; the map is the largest block in the prompt and changes a few times a day. The big stable thing goes in front of the thing that ticks.
The block itself therefore contains no date, no clock, no counter — and build()
refuses a header that does, at build time rather than through mysteriously slow turns
three weeks later.
It never hands over half a map
situation | what the agent gets |
all well | the map, between |
over | the whole map, and a warning in the JSON (never in the text — a banner is volatile content) |
over | a stub with no index lines, saying the map is missing rather than empty |
kura unreachable | an explicit note that the map is missing, never an empty string |
A truncated map is the worst artifact available: it looks complete, and every memory
below the cut appears not to exist. weave will shorten the fresh window to fit, but it
will never drop a line — and if no setting reaches the budget it says so, keeps the
better map, and tells you where the weight is.
Getting it into a host
host | mechanism |
DSH | native plugin — a |
Claude Code, VS Code, Goose | MCP |
Claude Desktop, claude.ai | ignore |
anything else |
|
The MCP instructions field is a MAY in the spec, and a 9,000-token index cannot
travel through a 2KB cap regardless, so this project does not pretend otherwise.
Pay it forward
Byte-stability makes a prefix cache hold; it does not make the first turn cheap.
After a re-weave changes the map, the next turn pays the whole cold prefill — and on a
slow mouth that is minutes, not milliseconds. A llama.cpp server started with
--slot-save-path can save a slot's KV to disk and load it back, so the cold turn can
be paid once, in the quiet hours, and kept across restarts:
kura pay-forward # every [[payforward.mouths]] entry; -s / --mouth narrow, --force re-bakesMeasured on one machine (a 320B pure-CPU llama.cpp mouth, 16,444-token map): the bake 796 s; the save 283 ms (1.5 GB on NVMe); the server killed, rebooted, and the restore 655 ms — after which the first turn reprocessed 18 prompt tokens. A 13-minute cold turn became a 0.7-second restore. The name is the film's: the cold turn is paid forward, so the next turn — whoever's it is — receives it warm.
The slot filename carries the map's etag (kura-<store>-<etag…>.bin), so the files are
content-addressed: a fresh etag is proven, not assumed — a restore shows the file
still exists, a one-token probe reads timings.prompt_n, small means warm — and exits
2, nothing to do; a changed etag tries the restore first anyway (a file left by a lost
state or a parallel runner is still the right bytes) and only then bakes, saves, and
records _still/payforward.json. A mouth that cannot be reached is a loud, labeled
skip, never a crash and never a state advance. Old slot files are not pruned — the
slots API can save and restore a filename but cannot list the directory — and they are
not small (KV width × map length: that 16k-token map was 1.5 GB), so sweep the
directory by hand. kura tend runs this as a track after each weave; the recipe,
including the systemd shape for mouth restarts, is in docs/OPERATING.md.
The shortest cue that still recognises (M4, shadow)
A trigger line is paid on every turn. The question is not how short it can be cut but how short it can be while the reader still thinks "ah, THAT one" — and not its neighbour. With
[prefill]
trigger_tokens = 24 # the legacy budget, still what production wears
adaptive_triggers = true # generate and judge shorter candidates (shadow)
adaptive_apply = false # nothing enters the cloth until a benchmark earns it
trigger_steps = [8, 12, 16, 24]kura weave also writes _still/adaptive.json: per memory, a candidate per rung,
the shortest one that passes every floor the production trigger passes (plus the ones
a shorter cue newly needs — a number tied to a different unit, a dropped negation or
retirement word, a cut identifier) AND is recognised alone by the recognizer with the
callsign pre-head and the body off, and why_not_shorter for each rung refused. A
verified callsign is judged by its receipt, not by word overlap. Candidates are
cached per memory; the verdicts are recomputed every time, because whether a cue is
ambiguous depends on every neighbour. Falling back to 24, or to the line itself, is a
measurement — that memory needs that many tokens.
Promotion is a benchmark's decision, not a flag's: kura bench worldline --resident canonical,woven --resident-file adaptive=<rendered map> under --routing agent-only, read per category, with no increase in wrong or obsolete branches and
no rise in remembered_but_unreachable. Until then the shadow only watches.
Modes: more than one kura
A single memory that serves both "help me build this" and "help me think this through" serves neither well: the recall that helps you debug is noise in a conversation about what to do next. So a store is a directory, and a mode maps to a store.
[stores.maker]
path = "~/kura/maker"
label = "maker mode — building things"
[stores.eq]
path = "~/kura/eq"
label = "EQ mode — talking things through"
[modes]
maker = "maker"
eq = "eq"Every route takes a selector, so one process serves them all:
curl -s -X POST localhost:8085/recall -d '{"question":"...","mode":"eq"}'
curl -s localhost:8085/index?store=maker
curl -s localhost:8085/s/eq/doctor # path form, for clients that only vary a base URLThe stores share no memories, no index, and no distiller watermark. Switching mode genuinely changes what is remembered — not the same memory in a different voice.
The room is chosen before the conversation. A mode is what the host sends — a DSH
preset, KURA_STORE in an MCP environment, -s on the CLI — and it is the whole
session's home. Nothing in this project reads a message and decides which store it
belongs to; a conversation that drifts from building into feeling stays where it
started, and the host may offer another room for the next session. An unknown
selector is an error at the door, never a quiet fall to the default.
One room, many tags. A memory lives in exactly one store and may carry several
tags that describe its character (decision, landmine, emotion-carried, …). Tags
are words, not weights: nothing ranks by them, nothing counts them, and a Develop
memory tagged emotion-carried is still a Develop memory. There is no command that
moves or copies a memory to another store, and a mode change affects only future
sessions. The same topic raised in two rooms yields two memories, each distilled from
that room's own evidence — Research's "what we learned" and Develop's "what we did"
are different facts, and nothing crosses the boundary to deduplicate them.
A wide room recalls a little softer. Narrow stores with fixed charters recognise
sharply. A store that accepts anything — a USER room that follows the person rather
than a purpose — is expected to be looser, and in exchange is the one whose
understanding may grow: a profile.md beside its charter, in sentences, read after
the charter, drafted from its own memories and applied by a person. Five such rooms,
with their charters and a config, are in examples/rooms/.
Independent as routing, not as confidentiality. The server has no authentication, so
any process that can reach its port can name any store it holds. Binding an agent keeps
a model in its lane; it does not keep a process out. One trust level per process —
docs/TRUST.md is short and worth reading before a private store goes
in. It also covers the two boundaries that are easy to miss: two stores drinking from
one journal root, and two stores behind one model endpoint.
With DeepSeek Harness
DSH switches persona and tools by agent preset. distill-kura switches memory by store. Bind them and one preset change moves the whole self:
# .agent-presets/eq/agent.cordis.yml
- id: kura-eq
name: distill-kura
config:
url: http://127.0.0.1:8085
store: eq # this preset's memory
readonly: true # the CLIENT's own switch: do not even offer a write tool
# (the store's own `write_policy` is the authority; this just keeps the tool
# out of the model's hands. Naming a store already binds the preset.)One dependency, and why it is a peer. The plugin imports defineTool from
@deepseek-ai/dsh-tools. A profile-local second copy can split the package's
module-local Symbol identity — even at the same version — and make the first tool
call fail on undefined.prepare. The plugin therefore declares the package as a
"*" peer so the profile supplies its copy without a version mismatch. A stale
physical duplicate can still require deduplication; see the install checks in
examples/dsh-presets/.
allowSwitch has no fixed default: it follows store. A preset that names a store is
bound to it — no kura_use, no drifting mid-conversation — and only an explicit
allowSwitch: true reopens that door. Name no store and the session is free: every
tool takes a store argument and kura_use switches for the session. Tools: kura_recall,
kura_read, kura_doctor, kura_list, kura_use, and kura_remember (only when the
store is writable). Full wiring, including the MCP bridge and the isolate realm rule
for service rows, is in examples/dsh-presets/.
Persona is the host's business, not ours. This project never renders or injects a
persona; it only records, per store, which persona file belongs with it, readable at
GET /profile?store=eq so the two halves can be kept in step by whoever owns the
preset. Agent instructions likewise stay with the host's AGENTS.md mechanism — see
AGENTS.md in this repo for the conventions an agent working on this
codebase should follow.
With any MCP host
{ "mcpServers": { "kura": {
"command": "python3", "args": ["-m", "distill_kura.mcp"],
"env": { "KURA_URL": "http://127.0.0.1:8085", "KURA_STORE": "eq", "KURA_READONLY": "1" }
}}}Leave KURA_STORE unset for free mode: the tools take an optional store argument and
kura_use switches for the session.
Models: one by default, upgrade a role at a time
Three roles, not three machines:
role | when it runs | wants |
| every recall | small and fast; must judge relevance by meaning |
| distilling: reads a whole batch of journal | context length and patience |
| distilling: writes the memory, then judges drafts | good prose in your language, judgement |
Declare only [models.thinker] and one model does all three — the model you talk
to is also the editor that writes and judges your memories. That is the default and
it is a fair one: a capable GPU model does the editor's work well enough in its idle
minutes, and kura tend stops it the moment you come back (see "Unattended" below).
The upgrade path is to give the editor its own seat — a bigger model, an online API,
or a CPU model that does not compete for the GPU at all, so maintenance can go on
while you are talking. The house this was built in runs a 1-trillion-parameter MoE
on CPU at about 3 tokens/second as the editor: slow, but it never touches the seat the
conversation uses, and the memories it wrote over five days are a third of the store
today. Upgrade either of the other roles independently — a bigger local model, or an online API (any OpenAI-compatible
/chat/completions; the key is read from an environment variable you name, never
stored in the config):
[models.thinker] # always-on, local, small
url = "http://127.0.0.1:8000/v1"
model = "local-small"
[models.scribe] # upgrade just the writing
url = "https://api.example.com/v1"
model = "big-model"
api_key_env = "EXAMPLE_API_KEY"Two things this handles for you: reasoning-effort dialects differ per model family
(reasoning_effort, thinking_effort, enable_thinking), so all of them are sent
— an unknown one is ignored by the template, while a model left on deep-thinking by
default can spend its whole budget reasoning and return nothing. And the charter text
is placed byte-identically at the head of every role's prompt, so on a slow local model
the three roles share one cached prefix instead of paying three prefills.
A slow editor needs the prefix. The charter sits byte-identically at the head of
every call, so a 3 tok/s CPU editor pays its prefill once per silence, not once per
draft; on llama.cpp, keep --cache-reuse 0 off the table for recurrent models and let
the server keep its slots warm (--slot-save-path). The editor's calls are the ones
that wait an hour (timeout=3600) on purpose.
If the thinker is down, recall does not go silent — it falls back to word overlap
and labels the answer how=words, which the tools surface as ⚠ degraded. Quiet
degradation is worse than degradation. And before either tier runs, a deterministic
recognizer ([fastpath], on by default) answers DIRECT questions — ones that name a
memory — in under a millisecond with how=fastpath, thinker up or not; anything it is
not sure of falls through unchanged, and every reply says what it did in
fastpath_verdict / fastpath_ms.
Unattended: kura tend
The distiller, the pourer and the loom are meant to run in the quiet hours, and the watcher that decides when that is needs no model:
kura distill catchup -s maker # first: start from today, do not drink a year of history
kura tend -s maker # stays resident; one process per store
kura tend -s maker --once # one tick, for a scheduler or a testRun catchup once when you point a distiller at a journal it has never seen —
otherwise its first act is to drink the whole history, which for a year-old journal
is days of model time spent re-learning what the store may already know. It only
moves the marks forward, so it can never lose progress.
"Quiet" is the newest journal file's mtime. After idle_min (10) of silence it drains
waiting drafts (the editor reads each one cold: pour / fix / toss), or runs one
distilling pass when there are none; when something was poured it re-weaves the
resident map once, then pays the fresh map forward into the registered mouths (a cheap
verified skip when the weave changed nothing); and it tidies the index once per
silence. A track that had nothing
to do exits 2 and rests for backoff_min (20), so an empty journal does not spin. It
counts work — poured, tossed, fixed, drafted — never launches. Every track's output is
kept in _still/tend.log. And it writes a heartbeat that kura doctor reads
(tending.alive), because a watcher that dies quietly is the one failure a watcher
must not have.
When the journal changes, a running track is stopped: the editor is usually the same
GPU you are about to talk to. With the editor on a separate seat — a CPU model, another
machine — set yield_on_return = false under [distill] and a verdict in flight is
left to finish. This is the watcher the house ran its CPU editor with for five days,
rebuilt with the lessons it taught; docs/OPERATING.md has the systemd unit.
What this project does not ship: an autonomous research loop that reads papers and grows the store on its own. The house has one; it needs a model that can be left alone for an hour per question, and its results are not evidence in the sense the gate uses. It stays on the house side of the line.
What a memory looks like
One file, one fact.
---
name: archive-on-slow-disk
description: the archive lives on the slow disk; the fast one stays scratch
metadata:
type: project # user | feedback | project | reference
tags: ["decision", "landmine"]
evidence_manifest: sha256:…
belongs_because: this store keeps how the machine is laid out and why
keep: which disk, and the reason
may_fade: the df figures from that afternoon
---
The archive goes on the slow disk. The fast disk is scratch space.
**Why:** the other way round burns write endurance for nothing.
**How to apply:** check which disk a target directory is on before writing there.
Related: [[disk-layout]]And one line in MEMORY.md:
- [Archive on the slow disk](archive-on-slow-disk.md) — the archive lives on the slow disk; the fast one stays scratchThat line is the only thing read every single time. It is a recognition trigger,
not a summary: proper nouns, numbers, ⚠️ landmines, the conclusion reached. If a line
could be swapped with another memory's line and still read fine, it is not doing its
job — kura distill tidy finds the mechanically detectable cases and rewrites them.
The four lines under metadata/at the top are curation, not facts: tags are
words about the memory's character — several is normal, and a memory written before
they existed simply has none — and the three sentences say why it belongs in this
store, what meaning must outlive any later thinning, and what detail need not. The
distiller proposes them against the store's charter; tags that claim something about
the human (entrusted, emotion-carried, recurred) are checked against the quotes
and the check is recorded in the manifest. recurred is written once, by the
distiller, when the human brings a topic up again from another session — it is a
property, not a counter, and there is no number behind it.
kura doctor reports counts, dead links, islands (memories nothing links to), index
drift, tag lines it cannot read, manifests a memory points at that are gone, the state
of the learned profile, and the store's capacity in four units side by side —
memories, index tokens, body tokens, bytes — with limit and pressure left None.
It is the eye the metabolism needs. What happens when a shelf is full is not decided
yet: see docs/DESIGN.md §8.
The HTTP surface
route | what it does |
|
|
|
|
|
|
| the raw index |
| the resident block, ready to inject ( |
| one memory in full, with its |
| health of one store ( |
| stores, modes, and which model fills each role |
| the store's charter, the learned profile with its state (absent / present / broken), and a pointer to its persona (never rendered here) |
| liveness |
Any route accepts ?store= / ?mode=, a store/mode field in the body, or the
/s/<name>/… path prefix. No authentication: bind to loopback, or put something in
front of it.
Design notes worth reading before you change things
docs/DESIGN.md — why recognition beats search, what the gate buys, and the failure that motivated each mechanism.
docs/OPERATING.md — running it resident, schedulers and exit codes, backups, what to watch.
docs/TRUST.md — what a store boundary is and is not, write policies, and the two boundaries that are easy to miss (shared journals, shared models). Read it before a private store goes in.
A few decisions that look odd until you hit the thing they prevent:
Reserve before drinking. The distiller claims a stretch of journal before reading it, under a lock, and watermarks only ever move forward. Two distillers each writing back their own snapshot erased each other's progress and re-drank the same water a dozen times.
Watermarks are per-adapter units. Byte offsets for append-only transcripts, sequence numbers for archives that get rewritten (a byte offset into a recompressed file is a lie).
Echo suppression. A quote that already exists in the store is not new material — it is the store reading itself back through a tool result. Without this, a memory system rediscovers and re-records its own contents forever.
The last gate is a model, not a human. If a person must approve every draft, the system has quietly made that person its bottleneck, and drafts pile up forever. Nothing in the loop may require someone who is not always present.
kura distill runexits 2 when there was nothing to do. A scheduler must be able to tell "did work" from "found nothing", or a watchdog spins on an empty queue and starves the steps that need the idle time.
Measuring it, instead of claiming it
Two questions get answered with one number and should not be.
How much smaller? store_ratio = tokens in the memories and index / tokens of raw
journal actually consumed. What was lost? That is a different measurement, and a
store that keeps one memory in a hundred scores beautifully on the first while being
useless.
kura bench compress # what this store cost, from the distiller's own metrics
kura bench compress --tokenizer-command "./count-tokens" # exact, not estimated
kura bench retention --questions bench/fixtures/questions.jsonMeasured here, with the shipped fixtures and the built-in estimator:
corpus |
|
| 0.18 |
| 1.14 |
The second one is not a bug. On material where nothing is filler, distilling does not compress — each memory adds its why and how to apply, and the store comes out slightly larger than the transcript. The ratio is a property of the corpus, not of this tool, which is why there is no headline number here and why the command reports what it counted with.
Retention is scored model-free: each planted fact carries a marker that must appear in
what recall returns, so the score is reproducible on someone else's machine. Distractors
invert — a fact marked must_not_store costs a point if the store kept it, because a
memory system is judged by what it declines as much as by what it keeps.
score 1.0 (10/10) decision 1/1 number 2/2 negation 1/1 reversal 1/1
conditional 1/1 landmine 1/1 returning 1/1 distractor 2/2That is ten planted facts in a synthetic fixture, distilled by a local Qwen3.8-27B
(NVFP4) as brain and scribe with max_items = 8, coverage_passes = 2, and scored with
the same model as thinker. A different model will give a different score: the score
measures a pipeline-plus-model, and the fixture exists so the model is the only thing
that varies. It measures whether a fact is findable, not whether the answer reads
well — judging prose needs a model, and then the benchmark stops being reproducible.
kura distill run writes one line per batch to _still/metrics.jsonl, which is where
the raw side comes from. The canonical side counts only memories whose evidence
manifest points at a recorded batch — dividing a whole store by the raw material of a
few batches is a number in the wrong direction by an order of magnitude, and the first
version of this command did exactly that. Memories that predate manifests are reported
as unattributed, not silently included. The raw side is always the distiller's
estimate at drink time, so with --tokenizer-command the ratio is labelled mixed.
What this runs against
requirement | |
Python | 3.11+ (no dependencies; |
Node | 20+, for the DSH plugin only |
| only to read DSH session archives |
model endpoint | anything answering |
"OpenAI-compatible" is narrower than "any provider." A vendor's native API needs an
OpenAI-compatible gateway in front of it; its own URL will not do. A strict service also
rejects unknown top-level fields, so set dialect = "openai" (or "generic") — the
default "vllm" sends chat_template_kwargs, which local servers want and a strict one
400s on. The client retries once with a plain body and records why a call failed
rather than collapsing every cause into a silent None.
Tests
python3 -m pytest tests -q # 346 tests, no model required
cd dsh-plugin && npm test # 24 more for the pluginThe gate is tested adversarially: every case is a way a real model actually tried to
smuggle something past it. test_containment.py is written the same way — every case is
an escape attempt, not a happy path — because it guards a hole that was real: a store
used to answer for any file whose path you could spell. The end-to-end test runs a full
distil→drain cycle against a scripted model server on a real socket.
License
MIT.
Available Tools
6 toolskura_doctorA
Health check of the kura: how many memories, resolved and dead [[links]], islands (memories nothing links to), index drift. Call when recall behaves oddly, or when asked about the memory system itself.
| Name | Required | Description | Default |
|---|---|---|---|
| store | No | Which kura. Omit for the current one. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It does well by naming the diagnostic nature ('health check'), the categories of data inspected, and the symptom-triggered use case. It does not explicitly state whether the tool is read-only or whether any repair/reindexing occurs, but 'health check' strongly implies diagnostic rather than mutating behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact: two sentences, front-loaded with the core purpose, followed by a concise list of diagnostic metrics and explicit usage triggers. Every clause earns its place, with no filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple tool with zero required parameters and a fully documented optional parameter, the description provides enough context to select and invoke the tool correctly. It explains what will be inspected and when to use it. The only gap is that it does not describe the return format or explicitly confirm the absence of side effects, but neither is critical for a diagnostic health-check tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, and the description provides no additional parameter-level meaning beyond what the schema already states. The optional 'store' parameter is fully documented in the schema, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb-resource pairing: 'Health check of the kura', and enumerates concrete diagnostic outputs (memory counts, resolved/dead links, islands, index drift). This clearly differentiates it from the sibling tools like kura_read or kura_recall, which focus on reading or retrieving memories rather than diagnosing system health.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly states when to use the tool: 'Call when recall behaves oddly, or when asked about the memory system itself.' This provides clear invocation context. However, it does not explicitly mention alternatives or when not to use it, so it stops just short of full routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
kura_listA
List the kura this server holds and which agent mode each one belongs to. Call it when you are unsure which memory you are speaking from, or before switching.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It states the operation (list) and the output (kura and agent modes) but does not mention side effects, read-only nature, performance, or any edge cases. Adequate but not rich for a tool without annotation safety signals.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, zero wasted words. The first sentence states the core function; the second provides immediate usage context. Information is front-loaded and every word serves a purpose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a no-parameter list tool with no output schema, the description explains what is listed and exactly when to invoke it. It could mention what 'kura' refers to or if results are ordered, but the definition is sufficient for an agent to decide and call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, and the schema coverage is trivially 100%. Per the rubric, 0 params earns a baseline of 4. The description does not need to add parameter details since none exist.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb 'list', the resource 'kura', and what is listed (which agent mode each belongs to). It implies a distinction from siblings like kura_read or kura_recall by focusing on the server's collection and mode mapping, though it does not explicitly name alternatives.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit when-to-use guidance: 'Call it when you are unsure which memory you are speaking from, or before switching.' It lacks explicit when-not-to-use or named alternative tools, but the usage context is concrete and helpful.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
kura_mapA
Show the whole index of the kura — every memory's one-line recognition trigger, in one answer. Use it when you need to see WHAT EXISTS rather than look something up: before claiming a topic was never discussed, when choosing which memory to open, or right after switching kura. It is a map, not the contents: open a memory with kura_read for the detail.
| Name | Required | Description | Default |
|---|---|---|---|
| store | No | Which kura. Omit for the current one. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are present, so the description carries the full burden. It clearly frames a display operation ('Show the whole index... in one answer') and sets output expectations ('map, not the contents'). It does not discuss permissions or exact formatting, but for a non-mutating index tool the core behavioral disclosure is present.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, each earning its place: the first states the action and scope, the second gives concrete use cases, and the third clarifies depth and names the sibling for details. It is front-loaded and free of filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a low-complexity tool with one optional parameter and no output schema, the description provides the essential context: what it returns, when to use it, what it is not, and where to go for detail. Nothing needed to call it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the single 'store' parameter is already documented as 'Which kura. Omit for the current one.' The description adds no parameter-level detail beyond that, so it meets the baseline but does not exceed it.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Show the whole index of the kura - every memory's one-line recognition trigger.' It also distinguishes itself from siblings by saying it is 'a map, not the contents' and explicitly directs details to kura_read, so an agent can tell it apart from the other tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives explicit when-to-use conditions ('when you need to see WHAT EXISTS... before claiming a topic was never discussed, when choosing which memory to open, or right after switching kura') and an explicit alternative ('open a memory with kura_read for the detail'). This is clear routing guidance rather than leaving usage to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
kura_readA
Read one whole memory from the kura by its slug (e.g. 'storage-doctrine'). Use after kura_recall when a summary is not enough and you need the full text.
| Name | Required | Description | Default |
|---|---|---|---|
| slug | Yes | Memory slug, without .md | |
| store | No | Which kura. Omit for the current one. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations available, the description carries the burden of behavior. It clearly indicates a read operation fetching the full memory text, which is sufficient for this simple tool. It could add detail about return format or failure behavior, but the core behavior is transparent enough.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with no filler. The primary behavior is front-loaded, and the usage guidance, alternative, and selection condition are packed efficiently into the second sentence.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read-by-slug tool with only two parameters, no output schema, and full schema coverage, the description is complete. It tells the agent what it does, when to use it, and provides an example. Nothing essential is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description adds a useful example slug ('storage-doctrine') but otherwise does not need to add much beyond what the schema already documents for 'slug' and 'store'.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Read'), the resource ('one whole memory'), and the retrieval key ('by its slug'), with a concrete example. It also names the sibling tool it complements, distinguishing it from kura_recall without needing to inspect the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly states when to use this tool: 'Use after kura_recall when a summary is not enough and you need the full text.' This directly routes the agent to the appropriate alternative and gives a clear selection condition.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
kura_recallA
Recall from the kura — long-term memory retrieved by MEANING rather than keyword, then following [[links]] between memories. Call it whenever the question touches past decisions, measurements, people, machines, or anything done before — prefer it over guessing. An empty result means it is simply not remembered yet: say so plainly and never fill the gap with invention. Pass store to reach a different kura, or call kura_use to switch for the session.
| Name | Required | Description | Default |
|---|---|---|---|
| hops | No | How many [[link]] hops to walk (default 1) | |
| store | No | Which kura to ask (store or mode name). Omit for the current one. | |
| question | Yes | What you want to remember, as a natural question |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations are absent, so the description carries the full burden of behavioral disclosure. It does reveal important behavior: it retrieves by meaning, follows links, and emphasizes not to fabricate when no result is found. However, it does not disclose potential side effects, rate limits, or what happens with the hops parameter (e.g., how deep the search goes and whether more hops affect latency). It also doesn't specify if the tool is read-only, though recall implies reading. Overall, it gives useful context but leaves some behavioral aspects unmentioned.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, dense paragraph but contains no filler. It front-loads the core purpose and usage guidance, then addresses edge cases and parameter use. It is appropriately sized for a tool with multiple interconnected behaviors (meaning-based recall, link-following, empty result handling). It could be slightly more structured (e.g., bullet points for handling empty results or parameter variants), but it is efficient and every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has 3 parameters (1 required), no output schema, and moderate complexity due to link-following and alternate stores, the description covers the essential aspects: what it does, when to use it, how to handle empty results, and how to use parameters. It does not explicitly describe return format (e.g., whether it returns text or a list of memories), but with no output schema, it could have provided more on that. Missing details on hops behavior and potential limitations are minor gaps. Overall, it is fairly complete for an agent to use it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema already documents all three parameters (question, hops, store) with descriptions, and coverage is 100%, so the baseline for parameter semantics is 3. The description adds value by explaining the semantics of 'store' ('reach a different kura') and hints that 'question' should be phrased as a natural question. It also implies the meaning of 'hops' (following [[links]]). It doesn't detail the format or constraints beyond the schema, but it enriches the meaning of each parameter sufficiently.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: 'Recall from the kura — long-term memory retrieved by MEANING rather than keyword, then following [[links]] between memories.' It differentiates itself from siblings by emphasizing meaning-based retrieval over keyword, which distinguishes it from tools like kura_read or kura_list. It also mentions specific use cases, making it easy for an agent to understand what it does.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly tells the agent when to use this tool: 'Call it whenever the question touches past decisions, measurements, people, machines, or anything done before — prefer it over guessing.' It also explains what to do on empty results ('An empty result means it is simply not remembered yet: say so plainly and never fill the gap with invention') and mentions alternatives or related actions ('Pass store to reach a different kura, or call kura_use to switch for the session'). This provides clear guidance on selection and behavior.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
kura_useA
Switch which kura the following calls read from, for the rest of this session. Use when the conversation moves to a different mode of work (for example from building things to talking things through). Has no effect on a bridge that was bound to a single kura at startup — it will say so.
| Name | Required | Description | Default |
|---|---|---|---|
| store | Yes | Store or mode name |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden of behavioral disclosure. It does this well by revealing that the switch persists for the session, affects subsequent calls, and has no effect on a single-kura bridge (and that it will say so). It does not cover every possible error or side effect, but the key stateful behavior is disclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three sentences with no wasted words. The main effect and scope are front-loaded, the usage condition follows, and the edge case is placed last. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple one-parameter state-switching tool with no annotations and no output schema, this description is complete enough. It explains what the tool does, when to invoke it, how long the effect lasts, and the one major exception an agent needs to know about.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% — 'store' is already described as 'Store or mode name.' The description adds some context by suggesting that stores correspond to modes of work, but it does not enumerate valid values or add meaningful parameter detail beyond the schema. This is the baseline for high schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Switch') and names both the resource ('which kura the following calls read from') and the scope ('rest of this session'). This clearly distinguishes it from the sibling read/analysis tools like kura_read and kura_map, which retrieve data rather than change the active context.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly says when to use it: 'Use when the conversation moves to a different mode of work,' with a concrete example. It also discloses a limitation about bridges bound to a single kura. It does not explicitly name sibling alternatives or state when not to use it, but the intended context is clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
6 tool updates
v0.1.0- First observed
kura_doctor - First observed
kura_list - First observed
kura_map - First observed
kura_read - First observed
kura_recall - First observed
kura_use
TDQS
Scored across 6 tools
Each tool has a distinct purpose: recall for meaning-based search, read for slug-based retrieval, map for index views, doctor for health checks, list for enumerating kuras, and use for switching. No overlap in functionality.
All tools follow the same `kura_` prefix with a verb-based second part (recall, read, map, doctor, list, use), creating a consistent and predictable naming scheme.
Six tools is a well-scoped number for a memory-management server, covering retrieval, inspection, maintenance, and context switching without excess or deficiency.
The set covers all read and management operations (recall, read, map, health, list, switch), but lacks an explicit write or delete capability, which is a minor gap for a full memory system. Agents can work around this by relying on existing memories.
Maintenance
Related MCP Connectors
Shared, governed long-term memory for AI agents across tools and sessions via MCP and REST.
- mem0OAuthio.github.mem0ai
Persistent memory for AI agents: add, search, update, and delete long-term memories.
Universal persistent memory and knowledge retrieval layer for AI agents and LLMs.
Persistent memory for AI agents. Search, store, and recall across sessions.
Related MCP Servers
- FlicenseNot gradedqualityDmaintenanceProvides AI agents with persistent, searchable memory that survives across conversations using semantic search, temporal versioning, and smart organization. Enables long-term context retention and cross-session continuity for AI assistants.14-
- AlicenseNot gradedqualityAmaintenanceProvides persistent, searchable memory for MCP-compatible agents, enabling recall by meaning, automatic decay, trust scoring, and cross-agent handoffs.5MIT

Hebbrix MCP Serverofficial
AlicenseAqualityAmaintenanceProvides long-term memory and a temporal knowledge graph for AI agents, enabling persistent memory and reasoning across sessions.331MIT
whimsicality-mcpofficial
AlicenseBqualityBmaintenanceProvides persistent memory for AI agents, including context storage, facts, plans, RAG search, code snippets, and conversation compaction, enabling state to survive across sessions and processes.1411 npm2MIT