memory-ultra-rag-mcp
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@memory-ultra-rag-mcpAdd today's discussion to the project memory."
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Memory UltraRAG MCP
Purpose: serve UltraRAG's memory to an agent over stdio MCP, in two kinds:
local memory belongs to the project this server is bound to and lives inside that repository, under
.memory-rag, so a project carries its memory with it and no other project reads it;global memory belongs to the account and lives in UltraRAG's shared storage tree, so every instance serving that tree reads the same memory. It is not a project's: who the user is, what they want remembered everywhere, and the durable rules that apply wherever they work.
Each kind is one file, memory.sqlite3, holding statements in UltraRAG's shape, and the two kinds take the same arguments. A standing instruction the user gave once belongs in global memory, where it is read at the start of every conversation, rather than anywhere local.
There are no dated rounds. Upstream keeps a day-per-file dialogue log beside the standing document; this server does not, because anything worth keeping is recorded as a statement. Every divergence from upstream is in the choice table below.
The behaviour, the tool names, the file names, and the byte formats are UltraRAG's, taken from servers/memory/src/memory.py at the pinned revision below and checked against files that revision produced. The difference this package makes is where a project's memory lives: UltraRAG keeps everything under one storage tree, and this server keeps each project's memory inside the project. What the extension is, and what it deliberately leaves alone, is decided in ADR 0001.
Reference revision | Declared version |
|
|
Requirements
CPython 3.11 or 3.12, Linux, uv, and a SQLite built with FTS5, which the search index needs. Run the probe to check that last one:
uv run --frozen python scripts/check_sqlite_fts5.pyWithout FTS5 the server refuses to start, because a read would have to pass over every file a memory holds, which is the cost the index exists to remove.
A lookup also runs a local sentence embedder on the CPU, so the first use fetches its weights once — about 65 MB — into this package's own cache under ~/.cache/memory-ultra-rag-mcp/models, and never again. No API is used at any point: the model runs in this process, and a machine that cannot fetch it still starts, records statements, and answers in words, saying semantic_available: false in every read so a caller is never left guessing why an answer is thinner. Which model it is takes one line, in a documented table.
No UltraRAG checkout is needed: the formats are verified against captured fixtures and, when a checkout is available, against the real server. The browser view is the shared ui-ultra-rag-mcp package, pinned by commit and installed with this one.
Related MCP server: Bilinc
Install
git clone https://github.com/AhmedKishki/memory-ultra-rag-mcp-server.git
cd memory-ultra-rag-mcp-server
uv sync --frozenConfigure an MCP client
{
"mcpServers": {
"memory-ultra-rag-mcp": {
"command": "/ABSOLUTE/PATH/TO/memory-ultra-rag-mcp-server/.venv/bin/memory-ultra-rag-mcp",
"args": [
"--project-root", "/ABSOLUTE/PATH/TO/my-project",
"--storage-root", "/ABSOLUTE/PATH/TO/UltraRAG/ui/storage"
]
}
}
}Flag | Meaning | Default |
| The repository whose local memory is served | required |
| Where the account's global memory lives, as a normal directory |
|
| Also serve a browser view of this memory on this loopback port | off |
The first root is inside the project and the second is in the account's home, and each can be moved without touching the other. memory-ultra-rag-ui takes the same --project-root and --storage-root, plus --host and --port.
One server instance serves one project, so --project-root binds the session: local memory always means that project's memory.
The four tools
Tool | Parameters | What it does |
|
| Records one statement in the project's or the account's memory. |
|
| Returns the statements from both memories that match |
|
| Removes the one statement whose text is exactly |
|
| Records this session's handoff for the next, replacing the previous one. |
Four verbs, named for what they do. The scope is an argument, not a second
tool: an agent chooses an operation and a memory, and never picks between two
tools that differ only by a word in their name. scope is "local" — this
project, inside the repository — or "global" — across projects, in the
account's memory — and it defaults to "local".
Which one a statement belongs in is read off what the user is asking for and how
far they mean it to reach. A fact about this repository is local whatever it is
called; a preference about how they want to be spoken to, or a rule about their
own work rather than this code, is the account's. That judgement is semantic and
is made on the substance rather than the wording, because most prompts carry no
marker of where a statement should go — so the absence of one is not a reason to
file local. The server names the two destinations and leaves the judgement to
the caller. recall_memory searches both
memories in one call, so one question gets one answer: the best statement wins
whichever memory it is in, and every statement names its own scope. Forgetting
searches both memories too, so a caller who has lost track of which one meant it
still gets the right one removed.
Record after recalling. The instructions say so, and the tools are built for it: a recall that returns something covering the same thing means a second copy would make the answer worse, and one that contradicts it means that statement should be forgotten first. Forgetting matches the words and never the meaning, so it removes that statement or nothing, and a text that matches more than one removes none and names them.
Settings
Every tunable this server reads is a setting, and the packaged file that declares
them all is default.toml, inside the
package. A layer above it names only what it changes, and the layers win per key,
lowest precedence first:
default.toml < ~/.config/memory-ultra-rag-mcp/config.toml (the global layer)
< <project>/.memory-rag/config.toml
< --config PATH
< MEMORY_ULTRARAG_* environment variables
< --set key=valueThe user layer is the one that applies globally. One file in the account's
config directory changes every project on this account, because a memory is
shared by every project on the account. A project may still narrow it for itself
with its own config.toml inside its own state, which is what the project layer
is for.
# What is in force, and where each value came from.
memory-ultra-rag-mcp --project-root ~/my-research --print-config
# Change it for one run, for a layer of your own, or for good.
memory-ultra-rag-mcp --project-root ~/my-research --set retrieval.recency_bonus=0
memory-ultra-rag-mcp --project-root ~/my-research --config ./memory.tomlA recall never shows the same statement twice. One filed in both memories, or two worded differently, is collapsed to the best-ranked of the pair and reported in collapsed_repetitions; retrieval.duplicate_cosine decides how alike two wordings have to be before they are the same statement, and no setting turns the check off.
A key the registry does not declare is an error in every layer, so a typo is loud
rather than silent, and a value out of range is refused by name. An empty string is
a value only where the setting gives it a meaning — dense.reranker_model = ""
turns the cross-encoder off, and an environment variable set to nothing is not a
layer saying anything, so that one is a file or a --set.
A setting is one of three kinds, and the kind says what changing it costs:
Kind | What it decides | What changes when you do |
| what a memory is made of | the statements are embedded again, rather than compared across two model spaces |
| where something lives | the cache moves or resizes; a statement is what it was |
| what a read admits and in what order | the next answer is ordered differently; the statements do not change |
Two things deliberately stay in code rather than in a file, and AGENTS.md says
why: the record's schema version, and the names and paths that make up a scope,
because those are what makes two UltraRAG servers agree on where a memory is.
The stack that merges those layers — the registry type, the merge, the coercion
every layer shares, the provenance, and the three path helpers — is the pinned
config-ultra-rag-mcp
library, which the research server uses too. This server keeps its own keys, its
own default.toml, its own MEMORY_ULTRARAG_* names, and its own account and
project directories, so a change to one server's tunables never reaches the other.
Browser view
Memory is worth looking at, so the same memory the tools serve can be read — and, if you want, added to — in a browser on this machine. Two commands do it:
# a view of its own, until you stop it
memory-ultra-rag-ui --project-root /ABSOLUTE/PATH/TO/my-project
# or inside the MCP server, reported on stderr and stopped with it
memory-ultra-rag-mcp --project-root /ABSOLUTE/PATH/TO/my-project --ui-port 5052The view shows the bound project's memory and the account's global memory together, as one block per scope:
each block names its scope, its directory, and whether its standing document is there;
that scope's standing memory, whole, with a copy button;
Edit standing memory on the block, which replaces the whole document only when it still matches the version the page read, so a write an agent made meanwhile is never overwritten.
The shared page still asks for dated exchanges, because it fetches them under the same capability that shows the memory view and an action that refused would take the whole view down. This adapter answers with an empty list, which is the true answer here; the page's own empty-state wording is the rough edge, and it belongs to ui-ultra-rag-mcp.
memory-ultra-rag-ui takes --project-root, --storage-root, --host (loopback only), and --port (default 5052). The page is the shared ui-ultra-rag-mcp package, pinned by commit, and this package supplies a thin adapter: the browser calls this server, the server reads and writes its own memory, and the view is handed no directory of its own.
Storage
<project-root>/.memory-rag/memory.sqlite3 the project's memory: the record
<storage-root>/memory/default/memory.sqlite3 the account's memory: the recordThe record is one file. Every statement is a row in
memory.sqlite3: its text, the kind it was filed under, its place in the
document, when it was added, how often it has been recalled, and one vector for
the meaning side. The same table is the FTS5 word index, so a statement is stored
once and found by words without a second copy to keep in step.
MEMORY.md is a rendering, not the memory. It is the same statements as
plain prose, newest first, and it is written only when it is asked for:
memory-ultra-rag-reindex --project-root ~/my-research --export
memory-ultra-rag-reindex --project-root ~/my-research --export-to ~/notes/memoriesA rendering is not read back as memory, and the next export replaces it. A
document left by an older version — when the document was the memory — is read
once for any statement the record lacks and is then removed, so a scope holding
both a document and a database is not a state this server leaves behind. One
that agrees with the record is left alone, because it is what somebody asked for;
--retire-document asks for it to go.
The browser view shows the same rendering, rendered from the record on every request, so the page can never show a document the memory has moved past. markdown
MEMORY
i am jack. i like LLMs.
**The newest statement is at the top**, so the file a person opens reads in the order things were remembered and a read can prefer what is new.
```markdown
# MEMORY
Always cite the commit that introduced a change.
The draft lives in docs/The file no longer writes a kind in front of a statement, and no longer carries
the seed. A file written by an older version is read once on upgrade — the
RULE: prefix becomes the row's kind and the words after it become the
statement — and the next export has neither.
memory.sqlite3 beside it is the one file, and it is the record: an FTS5 table of every statement — its text, its kind, and its position — with a vector per statement for the meaning side, and four columns on the row itself for what the document cannot hold: added_at, recalls, and last_recalled_at. A statement is one row and not a row plus a file, and a read counts a recall by updating the rows it just returned, in the transaction it read them in. Nothing is written beside the memory and no thread is involved, and the counting is proportional to the answer rather than to the file.
It is therefore not fully derived: deleting it costs a re-read of the document plus a re-embedding plus the dates and counts, and the last of those is the one thing no rebuild can restore. memory-ultra-rag-reindex collects the rows of statements that have since been reworded or removed, and the document itself is never in this file's power to lose. One file rather than three: the index, the vectors, and the history are all keyed by the same statement digest and all answer the same read, so one file means one schema version, one connection, and no way for two of them to disagree about one statement.
The storage root is where global memory lives, and its default is in the account's home:
Platform | Default storage root |
Linux |
|
macOS |
|
Move it with --storage-root, or set MEMORY_ULTRARAG_STORAGE_ROOT once for every project this account opens. Set ULTRARAG_UI_STORAGE_ROOT instead and global memory lands in UltraRAG's own UI storage tree, under the memory/ directory UltraRAG already reads — which is how a UltraRAG UI shows the same memory with no exchange format and no import step. The flag wins over the variables, and MEMORY_ULTRARAG_STORAGE_ROOT wins over ULTRARAG_UI_STORAGE_ROOT.
A project's memory is ordinary Markdown in that project: visible, diffable, and optionally committed, by the project's own choice.
Not yet
Recorded so an absence reads as a decision rather than an oversight, and so a later change extends this server instead of replacing what it stands on. ADR 0001 carries the reasoning.
A read understands a query, it does not rewrite it. No query expansion, no stemming, no synonym list, and no model asked to rephrase; a statement is found because its meaning is near the query's, which is close to but not the same as understanding the question.
A read cannot be paged or offset.
limitcaps the units; there is no "next page" and no date range.A hand edit is picked up on the next read, not before it. A statement you add is findable by its words at once and by its meaning once its vector is embedded; the read discloses both. Deleting a statement stops it answering immediately — a stale index can no longer serve one that is gone.
The semantic side is unmeasured here. This release adds it; the measurement that decides which model, and whether it beats words on this data, is the next phase. Nothing in this package claims a gain it has not measured.
Nothing removes or corrects a statement. A statement is appended; the standing document is replaced only through the browser view or by hand.
No session tier. Every statement is durable; nothing expires.
No typed or relational structure. The memory model is upstream's, deliberately.
No cross-scope read. One call reads one scope, so nothing can quietly sweep a project's memory into another's, and a project is reachable only by the server bound to it.
The choices this package makes
# | Choice | Reason |
1 | Local memory is kept inside the project, under | isolation is the point: a project's memory is private to that project and travels with it |
2 | The project is bound at startup, and the project tools take no project argument | one session means one project, so local memory cannot be written into the wrong one |
3 | Global memory keeps UltraRAG's layout in the storage tree | the shared memory stays where a UltraRAG UI already reads it |
4 | The file name and the scope layout are UltraRAG's, and the two kinds share one implementation | a project's memory and the global one have one shape, so either can be read the same way, and a UltraRAG UI finds the global memory where it looks. What is in that file is now an export rather than the record — choice 33 |
5 | The surface is four memory tools, and this server is not a UltraRAG pipeline node | UltraRAG's server base class registers a pipeline |
6 | The reference revision and its captured output are committed and tested | the format cannot drift quietly: a change fails |
7 | The browser view is the shared | interface code stays in one repository, and this package keeps no second UI |
8 | The page reaches memory only through this server: no directory is handed to the browser | one reader and one writer per file, and the page shows what an agent reads |
9 | Replacing the standing document is this package's own write, guarded by the digest the page read | upstream only creates that document from its template, and a plain overwrite could drop an agent's write |
10 | Global memory defaults to this account's home data directory, and moves by flag or variable | it belongs to the account rather than to a project, and the default follows the sibling repositories' pattern: one flag, one variable, a place in the home directory |
11 | The tools are named for what they do — record, recall, forget — and the memory is a | an agent should choose an operation and a memory, not pick between two tools that differ by a word in their name. Two verbs became four operations with one name each, and a read searches both memories at once because one question deserves one answer. Upstream's two tools are not re-exposed under any name |
12 | A read takes a query and answers it, and no tool returns a memory whole | a memory grows by appending; reading it whole spends the context window on the file instead of on the work, and would be unusable after a few weeks |
13 | The search runs over an FTS5 index built from the memory's own document, kept beside it and disposable | a read is then a lookup and a page of rows rather than a pass over the whole memory, so its cost does not grow with it; the document stays the record, so an index is never the thing that is true |
14 | The index follows the files instead of leading them, and no write path updates it | a memory edited by hand or by a UltraRAG UI is read correctly, because a read compares each file against what was indexed and re-reads only what changed |
15 | A statement is inserted at the top of the standing document, and nothing is attributed to a speaker | attributing a statement to the user or to an assistant would put a claim in the file that nobody made. Insertion rather than appending is choice 27 |
16 | Global memory has no | a project is identified by the repository it is bound to, so a user dimension would be a second identity for something already identified; the directory is left as upstream's so the layout, an existing store, and a UltraRAG UI are unaffected |
17 | A read works in words and in meaning, fused, with the model behind a seam | a memory is largely proper nouns, paths and identifiers, which the collection's own measurements say need words; a query that never used the caller's wording needs meaning |
18 | The model is local, pinned, listed, and run on the CPU in this process | no API is used at this stage, the weights are fetched once into this package's own cache, and a stronger or multilingual model is one line rather than a rewrite |
19 | Recording is asynchronous; a lookup is not | a write may cost an append and a queued unit and never fails because the lookup layer is missing; a read may only do constant work, so it never embeds units, walks the scope, or loads a model |
20 | A read keeps the word index in step with the Markdown, and | superseded by choice 26, which reverses the reasoning: the sweep this trade was measured against counted a few thousand dated files, and a scope here is one document — a read now fingerprints it and re-reads it only when it changed |
21 | A statement is filed under a kind the caller names, held as a column of the record and returned as a field of its own rather than as part of the text | a memory is a list of unrelated statements, and a category is what makes a read for one kind answerable; the kind is the caller's own and carries no meaning, so the server defines no vocabulary and stores the word as it was given. A single word keeps it separable from the statement, and the document stays plain prose |
22 | A scope is its standing document, with no dated exchange log beside it | anything worth keeping is recorded as a statement, and a date does not change what a statement is, so the log upstream keeps beside the document would be a second record of the same thing with a shape nothing here can answer a question from — the store, the index and the tools have one kind of unit, and the differential test asserts this side writes no dated file at all |
23 | A lookup drops English stop words before it searches | the word side asks FTS5 for every query term at once, so a term in no statement does not narrow the answer, it empties it, and the fallback then rescues the query by accident: "what is research-rag for" finds nothing, because no statement contains "what" |
24 | Reranking is on by default, at a depth of ten candidates | a memory's units are one line each, which is where a bi-encoder's vector is weakest, and a cross-encoder reads the pair rather than either side alone; measured here at 54 ms for ten candidates, which is what sets the depth against the 250 ms read ceiling. That it is more accurate is unmeasured, and nothing in this package claims it |
25 | This memory is English, and the models and the stop words say so | a memory is a record of what one person decided and asked to be kept, in one language; the embedder is |
26 | A read brings the document back in line with the record, and a hand edit to it is disclosed | a memory is something a person is meant to open, and a hand edit has to do something definite rather than nothing. With the record in the table, the definite thing is that the edit does not take: the document is written out from the record and the read reports |
| 27 | A new statement is inserted at the top of the standing document, and a read prefers new information by a bounded bonus | upstream appends, so its oldest statement is first, and a memory that only grows ends up answering with what it learned months ago as readily as what it learned an hour ago. Insertion on top makes the file a person opens read in the order things happened, and the bonus — at most ten per cent of the fused score, decaying with distance from the top — settles a tie towards the caller's current state without ever preferring a weaker match. The block format is untouched and a fresh scope is still byte-identical to upstream's (tests/test_fidelity.py) |
| 28 | Forgetting is a tool, and it matches the words exactly | a memory that cannot be emptied accumulates statements that are no longer true, and a later read returns them as confidently as the true ones. The match is on the exact text and never on meaning, and a query that matches several statements removes none of them, so one vague call cannot lose two decisions. The alternative — leaving removal to a person editing the Markdown — is what made the gap hard to notice |
| 29 | A handoff is one statement under a reserved kind, and setting it removes the previous one | yesterday's handoff and today's are both plausible and neither is current, which is worse than having none: a session cannot tell which one it is reading. Replacing by kind means the caller records a handoff and never thinks about the old one, and replaced reports what went so the loss is disclosed rather than silent. The handoff is a project's own state, so it has no global form |
| 30 | A statement's kind is optional and defaults to ITEM | a required kind is a decision every write has to make before it can make the write. A default that is a real kind means every statement is filed under something findable, and ITEM is exactly the right name for a statement nobody classified |
| 31 | When a statement was added, and how often it has been recalled, are a table in the same SQLite file as the word index | a statement that has been recalled fifty times and one that has never been recalled are not equally worth keeping, and neither fact is in the Markdown, whose bytes must stay the record and stay compatible with what upstream and a UltraRAG UI read. Putting the history in the file the read already has open means a recall is one UPDATE per statement answered, in the transaction the answer was read in: atomic without a temporary file, no background thread, and proportional to the answer rather than to the memory. The file is therefore no longer purely derived — deleting it costs a re-read and the counts, which is stated rather than assumed, and the document itself is never in its power to lose. A separate JSON file was the first shape of this and needed a background writer precisely because it was not in the transaction |
| 32 | A column that cannot vary is not kept, and the field is called kind | a kind column held statement for every row and said nothing, because a scope is one document of statements. A column that cannot vary is a claim the schema makes and does not keep. The field that does vary is the statement's own kind, and it is the column now (kind), indexed with the statement so a read can be narrowed to one |
| 33 | The record is one SQLite file, and MEMORY.md is a rendering of it written on request | two copies of the same statements is one thing too many, and the copy that is written by hand is the one that drifts. So the record is the table, and the document is rendered from it — by --export, and by the browser view on every request — and is never read back as memory. A memory written before this existed is recovered from the file it is in, and a rendering that a server wrote is left alone until someone asks for it to go. What is given up is stated rather than implied: a UltraRAG UI can no longer read the global memory out of the shared storage tree, and a memory is read with this server rather than with cat |
| 34 | A statement's kind is a column, not a prefix on its line | RULE: the draft lives in docs/ is a tagged line, and a file of them reads as data rather than as something a person wrote — which is what a memory is for. The kind is still one run of block letters and still the caller's, and it is still returned with every answer; what changed is that it is held with the statement instead of inside its text, so a read narrows to a category rather than searching for a word that is no longer there. A file written before this is read once on upgrade and its prefixes become kinds |
| 35 | The export holds the statements and nothing else | A file that carried the kinds, the dates, and the counts would be a second record to keep in step, and keeping two in step is how a memory loses a statement. So the file is a rendering: the words, in order. It is regenerated on every write, so it cannot drift; it is not read, so it cannot be wrong; and a memory lost with its database comes back as what was said rather than as how it was said, which is the honest limit of a file that was never a backup |
| 36 | A recall searches both memories and ranks the results together | a caller asking what is remembered is asking one question, and the answer is the best statement from either memory rather than the best from each in turn. The scores are the same measure in every scope, so a statement from the account's memory does not outrank a better match from the project's, and every statement names the scope it came from so the caller can say which memory it was quoting |
| 37 | scope is one argument with local as its default, rather than a tool per memory | a project fact written to the account's memory is wrong everywhere else and a standing instruction left in one project goes unnoticed in the rest. One tool with a defaulted argument makes the project the default that a caller has to choose against, and it leaves the two memories named in one place instead of across seven tool names |
| 38 | A rendering is written only when it is asked for, and is not read back | a file the server rewrites whenever the memory changes is a second record in everything but name, and one nobody asked for. So --export writes it, the page renders it, and nothing else does. A rendering is stale the moment a statement is recorded, which is what a rendering is, and the honest way to offer one is to say it is generated rather than to keep it silently in step |
| 39 | A document from an older version is read and then removed, and one this server wrote is left alone | a memory of an earlier version exists as a file, so the file has to be read once for anything the record lacks — and then it is gone, because a scope holding both a document and a database is the state a user recognises as two versions of one thing. A rendering this server wrote is a different case: it was asked for, it is not a record, and --retire-document is how it is asked for again |
| 40 | Every tunable is a declared setting, and a config.toml is an overlay rather than a replacement | a value written in two places is a value that will differ. So the registry in settings.py and the packaged default.toml are the two halves of one declaration, a layer may only set a key the registry declares, and every effective value reports the layer that supplied it. The user layer is the one that applies globally, because a memory is shared by every project on the account; the project layer is there for the project that differs |
| 41 | A setting is declared as identity, runtime, or retrieval | the three cost different things when they change, and a flat list of names does not say that. An identity setting is written beside the statements it produced, so changing it means they are embedded again rather than compared across two spaces; a retrieval setting reorders the next answer and changes nothing about what is remembered. A config file that cannot say which of the two a number is would let someone set one and expect the other |
| 42 | The layer stack is a pinned library rather than a second copy in this package | two servers in this collection were carrying the same nine functions, and the first copy refused the environment and --set layers because its coercion accepted only the types a TOML file produces. The stack is the same problem solved once: the registry, the merge, the coercion, the provenance, and the path helpers are shared behind a pinned commit, while the keys, the packaged default.toml, the environment names, and the two directory names stay here. Skew between this server's pin and the other server's is allowed, because a floating or vendored dependency is what would make one server's release depend on another's |
| 43 | A repetition is collapsed in the answer, and a write is never refused for it | a write does not wait for a vector, so deciding on a write decides with whatever has been embedded so far: the same statement would be refused once and accepted the next time, and a caller who cannot record is a caller with no memory. The read is where the whole answer is in hand, so that is where it is decided — the words first, which cost nothing, then the cosine within retrieval.duplicate_cosine for the statement that says the same thing in other words. The best-ranked of a repeated pair is the one shown, and collapsed_repetitions says what went and which test caught it, so an answer is never quietly short. What is given up: a memory may hold the same statement twice, and only a read reveals it |
Develop and test
uv sync --frozen
uv run --frozen ruff format --check .
uv run --frozen ruff check .
uv run --frozen pytest
uv buildThe suite runs without any UltraRAG checkout. It covers one scope's behaviour (template, create-on-read, the statement format, append, refusals, and that a scope is one document), the four tools over the real stdio server, the local reranker's call into its runtime, project isolation between two projects, the shared global tree, a byte-for-byte comparison against the standing document UltraRAG's own server produced (tests/fixtures/upstream/README.md records how it was captured), and the browser view: its scopes, its reads, its one write, the scope a request may name, and the routes a page can call.
The formats are also compared against a real checkout when you have one: point ULTRARAG_CHECKOUT at it and tests/test_upstream_differential.py starts that checkout's own servers/memory/src/memory.py over stdio, gives it the same inputs, and compares the bytes both wrote, the payload keys both returned, and the surface differences this package chose. It is the assurance a wrapper would have given by construction, without depending on a checkout at runtime.
ULTRARAG_CHECKOUT=/ABSOLUTE/PATH/TO/UltraRAG uv run --frozen pytest -m upstreamCredit and licensing
The memory model, the file names, and the file formats are UltraRAG's, from OpenBMB/UltraRAG at commit 3a709a2 (Apache-2.0), and the tools are its memory tools renamed and extended — see ADR 0001. Credit belongs to the UltraRAG team and contributors, including participants from THUNLP, NEUIR, OpenBMB, and AI9stars.
This package is licensed under the Apache License 2.0. It is an independent project: not an official release of UltraRAG, not affiliated with or endorsed by its maintainers, and the UltraRAG name describes what it serves.
The graph memory server (graph-memory-ultra-rag-mcp-server) is the collection's other memory project, which extends memory into a typed, relational store.
Available Tools
4 toolsforget_memoryA
Removes one statement that is no longer true.
Pass the exact text a recall returned. It matches words and never
meaning, so it removes that statement or nothing: if the text matches
more than one, nothing is removed and the answer names them.
Use it when a recalled statement is contradicted, and when a statement has simply stopped being true.
| Name | Required | Description | Default |
|---|---|---|---|
| text | Yes | The exact `text` of the one statement to remove, as a recall returned it: not a summary, not a fragment, and not the kind. The match is on those words and never on meaning, so it removes that statement or nothing. | |
| scope | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare a non-read-only, non-destructive, non-idempotent mutation; the description goes further by disclosing the matching rule ('never meaning'), the ambiguity behavior (nothing removed, answer names the matches), and the calling convention (pass text exactly as recall returned it). This is real context beyond the annotations, though it never states whether the removal is reversible.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the action, then the calling convention, then the matching rule. Efficient overall, though the final 'when to use' sentence slightly restates the same idea twice.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present and annotations covering the safety profile, the description covers purpose, matching semantics, and trigger conditions adequately. The one gap is the local/global scope decision, which is a subtle routing choice left entirely to the schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 50%: `text` is well documented in both places and the description usefully reinforces the exact-match rule, but `scope` (local vs global, an easily-misused default) is only covered by the schema and is never mentioned in the description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Removes one statement') with the exact scope of the action: exactly one statement, matched by words. An agent can immediately distinguish it from recall_memory and record_memory without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit when-to-use: 'Use it when a recalled statement is contradicted, and when a statement has simply stopped being true.' It even handles the edge case of matching more than one statement. It does not name a sibling as the alternative, but the recall-then-forget workflow is clearly implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
recall_memoryARead-onlyIdempotent
Returns what is remembered that matches these words.
It searches this project's memory and the account's memory together and ranks the results by how well they match, so the best statement wins whichever memory it is in. Each one names its scope.
Give it a few words, not a sentence. When nothing comes back, that is what the search found: try other words, or pass a kind.
| Name | Required | Description | Default |
|---|---|---|---|
| kind | No | Return only the statements filed under this kind, one word in block letters. A kind is a column of the record rather than part of a statement's words, so it is asked for here rather than searched for. | |
| limit | No | How many statements to return across both memories, at most 50. This caps matches, not the size of a file read. | |
| query | Yes | What to recall, as words. Answered from the statements that match, by their words and by their meaning. Required: this server answers a question rather than returning a memory whole. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly/idempotent/closed-world, but the description adds real behavior beyond them: results span two memory scopes, are ranked by match quality, each names its scope, and an empty return means 'not found' rather than an error. That empty-result semantics is valuable and not derivable from annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Short and front-loaded with the core action, then layered constraints. The prose style is somewhat ornate ('so the best statement wins whichever memory it is in'), but no sentence is truly wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present and annotations covering safety, the description needs only to convey scope and empty-result behavior, which it does. Minor gaps: no explicit sibling routing or pagination note, but nothing essential is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so each parameter is already documented. The description adds only marginal meaning — implying query is a few words and that 'kind' can be passed to widen a failed search — so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource — returns remembered statements matching given words — and clarifies it searches project and account memory together with ranking. The distinction from siblings (record_/forget_) is inferable from the verb but not called out explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives concrete usage advice: 'Give it a few words, not a sentence' and, on empty results, 'try other words, or pass a kind.' This tells the agent how to use and recover, though it never explicitly names alternative tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
record_handoffA
Records this session's handoff for the next one, replacing the last.
One statement at the top of this project's memory. The previous handoff
is removed as part of the same call, so a project holds one handoff
rather than a list of them, and replaced says how many went.
Write what the next session needs to pick the work up: what is done, what is in flight, and what to do first.
| Name | Required | Description | Default |
|---|---|---|---|
| content | Yes | This session's handoff, in a few sentences: what is done, what is in flight, and what the next session does first. It replaces the previous handoff in this project, which is a project's own state. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With annotations covering safety (readOnlyHint=false, destructiveHint=false, idempotentHint=false), the description adds real value by disclosing the replacement semantics — the prior handoff is deleted in the same call and a project holds one handoff, not a list. It also names the `replaced` count. It stops short of saying whether the removal is recoverable, which matters given destructiveHint=false despite a documented delete of prior content.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action, then the replacement rule, then what to write. Each sentence earns its place, though the third sentence partially duplicates the schema's parameter description.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists and the description already references the `replaced` return field, so return values need no further explanation. For a one-parameter write tool, everything an agent needs to invoke it correctly is present.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% for the single `content` parameter, and the schema already spells out the same guidance (a few sentences, replaces the previous handoff). The description restates rather than extends that, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb (records) and resource (this session's handoff), plus the critical scoping fact that it replaces the previous handoff. That single-slot semantics clearly separates it from record_memory, which appends discrete memories.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives clear context for when to call it ('what the next session needs to pick the work up: what is done, what is in flight, and what to do first'), which is actionable. It does not explicitly name record_memory as the alternative for non-handoff memories, so routing is implied rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
record_memoryA
Records one statement in this account's or this project's memory.
Call it for something meant to hold beyond this reply: a rule, a principle, a decision, a correction, a preference, a path. Not for a request that is finished when it is answered.
Recall first, with words covering the same thing. If a statement comes back that says the same, record nothing. If one comes back that contradicts it, forget that one first.
scope decides where it goes. "local" is this project, inside the
repository. "global" is across projects, in the account's memory. Choose
it from what the user means and how far it reaches — a fact about this
code, a path, a convention here is local; a preference about how they
want to be spoken to, or a rule about their own work rather than this
repository, is global. Prompts rarely say so, so judge the substance and
not the wording.
| Name | Required | Description | Default |
|---|---|---|---|
| kind | No | The kind to file it under, chosen from what the statement means. RULE for an instruction or a standing fact, PLAN for what the project is meant to be or achieve, PREFERENCE for what the user likes, CORRECTION for something to stop doing, and any other word in block letters if the statement fits one better. The same word for the same kind of thing, because a recall filtered by kind returns everything filed under it. If nothing fits, ITEM: it is a kind like any other, and a statement filed under it is found by asking for it. Nothing here interprets the word. | ITEM |
| scope | No | Where the statement belongs. "local" is this project, inside the repository, and is the default. "global" is across projects: the account's memory, shared by every project on this machine. Read it off what the user means and how far it reaches rather than off the wording, because most prompts carry no marker of where a statement should go. | local |
| content | Yes | The one statement to remember, in the user's own words where it is theirs. A line or a short paragraph. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=false, idempotentHint=false, and destructiveHint=false, so the agent already knows this is a non-idempotent write. The description goes beyond that by explaining the deduplication discipline (recall before record, skip duplicates, forget contradictory entries first) and by spelling out the scope semantics (local = this project/repository, global = the account's memory across projects). It does not disclose storage limits, rate limits, or failure modes, so it earns a 4 rather than a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-structured with front-loaded purpose, then eligibility, then workflow, then scope disambiguation. It is somewhat repetitive in its insistence that prompts rarely mark scope ('Prompts rarely say so, so judge the substance and not the wording') and the scope sentences are longer than needed. Still, nearly every sentence carries distinct operational guidance.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 3-parameter mutation tool with full schema coverage, annotations covering the safety profile, and an output schema (so return values need not be explained), the description supplies exactly the missing pieces: when to write, the recall-first workflow, and how to resolve scope ambiguity. Nothing an agent needs to invoke it correctly is absent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, with the schema itself carrying detailed semantics for kind, scope, and content. The description reinforces scope ('local is this project, inside the repository; global is across projects, in the account's memory') and gives guidance on choosing it from substance rather than wording, which adds meaningful interpretation beyond the raw schema. That interpretation is the valuable addition; the kind parameter's semantics are left entirely to the schema, so this lands at a 3.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource ('Records one statement in this account's or this project's memory') and clearly distinguishes what qualifies ('something meant to hold beyond this reply' vs 'a request that is finished when it is answered'). It does not directly name the sibling tools (recall_memory, forget_memory), so an agent must infer the relationship from the described workflow.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit when-not guidance ('Not for a request that is finished when it is answered') and a prescriptive prerequisite workflow: 'Recall first, with words covering the same thing. If a statement comes back that says the same, record nothing. If one comes back that contradicts it, forget that one first.' This names the alternatives (recall, forget) and the conditions that route to them, leaving little to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
4 tool updates
v0.8.0- First observed
forget_memory - First observed
recall_memory - First observed
record_handoff - First observed
record_memory
TDQS
Scored across 4 tools
The four tools have largely distinct roles: record_memory creates/overwrites a statement, recall_memory searches, forget_memory removes, and record_handoff manages session handoffs. There is minor overlap because record_handoff also writes memory, but the descriptions clearly differentiate it by its replace-last semantics and session-focused purpose.
All names follow a consistent snake_case verb_noun pattern: record_memory, recall_memory, forget_memory, record_handoff. The convention is predictable and easy to parse.
Four tools is well-scoped for a memory server, covering the core lifecycle operations (record, recall, forget) plus a specialized handoff tool. No tool feels redundant or missing for the stated purpose.
The surface covers create, read, delete, and session handoff, but lacks a direct update operation. Agents can work around this by forgetting the old statement and recording a new one, so the gap is minor but noticeable.
Maintenance
Related MCP Connectors
Read-only MCP server: let AI agents read your ORANO saved-video library, tasks, and memory.
Private-by-default, local-first memory/context/task orchestrator for MCP apps and agents.
Agent-native MCP server over the public saagarpatel.dev corpus. Read-only, stateless.
Shared memory and actions for Claude, Kiro, OpenAI, Cursor, and other MCP-compatible AI clients.
Related MCP Servers
AlicenseAqualityBmaintenanceExposes local memory vaults as 6 MCP tools over stdio, enabling clients like Windsurf, Zed, and Claude Desktop to read and write memory.6252 npm2MIT- AlicenseAqualityBmaintenanceHosted memory for coding agents, exposing commit_mem, recall, and status over stdio. Each write carries provenance, contradictions surface as inspectable state changes, and a bad run can be rolled back from a snapshot.8305 PyPI1Business Source 1.1
- AlicenseNot gradedqualityBmaintenanceLocal MCP server exposing the dek-memory hybrid memory engine over stdio, with optional ChatGPT tools for managing account memories via a browser profile.MIT
- AlicenseNot gradedqualityAmaintenanceProvides a drop-in MCP memory server over stdio using a local SQLite store, enabling agents to store, search, and audit memory with provenance, scoped access, and contradiction tracking without requiring API keys, LLMs, or embedding providers.551 npm16Apache 2.0