loreweave
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@loreweavewhat's the status of project atlas?"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Most knowledge tools are write-only. You capture diligently, the vault grows, and six months later you can't find the thing you know you wrote — because retrieval is keyword search over prose, nothing ever resurfaces on its own, and nothing notices when what you wrote last year stopped being true.
Loreweave is the layer that fixes that. Point it at a folder of markdown (Obsidian or plain) and it builds a knowledge graph, a bitemporal fact store, and a memory model over your notes — then hands them to you through a CLI and to your AI agents through MCP.
Three guarantees, enforced by the code rather than promised:
Your files win. User markdown is never mutated; the engine only appends, and only under
lore/. The vault is the source of truth — the index is a cache you can delete at any time and rebuild identically (there's a test for that).No LLM anywhere in the core. Indexing and retrieval use zero tokens and make zero network calls. Same vault, same query, same answer — forever.
Memory you can read. Every fact an agent stores is a markdown line you can open, edit, and
git diff.
Quickstart
cd ~/my-vault
npx loreweave init # creates .lore/
npx loreweave index # incremental sync; one changed note: 32 ms at 1k notes, 656 ms at 20k (see Scale)
npx loreweave search "why did we drop the queue design"
npx loreweave ask "what's the status of project atlas"
npx loreweave dream # what's duplicated, contradicted, stale, unlinkedZero configuration and no network: out of the box it runs on BM25 + knowledge-graph spreading activation. Add local embeddings when you want them:
// .lore/config.json
{ "embedding": { "provider": "ollama", "model": "mxbai-embed-large" } }Pick mxbai-embed-large (670 MB) for quality or nomic-embed-text (274 MB) when disk
and indexing speed matter more — required task prefixes are applied automatically for
both, and for the E5, BGE and Arctic families. The measured difference between the two
is in Benchmarks. No embedding provider means lexical + graph retrieval,
still fully functional.
Works with non-English vaults: Chinese, Japanese and Korean text is segmented per character so it is searchable at all, and other scripts index as written.
Related MCP server: Engram
What it does
1. Knowledge that has a timeline. Facts are bitemporal: when they were true in the
world (valid_from/valid_until) and when the system learned them (recorded_at).
Contradictions supersede rather than overwrite, so history stays queryable.
$ lore assert "Ledger Format" status draft --valid-from 2026-01-01
$ lore assert "Ledger Format" status final --valid-from 2026-08-01
✓ Ledger Format :: status :: final
superseded: "draft" (now valid until 2026-08-01)
journal: lore/journal/2026-08-01.md(lore and loreweave are the same binary — npm i -g loreweave gives you both;
npx loreweave works without installing.)
Both time axes are queryable, which is what makes it bitemporal rather than merely
historical. --as-of asks what was true then; --as-known-at asks what was
believed then. They disagree exactly when you learn something after the fact — which
is when you most need to reconstruct what a past decision was actually based on:
$ lore facts --subject Vendor --as-of 2024-06-01
Vendor :: reliability :: poor — outage postmortem (2024-01-01 → now)
$ lore facts --subject Vendor --as-known-at 2024-06-01
Vendor :: reliability :: good (2024-01-01 → 2024-01-01) [superseded]Which fact wins is decided deterministically (newest valid-time, provenance as tiebreak) — never by asking a language model which one looks fresher. And the whole history of anything is one command — every value change merged chronologically with the dated prose that mentions it:
$ lore timeline Project Atlas
2024-01-15 status: planning (until 2024-09-01)
2024-02-10 • [[Project Atlas]] kicked off with a three-person crew. [kickoff.md]
2024-09-01 status: planning → active
2025-06-20 • The [[Project Atlas]] midpoint review went long but well. [review.md]Temporal-graph systems build this by running an LLM over every document at ingestion. Here the supersede chain has been maintained all along, so it is a read-side join: no LLM, no network, same answer every time.
2. Retrieval that follows connections, not just words. Queries fuse BM25, dense similarity (when configured), and Personalized PageRank over the vault's own graph — wiki-links, shared entities, tags, co-occurrence. Two-hop neighbors surface even when they share no vocabulary with your query, and every result tells you why:
• data/glacier-dataset.md#@0 (0.0327) ⟨via amara osei⟩
The Glacier Dataset holds meltwater sensor readings from 2019-2024.3. Memory with dynamics. Every passage carries FSRS-style stability and retrievability — a power-law forgetting curve. Passages that actually get used (not merely retrieved) decay slower; important-but-fading knowledge gets surfaced for review instead of silently rotting. Nothing is ever deleted.
4. It dreams. lore dream is an idle-time consolidation pass that reviews the vault
and reports duplicate passages, contradicted facts, stale knowledge, missing links
between notes that clearly belong together, and orphans. With --apply it writes a
digest and a review queue — append-only, under lore/. It never rewrites your prose:
LLM-driven whole-file rewriting is a documented failure mode — each rewrite quietly
drops details until the file collapses to mush — so the architecture forbids it.
5. Questions retrieval can't answer. Counting, grouping, and date-range queries run as deterministic SQL over the fact store, not as vibes over embeddings:
$ lore count --predicate trip_to --since 2025-01-01 --until 2025-12-31
2 Japan
1 Kenya6. Facts come from your notes, not from a form. Nobody hand-writes
- [fact] X :: y :: z, so the extractor mines the conventions vaults already use:
status: shipped # frontmatter
- owner:: Priya # Dataview inline field
- [location] Hyderabad # Basic Memory observationOnly unambiguous field syntax is accepted automatically. Prose formatting like
- **Owner:** Priya is precise on entity notes and noisy on report notes, so it is
opt-in (facts.extract: "all") — or an agent can review candidates via
lore_propose_facts and assert the real ones. Judgement stays out of the index.
7. Time means when it happened, not when you saved the file. --since and
--until filter on content time — frontmatter dates, dated filenames
(2025-03-14-standup.md), or dates in the text — falling back to file mtime only when
a note carries no date of its own. lore watch keeps the index current so you never
have to remember to reindex.
8. Built for agents. An MCP server exposes 15 typed tools so Claude Code, Cursor,
or any MCP client can use your vault as durable memory. Session continuity is a query,
not a paraphrase — lore resume returns exactly what changed since the agent last
connected, computed from record time:
$ lore resume
since 2026-08-11 15:55
~ lore/journal/2026-08-11.md
+ Project Atlas :: status :: shipped (since 2026-08-11)
± Project Atlas :: status: active → shipped
$ lore resume
since 2026-08-11 15:56
nothing changedAlternatives that summarize the previous session with an LLM inject a paraphrase; this is a deterministic diff. Full setup in Agent memory.
The CLI
Command | What it does |
| create |
| incremental sync of vault → index |
| hybrid retrieval with provenance |
| extractive answer: current facts + top passages (no LLM needed) |
| query the fact store (200 rows unless |
| chronological history: fact changes merged with dated mentions |
| what changed since the last resume: notes, facts, supersessions |
| important-but-fading knowledge to revisit or archive |
| record a fact (journalled, supersedes) |
| close the current fact in a slot |
| aggregate over fact history |
| append a timestamped line to |
| consolidation pass + optional digest/review queue |
| reindex automatically as the vault changes |
| reinforce a passage that proved useful |
| export the graph |
| health check: broken links, integrity, coverage |
| vault statistics and top entities |
| start the MCP server on stdio |
Use it as agent memory (MCP)
// Claude Code: .mcp.json (or claude_desktop_config.json)
{
"mcpServers": {
"loreweave": {
"command": "npx",
"args": ["-y", "loreweave", "--vault", "/path/to/vault", "serve", "--mcp"]
}
}
}Tool | What the agent gets |
| hybrid retrieval with provenance and temporal filters |
| one-call session context: relevant passages + current facts + recent changes |
| full text of a note by path |
| write/close facts — journalled, superseding, never destructive |
| point-in-time fact queries ( |
| an entity's merged fact + prose chronology |
| exactly what changed since the agent last connected |
| important-but-fading passages worth resurfacing |
| count/group-by over fact history |
| append a timestamped line to the inbox |
| reinforcement signal: this passage actually helped |
| extraction candidates for the agent to review and assert |
| the consolidation report (duplicates, contradictions, stale, orphans) |
| trigger a reindex |
Facts asserted through MCP are written back to lore/journal/YYYY-MM-DD.md as readable
markdown, so an agent's memory is something you can open, edit, and git diff:
- [fact] Ledger Format :: status :: final {valid_from=2026-08-01, confidence=0.9, source=stated}Delete .lore/ and reindex — every fact and edge is reconstructed from those files.
- [fact] lines are replayed out of any note, but the source= attribute is read back
only under lore/journal/, which is the only path loreweave writes. The same line in an
ordinary note still records the fact, as extracted — a note can state something; it
cannot say that you stated it.
Compared to hosted memory services (Mem0, Zep): those run LLMs at write time to extract and summarize into their own store; loreweave runs no model in the core, keeps memory in your files under your version control, and makes every retrieval reproducible. The trade: they do abstractive summarization, this engine deliberately does not. Retrieval quality against their published benchmarks is below — with the caveats stated, because most published agent-memory numbers measure end-to-end QA with an LLM, which is a different quantity than retrieval.
Benchmarks
These are third-party benchmarks with relevance labels nobody here chose. All loreweave
numbers are retrieval metrics — it finds the evidence, it does not write the
answer — so they are not comparable to the end-to-end QA accuracy quoted by systems
that put a language model after retrieval. Reproduce any number:
docs/benchmarks.md. Field comparison with the traps called out:
docs/scoreboard.md.
LongMemEval_S (ICLR 2025) — 500 questions, each with ~50 sessions of chat history; find the sessions holding the evidence. Session-level recall, all 500 questions:
configuration | R@1 | R@5 | R@10 |
model-free (no embeddings, no network) | 0.552 | 0.899 | 0.943 |
+ embeddings & 8-turn chunking | 0.590 | 0.959 | 0.983 |
A third-party benchmark of the same task reports BM25 alone at 86.2% R@5, BM25+vector hybrid at 95.2%, and a vector-only system at 96.6%. One caveat before the comparison: their metric scores a question 1 if any gold session is retrieved, ours scores the fraction of gold sessions found — ours is the stricter definition, so treat cross-system gaps of under a point as noise. With that stated: at 95.9% loreweave is past the hybrid and just short of the vector-only system, with a local model, no network, and a lexical index it can fall back to.
BEIR / SciFact — 5 183 scientific abstracts, 300 claims, scored by nDCG@10 (ranking quality: rewards putting relevant documents nearer the top):
configuration | nDCG@10 | Recall@10 |
BM25 baseline (BEIR paper, Anserini) | 0.665 | — |
loreweave, model-free | 0.681 | 0.817 |
loreweave + nomic-embed-text | 0.727 | 0.865 |
loreweave + mxbai-embed-large | 0.742 | 0.884 |
LoCoMo — 10 long conversations, 1 982 evidence-labelled questions, turn-level recall. No comparable third-party retrieval number exists (published LoCoMo results are LLM QA accuracy), so these are offered as a target rather than a comparison:
R@1 | R@5 | R@10 | R@20 | |
model-free | 0.337 | 0.538 | 0.610 | 0.656 |
+ embeddings | 0.318 | 0.532 | 0.627 | 0.705 |
What the numbers say, plainly. Recall is strong and rank-1 is the weakness —
LongMemEval R@10 0.983 vs R@1 0.590 — which is why figures here are quoted at R@5 and
why an agent consuming these results should read a top-5 list, not trust rank 1. The
single biggest quality lever measured is the embedding model itself: mxbai over nomic
is worth more than every downstream tuning combined on document (BEIR) and session
(LongMemEval) retrieval — and measures flat on LoCoMo's single-turn passages, so it
is a scale effect, not magic. A cross-encoder reranker is available (rerank in config) but earns its keep only
for rank-1 consumers without embeddings — stacked on embeddings it loses recall on
both public benchmarks, so leave it off unless that trade is yours. Full analysis,
including the failed experiments and the defect that benchmarking caught:
docs/evaluation.md.
Loreweave also ships an internal regression benchmark (npm run eval) over three
purpose-built vaults — including a temporal-perturbation test where the shipped config
scores 100% consistency vs BM25's 0%, gated in CI on every push. Internal corpora
are good for regression and worthless as proof, so the details live in
docs/evaluation.md rather than here.
Scale
Measured, like the quality numbers — npm run scale reproduces this on your own
machine (synthetic vaults, 3 blocks per note, dense interlinking):
notes | blocks | entities | edges | full index | incremental | search p50 | p95 | dream | heap |
1 000 | 3 000 | 2 766 | 17 k | 1.2 s | 32 ms | 3 ms | 4 ms | 0.2 s | 63 MB |
5 000 | 15 000 | 13 766 | 85 k | 6.1 s | 165 ms | 9 ms | 12 ms | 1.0 s | 145 MB |
20 000 | 60 000 | 55 016 | 339 k | 22.5 s | 656 ms | 38 ms | 53 ms | 4.8 s | 333 MB |
Full index scales at 0.93× per note from 5 k to 20 k — linear or better, no
superlinear step hiding in the middle. "Incremental" is one changed note, which is what
lore watch actually does all day. Everything here is one process, one SQLite file, no
daemon — and search at 3-38 ms p50 is fast enough to sit inside an agent loop.
How it works
vault/*.md ──parse──▶ notes · blocks · wiki-links · tags · entities
│ (incremental: mtime + content hash)
▼
SQLite .lore/index.db ── disposable cache, rebuildable
│
┌───────────────────┼────────────────────┐
▼ ▼ ▼
graph (CSR) retrieval facts
blocks ∪ entities BM25 + dense + PPR bitemporal, supersession,
2-iteration PPR → weighted RRF deterministic freshness,
α = 0.5 → FSRS boosts aggregates
└─────────┬─────────┴──────────┬─────────┘
▼ ▼
dream (idle-time) CLI · MCPDesign rules the code enforces:
Files win. User markdown is never mutated. The engine only appends, and only under
lore/.Invariants in code, not prompts. Schema, migrations, graph construction, and supersession are typed, versioned, and tested — no LLM re-specifies them at runtime.
No LLM required anywhere in the core. Indexing and retrieval use zero tokens. Language models are consumers of this engine, not dependencies of it.
Everything is re-derivable. A full rebuild reproduces byte-identical derived state (there's a test for that).
Research lineage
Every significant choice traces to 2024-2026 literature; the survey lives in
docs/research/.
Choice | Source |
Dense-sparse fusion + PPR with dense reset probabilities | HippoRAG 2 (ICML 2025), 2502.14802 |
Shallow 2-iteration PPR, heterogeneous nodes | NodeRAG (2025), 2504.11544 |
Relation-free graph — no LLM triple extraction | LinearRAG (ICLR 2026), 2510.10114 |
No index-time community summarization | LazyGraphRAG (Microsoft, 2024) — same quality at 0.1% index cost |
Route/fuse instead of graph-everything | GraphRAG-Bench (ICLR 2026), 2506.05690 |
Bitemporal facts, invalidate-never-delete | Zep/Graphiti (2025), 2501.13956 |
Power-law forgetting, use-gated reinforcement | FSRS; RMM (ACL 2025), 2503.08026 |
Consolidation as idle-time work | Sleep-time compute (Letta, 2025), 2504.13171 |
Never let an LLM rewrite whole memory files | ACE (2025), 2510.04618 |
Computable facts for aggregation | User as Code (2026), 2606.16707 |
Fine-grained indexing + fact-augmented keys | LongMemEval (ICLR 2025), 2410.10813 |
Library use
import { openContext, indexVault, search, assertFact, queryFacts, dream } from 'loreweave';
const ctx = openContext('/path/to/vault');
await indexVault(ctx.store, ctx.root);
const hits = await search(ctx, 'streaming compaction', { k: 5 });
assertFact(ctx, { subject: 'Atlas', predicate: 'status', object: 'shipped', validFrom: '2026-08-01' });
const asOfMarch = queryFacts(ctx.store, { subject: 'Atlas', asOf: '2026-03-01' });
const report = dream(ctx);
ctx.close();Development
npm install
npm test # 440 tests
npm run eval # retrieval benchmark vs BM25 baseline
npm run typecheck
npm run buildRequires Node ≥ 20. Single native dependency (better-sqlite3). Tested in CI on
Linux, macOS and Windows across Node 20 and 22.
License
MIT © Ambuj Upadhyay
Available Tools
15 toolslore_aggregate_factsCount factsA
Deterministic aggregation over fact history — counts grouped by object/subject/predicate with date-range filters. Use for "how many X", "which Y most often" questions; similarity search cannot answer these reliably. Returns { groups, totalGroups, limit }: groups is the top limit (default 100), so read totalGroups for "how many distinct values are there" rather than counting groups, and raise limit if you need the tail.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | max groups returned (default 100); totalGroups always reports the real count | |
| since | No | ||
| until | No | ||
| groupBy | No | ||
| subject | No | ||
| predicate | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does well: it declares determinism, the return envelope {groups, totalGroups, limit}, and the critical interpretation rule that groups is truncated to limit while totalGroups is the true count. It omits cost/performance characteristics and any permission or auth notes, which keeps it short of a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, each earning its place: purpose first, then the routing rule, then the return-value caveat. The totalGroups pitfall is front-loaded where an agent will read it before writing code.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so describing the return envelope and the groups-vs-totalGroups distinction is exactly the right compensation, and it is complete for calling the tool correctly. The residual gap is filter parameter syntax (especially date formats), which is not addressed anywhere.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 17%, so the description must compensate, and it does so well for limit (default 100, raise it for the tail) and implies groupBy's role. However, since/until date format and the semantics of the subject/predicate filters remain undocumented in both schema and description, leaving several of the 6 parameters ambiguous.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('deterministic aggregation over fact history') plus the exact grouping dimensions (object/subject/predicate) and date-range filtering. It also implicitly distinguishes itself from similarity-based siblings, so an agent can tell it apart from lore_search without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives concrete trigger questions ('how many X', 'which Y most often') and an explicit exclusion: similarity search cannot answer these reliably. This routes the agent between this tool and the search family with no inference required.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
lore_assert_factAssert a factA
Record an atomic fact (subject :: predicate :: object) with bitemporal validity. Contradicting facts in the same slot are superseded automatically (never deleted; history stays queryable). The fact is journalled to lore/journal/ in markdown, so the vault remains the source of truth. Use for durable knowledge: decisions, states, preferences, relationships.
| Name | Required | Description | Default |
|---|---|---|---|
| object | Yes | ||
| subject | Yes | ||
| predicate | Yes | snake_case relation, e.g. works_at, status, lives_in | |
| validFrom | No | ISO date when it became true (default today) | |
| confidence | No | ||
| sourceType | No | stated: user said it · extracted: from a document · inferred: your deduction | |
| validUntil | No | ISO date when it stops being true, if known |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does well: it discloses automatic supersession of contradicting facts, that facts are never deleted and history remains queryable, and that the fact is journalled to lore/journal/ in markdown. It omits permissions, error cases and return behavior, but the key non-obvious side effects are surfaced.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four sentences, front-loaded with the operation and its side effects, then the usage cue. Dense but nearly every clause adds value; no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema and no annotations, yet the description covers the mutation semantics, supersession model, and persistence location that an agent most needs. It stops short of documenting half the optional parameters, but is otherwise complete for the call.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 57%, so the description should compensate more. It clarifies the subject/predicate/object triple and the bitemporal (validFrom/validUntil) intent, which helps, but says nothing about confidence or the stated/extracted/inferred sourceType beyond what the schema already documents.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb (record) and resource (an atomic subject :: predicate :: object fact), making the operation unambiguous. It does not name sibling tools like lore_propose_facts, but the tripartite structure is enough to identify the operation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The final sentence gives clear when-to-use context ('durable knowledge: decisions, states, preferences, relationships'). However, it offers no when-not guidance and never distinguishes this from the related sibling lore_propose_facts, leaving the routing to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
lore_captureCapture a noteA
Append a timestamped line to lore/inbox.md (or another vault note). Use for fleeting observations worth keeping that are not atomic facts. The captured text is searchable immediately. Never overwrites anything, and never writes outside the vault — a path that escapes it, including through a symlink, is refused.
| Name | Required | Description | Default |
|---|---|---|---|
| to | No | target .md path (default lore/inbox.md) | |
| text | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden and does so well: append-only, never overwrites, never writes outside the vault, symlink escapes refused, and capture is immediately searchable. That is exactly the safety and effect profile an agent needs to decide it can call this freely.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four short sentences, front-loaded with the primary action and default target, then usage, then guarantees. No filler or repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 2-parameter mutation tool with no annotations and no output schema, the description covers effect, safety, and semantics thoroughly. The only minor gap is that it never says what a successful call returns (path written, timestamp, etc.), which the absence of an output schema makes slightly more relevant.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 50% ('text' is undocumented in the schema), but the description compensates by explaining the append-and-timestamp behavior of the content and restating the default/alternate target for 'to'. It does not cover the maxLength bounds, but the intent of both parameters is clear.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Append a timestamped line to lore/inbox.md') with the default target and an allowed override. It also implicitly distinguishes itself from lore_assert_fact by scoping to 'fleeting observations ... not atomic facts', so an agent can route between the two without reading schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says when to use it ('fleeting observations worth keeping that are not atomic facts'), which is real selection guidance. It does not name a sibling tool outright as the alternative, so the routing is inferable rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
lore_context_packSession context packA
Progressive-disclosure primer: vault stats, top entities, recently modified notes, currently-valid facts, and (if topic given) top search hits. Call once at session start to orient; then drill down with lore_search / lore_read_note. Every list here is a sample: when one is cut, a truncated field names it with { shown, of, rest } and the tool to call for the remainder — so treat a missing item as "not in this sample", never as "not in the vault".
| Name | Required | Description | Default |
|---|---|---|---|
| topic | No | optional focus topic |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It does disclose an important behavioral trait — every list is a sample and truncated lists expose a { shown, of, rest } field pointing at the remainder tool — which is genuinely useful context about partial results. However, it says nothing about cost, latency, permissions, or whether the primer is a pure read, so the behavioral picture is incomplete rather than wrong.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the payload list, then usage, then the truncation caveat. Dense but every sentence carries distinct information; nothing restates the name or title.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and no annotations, the description must cover return contents and semantics — and it does, enumerating the sections and the truncation contract. An agent has everything needed to call it and interpret a partial result correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and only one optional parameter exists, so the baseline would be 3. The description earns one extra point by explaining the parameter's effect on output ('if topic given, top search hits'), which tells the agent what changes when topic is supplied — meaning beyond the schema's 'optional focus topic'.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a concrete verb and enumerates the exact payload: vault stats, top entities, recently modified notes, valid facts, and optional search hits. This is clearly distinguishable from siblings like lore_search (targeted hits) and lore_query_facts (fact lookup), which the description itself reinforces.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says when to use it ('call once at session start to orient') and names the follow-on alternatives ('then drill down with lore_search / lore_read_note'). No inference required about sequencing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
lore_dream_reportConsolidation reportA
Run the consolidation pass: duplicate passages, contradicting/recently-changed facts, stale knowledge needing review, suggested missing links, orphan notes. Leaves the vault untouched unless apply=true (which writes a digest + review queue under lore/); it does perform index maintenance either way, which changes no results. Findings are a summary — pass verbose:true for every one.
| Name | Required | Description | Default |
|---|---|---|---|
| apply | No | ||
| verbose | No | return every finding rather than a summary |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does well: it clarifies that the vault is untouched unless apply=true, that apply writes a digest and review queue under lore/, and that index maintenance happens either way but changes no results. It does not cover auth, rate limits, or the exact format of writes, but the key mutation and side-effect behavior is disclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, front-loaded with purpose, then mutation behavior, then output verbosity. Every sentence earns its place and there is no redundancy or filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema or annotations exist, so the description must provide context. It covers the tool's purpose, side effects, parameter effects, and output verbosity (summary vs. verbose). It lacks detail on the exact return format, but for a report tool the description is sufficiently complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 50% (apply has no schema description). The description compensates by explaining that apply=true writes a digest and review queue, and that verbose=true returns every finding rather than a summary. The verbose explanation duplicates the existing schema description, but the apply clarification adds real meaning beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Run the consolidation pass') and resource, then enumerates the exact finding types (duplicate passages, contradicting facts, stale knowledge, missing links, orphan notes). This clearly distinguishes it from most siblings, but it does not explicitly name an alternative tool (e.g., lore_review) or state when to prefer this over them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no explicit when-to-use or when-not-to-use guidance. It says 'Run the consolidation pass' but does not describe the scenario that should trigger this tool versus alternatives like lore_review or lore_search. Usage is only inferred from the purpose.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
lore_indexReindex the vaultA
Incrementally sync the markdown vault into the index. Call after writing files to the vault through anything OTHER than lore_* tools (an editor, another agent, plain fs writes) — the lore_* write tools index their own writes, so their content is searchable immediately without this.
| Name | Required | Description | Default |
|---|---|---|---|
| full | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It usefully discloses two traits: the operation is incremental, and lore_* writes are already indexed so this call is unnecessary after them. However, it says nothing about what happens with a full reindex, whether the operation is safe or idempotent, cost/latency, or any locking or permission requirements.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the action and immediately followed by the routing rule. Every clause earns its place, including the parenthetical examples that make the condition concrete.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple, single-parameter tool with no output schema, the description covers purpose and routing well. But it leaves the only parameter unexplained and provides no behavioral notes on what 'full' does, so an agent can invoke it correctly in the common case but not fully safely in the full-reindex case.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There is one parameter, 'full', with 0% schema description coverage, and the description never mentions it. The word 'Incrementally' faintly implies a non-full default, but the semantics of toggling 'full' (rebuild vs. incremental sync, cost, potential data effects) are left entirely undocumented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description pairs a specific verb ('incrementally sync') with a specific resource ('the markdown vault into the index'), and it names the sibling behavior it is not ('the lore_* write tools index their own writes'). An agent can tell exactly what this tool does and how it relates to the lore_* family without inspecting a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives an explicit trigger ('Call after writing files to the vault through anything OTHER than lore_* tools') and an explicit exclusion ('their content is searchable immediately without this'). The alternative cases (editor, another agent, plain fs writes) are enumerated, leaving nothing to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
lore_invalidate_factInvalidate a factA
Close the currently-valid fact in a (subject, predicate) slot without asserting a replacement — e.g. "no longer true". Journalled; history preserved.
| Name | Required | Description | Default |
|---|---|---|---|
| subject | Yes | ||
| predicate | Yes | ||
| validUntil | No | ISO date (default today) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It usefully discloses that the operation is journalled and history-preserving, but leaves key behavioral questions unanswered: what happens if no currently-valid fact exists, whether the close is idempotent, and what validUntil does retroactively or in the future.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single tight sentence plus a short tag; the disambiguating constraint ('without asserting a replacement') is front-loaded. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given a 3-parameter mutation tool with no annotations and no output schema, the description covers the core intent and the history-preservation trait but omits failure modes, validUntil semantics, and any return/confirmation behavior. Adequate but with visible gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 33%, so the description must compensate. It adds real value by framing subject/predicate as a slot identifier that selects the fact to close, but it adds nothing about validUntil beyond what the schema already says, leaving a third of the parameters undocumented in both places.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Names a specific verb (close/invalidate) and the exact resource scope: the currently-valid fact in a (subject, predicate) slot. The clause 'without asserting a replacement' cleanly distinguishes it from lore_assert_fact, so an agent can route between the two without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear usage context by contrasting with the replace/assert case ('without asserting a replacement — e.g. "no longer true"'), which implicitly defines when this tool is the right choice. It stops short of naming lore_assert_fact explicitly or stating exclusions, so it is clear but not fully prescriptive.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
lore_mark_usedMark passage as usedA
Reinforce passages that actually contributed to your answer (spaced-repetition signal: used memories decay slower). Call after citing a note.
| Name | Required | Description | Default |
|---|---|---|---|
| anchor | No | block anchor from search results; omit for whole note | |
| notePath | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations exist, so the description carries the full behavioral burden. It usefully discloses the effect of the call (used memories decay slower), which is real value beyond the name. However, it says nothing about idempotency, behavior on an unknown notePath/anchor, whether re-calling compounds the signal, or any auth requirement.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences with the action and its reinforcement signal front-loaded, then the invocation trigger. No filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple two-parameter, no-output-schema tool this covers the essentials, but it leaves an agent guessing about the unexplained notePath parameter and error/idempotency behavior — gaps the annotations cannot fill since none are provided.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 50% and the description explains neither parameter: notePath has no schema description and is not mentioned, and anchor is only implicitly hinted at by the word 'passages'. The description does not compensate for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Reinforce passages...') and clarifies the mechanism (spaced-repetition / slower decay), which distinguishes it from read/search/review siblings. It is clear what the tool does, though it never explicitly contrasts itself with lore_review, the nearest sibling.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit timing guidance — 'Call after citing a note' — which is exactly the moment an agent needs to invoke this. It stops short of stating when NOT to use it or naming the reinforcement alternatives (e.g., lore_review), so it is clear context without exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
lore_propose_factsPropose facts from a noteA
Returns candidate facts mined from a note's structure that are NOT yet in the fact store, for you to adjudicate. The engine only auto-accepts unambiguous field syntax (frontmatter, key:: value, - [key] value); prose formatting like - **Owner:** Priya is precise on entity notes and noisy on report notes, so it is surfaced here instead of assumed. Review these and call lore_assert_fact for the ones that are genuinely durable facts. This keeps judgement with you and out of the index.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| notePath | No | limit to one note; omit to sample the vault |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden and does substantial work: it discloses the engine's auto-accept rules (frontmatter, `key:: value`, `- [key] value`), why prose formatting is surfaced instead of assumed, and that only facts NOT yet stored are returned. It never explicitly states whether the call mutates the vault/index or requires auth, which is a gap for a tool feeding a write pipeline.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core purpose, then the mechanism, then the action. The closing line ('This keeps judgement with you and out of the index') is motivational rather than operational, so it burns a little space, but the rest is dense and informative.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-param read/surface tool with no output schema and no annotations, the description covers what is returned, why some content is surfaced rather than auto-accepted, and the follow-up action. Only the limit parameter and the explicit side-effect/auth profile are left unaddressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 50%: notePath is documented in the schema, limit is not. The description implies single-note vs. vault-wide sampling ('a note's structure' ... 'sample the vault') which aligns with notePath, but it says nothing about limit or how many candidates are returned by default. It partially compensates but does not close the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: returns candidate facts mined from a note's structure that are not yet in the fact store. It clearly distinguishes itself from the write-side sibling lore_assert_fact by framing itself as the adjudication/surfacing step.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly routes the agent: review the returned candidates and call lore_assert_fact for the ones that are genuinely durable facts. It also explains why prose formatting lands here rather than being auto-accepted, which tells the agent when this tool's output is worth acting on.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
lore_query_factsQuery factsA
Query the bitemporal fact store. Default: currently-valid facts. asOf answers "what was true on DATE", asKnownAt answers "what did we know on DATE"; includeHistory shows the full supersession chain. Prefer this over lore_search for factual slots (status, location, role, preference). Returns { facts } and, when the result is a sample, a truncated field with { shown, of, rest } — narrow by subject or raise limit before concluding a fact does not exist.
| Name | Required | Description | Default |
|---|---|---|---|
| asOf | No | ISO date: what was TRUE on this date | |
| limit | No | max facts to return (default 200) | |
| subject | No | ||
| asKnownAt | No | ISO date: what was KNOWN on this date. Facts recorded later are excluded however far back their validity was backdated — use it to reconstruct what a past decision was based on. Combine with asOf for "what was true then, as far as we knew then". | |
| predicate | No | ||
| includeHistory | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden; it discloses the default scope (currently-valid facts), supersession-chain behavior, and the return shape including the `truncated` sampling field with its { shown, of, rest } structure. It stops short of covering auth, rate limits, or ordering, but the sampling warning is a substantive behavioral disclosure few definitions include.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the default and the two temporal modes, then the sibling routing rule, then the return/truncation warning. Every sentence adds decision-relevant information with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a six-parameter read tool with no output schema and no annotations, the description supplies default scope, temporal semantics, return shape, and a false-negative guard ('narrow by subject or raise limit before concluding a fact does not exist'). Predicate semantics and result ordering are the only omissions.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With schema coverage at only 50%, the description compensates well: it explains asOf, asKnownAt (including the backdating exclusion rule), includeHistory, and the limit/subject narrowing. Only `predicate` goes unexplained, a minor gap against five well-contextualized parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Query the bitemporal fact store') and immediately scopes the default behavior, distinguishing it from lore_search by naming the exact domain it owns (factual slots). An agent can select this over siblings without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly names the alternative ('Prefer this over lore_search for factual slots') with the condition that selects it, and disambiguates the two temporal modes by their question form ('what was true on DATE' vs 'what did we know on DATE'), giving when-to-use guidance for each parameter mode.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
lore_read_noteRead a noteA
Read the raw markdown of a note by vault-relative path (as returned in search results). Never reads outside the vault — a path that escapes it, including through a symlink whose target lives elsewhere, is refused, and such a file is not indexed or searchable either. After reading a note that answered the question, call lore_mark_used to reinforce it.
| Name | Required | Description | Default |
|---|---|---|---|
| path | Yes | vault-relative path, e.g. projects/x.md |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral burden and does well: it discloses a security boundary (no reads outside the vault, symlink escapes refused) and a side effect that escaped files are not indexed or searchable. It omits error/return behavior details, so it falls short of full transparency.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, all front-loaded and information-dense with no filler. The path constraint and the follow-up action are distinct, earning their sentences, though the middle sentence is somewhat packed.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter, no-annotation, no-output-schema read tool, the description covers the addressing scheme, a security boundary, and the intended follow-up action. An agent has enough to call it correctly; only the failure/return shape is left implicit.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% with a single documented parameter, so baseline would be 3. The description adds value beyond the schema by specifying where the path comes from (search results) and constraining it to vault-relative, which is meaningful for correct invocation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Read) and resource (raw markdown of a note) plus the exact addressing scheme (vault-relative path). This clearly distinguishes it from lore_search or lore_context_pack, which locate rather than return raw content.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit context for the input ('as returned in search results') and a follow-up step ('after reading a note that answered the question, call lore_mark_used'). It does not, however, state when to prefer this over lore_search or lore_context_pack.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
lore_resumeResume a sessionA
What changed since this tool was last called: notes edited, facts asserted, and knowledge updates (slot: old → new). Call once at session start to continue where the previous session left off — the delta is computed from record time, so the same watermark always yields the same answer. Calling with no since consumes the delta (advances the watermark); pass an explicit since for a pure read that does not.
| Name | Required | Description | Default |
|---|---|---|---|
| since | No | explicit ISO watermark — pure read, does not advance the session boundary |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does a good job: it discloses that the call has a side effect (consuming the delta / advancing the watermark), that the result is deterministic ('the same watermark always yields the same answer'), and how to avoid the side effect. It does not describe error behavior, auth requirements, or the return envelope, so it is not exhaustive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three dense sentences, front-loaded with the return contents before the usage and parameter guidance. Every clause carries information; the determinism clause is the most expendable but still supports why the watermark contract matters.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With one optional parameter and no output schema, the description does the extra work of sketching the return contents ('notes edited, facts asserted, knowledge updates (slot: old → new)'). It omits error/failure behavior and whether the watermark is ever reset, which keeps it from being fully complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the schema already documents `since` as an ISO watermark that does not advance the boundary, so the baseline is 3. The description earns above baseline by explaining the consequence the schema omits: omitting `since` consumes the delta and advances the watermark, while passing it is a pure read.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb-plus-resource framing ('what changed since this tool was last called: notes edited, facts asserted, and knowledge updates') that tells an agent exactly what comes back. It implicitly separates itself from pure-read siblings ('pass an explicit `since` for a pure read'), but never names a sibling tool, so it falls short of full differentiation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives a concrete trigger ('call once at session start to continue where the previous session left off') and distinguishes the two modes of use, consuming vs pure read. No sibling alternatives are named, so it is a strong context statement rather than a full routing guide.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
lore_reviewWhat is fadingA
Important-but-fading knowledge: blocks whose retrievability has decayed below the threshold despite mattering, plus long-untouched open facts. This is the spaced-repetition loop made operable: review the list, then call lore_mark_used on anything still relevant — use is what reinforces stability. Deterministic, computed from the vault's own fitted forgetting curve.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | max items (default 20) | |
| threshold | No | retrievability below this counts as fading (default 0.3) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations, so the description carries the full burden. It usefully discloses that results are deterministic and computed from the vault's own fitted forgetting curve, and implies a read-only listing. It does not state pagination behavior, ordering, or the shape of returned items, leaving real gaps for a no-annotation tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, front-loaded with the definition of fading knowledge, then the loop mechanic, then the determinism note. No filler, though the spaced-repetition framing is slightly verbose for a list tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no annotations and no output schema, the description still conveys membership criteria, expected follow-up action, and the deterministic nature. Only the returned item shape and pagination remain unspecified, which is a modest gap for a filtered-list tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so both limit and threshold are already documented with defaults and bounds. The description restates the threshold concept (
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States precisely what the tool surfaces: blocks whose retrievability decayed below threshold plus long-untouched open facts. This distinguishes it from siblings like lore_search or lore_timeline by content type, though it reads more as a description of results than a clean verb+resource.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit workflow guidance: 'review the list, then call lore_mark_used on anything still relevant — use is what reinforces stability.' It names the follow-up tool and the condition for using it. No explicit when-not-to-use or alternative-tool exclusion, so it falls short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
lore_searchSearch the vaultA
Hybrid retrieval over the markdown vault: BM25 + knowledge-graph spreading activation (+ dense embeddings when configured). Returns one passage per note — the section that best covers your query — with the file it came from, how much of your query it matched, and any entity that linked it in. Use for any "what do my notes say about X" question, including multi-hop associations where the answer shares no words with the query. Pass verbose:true only if you need score internals.
| Name | Required | Description | Default |
|---|---|---|---|
| k | No | max results (default 8) | |
| tags | No | only notes carrying every listed tag; prefix a tag with "-" to exclude it | |
| query | Yes | natural-language query | |
| since | No | only content dated on/after this ISO date | |
| until | No | only content dated on/before this ISO date | |
| folder | No | only notes under this vault-relative folder, e.g. "projects/" | |
| verbose | No | include score breakdown and block anchors |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does well: it discloses the retrieval mechanism, that it returns one passage per note (the best-matching section), and what metadata accompanies it. It omits auth/permission needs and performance characteristics, keeping it out of 5 territory.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences with the mechanism front-loaded, then usage, then the param caveat. Dense but every clause carries information; minor cost from the parenthetical technical detail.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a search tool with no output schema and no annotations, the description tells the agent what it gets back (passage, source file, match proportion, linking entity) and when to escalate to verbose. Missing only operational details like auth or result limits.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the schema documents all seven parameters; baseline is 3. The only added meaning is the verbose hint ("only if you need score internals"), which is a mild addition over the schema's description of score breakdown.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource (hybrid retrieval over the markdown vault) and names the exact retrieval methods (BM25 + knowledge-graph spreading activation + dense embeddings). It also defines the return shape precisely. It does not explicitly name a sibling it differs from, so it falls short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit when-to-use trigger ("any 'what do my notes say about X' question") and even the harder case (multi-hop associations sharing no words with the query). It adds a conditional param rule for verbose, but names no alternative sibling or when-not-to-use condition.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
lore_timelineEntity timelineA
Chronological history of an entity in one call: every value change from the bitemporal fact store (with what each value replaced and when it stopped holding) merged with content-dated passages mentioning the entity. Use for "what happened to X", "what was X before it changed", "history of X" — instead of sampling repeated as-of fact queries and windowed searches.
| Name | Required | Description | Default |
|---|---|---|---|
| since | No | only entries on/after this ISO date | |
| until | No | only entries on/before this ISO date | |
| subject | Yes | entity name, e.g. "Project Atlas" |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations exist, so the description carries the full burden. It does disclose return composition well (each value change with what it replaced and when it stopped holding, merged with dated passages), which is genuinely useful. However, it says nothing about authorization, result size limits, pagination, or ordering guarantees for a tool that could return a large merged history.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core purpose, then examples, then the anti-pattern guidance — a logical order with little waste. The single long sentence with em-dash parentheticals is dense, slightly reducing readability.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description usefully explains what the return contains and how the two sources merge. Date scoping lives in the schema. Adequate for a 3-param read tool, though limits/pagination on the merged result remain unaddressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so since/until/subject are already documented in the schema. The description's mention of 'content-dated passages' hints at temporal semantics but adds no syntax or format detail beyond the schema's ISO-date notes. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Chronological history of an entity') and precisely defines the two merged data sources (bitemporal fact-store value changes plus content-dated passages). It is clearly distinguishable from siblings like lore_query_facts and lore_search.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly frames the trigger questions ('what happened to X', 'history of X') and names the naive alternatives to avoid — repeated as-of fact queries and windowed searches. This routes the agent away from lore_query_facts/lore_search with a stated condition.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
15 tool updates
v0.38.0- First observed
lore_aggregate_facts - First observed
lore_assert_fact - First observed
lore_capture - First observed
lore_context_pack - First observed
lore_dream_report - First observed
lore_index - First observed
lore_invalidate_fact - First observed
lore_mark_used - First observed
lore_propose_facts - First observed
lore_query_facts - First observed
lore_read_note - First observed
lore_resume - First observed
lore_review - First observed
lore_search - First observed
lore_timeline
TDQS
Scored across 15 tools
Each tool targets a distinct operation—orientation (context_pack/resume), retrieval (search/read), structured facts (query/aggregate/timeline/assert/invalidate), and maintenance (review/dream/mark_used). Descriptions explicitly guide when to prefer one over another, leaving no meaningful ambiguity.
All tools use the lore_ prefix and lower snake_case, which is highly consistent. A few names are noun phrases (context_pack, timeline, dream_report) rather than strict verb_noun, but no mixed conventions or camelCase appear.
15 tools sit at the top of the recommended range but each covers a distinct capability in a rich knowledge-management system. No tool appears redundant or tacked on; the set is well-scoped for its domain.
The surface covers ingestion, retrieval, fact lifecycle, review, and consolidation comprehensively. Direct note edit/delete tools are absent, but the design supports external vault edits via lore_index, so core workflows are not blocked.
Maintenance
Related MCP Connectors
One memory, every AI. A shared, user-owned markdown memory your AI clients read and write over MCP.
Markdown-based note-taking with a hosted MCP server. Your notes serve you and your AI.
Cloud-hosted MCP server for durable AI memory
- memnodeOAuthdev.memnode
Persistent, inspectable memory for AI agents with lineage, correction, and a hosted MCP endpoint.
Related MCP Servers
- AlicenseNot gradedqualityCmaintenanceA local-first MCP server that gives AI assistants long-term memory by storing, searching, and recalling notes as Markdown files on your machine.8 npmMIT
- AlicenseAqualityBmaintenanceA self-hosted MCP server that gives AI agents shared, long-term memory over a git-backed folder of markdown, enabling persistent knowledge search, read, and write without a database.1620 npm14MIT
- AlicenseNot gradedqualityBmaintenanceMCP server that turns a git-versioned Markdown vault into a queryable memory for AI agents, providing tools like brain_search, brain_read, and brain_neighbors.Apache 2.0
- AlicenseNot gradedqualityBmaintenanceMCP server providing persistent, local-first memory for AI agents via Markdown files in a git repo, with search, branching, and auditability.5 npm2MIT