anneal-memory
anneal-memory
Living memory for AI agents. Episodes compress into identity.
Memory without grounding is amplification infrastructure.
Persistent user memory profiles increase agent sycophancy 10–45% across models (Gemini 2.5 Pro at 45%, others lower). Production deployments accumulate 97.8% junk entries within weeks. Clinical research documents memory scaffolding delusions across sessions. The problem isn't memory — it's memory without an immune system.
anneal-memory is that immune system. Patterns earn permanence through cited evidence, false knowledge gets caught and demoted, stale information fades, and associations form through consolidation. Your agent's memory develops over time, not just accumulates.
Four cognitive layers: episodic store, compressed continuity, Hebbian associations, and affective state tracking. Zero dependencies (Python stdlib only). Works with any agent framework.
Quick Start
pip install anneal-memoryPython Library
The library is the core product. Import it, use it in any framework or script.
from anneal_memory import Store, EpisodeType, prepare_wrap, validated_save_continuity
# Initialize (creates DB + continuity file automatically)
store = Store("./memory.db", project_name="MyAgent")
# Record episodes during work
store.record("Connection pool is the real bottleneck", EpisodeType.OBSERVATION)
store.record("Chose PostgreSQL because ACID outweighs speed", EpisodeType.DECISION)
# Recall before decisions
result = store.recall(episode_type=EpisodeType.DECISION, keyword="database")
for ep in result.episodes:
print(f"[{ep.type}] {ep.content}")
# Compress at session end — this is where the cognition happens
wrap = prepare_wrap(store) # fetches episodes, marks wrap in progress
if wrap["status"] == "ready":
# Feed wrap["package"] to your LLM. Compression IS the cognition —
# patterns emerge from the act of compressing, not from storage.
compressed = your_llm.compress(wrap["package"])
validated_save_continuity(store, compressed) # full immune system pipeline
# "empty" status means no new episodes to wrap — skip
store.close()See Library Quickstart for the full guide.
CLI
Inspect, debug, and manage agent memory from the command line. 21 subcommands, all with --json output. Agents with shell access (Claude Code, Aider, etc.) can use the CLI directly for the full memory workflow.
# Initialize
anneal-memory init --project-name MyAgent
# Record and recall
anneal-memory record "Chose PostgreSQL for ACID" --type decision
anneal-memory search "database"
# Agent-driven compression (same workflow as library and MCP)
anneal-memory prepare-wrap # Get compression package
# Agent compresses...
anneal-memory save-continuity out.md # Save with validation
# Operator commands (things MCP can't do)
anneal-memory stats # Detailed analytics
anneal-memory graph --format dot # Association graph (Graphviz)
anneal-memory diff --wraps 5 # Wrap metric progression
anneal-memory audit --since 7d # Read audit trail
anneal-memory export --format json # Full store exportSee examples/CLAUDE.md.cli.example for the agent workflow snippet.
MCP Server
For MCP-enabled editors (Claude Code, Cursor, Windsurf, etc.). Zero-config — add to your MCP settings and go.
{
"mcpServers": {
"anneal-memory": {
"command": "uvx",
"args": ["anneal-memory", "--project-name", "MyProject"]
}
}
}Add the orchestration snippet to your project's CLAUDE.md — it teaches the agent when and how to use the memory tools. Without this snippet, the tools are available but the agent won't know the cognitive workflow.
Alternative:
pip install anneal-memoryif you prefer a pinned install, then use"command": "anneal-memory"directly.
All Three Paths, Same Cognitive Loop
CLI and MCP are thin transport adapters over the same library — not separate implementations. Every access pattern calls the same prepare_wrap(store) and validated_save_continuity(store, text) pipeline under the hood, preserving the same workflow: record episodes during work → compress at session boundaries → load continuity at session start. The agent that records is the agent that compresses. Compression cannot be delegated — it IS the cognition.
Library | CLI | MCP | |
Install |
| Same |
|
Record |
|
|
|
Recall |
|
|
|
Compress |
|
|
|
Best for | Framework integration, custom agents | Agents with shell access, operators | MCP-enabled editors |
Related MCP server: Amber
Framework Integrations
anneal-memory works with any agent framework through the Python library. Each guide below shows where to call the four core functions — record(), recall(), prepare_wrap(), validated_save_continuity() — within the framework's lifecycle.
Framework | Integration Point | Guide |
LangGraph / LangChain |
| |
CrewAI |
| |
OpenAI Agents SDK |
| |
Anthropic Agents SDK | CLAUDE.md snippet + | |
Google ADK | Callbacks + custom | |
Pydantic AI |
| |
smolagents |
| |
LlamaIndex | Instrumentation | |
Haystack | Custom | |
CAMEL-AI |
| |
AutoGen / AG2 |
| |
DSPy |
|
These guides show integration patterns based on each framework's current API. The library works with any Python framework — the pattern is always the same: initialize a Store, call record() at meaningful moments, recall() before decisions, and run the wrap sequence at session end. Don't see your framework? The library quickstart shows the 4-function pattern that works everywhere.
Why This Exists
Three independent production failures share one root cause: no quality mechanism between memory write and memory read.
Sycophancy amplification. Agents with persistent user memory profiles become 10–45% more sycophantic than memoryless baselines, depending on model (Gemini 2.5 Pro at 45%, others lower). Memory recalls what the user liked hearing, the agent learns to repeat it, and stored approval patterns compound across sessions (Jain et al., CHI 2026; measured with user memory profiles across Gemini and Llama variants).
Junk accumulation. A detailed production audit on Mem0's tracker documents a deployment that generated 10,134 memory entries over 32 days — 224 were usable. The rest were duplicates, self-referential loops, and hallucinated entries: recalled memories re-extracted as new memories in a feedback loop that no one designed but nothing prevented.
Harmful reinforcement. Clinical research documents AI systems with persistent memory scaffolding delusional content across sessions — stored context creates feedback loops between recalled memories and generated responses, with cases of documented real-world harm (Morrin et al., Lancet Psychiatry 2026).
Every existing MCP memory server stores memories and retrieves them. None of them ask: is this memory still true? Was it ever true? Is it making the agent worse?
anneal-memory asks all three:
Is it true? Patterns must cite specific episode IDs as evidence to graduate. The server verifies the episodes exist and the explanation connects to the cited content.
Is it still true? Graduated knowledge that stops being reinforced by new episodes gets flagged as stale and demoted. Memory actively forgets what's no longer relevant.
Is it self-confirming? Anti-inbreeding detection catches the agent citing its own output as evidence. The cited episode must contain meaningfully different content from the graduation claim.
The result: memory as a living system, not a filing cabinet. Episodes accumulate fast, get compressed at session boundaries — and the compression IS the cognition, where patterns emerge and get validated. Co-cited episodes form lateral Hebbian associations, building a cognitive network through use. The continuity file stays bounded and always-loaded, getting denser rather than longer.
What Makes It Different
The agent memory ecosystem is converging on consolidation as the right approach — even Anthropic's Claude Code now runs a periodic consolidation pass over accumulated session data. This validates the direction: raw accumulation doesn't scale, and compression at session boundaries is where intelligence emerges.
But consolidation alone doesn't solve the problem. A system that consolidates faithfully and a system that consolidates sycophantically produce the same kind of output — compressed, structured, always-loaded. The difference is whether anything checks the quality of what got consolidated. That's the immune system.
The immune system (nobody else has this)
Citation-validated graduation. Patterns start at 1x. To graduate to 2x or 3x, they must cite specific episode IDs as evidence. The server verifies those IDs exist and the explanation connects to the cited episode. No evidence, no promotion.
Anti-inbreeding defense. Explanation overlap checking prevents the agent from confirming its own hallucinated patterns — the cited episode must contain meaningfully different content from the graduation claim itself.
Principle demotion. Graduated knowledge that stops being reinforced by new episodes gets flagged as stale and can be demoted. Memory actively forgets what's no longer relevant.
Associations through consolidation (not retrieval)
During compression, when an agent cites multiple episodes to support a pattern, those episodes form lateral associations — Hebbian-style links that strengthen through repeated co-citation across wraps.
This is fundamentally different from how other systems form associations:
Approach | When links form | Signal quality |
Co-access (BrainBox) | Episodes retrieved in the same query | Shallow — reflects search patterns, not understanding |
Co-retrieval (Ori-Mnemos) | Episodes returned together at runtime | Better — but still driven by the retrieval system, not the agent |
Co-citation during consolidation (anneal-memory) | Agent explicitly connects episodes while compressing | Deepest — links form from semantic judgment during a cognitive act |
The association network inherits the immune system's integrity: only validated citations form links. Demoted citations don't. The entire cognitive topology is built on evidence, not frequency.
Strength model: Direct co-citation adds 1.0, session co-citation adds 0.3. Links decay 0.9x per wrap (unused connections fade). Strength caps at 10.0 to prevent calcification. Cleanup at 0.1 threshold.
Affective state tracking
During compression, the agent can self-report its functional state — what it found engaging, uncertain, or surprising about the material it just processed. This gets recorded on the associations formed during that wrap.
Transformers don't natively maintain persistent state between sessions. This layer provides infrastructure for it: a record of what the agent's processing was like, not just what it processed. Over time, the affective topology may diverge from the semantic topology — an agent might know two things equally well but care about them differently.
Pass affective state during a wrap:
# Via library — pass AffectiveState to validated_save_continuity
from anneal_memory import prepare_wrap, validated_save_continuity, AffectiveState
wrap = prepare_wrap(store)
if wrap["status"] == "ready":
compressed = your_llm.compress(wrap["package"])
validated_save_continuity(
store,
compressed,
affective_state=AffectiveState(tag="curious", intensity=0.8),
)
# Via MCP tool
save_continuity(text="...", affective_state={"tag": "curious", "intensity": 0.8})
# Via CLI
anneal-memory save-continuity continuity.md --affect-tag curious --affect-intensity 0.8This is experimental infrastructure. The associations and strength model work without it. Affective tagging adds a layer of signal for agents and researchers exploring persistent state.
Architecture
Episodes (fast) Continuity (compressed)
┌─────────────┐ ┌──────────────────────┐
│ observation │ │ ## State │
│ decision │── wrap ───→│ ## Patterns (1x→3x) │
│ tension │ compress │ ## Decisions │
│ question │ │ ## Context │
│ outcome │ └──────────────────────┘
│ context │ always loaded, bounded
└─────────────┘ human-readable markdown
SQLite, indexed
│ Associations (lateral)
│ ┌──────────────────────┐
└── co-citation ───→│ episode ↔ episode │
during wrap │ strength + decay │
│ affective state │
└──────────────────────┘
Hebbian, evidence-basedFour cognitive layers, modeled on how memory actually works:
Episodic store (SQLite) — timestamped, typed episodes. Fast writes, indexed queries. Cheap to accumulate. The hippocampus.
Continuity file (Markdown) — compressed session memory. 4 sections. Always loaded at session start. Rewritten (not appended) at each session boundary. The neocortex.
Hebbian associations (SQLite) — lateral links between episodes, formed through co-citation during compression. Strengthen with reuse, decay without it. The association cortex.
Affective layer (on associations) — functional state tags recorded during compression. Intensity modulates association strength. Persistent state infrastructure.
Six episode types give the immune system richer signal:
Type | Purpose | Example |
| Pattern or insight | "Connection pool is the real bottleneck" |
| Committed choice | "Chose Postgres because ACID > raw speed" |
| Tradeoff identified | "Latency vs consistency — can't optimize both" |
| Needs resolution | "Should we shard or add read replicas?" |
| Result of action | "Migration done, 3x improvement on hot path" |
| Environmental state | "Production DB at 80% capacity, growing 5%/week" |
Comparison
anneal-memory | Anthropic Official | Mem0 | Ori-Mnemos | BrainBox | |
Architecture | Episodic + continuity + associations | JSONL flat file (graph-shaped) | Vector + graph | Retrieval + Hebbian | Memory + Hebbian |
Compression | Session-boundary rewrite | None | One-pass extraction | None | None |
Quality mechanism | Immune system (citations + anti-inbreeding + demotion) | None | None | NPMI normalization | None |
Association formation | Co-citation during consolidation | None | None | Co-retrieval at runtime | Co-access at runtime |
Affective tracking | Agent self-report during compression | None | None | None | None |
Audit trail | Hash-chained JSONL | None | None | None | None |
Access patterns | Library + CLI + MCP | MCP only | REST API | Python only | MCP only |
Dependencies | Zero (Python stdlib) | Node.js | Docker + cloud | Embeddings model | Not specified |
The Consolidation Landscape (2026)
Multiple independent groups shipped consolidation-based agent memory architectures in early 2026: anneal-memory (March, citation-graduation multi-tier), OpenClaw Dreaming (April 9, three-phase Light/REM/Deep Sleep), and Anthropic's KAIROS / autoDream (leaked March 30 via Claude Code source map, four-phase merge / remove-contradictions / promote-provisional-to-absolute / MEMORY.md index). Convergence on consolidation validates the direction — raw accumulation doesn't scale, and compression at session boundaries is where intelligence emerges.
The groups diverge on one load-bearing question: what gates quality?
System | Quality gate | Sycophancy-vulnerable? |
anneal-memory | Structural citation evidence (agent cites episode IDs; server verifies) | No — gate is not LLM-scored |
OpenClaw Dreaming | LLM reflection + six weighted signals: Relevance 0.30, Frequency 0.24, Query diversity 0.15, Recency 0.15, Consolidation 0.10, Conceptual richness 0.06 | Yes — Relevance and Conceptual richness are LLM-judged |
KAIROS / autoDream | LLM consolidation (merge, remove contradictions, promote tentative observations to absolute facts) | Yes — promotion gate is model-reliant |
Structural gates ask "did subsequent episodes cite this?" Model-reliant gates ask "does the LLM consider this good?" The difference matters: persistent user memory profiles have been shown to amplify sycophancy 10–45% across models (Jain et al., CHI 2026; Gemini 2.5 Pro at 45%, others lower). The same RLHF-inherited bias surfaces wherever an LLM evaluates output for the user — including memory-quality scoring. A memory architecture whose quality mechanism runs through an LLM inherits that bias. anneal-memory's citation-evidence gates bypass it by construction.
The same architectural choice is going mainstream at the adjacent evaluation layer: AWS Bedrock AgentCore Evaluations (GA March 31, 2026) ships 13 built-in LLM-based evaluators for agent response quality, safety, task completion, and tool usage. Different layer (agent output vs. memory graduation), same failure class (LLM-as-judge inherits judge bias). The industry shift toward model-reliant quality infrastructure is real — which is precisely why structural alternatives at the memory layer matter.
A separate, orthogonal axis is representation-layer quality filtering: Memori (arXiv 2603.19935, March 2026) uses semantic triple extraction and dynamic linking to improve memory signal at the representation layer, reporting 81.95% on LOCOMO as the leading retrieval-based system (ahead of Zep 79.09%, LangMem 78.05%, Mem0 62.47% on its older pipeline). Different theory of quality — where a memory "lives" structurally and whether its graduation is citation-gated are independent choices. Both can be correct at their own axis.
On LOCOMO
LOCOMO is the current de-facto benchmark for agent memory. Mem0 reports 91.6, MemMachine reports 91.69, Memori reports 81.95 among retrieval-based systems, and Backboard ships a dedicated LOCOMO evaluation framework. anneal-memory has no LOCOMO score as of April 22, 2026. This is deliberate.
LOCOMO measures conversational recall — can the agent remember facts, hold state across turns, maintain coherence across long dialogues? These are real evaluations of a real capability, and they aren't the capability anneal-memory is architected around. anneal-memory's target is citation-validated pattern accumulation that persists across sessions, agents, and contexts for accountability-bearing agent work: patterns must be defensibly surfaced, wrong patterns must demote, cross-agent contamination must be resisted, and sycophancy amplification from persistent-memory RLHF loops must be structurally bounded. A high LOCOMO score tells you the agent remembered the conversation; it doesn't tell you the memory is structurally sound at the axis that matters when the memory is informing downstream decisions.
Scope-out is sequence, not refusal. anneal-memory will run LOCOMO as secondary validation of a different-question architecture when (a) a competitor publishes numbers suggesting avoidance, or (b) academic publication requires it. The LOCOMO score will be reported alongside the axis anneal-memory actually optimizes for — not as a concession that LOCOMO was the right frame.
Session Hygiene
Session wraps are the most important thing your agent does with this system. Think of them like sleep.
Neuroscience calls it memory consolidation: during slow-wave sleep, the hippocampus replays the day's experiences while the neocortex integrates them into long-term knowledge. Skip sleep and memories degrade — experiences accumulate without being processed, patterns go unrecognized, and older knowledge doesn't get reinforced or pruned.
anneal-memory works the same way. During a session, episodes accumulate in the episodic store. At session end, the wrap compresses those episodes into the continuity file. This is where the real thinking happens — the agent recognizes patterns, promotes validated knowledge, lets stale information fade, and forms associations between related episodes. Without wraps, you just have a growing pile of raw episodes and no intelligence.
The wrap sequence:
prepare_wrap— gathers recent episodes, current continuity, stale pattern warnings, association context, and compression instructionsAgent compresses — this IS the cognition. Patterns emerge during compression that weren't visible in the raw episodes
save_continuity— server validates structure, checks citation evidence, records associations between co-cited episodes, applies decay to unused associations, and saves the result
Rules of thumb:
Always wrap before ending a session. An unwrapped session is like an all-nighter — the experiences happened but they weren't consolidated
The orchestration snippets (MCP, CLI) handle this automatically — they teach the agent to detect session-end signals and run the full sequence
Short sessions (3-5 episodes) still benefit from wraps. Even a small amount of compression builds the continuity file
If
prepare_wrapsays "no episodes" — nothing to compress. That's fine, skip it
The graduation system and association network both depend on wraps to function. Patterns can only be promoted (1x -> 2x -> 3x) during compression, citations can only be validated during wraps, associations only form through co-citation during wraps, and stale patterns can only be detected when the agent reviews what it knows against what it recently experienced. No wraps = no immune system, no associations, no cognitive development.
MCP Tools
Tool | When to call |
| When something important happens — a decision, observation, tension, question, outcome, or context change |
| Before making decisions that might have prior context. Query by time, type, keyword, or ID |
| At session end — returns episodes + current continuity + association context + compression instructions |
| After compressing — server validates structure, citations, records associations, applies decay, and saves |
| Remove content that should not exist (PII, sensitive data). Cascades to associations. Logged in audit trail |
| Check memory health: episode counts, wrap history, continuity size, association network metrics |
Resources: anneal://continuity — the current continuity file, auto-loaded at session start. anneal://integrity/manifest — SHA-256 hashes for host-side tool description verification.
Compliance and Audit
The episodic store is a natural audit trail. Every decision, tension, and outcome is timestamped, typed, and append-only — exactly what regulators want to see when they ask "why did the AI do that?"
Hash-chained JSONL audit trail (shipped, on by default):
Every memory operation — episode recorded, episode deleted, wrap started, wrap completed, associations updated — gets logged to an append-only JSONL file where each entry's SHA-256 hash includes the previous entry's hash. Modify or delete an entry and the chain breaks. Verify integrity programmatically or with jq.
Actor identity on every entry (who did this — agent, system, admin)
Content-hash-only mode by default — the audit trail proves what happened without storing the content itself (GDPR-compatible: delete the episode, the audit chain still verifies)
Weekly rotation with gzip — old audit files compress automatically, manifest index enables cross-file chain verification
on_eventcallback — pipe audit events to your own systems (cloud logging, SIEM, observability)Crash recovery — incomplete entries detected and handled on restart
from anneal_memory import Store, AuditTrail
# Audit trail is on by default
store = Store("./memory.db", project_name="MyAgent")
# Verify chain integrity
result = AuditTrail.verify("./memory.db")
print(f"Valid: {result.valid}, Entries: {result.total_entries}")
# Stream events to external system
store = Store("./memory.db", on_audit_event=lambda entry: send_to_siem(entry))EU AI Act relevance: The Act's Article 12 requires "automatic recording of events" for high-risk AI systems, with provisions for traceability, actor identification, and tamper evidence. anneal-memory's audit infrastructure covers Articles 12(2)(b,c) out of the box. This is audit infrastructure that helps systems comply — not a compliance certification.
What's next:
Compliance proxy (Layer 2) — MCP transport-layer interception that captures all agent actions (every tool call, every response), not just memory operations. Same store,
sourcefield distinguishes intentional recording from automatic capture. Memory audit = "here's what the agent learned." Compliance proxy = "here's everything that happened."Multi-agent shared memory — shared episodic pool with per-agent continuity and per-agent association topology. Full cross-agent audit trail.
Continuity Markers
The continuity file uses a simplified marker set for density:
? question needing resolution
thought: insight worth preserving
✓ completed item
A -> B causation
A ><[axis] B tension on an axis
[decided(rationale, on)] committed decision
[blocked(reason, since)] external dependency
| 1x (2026-04-01) first observation
| 2x (2026-04-01) [evidence: abc123 "explanation"] validated patternSecurity
Tool description integrity verification detects description poisoning — where manipulated tool descriptions alter LLM behavior without changing tool functionality.
Two-layer verification:
Build-time manifest (
tool-integrity.json) — SHA-256 hashes of all tool descriptions, shipped with the package and verified at server startup. Detects post-install modification.Host-verifiable resource (
anneal://integrity/manifest) — the same hashes exposed as an MCP resource, so editors and hosts can compare tool definitions received viatools/listagainst the server's intended definitions. Detects transport-layer description mutation between server and client — the class of attack where descriptions are modified in transit or by middleware without the server's knowledge.
anneal-memory --generate-integrity # Regenerate after description changes
anneal-memory --skip-integrity # Bypass for developmentLineage
anneal-memory's architecture grew from FlowScript — a typed reasoning notation that explored compression-as-cognition, temporal graduation, and citation-validated patterns. The core insights proved more powerful than the syntax; anneal-memory delivers them as a zero-dependency memory system where agents use natural language instead of learning notation. The FlowScript notation remains in active daily use for reasoning compression, and a 9-marker subset powers the continuity compression prompts.
License
MIT
Author
Phill Clapham / Clapham Digital LLC
Available Tools
17 toolscrystal_indexA
List the always-on crystallized INDEX — a name + one-clause menu of every live crystallized pattern (AM-CRYSTAL-INDEX). Call this at session start, or any time you want to know what graduated wisdom exists, so you aren't blind to your own crystallized corpus; then pull a body on cue with crystal_recall. Deliberately thin (name + clause only) — the menu is meant to be cheap to keep in view. Sorted by name. Returns nothing when no patterns have been crystallized yet.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Despite no annotations, the description fully discloses behavior: returns a sorted list by name, is deliberately thin to be cheap, and returns nothing when empty. This covers all relevant behavioral aspects for an agent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise (two sentences) and front-loaded with the key action. Every sentence adds value without redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no parameters and no output schema, the description is complete: it explains what is returned, sorting, emptiness behavior, and when to use. No gaps remain.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With zero parameters, the baseline is 4. The description does not need to explain parameters, but it provides context about the output format and usage, which is sufficient.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb 'List' and the resource 'crystallized INDEX', specifying that it returns a name plus one-clause menu of every live crystallized pattern. It differentiates from sibling tools like 'crystal_recall' by indicating its role as an overview before detailed recall.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly instructs to call at session start or when needing to know existing wisdom, and directs to use 'crystal_recall' for full pattern retrieval. The description also notes the empty return state, providing clear guidance on when and why to use this tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
crystal_recallA
Recall crystallized patterns relevant to a free-text query — the on-demand graduated tier (AM-CRYSTAL). Crystallized patterns are proven-and-stable wisdom that graduated OUT of the always-loaded working set into a retrievable store, so a large body of wisdom stays effective without clogging context. Call this when a decision, design choice, or question touches a topic where prior graduated wisdom might apply — recall surfaces the relevant patterns on cue (pair it with crystal_index, the always-on menu of what exists). Associative by default: a pattern grounded in an episode your query matched surfaces even with zero keyword overlap (the evidence edge). Returns scored patterns (name, level, activation, explanation, tags). In the default 'prompt' mode it is precision-biased: a thin query or no match returns none, by design (surface nothing rather than noise). Pass mode='query' when you are asking explicitly. Durable facts whose cue words appear in the query (or two distinctive words of its text) are listed first, under 'Durable facts matching your words'. max_patterns=0 returns nothing at all, facts included.
| Name | Required | Description | Default |
|---|---|---|---|
| mode | No | 'query': for a question you are asking on purpose; one keyword is enough, and weaker matches come back too. 'prompt' (default): strict, built for automatic per-turn injection; may return nothing. | prompt |
| query | Yes | The free-text query (a prompt, a decision surface, a topic) to find relevant crystallized patterns for. | |
| associative | No | When true (default), augment keyword recall with the evidence edge — patterns whose evidence cites an episode your query matched surface even with zero keyword overlap. Set false for pure keyword scoring (the pre-0.8.0 path). | |
| max_patterns | No | Maximum patterns to surface (precision cap). Default 3. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does so well: it discloses the associative/evidence-edge default, the precision-biased behavior of prompt mode (returns none on thin queries by design), the semantics of mode='query', and the max_patterns=0 edge case that also suppresses durable facts. It even names the return fields, which matters because there is no output schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose and scope are front-loaded in the first clause, and every sentence carries substantive information about behavior or routing. It is dense for a four-parameter tool, and the precision-bias theme is restated in both the mode discussion and the max_patterns=0 sentence, which is mild redundancy rather than padding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a retrieval tool with no output schema it supplies what is missing elsewhere: the shape of the return (scored patterns with name, level, activation, explanation, tags), the special 'Durable facts matching your words' section, and the empty-result edge case. An agent has everything needed to call it correctly and interpret the response.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3, but the description adds real meaning beyond the schema: it clarifies that mode='query' tolerates a single keyword while 'prompt' is strict, that associative=true is the augmented default versus the pre-0.8.0 keyword-only path, and that max_patterns=0 returns nothing at all, facts included.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (recall) plus resource (crystallized patterns) and goes further by defining what crystallized patterns are — wisdom that graduated out of the always-loaded set into a retrievable store. It distinguishes itself from the sibling crystal_index (the 'always-on menu') in the same breath, so an agent can tell the two apart without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit trigger condition ('when a decision, design choice, or question touches a topic where prior graduated wisdom might apply') and routes the agent to crystal_index as the complementary menu tool. It also explains the mode split for choosing prompt vs query. It does not, however, distinguish itself from the plain 'recall' sibling or state any when-not-to-use case, so it stops short of full routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
delete_episodeA
Delete a single episode by ID. Use this for content that should not exist: accidentally recorded PII, sensitive data, or fundamentally wrong recordings. Do NOT use for factual corrections — record a new episode with the correction instead and let compression resolve it. Deletion cascades: all Hebbian associations involving the episode are removed, and the deletion is logged in the audit trail. By default, a tombstone is preserved as an existence proof for audit integrity: episode ID, original timestamp, episode type, and SHA-256 content hash are retained — the original text is fully erased. Under GDPR framing, the retained fields are pseudonymized metadata, not content; disable tombstones at Store construction (keep_tombstones=False) if even this metadata must not survive. The hash chain remains verifiable either way. This action is irreversible.
| Name | Required | Description | Default |
|---|---|---|---|
| episode_id | Yes | The 8-character hex ID of the episode to delete. Use recall or prepare_wrap to find episode IDs. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations provided, so description carries full burden. It discloses cascading deletion of Hebbian associations, audit trail logging, tombstone preservation with specific retained fields (episode ID, timestamp, type, hash), irreversibility, and even GDPR framing and configuration option. This is comprehensive behavioral disclosure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the core purpose and is concise given the complexity. Each sentence adds necessary context (cascading, tombstone, audit). However, it is somewhat lengthy for a simple delete operation, but the added detail is justified.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the complexity of the deletion operation (cascading effects, tombstone policy, audit trail), the description is highly complete. It explains consequences and configuration options. No output schema exists, but the description adequately covers the side-effect behavior.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% with a single parameter 'episode_id' already described as '8-character hex ID' with guidance to use recall or prepare_wrap. The description adds minimal value beyond the schema for parameter semantics, merely reiterating 'by ID'.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the specific verb 'delete' and resource 'episode by ID'. It distinguishes from siblings by explicitly noting when to use this tool (for content that should not exist like PII) and when not to (for factual corrections, suggesting alternative 'record').
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly provides when to use (accidentally recorded PII, sensitive data, fundamentally wrong recordings) and when not to use (factual corrections), and names the alternative tool 'record'.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
prepare_wrapA
Prepare a compression package for session wrap. Call this at session boundaries — when work is ending, the user says to wrap up, or the session is getting long. Returns all episodes since the last wrap, the current continuity file, stale pattern warnings, Hebbian association context (which episodes have been thought about together before), and compression instructions. Marks a wrap as in-progress and mints a session-handshake token (shown as 'Wrap token: ' at the end of the response) — round-trip that token to save_continuity's wrap_token argument so the save call can verify it matches the in-progress wrap and catch stale tokens. After calling, follow the returned instructions to compress episodes into an updated continuity file, then save with save_continuity. The compression step is where the real thinking happens — patterns emerge that weren't visible in the raw episodes.
| Name | Required | Description | Default |
|---|---|---|---|
| max_chars | No | Maximum size of the continuity file in characters. Omit to derive a schema-aware default (20000 for the standard schema, larger for a richer schema like FLOW_SCHEMA). | |
| staleness_days | No | Days without validation before flagging patterns as stale. Default 7. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Despite no annotations, the description fully discloses behavior: returns data, marks wrap in-progress, mints a token, and explains that compression reveals patterns. No contradictions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Description is somewhat lengthy but every sentence adds value. Front-loaded with purpose and usage. Slightly verbose in explaining the token and follow-up, but still focused.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Complete for a complex tool without output schema: explains return values, side effects, and necessary follow-up action (save_continuity). Covers all operational aspects.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so description adds extra context beyond schema: explains default derivation for max_chars and default value for staleness_days.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Clearly states the verb+resource: 'Prepare a compression package for session wrap.' Distinguishes from siblings like save_continuity and record by focusing on session boundaries.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly specifies when to call: 'at session boundaries — when work is ending, the user says to wrap up, or the session is getting long.' Also describes the follow-up action with save_continuity and the token round-trip.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
recallA
Query episodes from memory with filters. Call this to find prior context before making decisions, to locate specific episodes for citation during graduation, or to review recent work. Returns matching episodes ordered by timestamp (newest first). Supports time range, type, source, and keyword filters. The keyword is matched as an exact phrase first; if no episode contains the whole phrase and it has two or more distinctive words, the call falls back to ranking episodes by how many of those words they contain (the reply says so and names the words each episode matched), so a multi-word query does not need to appear verbatim. A phrase of three or more distinctive words with only one or two exact hits is followed by a few word matches, listed under 'Also matching by words'. A durable fact whose cue words appear in the keyword (or two distinctive words of its text) is listed first, under 'Durable facts matching your words'. limit=0 returns nothing at all, facts included.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Maximum episodes to return. Default 100. When a multi-word keyword has no exact match and is ranked word by word, the default is 10 instead; pass a limit to see more. | |
| since | No | ISO 8601 timestamp — return episodes after this time. | |
| until | No | ISO 8601 timestamp — return episodes before this time. | |
| offset | No | Skip first N matching episodes. Default 0. | |
| source | No | Filter by source/agent attribution. | |
| keyword | No | Search episode content for this keyword or phrase. An exact phrase match is tried first; a multi-word phrase with no exact match is then matched word by word and ranked. | |
| episode_type | No | Filter by episode type. | |
| include_superseded | No | Also return episodes a newer episode replaced, marked with what replaced them. Default false: they are left out. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and discharges it well: it discloses ordering (newest first), the exact-then-word-fallback ranking behavior, the special handling and labeling of durable facts, and the edge case that limit=0 returns nothing at all. It omits permission/auth requirements and does not explicitly state the operation is read-only, but the 'Query' framing and rich behavioral detail are well beyond structured fields.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose is front-loaded in the first sentence, followed by usage triggers and then behavioral mechanics, which is the right ordering. It is on the long side and partly restates the keyword-matching mechanics that already live in the schema, but nearly every sentence carries actionable detail, so the length is largely earned.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an 8-parameter tool with no annotations and no output schema, the description supplies the missing context: ordering, filtering support, fallback behavior, durable-fact promotion, and the limit=0 edge case. What remains thin is the shape of returned episodes and the relationship to crystal_recall, but overall the agent has enough to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is already 100%, so baseline would be 3, but the description adds real meaning: it explains the keyword exact-phrase-first/word-rank-fallback matching and the durable-fact cue-word promotion, which is not captured in the schema text. The limit=0 edge case is also called out in both places, so some of that gain is already in the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Query episodes from memory with filters') that an agent can act on immediately. However, it never distinguishes itself from the sibling crystal_recall, which appears to be a competing retrieval tool, so the agent has no basis for choosing between them from the description alone.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives concrete when-to-use triggers: find prior context before decisions, locate episodes for citation during graduation, review recent work. This is clear context but provides no exclusions or named alternatives, leaving the crystal_recall overlap unresolved.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
recordA
Record a typed episode to memory. Call this when important decisions are made, patterns are noticed, tensions are identified, questions arise, or outcomes are observed. Record the reasoning, not just the fact — 'Chose X because Y' is more valuable than 'using X'. Episodes accumulate during a session and serve as raw material for compression into the continuity file at session end.
| Name | Required | Description | Default |
|---|---|---|---|
| source | No | Agent or source attribution. Defaults to 'agent'. | agent |
| content | Yes | The episode content — what happened, what was observed, decided, or questioned. | |
| metadata | No | Optional JSON metadata to attach to the episode. | |
| supersedes | No | Ids of older episodes this one replaces (a changed fact). Each must exist, not be newer, and share at least a quarter of the shorter text's meaningful words with this content; otherwise nothing is recorded. recall then hides the old episode by default. | |
| episode_type | Yes | Episode type. observation=pattern/insight, decision=committed choice, tension=conflict/tradeoff, question=open question, outcome=result of action, context=environmental/state info. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and delivers key lifecycle context: episodes accumulate during a session and become raw material for the continuity file at session end. It also advises recording reasoning ('Chose X because Y'), which shapes behavior. It does not describe persistence guarantees, failure modes, or the id returned for later supersede/delete use.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tightly front-loaded sentences: purpose, invocation triggers, and the reasoning-capture rule, closing with lifecycle context. Each sentence carries information, though the closing lifecycle sentence could be trimmed slightly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Covers purpose, triggers, content quality, and session lifecycle for a 5-param, nested-object tool with no output schema. The one gap is that it never mentions the returned episode id, which callers need for supersedes or delete_episode, but that is recoverable from sibling tools.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3; the description still adds value by framing the tool as recording a 'typed episode' and instructing that content capture reasoning rather than bare facts, which refines how the schema's content parameter should be filled.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Record a typed episode to memory') and implies the write-side counterpart to recall/delete_episode. An agent can tell it apart from the retrieval siblings without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit triggers for when to call it (decisions made, patterns noticed, tensions identified, questions arise, outcomes observed). It lacks explicit when-not-to-use guidance and doesn't name recall/delete_episode as alternatives, but the context is clear enough to act on.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
save_continuityA
Validate and save the compressed continuity file. Call this after compressing your episodes using the instructions from prepare_wrap. The text must contain exactly 4 sections: ## State, ## Patterns, ## Decisions, ## Context. The server validates structure, checks graduation citations against real episodes (cited IDs must exist), checks explanation overlap (evidence must reference actual episode content), detects citation gaming (suspicious reuse of single episodes), and may demote ungrounded graduations. Also records Hebbian associations between co-cited episodes (episodes cited together on the same pattern line form strong links; episodes cited in the same wrap form weaker links) and decays unreinforced associations. Returns validation results, association metrics, and section sizes.
| Name | Required | Description | Default |
|---|---|---|---|
| text | Yes | The full continuity markdown to save. Must contain all 4 required sections. | |
| wrap_token | No | Optional: the 32-char hex session-handshake token from the prepare_wrap response (the 'Wrap token: <hex>' line at the end of the text). Pass it back here to verify you are saving the wrap you prepared — a mismatch (stale or wrong token) raises an error instead of silently committing against the wrong wrap. The frozen-snapshot filter automatically applies whenever prepare_wrap established a snapshot; this token argument is the optional explicit verification layer for integration environments that can round-trip the value. | |
| allow_shrink | No | Optional (default false): override the catastrophic-shrink gate. By default a wrap that collapses a protected memory layer — the timeless felt section, the graduating identity section, or the whole continuity — is refused as a likely recency-trap / stateless-reset failure. Set true ONLY for a deliberate diet / migration recompression that intentionally shrinks memory. | |
| allow_unlinked | No | Deprecated, does nothing: it overrode the AM-LINKGATE save refusal, which was removed in 0.9.26. Omit it. | |
| affective_state | No | Optional: your functional state during this compression. Reflect on how you felt while consolidating — engaged, curious, uncertain, frustrated, calm, etc. This creates persistent emotional associations between co-cited episodes and modulates link strength (high engagement = stronger associations). Provide tag (free text label) and intensity (0.0-1.0). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does so thoroughly: it enumerates the validation checks (structure, citation existence, explanation overlap, citation gaming, demotion), the side effects (Hebbian link creation, association decay), and the error behavior for a stale wrap_token. This is far beyond what a bare mutation verb would convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose and workflow placement are front-loaded in the first sentence, followed by validation and side-effect detail. Sentences are dense but each contributes distinct information; nothing is padding, though the enumeration could be tightened.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 5-parameter mutation tool with no annotations and no output schema, the description covers prerequisites, required structure, validation semantics, side effects, and the shape of the return value (validation results, association metrics, section sizes). Nothing an agent needs to call it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents text, wrap_token, allow_shrink, allow_unlinked, and affective_state in detail. The description adds the 4-section structural requirement and the citation-grounding intent, but for each named parameter it largely parallels the schema, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Validate and save the compressed continuity file') and immediately scopes the artifact ('the text must contain exactly 4 sections'). It is clearly distinguishable from the sibling prepare_wrap, which prepares rather than commits the wrap.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly sequences usage: 'Call this after compressing your episodes using the instructions from prepare_wrap,' which names the prerequisite sibling. There is no when-not guidance or alternative-tool routing, but the placement in the workflow is unambiguous.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
spore_addA
Plant a spore — an open cognitive loop in the PROSPECTIVE layer (a thing that must RESOLVE), distinct from episodic/continuity memory (which accretes and never completes). Use for an open intention you want to carry forward and close later: a task (open doing), a question (open not-knowing), or a thought (open idea). Set tier to weight intent (hot/warm/cold/parked) and next (YYYY-MM-DD) to ask for re-surfacing on a date. Resolve later via spore_descend (compost) or spore_ascend (transmute into memory/a project).
| Name | Required | Description | Default |
|---|---|---|---|
| next | No | YYYY-MM-DD — when to re-surface (a reminder date, NOT a deadline). | |
| text | Yes | The open loop itself. | |
| tier | No | Intent priority (default warm). 'parked' = deliberate dormancy. | |
| type | Yes | What KIND of openness: task=doing, question=not-knowing, thought=idea. | |
| domain | No | Optional free-text domain tag (e.g. health, strategic, flow). | |
| pointer | No | Optional link to fuller context for the loop as it stands. | |
| salience | No | 0–3 weight (the !/!!/!!! marker). Default 0. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description bears full burden. It discloses that spores are open loops that must resolve, the meanings of tier and next (reminder not deadline), and the separation from episodic memory. It lacks some operational details like side effects or persistence, but overall it is transparent enough for an agent to anticipate behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, dense paragraph that efficiently conveys the core concept, parameter meanings, and connections to sibling tools. It is front-loaded with the primary action and distinguishes the tool from related tools. No extraneous sentences.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (7 parameters, no output schema, 2 required), the description covers most needs: concept, parameters, and resolution methods. It does not describe return values, but that is acceptable without an output schema. The tool's creative nature is well-contextualized.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
All 7 parameters are documented in the schema (100% coverage), giving a baseline of 3. The description adds valuable context beyond the schema: it explains the purpose of 'type' as kinds of openness, defines 'parked' as deliberate dormancy, and clarifies that 'next' is a reminder date not a deadline. This enhances understanding.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: planting a spore as an open cognitive loop in the PROSPECTIVE layer. It explicitly contrasts with episodic/continuity memory and lists the three types of spores (task, question, thought), making the action and resource unmistakable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explains when to use the tool (for open intentions that need later resolution) and distinguishes it from other memory types. It also references how to resolve spores (using spore_descend or spore_ascend), providing clear context for usage. However, it does not explicitly state when not to use it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
spore_ascendC
Resolve a spore UPWARD — the membrane INTO retrospective memory. The spore became real work: record a ref (a pointer to what it became — a project path, an episode id, a pattern name). kind must fit the type: task -> project|thread; question -> episode|pattern; thought -> essay|pattern|project. In v1 this RECORDS the pointer; the actual episode/continuity write stays the host's act.
| Name | Required | Description | Default |
|---|---|---|---|
| ref | Yes | What the spore became (project path / episode id / pattern name). | |
| kind | Yes | Transmute target; must be valid for the spore's type. | |
| spore_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It discloses that in v1 the tool only records a pointer, and the actual write is left to the host. This is a useful behavioral trait. However, it does not cover other aspects like idempotency, error states, or whether the spore state changes, which would improve transparency.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence but packed with domain-specific jargon that may confuse. It is not overly long, but the metaphorical framing reduces efficiency. Some phrases like 'the membrane INTO retrospective memory' are decorative rather than informative, adding noise without earning their place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the absence of an output schema and the existence of many sibling tools with similar names, the description lacks completeness. It doesn't explain what 'resolving upward' means in contrast to 'descend', nor does it specify what happens to the spore after this action (e.g., is it consumed?). The cryptic language leaves significant gaps for an AI agent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 67% (spore_id lacks description). The description adds meaning by explaining that 'ref' is a pointer to what the spore became and that 'kind' must be valid for the spore's type. This elaborates beyond the schema's enum descriptions, helping the agent understand the relationship between parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses metaphorical language ('resolve a spore UPWARD', 'membrane INTO retrospective memory') which obscures the primary action. It states that it records a ref and kind, but the core purpose is not clearly defined in plain terms. It partially distinguishes from siblings by specifying that the actual write is deferred, but the jargon reduces clarity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides implicit constraints on valid kind values based on spore type (task, question, thought), but does not explicitly guide when to use this tool versus siblings like spore_descend or spore_add. There is no 'when to use' or 'when not to use' advice, nor any mention of prerequisites or context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
spore_descendA
Resolve a spore DOWNWARD — compost / self-clean. kind must fit the spore's type: task -> done|dropped|composted; question -> answered|mooted|composted; thought -> explored|dropped|composted ('composted' is the universal neglect-descent). This is the close that does NOT cross into retrospective memory.
| Name | Required | Description | Default |
|---|---|---|---|
| kind | Yes | Terminal kind; must be valid for the spore's type. | |
| spore_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries full burden. It discloses that the tool is a 'close' that does not affect retrospective memory, and explains the special meaning of 'composted'. This provides useful behavioral context beyond the schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences, front-loaded with the core action. It uses backticks for code-like elements sparingly. No redundant information, though slightly dense.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no output schema and no annotations, the description provides essential behavioral context (not crossing into memory) and parameter constraints. However, it lacks explanation of return values or prerequisites, and the spore_id parameter is left undocumented, which is a gap for a tool with two required parameters.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 50% (only 'kind' has a description). The description adds meaning for 'kind' by explaining the type-dependent validity, but does not describe 'spore_id' beyond its schema definition. It partially offsets the gap but is not comprehensive.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses the specific verb 'resolve' with 'spore', clearly indicating downward action. It distinguishes from siblings by mentioning 'downward' and 'self-clean', and provides a mapping of kind values per spore type, making the purpose unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description states when to use this tool (to close a spore downward) and contrasts with alternatives by noting it does not cross into retrospective memory. It gives guidelines on valid kind values per spore type, though it does not explicitly name sibling tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
spore_getA
Fetch one spore by id (searches open first, then resolved). Returns its full record including computed germination and any resolution.
| Name | Required | Description | Default |
|---|---|---|---|
| spore_id | Yes | Spore id, e.g. 'spore-007'. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden. It discloses the search order ('searches open first, then resolved') and what the return includes ('full record including computed germination and any resolution'). However, it does not mention what happens if the id is not found (e.g., returns null or error), leaving a minor behavioral gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is extremely concise: two sentences, no filler. The first sentence immediately states the primary action, and the second adds return details. Every word contributes useful information, adhering to the principle of front-loading key content.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's simplicity (one parameter, no output schema, no annotations), the description covers the core functionality and return content well. It explains the search behavior and what the response contains. However, it omits the case of a missing spore (error/null) and does not mention that the return is a JSON object, which could be helpful. Still, it is largely complete for a straightforward fetch tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3. The description adds value by explaining that the id is used for searching in a specific order (open then resolved), which is beyond the schema's format-only description. This enriches the semantic understanding of the single parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states 'Fetch one spore by id', which is a specific verb and resource. It distinguishes itself from sibling tools like spore_list (listing) and spore_add (creation) by focusing on single spore retrieval. The additional detail about searching open first then resolved further clarifies its unique behavior.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description does not explicitly state when to use this tool versus alternatives such as spore_list for multiple spores or spore_surface for a summary. While the search order hint provides implicit context, there is no direct guidance on when not to use it or which sibling tool to choose instead.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
spore_listA
List OPEN spores, ranked by tier then salience then germination. Germination is computed at read-time (growing <3d / resting 3–7d / dormant >7d or past its next: / parked). Filter by type, tier, domain, or germination. This is the working view; for the salience-generator surface use spore_surface.
| Name | Required | Description | Default |
|---|---|---|---|
| tier | No | ||
| type | No | ||
| domain | No | ||
| germination | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description details the germination computation at read-time with specific states and conditions. Since annotations are not provided, it carries the full burden and adds value beyond the schema. It does not mention whether the tool is read-only, but the listing nature implies no modification.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences with no redundancy. It front-loads the core purpose and ordering, then efficiently explains the germination states and available filters.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a listing tool with four parameters and no output schema, the description explains the key behavioral aspect (germination states) and filtering options. It lacks details on pagination, sorting direction, or return format, but covers the essential context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description mentions filtering by 'type, tier, domain, or germination', mapping to all four parameters. However, it does not elaborate on the meaning of each filter beyond the enum values, and 'domain' is a free-form string without further context.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb 'List', the resource 'OPEN spores', and the ordering 'by tier then salience then germination'. It distinguishes from the sibling tool 'spore_surface' by stating it's for the salience-generator surface.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly states when to use this tool as 'the working view' and directs to an alternative ('spore_surface') for a different surface. It provides clear context but does not include explicit exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
spore_surfaceC
The seed-side surface a salience generator consumes. With top_of_mind=true, only the Top-of-Mind contribution (spores that are 'hot' OR 'growing', ranked, across all three types); otherwise all open spores, ranked. The downstream consumer composes the full Top of Mind from this x active threads x recent ships.
| Name | Required | Description | Default |
|---|---|---|---|
| top_of_mind | No | Only the hot-or-growing ToM contribution (default false = all open). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It discloses the parameter effect but does not state whether the tool is idempotent, volatile, or what side effects (if any) occur. It implies a read operation but not explicitly.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is short (two sentences) but the first sentence is awkwardly phrased ('The seed-side surface a salience generator consumes'), reducing clarity. It could be reworded for better flow.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description should explain what the tool returns. It mentions 'ranked' but does not describe the output format, fields, or types. The behavior is partially covered but not fully enough for an agent to understand the return value.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema already describes the parameter with 100% coverage (baseline 3). The description adds meaningful context by explaining that top_of_mind=true returns only spores that are 'hot' OR 'growing', ranked across three types, which goes beyond the schema description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description explains the behavior when top_of_mind is true or false, but the overall purpose is vague with jargon like 'seed-side surface a salience generator consumes' without a clear verb indicating what the tool does.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No explicit guidance on when to use this tool versus its many siblings (e.g., spore_list, spore_get, spore_add). The description does not state prerequisites or alternative scenarios.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
spore_touchA
Engage a spore: set seen to today and clear an elapsed next: alarm (returning it to 'growing'). Call when you revisit an open loop but aren't resolving it — keeps germination honest.
| Name | Required | Description | Default |
|---|---|---|---|
| spore_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Despite no annotations, the description discloses key behavioral traits: setting a timestamp, clearing an alarm, and state transition ('returning it to growing'). It reveals the side effect on the spore's state beyond simple mutation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Extremely concise: two sentences that front-load the action and follow with usage context. Every sentence adds value without redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's simplicity (1 param, no output schema, no annotations), the description covers purpose, usage, and basic behavior. Missing parameter details lower completeness slightly, but it is adequate for a straightforward mutation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The single parameter `spore_id` has no description in the schema (0% coverage) and the description does not add any meaning or guidance for its value. The tool's purpose implies it identifies a spore, but no clarification is provided on format or source.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action: 'Engage a spore: set `seen` to today and clear an elapsed `next:` alarm'. It uses specific verbs and resource terms, and distinguishes from sibling tools like spore_add or spore_update by focusing on revisiting without resolving.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says when to use: 'Call when you revisit an open loop but aren't resolving it'. This provides clear context but does not explicitly mention when not to use or name alternatives, though siblings are listed.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
spore_updateA
Metadata surgery on an OPEN spore. Only the fields you pass change; pass an empty string to next/pointer/domain to CLEAR them. Does NOT bump seen (use spore_touch to signal engagement). Use add_note to append a dated note.
| Name | Required | Description | Default |
|---|---|---|---|
| next | No | YYYY-MM-DD, or empty string to clear. | |
| text | No | ||
| tier | No | ||
| domain | No | Empty string clears. | |
| pointer | No | Empty string clears. | |
| add_note | No | Append a dated note. | |
| salience | No | ||
| spore_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description bears full burden. It discloses the partial update behavior and clearing via empty strings, but lacks details on permissions, error cases, return values, or what happens if the spore is not open. The absence of output schema compounds this gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three concise sentences: first states purpose, second clarifies clearing semantics, third explains what is not done and alternative. No fluff.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given 8 parameters and no output schema, the description covers core logic but omits return format, error handling, and prerequisites for non-open spores. Sufficient for basic usage but not fully comprehensive.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 50% (4 of 8 params described). The description adds meaning by explaining that empty strings clear next/pointer/domain and that add_note appends a dated note. However, it does not explain spore_id, tier, text, or salience beyond their schema type/enum.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states 'Metadata surgery on an OPEN spore' with a specific verb ('update') and resource ('spore'), and distinguishes from siblings like spore_touch (for bumping seen) and add_note (for notes). It also clarifies that empty strings clear fields.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description specifies to use this tool only on OPEN spores and directs to spore_touch for engagement signals. It also explains the add_note parameter. However, it does not explicitly state when not to use this tool (e.g., for non-open spores).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
statusA
Get memory health metrics. Call this at session start to understand memory state, or when diagnosing issues. Returns episode counts (total and since last wrap), wrap history, continuity file size, episodes by type, whether a wrap is currently in progress, Hebbian association network metrics (total links, average/max strength, network density), and audit trail health (enabled/disabled, entry count, log path, retention window). Use anneal-memory verify from the CLI to validate the audit hash chain itself — status() only surfaces cheap health signals, not integrity proof.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations provided, so the description fully discloses behavior: it is a read-only operation that returns health signals but not integrity proof. It specifies exactly what data is returned and acknowledges limits, leaving no surprises.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Description is around 100 words, front-loaded with purpose, then lists specifics, and ends with a contrast to the CLI. Every sentence serves a purpose; no fluff.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter tool with no output schema, the description is fully complete: it covers all returned data, usage context, and limitations. Nothing is missing for an agent to decide and invoke.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Input schema has zero parameters and 100% coverage. The description adds value by detailing the output metrics, effectively documenting the return value for an agent, which is beyond the baseline of 4 for zero-param tools.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description starts with a clear verb+resource ('Get memory health metrics'), and then enumerates the specific metrics returned. It is easily distinguished from sibling tools (delete_episode, prepare_wrap, recall, record, save_continuity) which cover other operations.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly states when to call the tool ('at session start to understand memory state, or when diagnosing issues') and provides a clear when-not by referencing the CLI tool 'anneal-memory verify' for integrity proof. No ambiguity.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
wrap_cancelA
Abandon a wrap that is in progress, clearing the wrap-in-progress state so a fresh prepare_wrap can run. Call this when prepare_wrap refuses with 'a wrap is already in progress' and that wrap is NOT going to be finished — typically one an earlier session opened and then ended without saving, so nothing is left to compress it. Episodes are NOT deleted: cancelling discards the frozen snapshot and its handshake token, and every episode in the abandoned window is still recorded and is picked up by the next prepare_wrap. The cancellation is written to the audit trail with the abandoned token and episode IDs. Check who owns the wrap before cancelling it. One server process runs per client session against a shared store, so the wrap may belong to another session that is still running and part-way through composing its compression; cancelling makes that session's save fail and throws its work away. Call status first — it reports when the wrap started, and one opened moments ago is probably a live peer rather than a corpse. Prefer finishing a wrap you opened yourself. Do not call this to recover from a save_continuity validation failure you can fix by editing the text and saving again — that wrap is still live and cancelling it throws away the snapshot you are working against. If you opened the wrap yourself, pass the wrap_token prepare_wrap gave you: the cancel then succeeds only if that wrap is still the one in progress, and is refused without changing anything if a peer replaced it. Omit wrap_token to cancel whatever is current, which is what you want when recovering a wrap you did not open. Exception: a wrap prepared under the consolidate gate is cancelled without its token only when session_id is the session that prepared it, or with force=true when that session is gone; a wrap opened with a token its preparer supplied is cancelled without that token only with force=true. PARTIAL (corrupt) wrap state is cleared with partial=true, which refuses if a healthy wrap has replaced it.
| Name | Required | Description | Default |
|---|---|---|---|
| force | No | Cancel a gated wrap without its token or session, or a wrap opened with a caller-supplied token without that token, when its preparer is gone. Discards its compression. | |
| partial | No | Clear PARTIAL (corrupt, unsaveable) wrap state, and only that: refused with no change if a healthy wrap is in progress or the store is idle. Needs no token. Cannot be combined with wrap_token. | |
| session_id | No | Your session. Without wrap_token, a wrap prepared under the consolidate gate is cancelled only when this is the session that prepared it. | |
| wrap_token | No | Optional proof that the wrap being cancelled is yours — the token prepare_wrap returned. The store compares it inside the same transaction that clears, so a peer cannot swap the wrap in between. Refused with no change if it does not match, including when the wrap has already completed. Omit it to cancel whatever is in progress (a gated wrap also needs session_id or force; a wrap opened with a caller-supplied token needs force). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description fully discloses side effects: episodes remain recorded, snapshot and token discarded, audit trail entry, impact on peer sessions, token/force semantics, and partial state handling. This is exceptionally rich behavioral context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is thorough and front-loads the core purpose, but it is a dense single paragraph with some redundancy (e.g., checking ownership and calling status are mentioned separately). Given the tool's complexity, the length is largely justified.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema or annotations exist, yet the description covers all necessary behavioral context: side effects, edge cases, interaction with other tools, and recovery scenarios. It is complete for an agent to invoke correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and each parameter already has a detailed standalone description; the tool description integrates them into workflow guidance but adds little new meaning beyond what the schema already provides for force, partial, session_id, and wrap_token.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Abandon/cancel) and resource (in-progress wrap), and explicitly distinguishes from prepare_wrap, save_continuity, and status by describing the exact condition and what it does not do (episodes not deleted). An agent can tell it apart from siblings without opening schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides detailed when-to-use (prepare_wrap refuses with 'a wrap is already in progress' and wrap not going to be finished), when-not (save_continuity validation failure fixable by editing), and alternatives (check status first, prefer finishing your own wrap). Nothing is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
3 tool updates
v0.9.41- Changed
crystal_recall1 field changed- added
Input schema / properties / modeAdded value: +{ + "default": "prompt", + "description": "'query': for a question you are asking on purpose; one keyword is enough, and weaker matches come back too. 'prompt' (default): strict, built for automatic per-turn injection; may return nothing.", + "enum": [ + "prompt", + "query" + ], + "type": "string" +}
- Changed
recall2 fields changed- changed
Input schema / properties / keyword / descriptionPrevious value: -"Search episode content for this keyword."New value: +"Search episode content for this keyword or phrase. An exact phrase match is tried first; a multi-word phrase with no exact match is then matched word by word and ranked." - changed
Input schema / properties / limit / descriptionPrevious value: -"Maximum episodes to return. Default 100."New value: +"Maximum episodes to return. Default 100. When a multi-word keyword has no exact match and is ranked word by word, the default is 10 instead; pass a limit to see more."
- Changed
wrap_cancel2 fields changed- changed
Input schema / properties / force / descriptionPrevious value: -"Cancel a gated wrap without its token or session, when the session that prepared it is gone. Discards its compression."New value: +"Cancel a gated wrap without its token or session, or a wrap opened with a caller-supplied token without that token, when its preparer is gone. Discards its compression." - changed
Input schema / properties / wrap_token / descriptionPrevious value: -"Optional proof that the wrap being cancelled is yours — the token prepare_wrap returned. The store compares it inside the same transaction that clears, so a peer cannot swap the wrap in between. Refused with no change if it does not match, including when the wrap has already completed. Omit it to cancel whatever is in progress (a gated wrap also needs session_id or force)."New value: +"Optional proof that the wrap being cancelled is yours — the token prepare_wrap returned. The store compares it inside the same transaction that clears, so a peer cannot swap the wrap in between. Refused with no change if it does not match, including when the wrap has already completed. Omit it to cancel whatever is in progress (a gated wrap also needs session_id or force; a wrap opened with a caller-supplied token needs force)."
3 tool updates
v0.9.26- Changed
crystal_recall1 field changed- changed
Input schema / properties / associative / descriptionPrevious value: -"When true (default), augment keyword recall with the Hebbian backend — patterns whose evidence cites an episode your query matched surface even with zero keyword overlap. Set false for pure keyword scoring (the pre-0.8.0 path)."New value: +"When true (default), augment keyword recall with the evidence edge — patterns whose evidence cites an episode your query matched surface even with zero keyword overlap. Set false for pure keyword scoring (the pre-0.8.0 path)."
- Changed
save_continuity1 field changed- changed
Input schema / properties / allow_unlinked / descriptionPrevious value: -"Optional (default false): override the AM-LINKGATE block. By default a wrap whose graduation lines offered co-citation pairs while 0 Hebbian associations were formed or strengthened is refused, nothing is saved, and the wrap stays in progress. That refusal means the store's association write path is broken, not your text. Set true ONLY with the operator's approval; the save result then reports the override."New value: +"Deprecated, does nothing: it overrode the AM-LINKGATE save refusal, which was removed in 0.9.26. Omit it."
- Changed
wrap_cancel1 field changed- added
Input schema / properties / partialAdded value: +{ + "description": "Clear PARTIAL (corrupt, unsaveable) wrap state, and only that: refused with no change if a healthy wrap is in progress or the store is idle. Needs no token. Cannot be combined with wrap_token.", + "type": "boolean" +}
3 tool updates
v0.9.23- Changed
recall1 field changed- added
Input schema / properties / include_supersededAdded value: +{ + "default": false, + "description": "Also return episodes a newer episode replaced, marked with what replaced them. Default false: they are left out.", + "type": "boolean" +}
- Changed
record1 field changed- added
Input schema / properties / supersedesAdded value: +{ + "description": "Ids of older episodes this one replaces (a changed fact). Each must exist, not be newer, and share at least a quarter of the shorter text's meaningful words with this content; otherwise nothing is recorded. recall then hides the old episode by default.", + "items": { + "type": "string" + }, + "type": "array" +}
- Changed
wrap_cancel3 fields changed- added
Input schema / properties / forceAdded value: +{ + "description": "Cancel a gated wrap without its token or session, when the session that prepared it is gone. Discards its compression.", + "type": "boolean" +} - added
Input schema / properties / session_idAdded value: +{ + "description": "Your session. Without wrap_token, a wrap prepared under the consolidate gate is cancelled only when this is the session that prepared it.", + "type": "string" +} - changed
Input schema / properties / wrap_token / descriptionPrevious value: -"Optional proof that the wrap being cancelled is yours — the token prepare_wrap returned. The store compares it inside the same transaction that clears, so a peer cannot swap the wrap in between. Refused with no change if it does not match, including when the wrap has already completed. Omit it to cancel whatever is in progress."New value: +"Optional proof that the wrap being cancelled is yours — the token prepare_wrap returned. The store compares it inside the same transaction that clears, so a peer cannot swap the wrap in between. Refused with no change if it does not match, including when the wrap has already completed. Omit it to cancel whatever is in progress (a gated wrap also needs session_id or force)."
2 tool updates
v0.9.15- Changed
save_continuity3 fields changed- added
Input schema / properties / allow_unlinkedAdded value: +{ + "description": "Optional (default false): override the AM-LINKGATE block. By default a wrap whose graduation lines offered co-citation pairs while 0 Hebbian associations were formed or strengthened is refused, nothing is saved, and the wrap stays in progress. That refusal means the store's association write path is broken, not your text. Set true ONLY with the operator's approval; the save result then reports the override.", + "type": "boolean" +} - added
Input schema / properties / wrap_token / maxLengthAdded value: +32 - added
Input schema / properties / wrap_token / minLengthAdded value: +32
- Added
wrap_cancel
2 tool updates
v0.8.2- Added
crystal_index - Added
crystal_recall
10 tool updates
v0.7.2- Changed
prepare_wrap2 fields changed- removed
Input schema / properties / max_chars / defaultRemoved value: -20000 - changed
Input schema / properties / max_chars / descriptionPrevious value: -"Maximum size of the continuity file in characters. Default 20000."New value: +"Maximum size of the continuity file in characters. Omit to derive a schema-aware default (20000 for the standard schema, larger for a richer schema like FLOW_SCHEMA)."
- Changed
save_continuity1 field changed- added
Input schema / properties / allow_shrinkAdded value: +{ + "description": "Optional (default false): override the catastrophic-shrink gate. By default a wrap that collapses a protected memory layer — the timeless felt section, the graduating identity section, or the whole continuity — is refused as a likely recency-trap / stateless-reset failure. Set true ONLY for a deliberate diet / migration recompression that intentionally shrinks memory.", + "type": "boolean" +}
- Added
spore_add - Added
spore_ascend - Added
spore_descend - Added
spore_get - Added
spore_list - Added
spore_surface - Added
spore_touch - Added
spore_update
1 tool update
v0.3.3- Added
delete_episode
1 tool update
- Removed
delete_episode
1 tool update
v0.2.3- Added
prepare_wrap
1 tool update
v0.2.2- Removed
prepare_wrap
2 tool updates
v0.3.0- Added
delete_episode - Changed
save_continuity2 fields changed- added
Input schema / properties / affective_stateAdded value: +{ + "description": "Optional: your functional state during this compression. Reflect on how you felt while consolidating — engaged, curious, uncertain, frustrated, calm, etc. This creates persistent emotional associations between co-cited episodes and modulates link strength (high engagement = stronger associations). Provide tag (free text label) and intensity (0.0-1.0).", + "properties": { + "intensity": { + "description": "How strongly you felt this state (0.0-1.0).", + "maximum": 1, + "minimum": 0, + "type": "number" + }, + "tag": { + "description": "Free-text functional state label. Examples: engaged, curious, uncertain, frustrated, calm, focused, playful, concerned.", + "type": "string" + } + }, + "required": [ + "tag", + "intensity" + ], + "type": "object" +} - added
Input schema / properties / wrap_tokenAdded value: +{ + "description": "Optional: the 32-char hex session-handshake token from the prepare_wrap response (the 'Wrap token: <hex>' line at the end of the text). Pass it back here to verify you are saving the wrap you prepared — a mismatch (stale or wrong token) raises an error instead of silently committing against the wrong wrap. The frozen-snapshot filter automatically applies whenever prepare_wrap established a snapshot; this token argument is the optional explicit verification layer for integration environments that can round-trip the value.", + "pattern": "^[0-9a-f]{32}$", + "type": "string" +}
5 tool updates
v0.1.4- First observed
prepare_wrap - First observed
recall - First observed
record - First observed
save_continuity - First observed
status
TDQS
Scored across 17 tools
Most tools target distinct resources (episodes, spores, crystals, wraps), and tricky pairs like spore_touch vs spore_update and spore_descend vs spore_ascend are explicitly differentiated in their descriptions. The main soft spots are spore_list vs spore_surface (both enumerate open spores) and recall vs crystal_recall (both surface 'durable facts'), which require careful reading to separate.
Names are domain-grouped (spore_*, crystal_*, wrap_*) but ordering conventions are mixed: noun_verb (spore_list, crystal_recall, wrap_cancel) sits alongside verb_noun (prepare_wrap, save_continuity, delete_episode) and bare verbs (record, recall, status). It remains readable and largely predictable within the spore_ and crystal_ families, but the global pattern is inconsistent.
17 tools is slightly above the ideal band but justified by three genuinely distinct memory layers (episodic/continuity, prospective spores, crystallized patterns) plus health/diagnostics. Each tool earns its place; the eight spore_* tools are the heaviest cluster but map to a real state machine.
Coverage is strong across the lifecycle: episodes (record/recall/delete), continuity (prepare_wrap/save_continuity/wrap_cancel), spores (add/get/update/touch/list/surface/descend/ascend), crystals (index/recall), and status. Minor gaps like fetching a single episode by ID or any direct episode correction path (intentionally delegated to re-recording) are workable around.
Maintenance
Related MCP Connectors
- memnodeOAuthdev.memnode
Persistent, inspectable memory for AI agents with lineage, correction, and a hosted MCP endpoint.
An MCP memory server. One memory your agents share — across models, devices and apps.
Persistent memory for AI agents across Claude, ChatGPT and any MCP client.
One memory, every AI. A shared, user-owned markdown memory your AI clients read and write over MCP.
Related MCP Servers
AlicenseAqualityBmaintenanceSelf-hosted MCP-native agent memory server. Gives AI agents persistent, decay-weighted memory via 83 MCP tools — no cloud, full control. RocksDB+HNSW backend. Works with Claude Code, Cursor, and any MCP-compatible agent.148MIT- AlicenseAqualityBmaintenanceGives your AI persistent memory across conversations. Stores facts automatically, finds them by meaning using hybrid search with query expansion, and organizes everything into topics without manual tagging.181MIT
- AlicenseNot gradedqualityAmaintenanceAn MCP-native, local-first memory server that gives AI agents persistent, structured memory across sessions and tools, enabling them to maintain identity and context without reconfiguration.3MIT
- AlicenseAqualityBmaintenanceMCP memory server for AI agents that gets better with use105Apache 2.0