CPersona
OfficialCPersona is a persistent memory server for AI agents over MCP, storing and retrieving memories in a local SQLite file with hybrid search.
Store and manage memories:
store,update_memory,delete_memory,lock_memory/unlock_memory, andlist_memoriesfor CRUD and dashboard-style listing.Recall with hybrid search:
recall(vector + FTS5 + keyword),recall_with_context(auto-merge with conversation history),reconstruct(assemble evidence-backed items from candidate rows), andget_contents(expand preview-tier rows to full text).Manage episodes:
archive_episode,list_episodes, anddelete_episode.Maintain agent profiles:
get_profileandupdate_profile.Declare and traverse associative memory:
declare_associations(entities, aliases, relations) andtraverse(graph neighborhood exploration).Tune retrieval behavior:
calibrate_threshold(auto-calibrate vector gate) andset_recall_precision/get_recall_precision(strict/balanced/lenient specificity).Operate and maintain:
check_health(33-check registry with auto-repair),deep_check(heuristic data-quality analysis),get_queue_status, andmigrate_channel_axis.Backup and portability:
export_memories,import_memories,merge_memories(agent-to-agent), anddelete_agent_data.Session/process control:
pause_persistence,resume_persistence, andpersistence_statusfor opt-in no-write windows.Read operator doctrine:
get_operating_contextfor the server-served operating context.Self-update:
check_updateto check for and (for pip/checkout installs) apply newer releases.
CPersona
MCP Memory Server
Persistent memory for AI agents, over MCP. One SQLite file you own. No LLM in the loop. Honest when recall degrades. CPersona is an MCP persistent memory server that stores AI agent memories in a local SQLite file.
Documentation · Getting Started · Architecture · Tools · PyPI · Zenn Book (JP)
Standalone repository — This is the standalone version for use with Claude Desktop, Claude Code, Codex CLI, Cursor, VS Code, and any other MCP client (registration table). If you are a ClotoCore user, install CPersona from the in-app marketplace (ClotoHub) instead — it distributes this same repository.
Project status — 2.4.x is Stable; 2.5.x is Current, an internal stabilization line where all fixes land, pending production-soak certification. The DB schema is preserved across the line. Additive, rollback-safe features may land here as well (lifecycle standard §2.6); a change that cannot be rolled back waits for 2.6. Which version to run, and how long each line keeps receiving fixes: SUPPORT.md. Where the lines are heading: the roadmap.
Upgrading from 2.5.2 or earlier? Two things need a decision from you. v2.5.3 will not start the HTTP transport without
CPERSONA_AUTH_TOKEN, wherever it binds — set one, or opt out withCPERSONA_ALLOW_UNAUTHENTICATED_HTTP=true(why; stdio is unaffected). v2.5.2 changed tool response shapes — branch onok is false, and treat any response carryingerroras a failure whether or notokis present (contract §10).
The Problem
Claude forgets everything between sessions. Every conversation starts from zero — no context about your project, your preferences, or what you discussed yesterday.
cpersona fixes this. It's an MCP server that stores memories in a local SQLite file and retrieves them through hybrid search. Claude remembers you. It runs against any MCP-compatible host — Claude Desktop, Claude Code, ClotoCore (the AI agent platform where cpersona originated, and whose memory layer it is), or a client of your own.
Related MCP server: mcp-memory-graph
Quick Start
Setting this up as an AI agent? Read the Agent Skill —
skills/cpersona-memory/SKILL.mdin a clone of this repository — and follow it end to end. It is the complete procedure: install, embedding server, registration, and the two steps a quick start tends to skip — verifying from a new session, and persisting the memory policy so the next session knows whichagent_idholds the memories.Using Claude Code yourself? The same skill ships inside the wheel. Once
cpersonais installed, copy it in and say "Set up CPersona."python -c "import cpersona,pathlib,shutil; s=pathlib.Path(cpersona.__file__).parent/'skills'/'cpersona-memory'; shutil.copytree(s, pathlib.Path.home()/'.claude/skills/cpersona-memory', dirs_exist_ok=True)"
1. Install — Python 3.11+, and uv for the one-command path.
uvx cpersona # run directly, no install step
pip install cpersona # or install it2. Run an embedding server — strongly recommended; it powers the vector layer
uvx --from "cembedding[onnx]" cembedding-download-model --model jina-v5-nano
EMBEDDING_PROVIDER=onnx_jina_v5_nano uvx --from "cembedding[onnx]" cembedding # serves http://127.0.0.1:8401/embedThe reference server's lifetime is bound to its stdin. Started with stdin closed — by a service manager, by nohup … </dev/null, or from an agent's background shell — it binds the port and exits within the same second with status 0. Give it a stdin that stays open (sleep infinity | cembedding); Getting Started has the details.
Any endpoint implementing the embedding contract works and is equally recommended; CEmbedding is the reference implementation. The choice of backend is yours — the recommendation is to connect one, not to connect that one.
Without a backend, cpersona still runs — FTS5 + keyword search, and it says on every recall that it is degraded rather than quietly returning less. That is a supported fallback, not a recommended way to run: recall then matches on shared words, so a memory phrased differently from your question can be missed, and so can an older one.
3. Register it with your MCP client
claude mcp add-json cpersona '{"type":"stdio","command":"uvx","args":["cpersona"],"env":{"CPERSONA_DB_PATH":"/home/you/.claude/cpersona.db","EMBEDDING_MODE":"http","EMBEDDING_HTTP_URL":"http://127.0.0.1:8401/embed"}}' -s user4. Verify from a new session — ask the agent to store something, then recall it in a fresh session. Surviving the session boundary is the whole point.
5. Make it stick — registration gives the agent the tools. It does not tell the next session to use them, or which agent_id holds the memories: recall is scoped to an exact agent_id, so a session that guesses the wrong one gets nothing back. Persist the short policy block into the file your client loads every session (~/.claude/CLAUDE.md, AGENTS.md, …) — Getting Started §5.
At startup the server asks pypi.org whether a newer release exists and tells the
calling agent through recall; set CPERSONA_UPDATE_CHECK=false to turn that
off. Updating is never automatic.
Claude Desktop config, Windows paths, installing from source and the full walkthrough: Getting Started.
What You Get
Hybrid search — vector (the layer an embedding server powers), FTS5 (trigram, so it works on Japanese and other space-less scripts) and keyword, fused by rank or relative score. The FTS and keyword layers rescue what vectors miss: identifiers, error strings, exact names.
Three memory types — facts, session summaries and an accumulated profile.
Zero LLM dependency — cpersona never calls a generative model; your agent summarizes and hands over the result. Recall is deterministic given a calibrated gate, but the gate is sampled, so two installs on identical data can settle differently.
Single-file SQLite — no external database;
sqlite3 .backupcopies the corpus (the calibration sidecar beside it needs copying too).Operable — auto-calibrated thresholds, a health check with auto-repair, an advisory when the embedding layer dies, JSONL export/import, agent-to-agent merge.
Isolation —
agent_id,project_idandchannellet several agents and projects share one database without bleeding into each other.
How it fits together: Architecture · what the tools do: Tools · what you may rely on: Behavior Contracts.
Benchmarks
Measured on LMEB (Long-horizon Memory Embedding Benchmark, arXiv:2603.12572) — 22 datasets subsuming LoCoMo and LongMemEval, measured here as 22 retrieval tasks. The metric is Mean NDCG@10 across all 22 tasks. Track A is the raw embedding model alone; Track B routes the same embeddings through cpersona's real store/recall code paths (SQLite + FTS5 + RRF fusion + per-agent auto-calibration).
Embedding Model | Params | Dim | Track A (raw) | Track B (cpersona) | Δ |
all-MiniLM-L6-v2 | 22M | 384 | 43.67 | 50.10 | +6.43 |
bge-m3 | 568M | 1024 | 56.83 | 57.66 | +0.83 |
Track B lands at or above Track A on both models: the fusion layers add signal rather than merely persisting vectors, and a weaker embedding gains more because the FTS5/keyword layers rescue what its vectors miss. How to read the deltas, the noise envelope, the measurement harness and the reproduction regime: benchmarks/.
Documentation
cloto-dev.github.io/CPersona is canonical — when this README disagrees with it, the site wins.
Install, embedding server, client registration, verification | |
What you may rely on: recall ordering, dedup, scan window, response shapes | |
Every tool, grouped by what you reach for it for | |
Storage, the retrieval pipeline, isolation axes | |
What each release line is for and may break; planned retrieval features and the scale ladder | |
Backup, degradation detection, tuning, CJK guidance, corpus sync | |
Every environment variable and its default | |
How a release is gated: audits, the bug ledger, structural and mutation gates | |
Short answers to the questions operators actually ask |
Japanese translations are in the language selector (English is canonical) and
agents can read llms.txt.
Longer reads in Japanese: a book
on the design and setup, and an article
on the token economics of session-end → /clear → recall.
Quality Assurance
Every release is gated by a machine-verifiable process: multi-agent audit rounds with adversarial verification, a bug ledger that fails CI if a fix marker disappears or a removed defect returns, structural gates for invariants a plain test cannot express, a mutation proof that those gates go red when the invariant is broken, and gates holding the documented counts, defaults and version claims to the source that defines them.
Behind it: ~2,572 test functions across ~183 test modules (~3,309 cases parametrised, more test code than server code), on Schema v17 — how a release is gated.
Support
Three tiers — Stable (production-certified, critical fixes only), Current (newest line, all fixes land here) and Experimental (opt-in pre-releases). A superseded line keeps critical-fix support for 30 more days. Read SUPPORT.md § Known issues before pinning a version — some of them change what you should run.
Found a bug, or something the docs do not explain? Open a bug report or feature request, even when you are not certain — a configuration problem mistaken for a bug means the documentation was unclear, which is a defect of its own. Report security vulnerabilities privately via SECURITY.md.
Sponsorship
CPersona is MIT-licensed and stays fully usable whether or not anyone sponsors it. Sponsorship buys no feature, no release tier and no position in the issue queue — issues are triaged by impact, reproducibility and safety, and that does not change for anyone.
If CPersona has earned a place in your workflow and you would like the work to continue, you can sponsor Cloto-dev on GitHub. The same page covers CPersona, ClotoCore and the other projects published under that account; sponsorship goes toward development time, testing and infrastructure, documentation and maintenance.
Money is not the only thing that helps, and it is not the thing this project needs most. Starring the repository, saying which part of the setup was confusing, filing a reproducible issue, or correcting a sentence in the documentation all move it forward.
License
MIT — free to use from any MCP host without restriction.
Available Tools
34 toolsarchive_episodeA
Archive a conversation episode with pre-computed summary, keywords, and resolved status. All LLM processing is performed by the caller. A summary that runs past the embedding window adds nodes:{status:'queued'} to the response, as on store.
| Name | Required | Description | Default |
|---|---|---|---|
| channel | No | v2.4.22 conversation-channel tag (e.g. a Discord channel id). Default '' (= unscoped). Channel-scoped recall returns episodes whose channel matches; this powers the per-channel episodic loop. | |
| history | No | Original conversation messages (used for start/end timestamp extraction; the episode embedding is computed from summary) | |
| summary | Yes | Episode summary (pre-computed by caller) | |
| agent_id | Yes | Agent identifier | |
| keywords | No | Space-separated keywords (pre-computed by caller) | |
| resolved | No | Whether the topic was completed/concluded | |
| project_id | No | v2.4.17 isolation axis. Omit or pass '' for the global pool. v2.5.1: pass '@auto' to resolve this agent's default from the server's operating context (the resolution is echoed as resolved_project_id; an unmapped agent yields operating_context_warning). bug-186: resolution requires a configured operating context. With none — the default, and equally the outcome of a sidecar that fails to parse — the sentinel is NOT resolved: it is stored and filtered as the literal project_id '@auto', resolved_project_id echoes '@auto', and no warning is raised. Read resolved_project_id before relying on the resolution. | |
| session_key | No | Opaque session identity you declare: a partition hint, not authentication and not a data filter. Selects which no-persist pause applies to this call. Omit to share one bucket with every caller that omits it. Full text on recall. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only state non-readonly, non-idempotent, non-destructive; the description adds the meaningful edge case that an oversized summary results in nodes:{status:'queued'} in the response, and that the caller owns all LLM processing. It does not contradict the annotations and provides useful behavioral context beyond them.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences with no filler. The core operation is front-loaded, the caller-responsibility constraint is stated clearly, and the edge-case behavior is given in a single efficient sentence.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with 8 fully-documented parameters and no output schema, the description covers the essential operation, caller obligations, and a notable response anomaly. It relies on 'as on store' for one behavior, which is acceptable but slightly indirect; otherwise an agent has enough context to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description's first sentence restates the 'pre-computed' nature of summary, keywords, and resolved status, but adds no parameter-level meaning beyond what the schema already documents.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb ('Archive'), a specific resource ('conversation episode'), and the key payload elements ('pre-computed summary, keywords, and resolved status'). It further distinguishes from siblings by stating 'All LLM processing is performed by the caller,' which clarifies that this tool is not the one that generates summaries or embeddings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The condition for using this tool is implied: the caller must already have a pre-computed summary, keywords, and resolved status. However, it never explicitly names an alternative or says 'use store instead if you need LLM processing.' The reference to 'as on store' hints at sibling behavior but does not provide concrete when-to-use vs. when-not-to-use guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
calibrate_thresholdA
Auto-calibrate the vector search threshold from the null (random-pair) cosine distribution. Samples random memory pairs and places the threshold ABOVE the null mean so unrelated pairs are rejected. method='separation' (default) learns the operating point from two populations — null pairs vs temporally-adjacent same-session positives (nearest-neighbour fallback when too few exist); method='percentile' uses a quantile of the null distribution (robust to anisotropic models such as bge-m3); method='zscore' uses mean + z*std. No labels used, purely statistical. Adapts to both embedding model and corpus characteristics.
| Name | Required | Description | Default |
|---|---|---|---|
| method | No | 'separation' (default; two-population — learns the operating point from null pairs vs temporally-adjacent same-session positives, falling back to nearest-neighbour when too few exist), 'percentile', or 'zscore' | |
| agent_id | Yes | Agent ID whose memories to sample | |
| z_factor | No | Z-score multiplier for method='zscore' (default: 1.0, higher = stricter) | |
| percentile | No | Null-distribution quantile for method='percentile' (default: 0.95, higher = stricter) | |
| sample_size | No | Number of embeddings to sample (default: 200) | |
| session_key | No | Opaque session identity you declare: a partition hint, not authentication and not a data filter. Selects which no-persist pause applies to this call. Omit to share one bucket with every caller that omits it. Full text on recall. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only say the operation is not read-only, not idempotent, and not destructive. The description adds meaningful behavioral detail: it samples memory pairs, places the threshold above the null mean, uses no labels, and adapts to model and corpus. It stops short of explicitly stating whether the calibration result is persisted or how it affects future recalls, but the added algorithm-level transparency is strong.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but well-structured: purpose first, then method details. Every sentence adds meaningful information, though the method explanations are slightly verbose and overlap with the schema descriptions.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers the algorithm and method choices well, but it does not state what the calibration changes in persisted state, whether there are prerequisites (e.g., existing memories), or what the return value is. Given the absence of an output schema and the mutating nature of the tool, these gaps leave the description slightly incomplete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3. The description goes beyond the schema by explaining the statistical rationale for each method, including when percentile is robust (anisotropic models like bge-m3) and how separation falls back to nearest-neighbour. This helps an agent choose sensible parameter values.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states a specific action ('Auto-calibrate') on a specific resource ('the vector search threshold'), grounded in a concrete statistical basis (the null cosine distribution). It also distinguishes the different calibration methods, making the tool's intent unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explains how the tool works and the method options, but it does not say when to use this tool versus alternatives like set_recall_precision or get_recall_precision. It gives no explicit usage conditions, prerequisites, or exclusion criteria.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
check_healthA
Check memory database health (33-check registry, each issue tagged with severity critical/warn/info). Detects contamination, duplicates, oversized content, embedding issues, FTS integrity (count + content-level), schema version/object drift (missing UNIQUE indexes or FTS triggers), SQLite file integrity, project_id naming drift, invalid JSON/timestamps, timestamp format drift, stale tasks, missing profiles, empty content, invalid/anonymous sources. Returns storage stats incl. project_id/channel distributions. Set fix=true to auto-repair (agent-scoped, locked-safe); the one exception is dedup_msg_id_index, whose repair blanks colliding msg_id values under every agent because the UNIQUE index it restores is a global schema object — with an ACL configured that repair demands read-write on '*', so exclude it via checks to stay agent-scoped. critical file-integrity findings are report-only. Two repairs are lossy and irreversible, each against its own cap: oversized memories are cut to CPERSONA_MAX_CONTENT_LENGTH (default 16000 since 2.5.4a2) and the agent's profile row to CPERSONA_MAX_PROFILE_LENGTH (default 2000), keeping the start. Lower either cap and a fix run shortens rows that were within the old one. Some repairs are bounded per run (source canonicalisation classifies at most 10000 rows); a fix response carrying remaining > 0 with a re-run hint has NOT converged — run fix again until remaining stops decreasing. Use checks parameter to run a subset — an unknown name is rejected (ok=false) rather than silently running nothing, and every response echoes checks_run. The verdict is status: healthy / degraded / unhealthy, derived from severity counts (info never degrades). The pre-2.5.2b1 healthy boolean (len(issues) == 0) is gone — it reported False for an info-only database that status called healthy; read issues / severity_summary for the underlying counts. Read status as a verdict on what is IN the database, not on whether the pipeline that fills it is working: a corpus where every embedding is NULL is internally consistent, so it scores healthy while semantic recall is dead. Nothing here contacts the embedding backend unless fix=true — on a report-only run the liveness findings cannot appear at all, and their absence is not evidence the backend answered. The null_embedding finding carries the reason its repair cannot run; read that before reading status.
| Name | Required | Description | Default |
|---|---|---|---|
| fix | No | Auto-fix detected issues | |
| checks | No | Registry check names to run (empty = all). See cpersona.checks.HEALTH_CHECK_NAMES. | |
| agent_id | No | Agent ID to check (empty = all agents) | |
| session_key | No | Opaque session identity you declare: a partition hint, not authentication and not a data filter. Selects which no-persist pause applies to this call. Omit to share one bucket with every caller that omits it. Full text on recall. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With only readOnlyHint=false in annotations, the description carries the full load, and it delivers extensively: irreversible lossy repairs with caps, per-run bounds, unknown check names rejected via ok=false, status semantics, removed healthy boolean, no embedding-backend contact unless fix=true, and the null_embedding repair caveat. Nothing contradicts the non-read-only annotation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long but dense, with almost no filler; the first sentence front-loads the purpose and every subsequent sentence adds a concrete caveat or behavior. It could be better organized into paragraphs or bullet-like sections, but each sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity, the absence of an output schema, and the many edge cases around fixes and status interpretation, the description is remarkably complete. It explains return semantics (`status`, `remaining`, `checks_run`, `issues`, `severity_summary`), dangerous repair behavior, and what absence of findings does and does not mean.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds real value beyond the schema: fix=true triggers auto-repair with specific behavioral caveats, the checks parameter has rejection/echo semantics and scoping implications, and the relationship between fixes and convergence is explained. It doesn't add extra meaning to agent_id or session_key, but those are already well described in the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource ('Check memory database health') and enumerates the exact classes of issues detected, making the tool's function unmistakable. It does not explicitly differentiate this tool from the sibling deep_check, so it stops short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives rich behavioral guidance: when to use fix=true, how to stay agent-scoped by excluding dedup_msg_id_index, when re-runs are required until `remaining` stops decreasing, and the fact that report-only runs cannot produce liveness findings. It does not explicitly contrast check_health with sibling alternatives such as deep_check, so it lacks an explicit when-not/alternative statement.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
check_updateA
Report whether a newer release of this server exists, and — only if you ask — install it. The check itself runs ONCE per process start, in a background task that nothing waits on, and its verdict is cached for 24h (CPERSONA_UPDATE_CHECK_INTERVAL_SECONDS) in a file beside the database; a bare call here reads that verdict and reaches neither the network nor the disk. state is one of: ok (running the newest release) / newer (a newer final release exists — pre-releases are never proposed) / yanked (every file of the RUNNING version has been withdrawn on PyPI; reason carries the publisher's text) / unlisted (this version is not on the index at all — a development checkout; not a defect) / unknown (no check has completed, e.g. no network) / disabled. install names how this process was installed (uvx / pip / checkout / unknown) and the exact command that would update it. refresh=true performs the fetch now (3s budget) and updates the cache. apply=true runs that command as an argv list (never a shell), returning exit_code and the last 40 lines of output — supported for pip and checkout installs only; under uvx the environment is a cache entry keyed by the launch arguments, so the update belongs in your client's config (uvx cpersona@latest), and an install here would be discarded on the next launch. A checkout parked at a tag (detached HEAD) is likewise refused before anything runs, and answers with the git fetch --tags && git checkout <tag> form to use instead. Updating is NEVER automatic and never a side effect of any other call. A RESTART IS ALWAYS REQUIRED afterwards: this process keeps serving the old code until it is replaced. Unaffected by pause_persistence — an install writes no memory row, so a no-persist session can still repair a withdrawn version. Set CPERSONA_UPDATE_CHECK=false to disable the feature entirely: no fetch, no cache, no notice on recall or check_health, and this tool answers state=disabled.
| Name | Required | Description | Default |
|---|---|---|---|
| apply | No | Run the detected update command (pip / checkout installs only). Off by default; a restart is required afterwards. | |
| refresh | No | Fetch the package index now instead of reading the cached verdict (3s budget; a failure answers state=unknown). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description extensively discloses behavior beyond the minimal readOnlyHint:false annotation: the once-per-process background check, 24h cache, no network/disk on bare calls, exact state meanings, restart requirement, no automatic updates, install-method constraints, and independence from pause_persistence. This is exemplary transparency for a tool with side-effecting capabilities.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The content is dense and every major point is relevant, but it is delivered as one long monolithic paragraph with no bullets or headings, making it harder to parse. It is somewhat longer than necessary, with minor repetitions around 'never automatic' and 'only if you ask'.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description fully covers the return contract: state values, reason, install command, exit_code, and output truncation. It also addresses edge cases such as detached HEAD, uvx cache behavior, and the CPERSONA_UPDATE_CHECK disable path, so an agent has everything needed to invoke the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, but the description adds substantial meaning: refresh has a 3s budget and failure yields state=unknown; apply returns exit_code and last 40 lines, never uses a shell, and is restricted to pip/checkout installs. The description enriches both boolean parameters well beyond their schema definitions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a precise verb and resource: 'Report whether a newer release of this server exists' and explicitly scopes the install capability as opt-in. It clearly separates check, refresh, and apply behaviors, making the tool's purpose unambiguous even among many siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context for when a bare call, refresh=true, and apply=true are appropriate, and explicitly states when apply is refused (checkout at tag, uvx). It does not explicitly name alternative sibling tools, but the usage boundaries are otherwise detailed enough for an agent to decide correctly.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
declare_associationsAIdempotent
Declare associative memory after the fact: entities with aliases, and subject–predicate–object relations, recorded verbatim and walked by reconstruct. The server extracts nothing and infers nothing — coverage is exactly what was declared. Names, aliases and predicates are compared after normalization (NFKC, case-folded, whitespace collapsed), so two declarations that normalize alike are one entity. An alias resolves to at most one entity per scope; a second claim on it is dropped. anchor_ref names the record (mem:<id> / ep:<id>) the declaration is evidenced by: every entity named is recorded as mentioned by it and every relation carries it. A relation's endpoint is an entity name (registered if new) or a record ref of this agent. Malformed items are reported in dropped and skipped; nothing else in the call is refused for them. retract removes relations by id and mentions by {entity, ref} — the only way a declaration leaves the store. Response: {ok, result:'declared', entities:[{id, name, created}], mentions, relations:[ids], dropped:[{item, reason}], retracted?:{relations, mentions}}. Under pause_persistence nothing is written (result:'skipped', persisted:false).
| Name | Required | Description | Default |
|---|---|---|---|
| channel | No | Memory channel the declaration belongs to. Default: '' (shared). | |
| retract | No | Declarations to remove: {relations: [relation ids], mentions: [{entity: <entity id>, ref: 'mem:<id>'}]}. Only this agent's rows are touched. | |
| agent_id | Yes | Agent identifier | |
| anchor_ref | No | The record this declaration is evidenced by: 'mem:<id>' or 'ep:<id>' of this agent. Optional. | |
| project_id | No | Project the declaration belongs to. Optional — omit or pass '' for the global pool. v2.5.1: pass '@auto' to resolve this agent's default from the server's operating context (the resolution is echoed as resolved_project_id; an unmapped agent yields operating_context_warning). bug-186: resolution requires a configured operating context. With none — the default, and equally the outcome of a sidecar that fails to parse — the sentinel is NOT resolved: it is stored and filtered as the literal project_id '@auto', resolved_project_id echoes '@auto', and no warning is raised. Read resolved_project_id before relying on the resolution. | |
| session_key | No | Opaque session identity you declare: a partition hint, not authentication and not a data filter. Selects which no-persist pause applies to this call. Omit to share one bucket with every caller that omits it. Full text on recall. | |
| associations | No | Associative memory to declare alongside this call: entities the text mentions, with their aliases, and subject–predicate–object relations. Stored verbatim; the server extracts nothing and infers nothing. On store, the stored memory is recorded as mentioning every entity named here and anchors every relation. Malformed items are reported in the response's associations.dropped and skipped; the memory is stored regardless. Optional. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description goes well beyond the annotations by disclosing normalization rules, alias conflict resolution, dropped-item handling, pause_persistence behavior, retract semantics, and the exact response shape. It provides rich behavioral detail without contradicting the readOnlyHint, idempotentHint, or destructiveHint annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense and front-loaded with the core purpose, but the long single-paragraph format makes it harder to scan. It is appropriately detailed for a complex tool with nested parameters and no output schema, though it could benefit from structured breaks or bullet points.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity, nested schema, and absence of an output schema, the description is remarkably complete. It covers response fields, failure handling, retraction, persistence-pause behavior, edge cases around project_id resolution, and interaction with reconstruct. Nothing essential is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Although the schema already covers 100% of parameters, the description adds substantial meaning: anchor_ref semantics, relation endpoint rules, alias deduplication, project_id '@auto' resolution caveats, session_key partitioning, and retract scope. This is far beyond the baseline expected for fully documented schemas.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Declare associative memory after the fact' and immediately defines entities with aliases and subject–predicate–object relations. It also distinguishes itself from sibling tools by stating that the server 'extracts nothing and infers nothing' and that coverage is exactly what was declared, making its role clear.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context: declarations are evidenced by an anchor_ref, stored verbatim, and later walked by reconstruct, while retract is the only way to remove them. It does not explicitly name alternative tools or exclusion conditions, but the intended usage is evident from the behavior described.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
deep_checkA
Deep heuristic analysis of memory data quality. Detects issues requiring recovery or judgment (anonymous sources, short/trivial content, stale profiles, orphaned episodes, stale threshold calibration, embedding-space near-duplicate pairs as merge candidates). fix=true applies repairs for: anonymous_source, short_content. Report-only (fix is accepted and ignored): stale_profile, orphaned_episodes, calibration_staleness, near_duplicate, unnormalized_content, embedding_norm — apply those decisions via merge_memories / delete_memory / calibrate_threshold / update_profile. Use checks parameter to select specific checks.
| Name | Required | Description | Default |
|---|---|---|---|
| fix | No | Apply repairs (default: dry-run preview only) | |
| checks | No | Checks to run (empty = all). Options: anonymous_source, short_content, stale_profile, orphaned_episodes, calibration_staleness, near_duplicate, unnormalized_content, embedding_norm | |
| agent_id | Yes | Agent ID to check (required) | |
| session_key | No | Opaque session identity you declare: a partition hint, not authentication and not a data filter. Selects which no-persist pause applies to this call. Omit to share one bucket with every caller that omits it. Full text on recall. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With only readOnlyHint=false in annotations, the description carries the burden of disclosure. It clearly states that fix=true mutates data for two check types, that fix is accepted but ignored for other checks, and that the default is dry-run preview. This is honest and useful, though it does not detail exactly what the repairs do.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is information-dense with no filler. Each sentence earns its place: opening purpose, list of detected issues, fix/report behavior, and routing to relevant tools. The critical behavioral caveat about fix being ignored for report-only checks is clearly stated.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers the tool's purpose, fix behavior, check selection, and follow-up actions, which is substantial for a four-parameter tool with no output schema. It does not describe the report/return format, but given the detailed check enumeration and routing instructions, the agent has enough context to call and interpret results reasonably.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3. The description adds real meaning by explaining that fix=true means 'apply repairs' versus dry-run, that certain checks ignore the fix flag, and that checks is an array where empty means all. This goes beyond the raw parameter descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'Deep heuristic analysis of memory data quality.' It enumerates the exact issues it detects (anonymous sources, short/trivial content, stale profiles, etc.) and links follow-up actions to sibling tools, making its scope clear and distinguishable from related tools like check_health.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear usage semantics: fix=true applies repairs only for anonymous_source and short_content, while other checks are report-only and route to merge_memories/delete_memory/calibrate_threshold/update_profile. It also explains the checks parameter for selecting specific checks. It does not explicitly contrast with check_health, but it gives enough operational guidance to choose and act on this tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
delete_agent_dataADestructiveIdempotent
Delete ALL data (memories, profiles, episodes) for a specific agent. Used by kernel during agent deletion.
| Name | Required | Description | Default |
|---|---|---|---|
| agent_id | Yes | Agent ID whose data should be purged | |
| session_key | No | Opaque session identity you declare: a partition hint, not authentication and not a data filter. Selects which no-persist pause applies to this call. Omit to share one bucket with every caller that omits it. Full text on recall. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, idempotentHint=true, and readOnlyHint=false. The description adds meaningful context beyond these flags by specifying exactly what data domains get destroyed (memories, profiles, episodes) and that it is scoped to a single agent, clarifying the blast radius.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences with no filler. The core action and scope are front-loaded, and the usage context is given in the second sentence. Every word earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the simple two-parameter schema and annotations that already cover destructive/idempotent behavior, the description sufficiently explains the tool's purpose and scope. Minor omissions like return values or preconditions are not critical for this kind of cleanup operation. The note about being used by the kernel adds adequate context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Both parameters (agent_id and session_key) have detailed descriptions in the input schema, including the nuanced behavior of session_key as a partition hint rather than a data filter. The tool description itself adds no param information beyond what the schema provides, so a baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses the specific verb 'Delete' and names the exact resource: ALL data (memories, profiles, episodes) for a specific agent. It clearly distinguishes this from sibling memory-manipulation tools by emphasizing it purges everything for one agent during deletion.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly states when to use this tool: 'Used by kernel during agent deletion.' It does not provide explicit alternatives or when-not-to-use guidance, but the context of a cleanup/teardown operation is unambiguous enough for an agent to decide.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
delete_episodeADestructiveIdempotent
Delete a single episode by ID. Ownership is enforced when agent_id is provided.
| Name | Required | Description | Default |
|---|---|---|---|
| agent_id | No | Agent ID for ownership verification (injected by kernel) | |
| episode_id | Yes | Episode ID to delete | |
| session_key | No | Opaque session identity you declare: a partition hint, not authentication and not a data filter. Selects which no-persist pause applies to this call. Omit to share one bucket with every caller that omits it. Full text on recall. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already indicate destructive and non-read-only behavior, and idempotentHint=true is present. The description adds value beyond the annotations by revealing that ownership is enforced when agent_id is provided, which is key behavioral context for callers. It does not contradict the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, no filler, and the core action is front-loaded. Every word earns its place by clarifying scope or the ownership condition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple single-ID deletion tool, the description covers the core operation and the ownership nuance. The destructive nature is already communicated by annotations, and no output schema exists, so return-value details are not required. It could mention the alternative of archiving, but that is not necessary for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents episode_id, agent_id, and session_key. The description largely restates what the schema says: delete by ID and ownership enforcement for agent_id. It does not add meaningful parameter-level detail beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'Delete a single episode by ID.' It clearly scopes the operation to one episode, distinguishing it from broader deletion tools like delete_agent_data or archiving tools like archive_episode. The ownership enforcement sentence adds further precision about the operation's conditions.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies a clear usage context: use when you need to delete exactly one episode by its ID. It also notes an important condition around agent_id and ownership verification, but it does not explicitly name alternatives or explain when not to use this tool versus archive_episode or delete_agent_data.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
delete_memoryADestructiveIdempotent
Delete a single memory by ID. Ownership is enforced when agent_id is provided.
| Name | Required | Description | Default |
|---|---|---|---|
| agent_id | No | Agent ID for ownership verification (injected by kernel) | |
| memory_id | Yes | Memory ID to delete | |
| session_key | No | Opaque session identity you declare: a partition hint, not authentication and not a data filter. Selects which no-persist pause applies to this call. Omit to share one bucket with every caller that omits it. Full text on recall. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish destructive and read-only behavior, so the description adds value by explaining ownership enforcement when agent_id is provided. It does not contradict the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two concise, front-loaded sentences with no filler. The core action is stated first and the ownership caveat follows.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple single-item deletion tool, the description plus annotations and full schema coverage are largely sufficient. It could additionally clarify not_found behavior or relationship to bulk deletion tools, but these are minor gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so parameters are already fully documented. The description does not add meaningful parameter details beyond what the schema provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb (Delete), a specific resource (memory), and a precise scope (single memory by ID). This clearly distinguishes it from sibling tools such as delete_episode and delete_agent_data.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The use case is implied: use this when deleting one memory by its ID. However, it does not explicitly mention when_not_to use it or name alternatives like delete_agent_data for broader deletion, leaving some routing to the agent's inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
export_memoriesADestructive
Export memories, episodes, and profiles to a JSONL file for backup or portability.
| Name | Required | Description | Default |
|---|---|---|---|
| agent_id | Yes | Agent identifier (empty string to export all agents) | |
| output_path | Yes | File path for the JSONL output | |
| include_embeddings | No | Include embedding BLOBs as base64 (default false, usually not needed) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description describes an export operation, which is typically non-destructive. However, annotations set destructiveHint to true, implying the tool may have destructive side effects (e.g., file overwrite). The description does not disclose this, contradicting the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, straightforward sentence with no wasted words. It is front-loaded with the action and purpose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description mentions the data types exported (memories, episodes, profiles) and the format (JSONL), but does not address potential side effects like file overwriting despite the destructiveHint annotation. Given no output schema, more detail on the return value or behavior would improve completeness.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema covers all three parameters (agent_id, output_path, include_embeddings) with descriptions. The tool description does not add additional meaning beyond what the schema provides, so baseline score of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool exports memories, episodes, and profiles to a JSONL file for backup or portability. It uses a specific verb and resource, and distinguishes from siblings like import_memories.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides context for when to use the tool (backup or portability) but does not explicitly state when not to use it or mention alternatives. Sibling list makes the purpose clear enough.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_contentsARead-only
Fetch full, untrimmed content for recall preview refs. Use after a preview-tier recall to expand only the rows that matter instead of opting the whole recall out with full_content=true. Bounded twice: at most 20 refs per call, and a 40,000-character budget across the batch (2.5.4a2) that does not move when CPERSONA_MAX_CONTENT_LENGTH does. Rows are never cut to fit — when the budget is spent the remaining refs come back in deferred (absent otherwise) alongside budget_chars; re-fetch them in a second call. A single row larger than the budget is still returned in full, because this tool is the only path back to a row's complete text. RANGES: a ref may instead be an object that names part of its record -- {ref, node: i} or {ref, node: [first, last]} (inclusive) for overflow-tree nodes, e.g. the node.index of a reconstruct quote and its neighbours, or {ref, span: [start, end]} for characters. Offsets are in the stored text (a memory's content, an episode's summary without the '[Episode] ' label). The item then carries that slice as content and range = {span, content_len, and node + of when nodes were named}; a span end past the text is clamped and range.span says what was served. A range that cannot be served exactly is never widened to the whole row: it comes back in unresolved (absent otherwise) as {ref, reason}, reason one of invalid_range, no_current_nodes (the record has no complete node set -- short records have none, and a new long one gets them shortly after store), node_out_of_range, span_out_of_range. Only the slice counts against the budget.
| Name | Required | Description | Default |
|---|---|---|---|
| refs | Yes | Refs from recall messages ('mem:<id>' / 'ep:<id>'), or range objects such as {'ref': 'mem:<id>', 'node': [2, 3]} / {'ref': 'ep:<id>', 'span': [0, 800]} (max 20 per call) | |
| agent_id | Yes | Agent identifier (ownership check — another agent's refs come back in `missing`) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Although readOnlyHint=true is annotated, the description goes far beyond that, detailing the bounded nature (20 refs, 40k char budget), handling of over-budget refs via 'deferred', behavior for rows larger than budget, range serving with clamping, range failure modes ('unresolved' reasons), and budget accounting. This is rich behavioral context that annotations cannot convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long but every sentence contributes new information. It is front-loaded with the core purpose and usage, then details bounds, range behavior, and failure modes. It is dense but efficient, with no filler or repetition. Given the complexity, this length is justified.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has no output schema, so the description must explain return values, which it does thoroughly: 'deferred', 'budget_chars', 'content', 'range' with {span, content_len, node}, and 'unresolved' with reasons. It covers edge cases and provides example range objects. For a complex tool with two parameters and no output schema, this is complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the schema already provides descriptions for both parameters, so baseline is 3. However, the description adds significant semantic depth: explains the meaning of 'mem:' and 'ep:' prefixes, how range objects work, the exact structure of node and span, inclusive ranges, offsets in stored text, and the 'deferred' and 'unresolved' fields, which the schema doesn't mention. This is beyond baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool fetches full, untrimmed content for recall preview refs, distinguishing it from the recall tools that provide previews. It specifies the verb, resource, and the exact use case, and is unique among siblings (e.g., 'reconstruct', 'recall') as the only path to full row text.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides explicit context: 'Use after a preview-tier recall to expand only the rows that matter instead of opting the whole recall out with full_content=true.' It also clearly states the alternative (recall with full_content=true) and when not to use the tool, giving an agent clear decision criteria.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_operating_contextARead-only
Read the server-served operating context (v2.5.1): the operator-owned doctrine distributed to every connected client. Without arguments returns the preview tier — context_revision, instructions_summary, project_id registry (+ enforce mode), @auto defaults, and doctrine section names. Pass section to fetch one section's full body. Read-only: the context is edited by the operator on the filesystem (~/.cpersona/operating-context.toml), never via MCP.
| Name | Required | Description | Default |
|---|---|---|---|
| section | No | Doctrine section name to fetch in full (from doctrine_sections). Empty = preview tier. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description not only aligns with the readOnlyHint annotation but adds significant context: the context is edited on the filesystem (~/.cpersona/operating-context.toml), never via MCP. This discloses the source of truth and mutation path, which annotations alone do not provide.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three concise sentences, front-loaded with the core purpose, followed by mode details and behavioral note. No wasted words. Structure is logical: what, how, important note.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given one optional parameter, no output schema, and full annotations, the description covers all necessary aspects: return types (preview tier components, full section body), usage modes, and behavioral constraints (read-only, filesystem editing). It is sufficient for an AI agent to correctly invoke this tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% with a description for the 'section' parameter. The description adds value by explaining the default behavior (preview tier) and that the section is from 'doctrine_sections'. It clarifies the parameter's effect beyond the schema's own description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool reads the 'server-served operating context' with a specific version (v2.5.1). It identifies the resource and its nature as 'operator-owned doctrine'. This is specific and distinct from sibling tools like 'get_profile' or 'get_contents'.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly states the two modes: without arguments returns the preview tier, and with a 'section' argument returns the full body. It mentions read-only and that editing is done via filesystem, not MCP. While it doesn't contrast with siblings, it provides clear usage context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_profileCRead-only
Get the current profile for an agent.
| Name | Required | Description | Default |
|---|---|---|---|
| agent_id | Yes | Agent identifier |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Description aligns with readOnlyHint by stating 'Get', but adds no further behavioral details such as error handling for missing agents, return format, or scope of the profile.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Single sentence, concise and front-loaded. Could include more detail without being overly long, but the brevity aids readability.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Adequate for a simple 1-parameter tool with no output schema. However, lack of return value description may leave the agent uncertain about the output format.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% with a clear description for 'agent_id'. The tool description does not add any additional meaning beyond the schema, so baseline score applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description 'Get the current profile for an agent' clearly states the action (get) and resource (profile) with specifier 'for an agent'. It distinguishes from sibling 'update_profile' but is slightly redundant with the tool name.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this tool vs alternatives, no prerequisites or limitations mentioned. The only implicit guidance is from the readOnlyHint annotation.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_queue_statusARead-only
Get the status of the background task queue (pending tasks, retry config).
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, so the description's addition of 'pending tasks, retry config' provides some context. However, it could be more transparent about the return format or any other behavioral traits.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence of 13 words. Every word is purposeful and concise.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple, parameterless tool with readOnlyHint annotation, the description is fairly complete. It could benefit from specifying the output structure, but given no output schema, it's adequate.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has zero parameters, so baseline is 4. The description does not add parameter details, which is acceptable since there are none.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states a specific verb 'Get' and resource 'background task queue status', and mentions what is included (pending tasks, retry config). This distinguishes it from sibling tools like check_health or list_episodes.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for checking queue status but provides no explicit guidance on when to use it versus alternatives, nor when not to use it. No sibling comparisons mentioned.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_recall_precisionARead-only
Read an agent's effective recall precision (knob 3) — the read-back companion to set_recall_precision. Returns the resolved specificity weight (beta) and its named precision level (strict / balanced / lenient, or 'custom' for a raw beta), and flags whether the value is a per-agent override or the global CPERSONA_RECALL_PRECISION default (overridden + global_precision / global_beta). Read-only: it never recalibrates and never persists, so a UI can load the current setting, let the user edit it, and write it back instead of the control being write-only.
| Name | Required | Description | Default |
|---|---|---|---|
| agent_id | Yes | Agent whose precision to read |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true. The description goes further by stating it never recalibrates or persists, and details the returned fields (beta, precision level, override flags). This adds behavioral context beyond the annotation, though it doesn't cover all edge cases.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single paragraph that front-loads the main purpose and then provides additional details. It is reasonably concise, though some sentences could be tightened. Every sentence adds value.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the simple input (1 param, no output schema), the description thoroughly explains the return value and its relationship to the global default and override behavior. It is complete for a read-only tool of this complexity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema covers the single required parameter (agent_id) with a description. The tool description does not add meaning beyond that, but since schema coverage is 100%, a baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool reads an agent's effective recall precision, identifies it as the read-back companion to set_recall_precision, and specifies it is read-only. This distinguishes it from its sibling and provides a specific verb-resource combination.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explains when to use this tool (as a read-back companion to set_recall_precision, for UI loading before editing) and implies it should be used before writing. However, it does not explicitly mention alternatives or when not to use it, though the sibling set tool is clearly the counterpart.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_session_findingsARead-only
Pull the storage-integrity findings on demand (SuperAuditor v1 pull contract, docs/SUPERAUDITOR_STANDARD.md) instead of reading them off check_health. Same detector as check_health(fix=false) over the WHOLE database, delivered as findings: each carries kind (the finding's name: a check registry name, or an escalation tier this seam mints for a runner that grades its own severity, e.g. null_embedding_pipeline_down — a tier is NOT a registry name), check (the registry name that produced it, so check_health(checks=[finding['check']]) re-runs exactly that probe) and a static per-kind severity (critical = the read contract is broken now / warn = two stored facts contradict / info = an observation). check_health's own instance verdict rides along as health_severity; a probe that raised is reported as kind check_crashed instead of failing the pull, so a partial result says which probe is missing. Read-only, never repairs. NOT free, though: the registry runs unfiltered, which includes two whole-database reads (the FTS5 integrity-check over both indexes, and PRAGMA quick_check over the file), so every pull is O(database) on a channel meant to be pulled once a session — budget it by call frequency. There is deliberately no cheap subset: choosing which probes run would be choosing which forgotten state stays forgotten. Findings are NOT filtered by agent_id or project_id — the channel surfaces forgotten state, and slicing it by the caller's bucket would hide exactly the rows that were forgotten (scope a repair with check_health(agent_id=...)). Honest caps: findings holds at most per_kind_limit rows per kind, capped_kinds names every kind that had more (observed, not inferred from count == limit), total and the counts describe the RETURNED set only, and per_kind_limit echoes the limit applied. summary restates the same trimmed set in prose (pass include_summary=false to skip paying for it). On a shared remote transport with no session_key declared the response carries identity_shared: true — this server has no session-scoped probes, so the key is a partition hint, not a filter. _meta.server_version identifies the running instance.
| Name | Required | Description | Default |
|---|---|---|---|
| session_key | No | Opaque client-declared session identity (partition hint, not authentication). Empty on a non-stdio transport marks the response identity_shared. | |
| per_kind_limit | No | Maximum findings returned per kind (default 5, minimum 1). Kinds that hit it are listed in capped_kinds. | |
| include_summary | No | Include the human-readable `summary` rendering (default true). It restates `findings` in prose — set false when machine-reading. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description goes far beyond the readOnlyHint annotation by disclosing the O(database) cost, the unfiltered registry, the check_crashed degradation behavior, the cap semantics, and the identity_shared case. It also explicitly states it is read-only and never repairs, matching the readOnlyHint annotation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The text is long and dense, but each sentence carries meaningful caveats or field semantics and the purpose is front-loaded. It loses a point for being a single wall of text; light formatting or grouping would improve scannability without dropping content.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, yet the description covers the shape of findings (kind, check, severity, health_severity), the special check_crashed finding, the caps/totals semantics, identity_shared, and _meta.server_version. An agent has everything it needs to invoke and interpret the result correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Although schema coverage is 100%, the description adds essential semantics: session_key is a partition hint rather than a filter, cap-related fields describe the returned set only, and include_summary=false can skip the prose. These details materially change how an agent would set and interpret the parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific action ('Pull the storage-integrity findings on demand') and explicitly contrasts it with check_health, calling out the same detector over the whole database. This clearly differentiates the tool from its closest sibling and leaves no ambiguity about the resource it operates on.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It tells the agent when to use the tool ('instead of reading them off check_health'), when to prefer the sibling ('scope a repair with check_health(agent_id=...)'), and what to avoid ('NOT free', 'budget it by call frequency', 'deliberately no cheap subset'). The when/when-not guidance is explicit and concrete.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
import_memoriesADestructiveIdempotent
Import memories, episodes, and profiles from a JSONL file. Idempotent: memories deduplicate on msg_id (and on content within a project/channel), episodes on their summary within a project/channel.
| Name | Required | Description | Default |
|---|---|---|---|
| dry_run | No | Count records without writing to DB (preview mode) | |
| input_path | Yes | Path to the JSONL file to import | |
| session_key | No | Opaque session identity you declare: a partition hint, not authentication and not a data filter. Selects which no-persist pause applies to this call. Omit to share one bucket with every caller that omits it. Full text on recall. | |
| target_agent_id | No | Remap all records to this agent ID (empty to use original agent_id from file) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description adds useful behavioral detail beyond annotations by specifying deduplication keys (msg_id/content for memories, summary for episodes). However, destructiveHint is true and the description does not disclose what destructive effect may occur, such as overwriting existing records, which leaves an important behavioral gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences with no filler: the first states the operation and scope, and the second adds essential idempotency and deduplication semantics. Information is front-loaded and every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutating tool with destructiveHint, the description is reasonably complete but missing return-value behavior and clarification of the destructive/overwrite effect. The schema and annotations cover parameters and high-level safety, but the description does not fully bridge those gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description adds no parameter-specific meaning beyond saying the source is a JSONL file; it does not elaborate on dry_run, session_key, or target_agent_id, though the schema already captures those.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb (Import), the resource (a JSONL file), and the data kinds (memories, episodes, and profiles). This clearly differentiates it from siblings like export_memories, store, or merge_memories.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The tool's intended use as a bulk importer from a JSONL file is implied, and the idempotency note suggests it is safe for repeated imports. However, it does not explicitly say when to use this instead of alternatives such as store or merge_memories, nor does it state exclusions or prerequisites.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_episodesARead-only
List archived episodes for an agent (for dashboard display). bug-385: limit is clamped to 200 rows and the response carries no marker when the clamp bit, so a listing of exactly that many rows may be a truncated one rather than the end of the data — reach the rest through export_memories or a narrower filter, not a larger limit. bug-255: within that cap the response holds an 800,000-character budget across summary and keywords together, with the same degradation and ceiling semantics as list_memories — rows past the budget that exceed the preview cap carry pure prefixes plus summary_truncated/summary_len and keywords_truncated/keywords_len; budget_chars appears iff at least one row was degraded. Their ref expands the summary via get_contents (under the row's own agent_id); a full keywords string is only available through export_data.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Max episodes to return | |
| agent_id | No | Agent identifier (empty for all agents) | |
| project_id | No | v2.4.17 γ filter. Same semantics as list_memories. v2.5.1: pass '@auto' to resolve this agent's default from the server's operating context (the resolution is echoed as resolved_project_id; an unmapped agent yields operating_context_warning). bug-186: resolution requires a configured operating context. With none — the default, and equally the outcome of a sidecar that fails to parse — the sentinel is NOT resolved: it is stored and filtered as the literal project_id '@auto', resolved_project_id echoes '@auto', and no warning is raised. Read resolved_project_id before relying on the resolution. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations provide only readOnlyHint=true, so the description carries the burden of rich behavioral disclosure. It goes far beyond that: explains the 200-row clamp, the 800,000-character budget, degradation semantics, presence of summary_truncated/summary_len and keywords_truncated/keywords_len, and budget_chars. It also describes how `ref` expands via get_contents and that full keywords require export_data. This gives an agent a precise mental model of response behavior and edge cases.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but highly relevant. It starts with the core purpose and then layers in bug-specific behavioral details, each earning its place by preventing misinterpretation. The structure with bug-385 and bug-255 is logical but could be slightly tightened; still, the level of detail is justified given the complexity and the absence of an output schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description must explain return behavior, and it does thoroughly. It covers truncation markers, character budgets, degraded rows, resolution echoes, and fallback semantics. It also references list_memories semantics for consistency. For a tool with this complexity, the description leaves no significant gap in what an agent needs to know to call it correctly and interpret results.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, which would set a baseline of 3, but the description adds substantial meaning, especially for project_id. It explains the gamma filter, '@auto' resolution, resolved_project_id, operating_context_warning, and the bug-186 sentinel fallback behavior. It also clarifies that limit is clamped to 200 rows, extending the schema's 'Max episodes to return.' These additions are critical for correct invocation and output interpretation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description begins with a clear, specific statement: 'List archived episodes for an agent (for dashboard display).' It names the verb (list), the resource (archived episodes), and the intended use (dashboard display), which cleanly distinguishes it from sibling tools like list_memories and archive_episode. The phrasing leaves no ambiguity about what the tool does.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit guidance on when to use alternatives: 'reach the rest through export_memories or a narrower filter, not a larger limit' when the clamp is hit, and 'a full keywords string is only available through export_data.' It also implies primary usage via 'for dashboard display.' However, it does not explicitly contrast with list_memories or state when not to use the tool beyond the truncation edge case, so some inference remains.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_memoriesARead-only
List recent memories for an agent (for dashboard display). bug-385: limit is clamped to 500 rows and the response carries no marker when the clamp bit, so a listing of exactly that many rows may be a truncated one rather than the end of the data — reach the rest through export_memories or a narrower filter, not a larger limit. bug-255: within that cap the response holds a 1,000,000-character content budget. Rows are returned newest-first and none is dropped by the budget; once it is spent, later rows LONGER than the preview cap (CPERSONA_RECALL_PREVIEW_CHARS, default 500) degrade to a pure prefix with content_truncated/content_len and a ref that get_contents expands under the row's own agent_id (in an all-agents listing, pair the ref with the row's agent_id field). budget_chars appears iff at least one row was degraded. The effective ceiling is the budget plus one whole row plus the degraded rows' prefixes, so it scales with the preview cap; preview cap 0 disables trimming and the budget with it.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Max memories to return | |
| agent_id | No | Agent identifier (empty for all agents) | |
| project_id | No | v2.4.17 γ filter. Omit → no filter; '' → global pool only; 'X' → 'X' ∪ global pool. v2.5.1: pass '@auto' to resolve this agent's default from the server's operating context (the resolution is echoed as resolved_project_id; an unmapped agent yields operating_context_warning). bug-186: resolution requires a configured operating context. With none — the default, and equally the outcome of a sidecar that fails to parse — the sentinel is NOT resolved: it is stored and filtered as the literal project_id '@auto', resolved_project_id echoes '@auto', and no warning is raised. Read resolved_project_id before relying on the resolution. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description is exceptionally transparent about edge-case behaviors: clamps at 500 rows, content budget of 1,000,000 characters, degradation to prefixes with truncation indicators, and the interplay with preview cap, including the case where preview cap is 0. It discloses what happens when the clamp is hit (no marker) and how to access full content via get_contents with proper agent_id pairing. This far exceeds what annotations (readOnlyHint) provide and adds critical behavioral nuances that affect tool invocation and result interpretation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense and technically precise, with each sentence carrying essential details about limits, budgets, and degradation. It is not front-loaded with a summary line, but the structure is logical. While lengthy, it earns its length due to the complexity of edge cases; a shorter version would lose critical information. Minor deduction for lack of a high-level overview at the start.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with no output schema, the description comprehensively covers what an agent needs to know: return order, potential truncation, how to detect truncation (budget_chars), how to access full content via get_contents, and how to handle all-agents listings. It addresses edge cases like preview cap 0 and unmapped agents, ensuring the agent can correctly interpret results and avoid misreading truncated data as complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Although the schema already covers 100% of parameters with descriptions (raising the baseline to 3), the description adds substantial value beyond the schema for the project_id parameter: it details the '@auto' sentinel behavior, error conditions, and resolved_project_id echoing. For limit, it adds the clamping behavior and budget context not mentioned in the schema. This extra context significantly enhances parameter understanding despite high schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool lists recent memories for dashboard display, which is specific and actionable. However, it does not explicitly differentiate from sibling tools like recall or export_memories beyond mentioning export_memories as an alternative, so it's clear but not fully distinguishing.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides explicit guidance on when to use this tool (for dashboard display) and when to use alternatives (export_memories for full data access). It mentions using narrower filters instead of larger limits, which guides behavior, but it doesn't explicitly say when NOT to use this tool beyond those points.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
lock_memoryAIdempotent
Lock a memory to prevent deletion and editing. Ownership enforced when agent_id provided.
| Name | Required | Description | Default |
|---|---|---|---|
| agent_id | No | Agent ID for ownership verification | |
| memory_id | Yes | Memory ID to lock | |
| session_key | No | Opaque session identity you declare: a partition hint, not authentication and not a data filter. Selects which no-persist pause applies to this call. Omit to share one bucket with every caller that omits it. Full text on recall. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description adds behavioral context beyond the annotations by explaining that locking prevents deletion and editing. It also clarifies the conditional ownership enforcement tied to agent_id, which is not captured in the readOnlyHint, idempotentHint, or destructiveHint flags. No contradiction with annotations exists.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two short sentences with no filler. It front-loads the primary purpose and immediately follows with the key ownership condition, making it efficient and easy to parse.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple locking tool with a single required parameter, the description covers the core behavior and the important ownership nuance. No output schema exists, but little is needed because the tool's effect is straightforwardly described.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, and the description adds value by clarifying that ownership is only enforced when agent_id is provided. This is a meaningful semantic qualifier beyond the schema's generic 'Agent ID for ownership verification' phrasing.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Lock') and resource ('a memory') and clearly states the intended effect: preventing deletion and editing. It is distinguishable from sibling tools like unlock_memory and delete_memory, so there is no ambiguity about what this tool does.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies the tool is used when a memory must be protected from deletion or editing, but it does not explicitly state when to choose this over alternatives or when not to use it. The ownership note ('Ownership enforced when agent_id provided') gives conditional usage context, but no direct comparison to siblings is provided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
merge_memoriesADestructiveIdempotent
Merge memories, episodes, and profiles from one agent into another. Atomic one-shot equivalent of export→import without intermediate files. Strategy 'skip' deduplicates by msg_id (memories) and summary (episodes).
| Name | Required | Description | Default |
|---|---|---|---|
| mode | No | Merge mode: 'copy' (preserve source) or 'move' (delete source after merge) | copy |
| dry_run | No | Preview merge without writing to DB | |
| strategy | No | Merge strategy: 'skip' (default) — skip duplicates, keep target's version | skip |
| session_key | No | Opaque session identity you declare: a partition hint, not authentication and not a data filter. Selects which no-persist pause applies to this call. Omit to share one bucket with every caller that omits it. Full text on recall. | |
| source_agent_id | Yes | Agent ID to merge FROM | |
| target_agent_id | Yes | Agent ID to merge INTO |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already mark the tool as non-read-only, idempotent, and destructive. The description adds useful behavioral context beyond those flags: atomicity, lack of intermediate files, and the deduplication behavior for strategy 'skip' by msg_id and summary. It does not contradict the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three focused sentences, each adding distinct value: the operation and scope, the atomic/one-shot nature, and the deduplication rule. No fluff or repetition of schema content.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the 6-parameter schema and destructive annotations, the description plus schema provide enough for an agent to select and invoke the tool correctly. It could go slightly further by describing what the call returns, especially since there is no output schema, but that is a minor gap for a mutating merge operation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema fully documents all six parameters. The description adds minor value by detailing how the 'skip' strategy deduplicates, but it does not need to compensate for missing parameter documentation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific operation ('Merge memories, episodes, and profiles from one agent into another') and the exact resource scope. The phrase 'Atomic one-shot equivalent of export→import without intermediate files' strongly distinguishes it from sibling tools like export_memories and import_memories.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly implies when to use this tool: when a single atomic merge between agents is desired, avoiding the multi-step export→import flow. It does not explicitly give exclusion criteria or name alternatives, but the contrast with export/import is enough to guide selection among related siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
migrate_channel_axisA
Re-channel bridge-type memories to their concrete channel (knob2 v2 default flip prep). Memories the kernel filed under the bridge type ('discord') are rewritten to the concrete channel recovered from the stored session_id ('{channel_id}:{user_id}:{chunk}' | '{channel_id}:shared' → channel_id), so per-channel recall can match them. Non-destructive (only the channel column changes) and idempotent (re-running is a no-op once moved). dry_run=true (default) reports the recoverable count, the channels that would be recovered, and an unrecoverable bucket (channel='discord' rows with no snowflake session_id) without mutating. globalize_unrecoverable=true moves the unrecoverable bucket to channel='' (global, matched by every channel-scoped recall) so the flip orphans nothing; default false (report only).
| Name | Required | Description | Default |
|---|---|---|---|
| dry_run | No | Preview counts only, no mutation (default: true) | |
| agent_id | No | Agent ID to migrate (empty = all agents) | |
| session_key | No | Opaque session identity you declare: a partition hint, not authentication and not a data filter. Selects which no-persist pause applies to this call. Omit to share one bucket with every caller that omits it. Full text on recall. | |
| globalize_unrecoverable | No | Also move channel='discord' rows with no snowflake session_id to channel='' (global). Default false. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only provide readOnlyHint=false, which is minimal. The description compensates fully by disclosing that the operation is non-destructive (only the channel column changes), idempotent (re-running is a no-op), that dry_run avoids mutation, and that globalize_unrecoverable moves unmatched rows to the global channel. It also explains the unrecoverable bucket behavior. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense and long, but every clause contributes necessary operational detail: the transformation rule, the session_id format, non-destructiveness, idempotency, dry_run behavior, and the globalize option. It is front-loaded with the core purpose. A more structured list might improve readability, but the content is warranted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a migration tool with no output schema, the description covers all essential context: what gets rewritten, the exact session_id format, default behavior, the unrecoverable case, and the optional globalize behavior. An agent can correctly decide whether to call this tool and how to set the parameters without additional information.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description adds value beyond the schema by explaining what dry_run=true reports (recoverable count, channels, unrecoverable bucket) and what globalize_unrecoverable does to the unrecoverable bucket. This goes beyond the simple boolean descriptions in the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action: 'Re-channel bridge-type memories to their concrete channel', naming the resource (bridge-type memories) and the transformation (rewriting channel from session_id). It clearly distinguishes this migration utility from the sibling memory tools, which operate on individual memories or recall rather than performing a schema migration.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear operational guidance: dry_run=true is the default and reports without mutating, while globalize_unrecoverable=true is an opt-in escalation. It also clarifies idempotency and non-destructive behavior. It does not explicitly name alternative tools or exclusion conditions, but there is no obvious sibling alternative for this migration, and the flag semantics effectively tell the agent when to use which mode.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
pause_persistenceAIdempotent
Pause write operations on this MCP server for an opt-in TTL window. While paused, every write tool — store, declare_associations, archive_episode, update_memory, delete_memory, delete_episode, delete_agent_data, lock_memory, unlock_memory, update_profile, import_memories, merge_memories, calibrate_threshold, set_recall_precision — returns a no-op response carrying persisted: false, dry_run: true and a reason (with the TTL remaining) instead of writing to the database. persisted: false is the authoritative signal: branch on it, not on an id. Where the success shape has an id, it reads "no-persist" (store, archive_episode); action-specific id keys (deleted_id / updated_id / locked_id / unlocked_id / episode_id) are blanked to null so a truthy echo cannot read as success. migrate_channel_axis is gated differently — it is forced to dry_run and reports repairs_skipped rather than returning a skipped-response, so it carries no persisted key. check_health and deep_check are not blocked but downgrade to fix=false (they answer with repairs_skipped: true). Read tools (recall, list_*, get_profile, etc.) still answer normally, except that recall suppresses its recall_count / last_recalled_at bump — a write that would otherwise move ranking state during a paused session. Blast radius follows session_key (response scope). Pass the same session_key here and on your write calls and the pause covers that key alone (scope: "session"): a session that sends a different key is neither silenced by it nor able to clear it. The key is a partition hint, not a credential — it is compared, never verified — so anyone who sends the same string shares the pause. Omit it and you arm the bucket every keyless caller shares (scope: "process") — on a streamable-HTTP deployment a single process serves every connected client, so a keyless pause silences writes for every other keyless session until resume or TTL elapse, and those sessions get no signal. Under stdio (one process per client) that bucket is the session. This affects only this MCP server (cpersona); call cscheduler's pause_persistence too if you want both paused. Use for benchmarking, AB testing, or ephemeral exploration where memory contamination must be avoided. Default TTL: 1800 seconds (30 minutes); upper bound: 86400 seconds (1 day).
| Name | Required | Description | Default |
|---|---|---|---|
| session_key | No | Opaque session identity you declare: a partition hint, not authentication and not a data filter. Selects which no-persist pause applies to this call. Omit to share one bucket with every caller that omits it. Full text on recall. | |
| ttl_seconds | No | TTL until automatic resume. Min 1, max 86400 (clamped). Default 1800. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Despite annotations already declaring idempotentHint=true, the description goes far beyond: it discloses the exact no-op response shape, the authoritative `persisted:false` signal, the special gating of migrate_channel_axis and check_health/deep_check, the recall side effect, and the full session_key partition semantics including the lack of verification and process-wide scope on streamable-HTTP. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long but densely packed and logically ordered: core effect, authoritative signal, special cases, session semantics, usage. Every sentence adds necessary context. It is front-loaded with the main purpose. The length is justified by the complexity, but it is slightly verbose and might benefit from tighter phrasing.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the full burden of explaining return values, and it does so comprehensively: the no-op response fields, the id blanking, the exceptions, the read-tool behavior, the session_key scope, and the TTL. For a tool of this complexity (many affected tools, deployment nuances), nothing an agent needs to call it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3, but the description adds substantial meaning beyond the schema. It explains session_key is a partition hint, not a credential, and details the consequences of omission (shared bucket, process-wide silencing). It also clarifies TTL bounds and the default. The schema gives the facts; the description gives the operational implications.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a precise verb ('Pause write operations') and resource (this MCP server), enumerates every affected tool, and contrasts with cscheduler's separate pause_persistence. It is unmistakable what this tool does and how it differs from siblings like resume_persistence and persistence_status.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly gives use cases (benchmarking, AB testing, ephemeral exploration) and tells when to call the sibling (cscheduler's pause_persistence) for broader coverage. The session_key semantics include when to omit it and the deployment-specific blast radius, providing thorough guidance on selecting and invoking the tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
persistence_statusARead-only
Report whether persistence is currently paused and the TTL remaining (in seconds). It reports the bucket session_key selects (response scope), not the server as a whole: with a session_key it answers for your session only, so paused: false here does not mean no other session is paused. Without one it reflects the bucket every keyless caller shares, which on a streamable-HTTP deployment means paused: true may have been armed by a different keyless session.
| Name | Required | Description | Default |
|---|---|---|---|
| session_key | No | Opaque session identity you declare: a partition hint, not authentication and not a data filter. Selects which no-persist pause applies to this call. Omit to share one bucket with every caller that omits it. Full text on recall. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare `readOnlyHint: true`, lowering the burden. The description still adds valuable behavioral nuance: the status is scoped to a bucket selected by `session_key`, not the whole server, and in streamable-HTTP deployments a keyless caller may see `paused: true` caused by another session. This meaningfully goes beyond the annotation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Although the description is longer than typical, every sentence earns its place: the main question is front-loaded, and the subsequent sentences clarify non-obvious scope semantics. No filler or redundant restatement of the tool name or schema exists.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema present, the description supplies the needed outcome expectations: paused state, TTL remaining in seconds, and scope. It also covers the one optional parameter and the tricky multi-session semantics. Nothing needed to call and interpret this tool correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema already documents `session_key` as an opaque partition hint with 100% coverage, so baseline is 3. The description adds extra meaning by connecting the parameter to response scope, explaining that omitting it shares a bucket with other keyless callers, and clarifying that it is not authentication or a data filter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Report whether persistence is currently paused and the TTL remaining (in seconds).' It clearly distinguishes itself from broader server-wide status by emphasizing the returned `scope` is tied to a session key, so it is not ambiguous against sibling status or mutating tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear usage context: it explains what happens when `session_key` is provided vs omitted, and warns that `paused: false` does not mean no other session is paused. It stops short of explicitly naming alternative tools or saying 'use this when X, not when Y,' but the inclusion criteria are strongly implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
recallARead-only
Recall relevant memories using multi-strategy search (vector + FTS5 + keyword). Message content is returned as a preview tier by default — expand selected rows with get_contents(refs), or opt out wholesale with full_content=true. full_content is itself budgeted (200k chars per response, bug-211): rows past the budget degrade to the preview tier and the response carries full_content_budget_chars (absent when the budget never bites). v2.5.2 additive: each scored message carries match_reason={signal, score, ...} where signal is the branch the ranking / quality gate keyed on (confidence > rsf > cosine > rrf) and the remaining keys (cosine / rrf / rsf) surface the internal per-retriever contributions present on that row. Unscored rows (cascade FTS/keyword) omit match_reason. A response carrying gate_fallback=true (absent otherwise) means every candidate fell below the quality gate and the below-gate lexical matches were returned instead of an empty result — treat them as low-confidence.
| Name | Required | Description | Default |
|---|---|---|---|
| deep | No | Deep recall — halves the quality gate (and the calibrated fused gate), so weaker matches are admitted. It also disables time and completion decay, which are inert unless CPERSONA_CONFIDENCE_ENABLED=true, and it does NOT widen the scan window (CPERSONA_MAX_MEMORIES) — deep is about how weak a match may be, not how far back the search reaches. | |
| limit | No | Per-retriever search depth, not a pure response cap: the value is handed to each retrieval channel (vector / episode FTS / keyword) as its top-K, so lowering it shrinks the candidate pool itself — rows beyond the depth are unreachable at any gate value, and score normalization / autocut operate on the smaller pool, which can also reorder what remains. Fewer rows than this may be returned. (Agent-facing cap; the library layer accepts up to the scan window for direct callers.) | |
| query | Yes | Search query (empty returns recent memories) | |
| channel | No | Filter memories by channel (e.g. 'chat', 'discord'). Default: '' (all channels). | |
| agent_id | Yes | Agent identifier | |
| source_id | No | v2.4.20 per-user source filter. Empty (default) = no filter. Non-empty = prefix match against json_extract(source, '$.id'), e.g. 'discord:12345' to restrict to one Discord user, or 'discord:' to scope to all Discord-sourced memories. Episodes carry no per-user source tagging, so they are skipped when this is set — UNLESS channel is also set, which scopes episodes to one conversation and re-admits them. | |
| project_id | No | v2.4.17 γ filter. Omit → no filter (all projects). '' → global pool only. 'X' → 'X' bucket ∪ global pool. Threaded through cascade / RRF / vector / FTS / keyword paths. v2.5.1: pass '@auto' to resolve this agent's default from the server's operating context (the resolution is echoed as resolved_project_id; an unmapped agent yields operating_context_warning). bug-186: resolution requires a configured operating context. With none — the default, and equally the outcome of a sidecar that fails to parse — the sentinel is NOT resolved: it is stored and filtered as the literal project_id '@auto', resolved_project_id echoes '@auto', and no warning is raised. Read resolved_project_id before relying on the resolution. | |
| session_key | No | Opaque session identity you declare — a partition hint, NOT authentication. It scopes this process's per-session state: the degraded-recall advisory's "already told you" memory, and which no-persist pause applies to this call. It does NOT filter stored data (use agent_id / project_id / channel for that), and it never reaches the database. Omit it to share one bucket with every other caller that omits it, which is the behaviour that predates this parameter. | |
| full_content | No | v2.5.0 preview tier opt-out. By default message content longer than the preview cap (CPERSONA_RECALL_PREVIEW_CHARS, default 500) is returned as a pure prefix with content_truncated/content_len markers; each message's `ref` expands via get_contents. true returns full text. | |
| exclude_contents | No | Normalized content strings to exclude from results (starts-with match). Used to prevent duplication with conversation context already known to the caller. bug-399: the starts-with rule holds only at or above 32 characters. A shorter entry has to EQUAL the stored content (after the normalization this parameter already asks for: stripped and lower-cased) — a short prefix would otherwise suppress every memory beginning with it, inside the retrievers and with nothing in the response reporting the exclusion. Size entries at or above that length when you mean a prefix. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description richly discloses output behaviors beyond the readOnlyHint annotation: the preview tier, full_content budget and degradation, match_reason structure with signal branches, omission on unscored rows, and the gate_fallback fallback. This gives the agent a detailed model of what to expect in the response, exceeding the bare annotation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single dense paragraph but is front-loaded with the core purpose and each sentence adds necessary behavioral or alternative-action detail. It is longer than ideal and lacks bullet or section structure, yet it contains no fluff or repetition; every sentence covers a distinct facet of behavior.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema present, the description carries the full burden of explaining return behavior. It explains the preview tier, budget marker, match_reason fields, unscored row behavior, and gate_fallback semantics. It also references the operating context via project_id auto-resolution and excludes_contents edge cases. For a 10-parameter tool with complex behavior, this is a complete and thorough description.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, with each parameter already having a detailed description. The tool-level description adds extra context for full_content by specifying the 200k char budget (bug-211) and the degradation behavior, which is not present in the schema. For other parameters, the description does not repeat schema content, but the budget detail lifts it above the baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with 'Recall relevant memories using multi-strategy search (vector + FTS5 + keyword)', giving a specific verb and resource. It also references get_contents explicitly for expanding previews, distinguishing the tool's primary function from content expansion. The multi-strategy detail adds useful specificity that separates it from a plain 'list memories' utility.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description offers clear guidance on content handling: use get_contents(refs) for full content or full_content=true to opt out of previews. It also explains the budget behavior and when gate_fallback and match_reason appear. However, it does not explicitly compare recall with sibling recall_with_context or list_memories, so exclusions are mostly implied rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
recall_with_contextARead-only
Recall memories and merge with external conversation context. Automatically deduplicates, sorts chronologically, and returns a unified list. Replaces separate recall + manual merge in the caller. Content is preview-tiered by default — see recall's full_content / get_contents (full_content shares recall's 200k-char response budget, bug-211). Every external_context entry's content filters the recall (the caller already holds that text), but only role=user / role=assistant entries are merged into messages. When entries of other roles are present the response carries context_filter_only={roles:[...]} — those entries filtered the recall without appearing in the output, whether or not they dropped a memory this time. context_field_issues={entries:[{index, fields}]} (absent otherwise) names entries whose declared field was not a string: it was read as absent and the entry merged without it. CPERSONA_EXTERNAL_CONTEXT_MODE=reject refuses such a call instead. gate_fallback=true (absent otherwise) is forwarded from the underlying recall: every candidate fell below the quality gate and the below-gate lexical matches were returned instead of an empty result — treat them as low-confidence.
| Name | Required | Description | Default |
|---|---|---|---|
| deep | No | Deep recall — same semantics as in `recall`: halves the quality gate so weaker matches are admitted. | |
| limit | No | Per-retriever search depth for the underlying recall, not a pure response cap — same semantics as recall's limit: lowering it shrinks the candidate pool itself, not just the rows returned. (Agent-facing cap; the library layer accepts up to the scan window for direct callers.) | |
| query | Yes | Search query | |
| channel | No | Memory channel filter | |
| agent_id | Yes | Agent ID | |
| source_id | No | v2.4.20 per-user source filter — passed through to recall. Same semantics as in `recall`. | |
| project_id | No | v2.4.17 γ filter — passed through to recall. Same semantics as in `recall`. v2.5.1: pass '@auto' to resolve this agent's default from the server's operating context (the resolution is echoed as resolved_project_id; an unmapped agent yields operating_context_warning). bug-186: resolution requires a configured operating context. With none — the default, and equally the outcome of a sidecar that fails to parse — the sentinel is NOT resolved: it is stored and filtered as the literal project_id '@auto', resolved_project_id echoes '@auto', and no warning is raised. Read resolved_project_id before relying on the resolution. | |
| session_key | No | Opaque session identity you declare — a partition hint, NOT authentication. It scopes this process's per-session state: the degraded-recall advisory's "already told you" memory, and which no-persist pause applies to this call. It does NOT filter stored data (use agent_id / project_id / channel for that), and it never reaches the database. Omit it to share one bucket with every other caller that omits it, which is the behaviour that predates this parameter. | |
| full_content | No | v2.5.0 preview tier opt-out — same semantics as in `recall`. | |
| external_context | No | Conversation history entries [{role, content, name?, user_id?, timestamp?}, ...]. Every declared field is a string; one that is not is read as absent and reported in context_field_issues (CPERSONA_EXTERNAL_CONTEXT_MODE). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With readOnlyHint=true already provided, the description still adds substantial behavioral context: automatic deduplication and chronological sorting, preview-tiering by default, role-based filter/merge behavior, context_filter_only, context_field_issues, the CPERSONA_EXTERNAL_CONTEXT_MODE reject option, and gate_fallback semantics. This is far beyond what the annotation conveys.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long but dense, and every sentence earns its place given the tool's complexity. It front-loads the core purpose and key behaviors before diving into edge cases and response fields. It could be slightly tightened by moving some low-level bug references out, but nothing is redundant.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex 10-parameter tool with no output schema, the description is remarkably complete. It explains the unified return list, preview-tiering and response budget, the context_filter_only container, the context_field_issues diagnostic, gate_fallback low-confidence results, and project_id '@auto' resolution. An agent has enough information to invoke it correctly and interpret unusual outcomes.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description adds meaning beyond the schema, particularly for external_context entries: which roles are merged, how non-string fields are handled, how timestamps affect ordering, and how gate_fallback affects result confidence. It does not add new semantics for every parameter, but the added detail is genuinely useful.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Opens with a specific verb-plus-resource statement: 'Recall memories and merge with external conversation context.' It also differentiates from the sibling recall by stating it 'Replaces separate recall + manual merge in the caller.' An agent can clearly tell this tool apart from recall without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context: use this when you have external conversation context that should filter and merge with recalled memories, since it replaces a manual two-step process. It references recall and get_contents as related alternatives, but does not explicitly state the when-not case (e.g., 'if you have no external_context, call recall directly').
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
reconstructARead-only
Assemble recall ITEMS from the candidate rows a recall produces: units of memory, each traceable to the canonical rows that support it. Reconstruction means select, order and assign roles -- never compose. No model is called and nothing is summarised: content is a verbatim excerpt of the item's head claim, cut the way the recall preview tier cuts, and expandable through head_ref via get_contents. HEAD CLAIM: the most relevant row in the item; if newer versions of that record (same message id in the same stored project) are present, their latest version. When max_evidence cuts an item, the head is kept and the most relevant remaining rows fill the rest. Stored rows are never modified. COUNT IS A CEILING, NOT A FILL TARGET AND NOT A SEARCH DEPTH: base = forced ?? requested ?? server default, effective = min(base, maximum), and 0 <= returned <= effective. Every response states effective_count and returned_count. A RESPONSE SAYS MORE ONLY WHEN THE SERVER DID SOMETHING OTHER THAN WHAT WAS ASKED: requested_count + count_policy {source, clamped, reason} when the count was clamped or operator-forced; requested_budget + budget_policy when the budget was clamped, raised or forced; effective_budget + used_budget when the budget withheld an item or an excerpt; bounds when a bound dropped rows, was reached, or was lowered by the library ceiling; reconstruction.excluded_without_provenance when rows were excluded. A response without them was served as asked. trace=true returns the full audit every time. Fewer items than the window is a NORMAL result and carries shortfall_reason (no_relevant_evidence / below_quality_threshold / exhausted_candidates); a shortfall is never padded with duplicates, fragments, or a cluster split in two. BREADTH IS SEPARATE FROM COUNT: top_k (candidate depth), max_hops (relation hops) and max_evidence are declared independently and none is derived from count -- changing count alone does not move the candidate id set. WHAT THE RESPONSE ADMITS: bounds.omitted names a bound that DROPPED rows the tool held (max_evidence -- each cut item also counts them in claims_omitted -- or max_hops: a declared relation was left unfollowed); bounds.reached names a bound that was only MET (top_k: retrieval returned as many rows as it was allowed; max_evidence: an entity the walk reached is mentioned by more records than were read -- whether more lay beyond is not known). Both are absent when empty. quote_selection: lexical_only appears when no query embedding was available and nodes were ranked by shared trigrams alone; an item whose cut quote is merely the start of its record carries node_unavailable (no_nodes, or not_current when nodes exist but are partial or another model's). ABSENCE IS NOT A VERDICT: a response without these fields does not say its items suffice to answer, that the whole store was searched, or that the rows were checked for contradiction -- conflicts detects one narrow case only. BREADTH BEFORE DEPTH: budget bounds the characters of quoted text -- each item's content and its excerpts -- where count bounds how many items. The quoted text is one fixed sequence: every head in item order, then each item's most relevant remaining excerpt, then the next, and the response is its longest prefix that fits. An excerpt the budget cannot carry is omitted (counted in excerpts_omitted, absent when zero; its claim and ref stay); an item is dropped only when its head does not fit, with shortfall_reason budget_exhausted. Raising the budget alone never removes an item or an excerpt. When budget is omitted the default is the configured default or one quote per item of the window, whichever is more, so a count you name is not cut by a budget you did not set; a budget you do name is taken as given. QUOTES: content quotes the head claim and each excerpts[] entry quotes another retained claim, most relevant first; all are verbatim and cut as the preview tier cuts. A long record with overflow-tree nodes is quoted from the node that best matches the query (rank by embedding similarity and by shared character trigrams, fused), and node gives its index, node count and character span in the stored text; a record without nodes is quoted from its start. READ FURTHER IN STEPS, SMALLEST FIRST: a node quote is the start of a node several times its length, and an item whose quote was cut carries expand -- pass it to get_contents as it is to read the rest of that node. If that is not enough, read its neighbours with {ref, node: [index - 1, index + 1]}. Pass the bare ref, the whole record, only when the parts did not answer: a record can be tens of times a node. Nodes are read after items are chosen, so they never change which items come back or their order. No relevance score is returned. ITEM SHAPE (the same for every item): claims carries one entry per retained row, newest first, each with ref, as_of, why (the key that admitted the row; relation:<predicate> when a declared relation did), hops when the relation walk reached the row, and roles when it has any -- sort by as_of for a chronological view. excerpts and excerpts_omitted are absent when empty. trace=true adds reconstruction (policy, candidate / cluster / selected counts), candidate refs, clusters and, for each record quoted by node, node_order -- its best few node indices, best first, as places to read next (an order, not a confidence). Gate fallback remains visible even when count is filled; zero count states count_zero. Retrieval degradation and update notices are delivered unchanged. If the library ceiling clamps top_k, bounds.effective_top_k reports the applied bound, including when the candidate pool is empty. independence_reason says why this is a separate item; conflicts appears only when two rows cannot be ordered. ROLE DIRECTION: roles[].role names what the REFERENCED row is to this claim (the ref is the subject, the claim is the object): the referenced episode SUPPORTS this claim, the referenced newer row SUPERSEDES it. The vocabulary is fixed at supports / supersedes / corrects / qualifies / contradicts / temporal_predecessor. The server derives supersedes (same message id, time order) and supports (episode span containment); any role word can also be DECLARED as a record -> record relation (declare_associations), whose subject is the ref. Ignore a role you do not know. BUNDLING KEYS are deterministic and never semantic: same message id within the same stored project (unknown project context cannot establish identity), containment in a candidate episode's time span, and adjacent timestamps FROM THE SAME SOURCE within the same project and channel, with the entire burst bounded by the time window (source alone is not a key -- in a single-agent store it is constant and would fold the whole pool into one item), and a declared record -> record relation between two candidates. Sharing a declared entity does not bundle. DECLARED ASSOCIATIONS (declare_associations, or associations on store) are read here and nowhere else, and change nothing when none apply: the names and aliases of entities the query mentions are added to the LEXICAL search only (the query's meaning, and so the vector search, is unchanged; the extra match is a vote, not a pass through the quality gate); and from each item's candidates the relation walk follows declared entity -> entity relations, either direction, up to max_hops, adding records that mention an entity it reached as evidence inside that item -- never as an item, never twice in one response, kept fewest hops first, then most recently declared relation, then lowest record id. This tool is additive: the recall contract is untouched.
| Name | Required | Description | Default |
|---|---|---|---|
| deep | No | Deep recall for the candidate stage -- same semantics as in `recall`. | |
| count | No | Ceiling on recall items returned -- not a fill target, not a search depth. Declare per call; omit to take the server default (1 unless configured). An operator-forced value overrides both. Requesting 5 with only 2 valid items returns 2; neither setting requires filling the window. Clamped to the server maximum, and the clamp is reported in count_policy rather than applied silently. | |
| query | Yes | Search query (empty returns recent memories) | |
| top_k | No | Candidate depth: how many rows the retrieval hands to bundling. This is the breadth knob; it is independent of `count` and is what to raise when items are missing evidence. | |
| trace | No | Include candidate refs and cluster membership for local diagnosis; no full text is added. | |
| budget | No | Payload budget: characters of quoted text (item `content` plus `excerpts`) the response may carry. Bounds depth, where `count` bounds breadth, and breadth wins: excerpts are omitted before any item is. Omit for the server default, which is never less than one quote per item of the window; an operator-forced value overrides both; clamped to the server maximum and raised to one preview-tier excerpt, and budget_policy says which. | |
| channel | No | Memory channel filter | |
| agent_id | Yes | Agent identifier | |
| max_hops | No | Relation hops the walk may follow from an item's candidates through declared entity -> entity relations. 0 adds no walked evidence. A relation left unfollowed at the bound is named in bounds.omitted. | |
| source_id | No | Per-user source filter -- same semantics as in `recall`. | |
| project_id | No | γ filter -- same semantics as in `recall`, including the '@auto' sentinel, which resolves this agent's default from the server's operating context and echoes the resolution as resolved_project_id. With no configured operating context the sentinel is NOT resolved: it is filtered as the literal project_id '@auto'. Read resolved_project_id before relying on the resolution. | |
| session_key | No | Opaque session identity you declare: a partition hint, not authentication and not a data filter. Forwarded to the candidate recall. | |
| max_evidence | No | Maximum retained rows per item, bounding its claims and role targets. A cut is named in bounds.omitted and counted in the item's claims_omitted. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description is consistent with the readOnlyHint=true annotation — "Stored rows are never modified" — so there is no contradiction, and it goes far beyond the annotation. It discloses the non-obvious traits the annotation can't: count is a ceiling not a fill target, ``0 <= returned <= effective`` with explicit effective_count/returned_count reporting, the absence-is-not-a-verdict caveat, bounds.omitted versus bounds.reached semantics, and the guarantee that no relevance score is returned. This is exceptional disclosure of behavioral nuance the annotation alone cannot convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is extremely long — effectively a full API specification covering policy reporting, bounds semantics, quote sequencing, role direction, bundling keys, and declared associations. It is well organized into CAPITALIZED sections and front-loads the core purpose, but it is far beyond what an agent should hold in context; much of the response-contract detail (item shape, role vocabulary, bounds semantics) belongs in an output schema or separate docs. There is also redundancy with the schema (top_k as the knob to raise for missing evidence appears in both). Density means real conciseness cost.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given 13 parameters, no output schema, and a genuinely complex response contract (bounds, count/budget policies, shortfall_reason, item shape, role vocabulary, trace additions), the description carries the full burden of documenting returns — and it does so exhaustively. It specifies the item shape, the fixed role vocabulary (supports / supersedes / corrects / qualifies / contradicts / temporal_predecessor), what fields appear only under error or clamp conditions, and what trace=true adds. Nothing an agent needs to correctly call and interpret this tool is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, placing the baseline at 3, and the description adds real meaning beyond the schema: it frames count as "a ceiling, not a fill target, not a search depth," clarifies that budget bounds characters where count bounds items ("breadth wins"), and names top_k as "the breadth knob... what to raise when items are missing evidence" — the same key the schema flags. The description enriches parameter intent beyond the schema's bare minimums, justifying a 4 rather than the baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening sentence states a specific verb and resource: "Assemble recall ITEMS from the candidate rows a recall produces: units of memory, each traceable to the canonical rows that support it." The core distinction is nailed precisely with "select, order and assign roles -- never compose," separating it cleanly from recall (which produces candidate rows) and from summarization tools. The tool's identity is unambiguous and differentiated from siblings like recall_with_context and get_contents.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is rich internal guidance — "READ FURTHER IN STEPS, SMALLEST FIRST" tells the agent to pass node quotes to get_contents before falling back to the bare ref, and the breadth-vs-depth discussion explains when to raise top_k versus count. However, explicit tool-selection routing among siblings is absent: the description never states when to choose reconstruct over recall or recall_with_context, and the contrast with get_contents is implied via the expand parameter rather than stated as a decision rule. The when-to-use-the-tool guidance is present but implicit in the behavioral detail.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
resume_persistenceAIdempotent
Re-enable persistence immediately, clearing this caller's active no-persist TTL. Returns was_active=true if THIS bucket was paused before the call. It clears only the bucket session_key selects (response scope): with a session_key, your own pause and no other session's; without one, the shared keyless bucket, which on a streamable-HTTP deployment re-enables writes for every other keyless session too.
| Name | Required | Description | Default |
|---|---|---|---|
| session_key | No | Opaque session identity you declare: a partition hint, not authentication and not a data filter. Selects which no-persist pause applies to this call. Omit to share one bucket with every caller that omits it. Full text on recall. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description goes well beyond the annotations by disclosing the non-obvious shared-bucket side effect: without a session_key, resuming can 're-enable writes for every other keyless session too.' It also explains what `was_active` means and that only the selected bucket is cleared. This is exactly the kind of behavioral nuance annotations do not capture, and there is no contradiction with the readOnlyHint/idempotentHint/destructiveHint flags.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded: the first sentence states the action and return value, and the second sentence details the critical scope side effect. No sentence is wasted, and the bolded scope warning earns its place because incorrect usage could affect other sessions.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter tool with no output schema, the description covers the essential agent decision points: what operation occurs, which bucket is affected, what happens when session_key is omitted, and what return value to expect. The behavior is fully specified enough for correct invocation, especially with the schema's additional session_key clarification.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema already documents session_key well, including that it is a partition hint rather than authentication. The description adds value by explaining the real-world consequence of each choice: with a session_key, only your own pause is cleared; without one, the shared keyless bucket is affected. This goes beyond the schema's neutral wording and helps the agent choose correctly.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action: 'Re-enable persistence immediately, clearing this caller's active no-persist TTL.' It clearly identifies the resource (the caller's no-persist pause bucket) and distinguishes this tool from siblings like pause_persistence and persistence_status by its reversal semantics. It also names the return value, giving the agent a concrete outcome.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description makes the usage context clear: this tool undoes a pause and works on either the session-specific bucket or the shared keyless bucket depending on session_key. It explains the behavioral difference between providing and omitting session_key, but it does not explicitly state 'use this instead of X' or list conditions when not to use it. This is clear context without formal exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
set_recall_precisionAIdempotent
Set an agent's recall precision (knob 3) and recalibrate its quality gate. precision = strict | balanced | lenient maps to a specificity weight beta of 2.0 / 1.0 / 0.5 in the gate separation objective (sensitivity + beta*specificity): strict sits the gate higher (fewer contaminants, more misses), lenient lower (fewer misses, more contaminants). A raw beta > 0 overrides the named level; an empty precision with beta <= 0 clears the per-agent override and returns the agent to the global CPERSONA_RECALL_PRECISION default. The gate is recalibrated at the new beta immediately and persisted, so the change is live without a restart. Precision is a per-agent setting, not a per-recall argument: the gate threshold is precomputed on the separation curve at a fixed beta, so this tool recalibrates once instead.
| Name | Required | Description | Default |
|---|---|---|---|
| beta | No | Raw specificity weight; overrides the named precision when > 0. | |
| agent_id | Yes | Agent whose precision to set | |
| precision | No | strict / balanced / lenient. Empty (with beta <= 0) clears the override. | |
| session_key | No | Opaque session identity you declare: a partition hint, not authentication and not a data filter. Selects which no-persist pause applies to this call. Omit to share one bucket with every caller that omits it. Full text on recall. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the annotations (readOnly=false, idempotent=true, destructive=false), it explains that the gate is recalibrated immediately, the change is persisted, it is live without restart, and how named levels, raw beta, and empty values resolve. No contradiction with the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Dense but organized: purpose first, then parameter semantics, then scope. It is longer than average, but every sentence carries behavioral information; minor jargon like 'knob 3' and 'separation curve' keeps it from being perfectly crisp.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Complete for a mutation tool of this complexity: it covers when to use it, all parameter interactions, reset behavior, persistence, and live effect. No output schema exists, but a setter's return value is not needed to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema already covers all four parameters (100% coverage), so the baseline is 3. The description adds real value by defining the precision-to-beta mapping (strict/balanced/lenient -> 2.0/1.0/0.5), beta override precedence, and clearing behavior, which go beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb+resource ('Set an agent's recall precision ... and recalibrate its quality gate') and later adds the distinguishing scope ('per-agent setting, not a per-recall argument'). This separates it clearly from recall-scoped siblings and from get_recall_precision.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives a clear exclusion — this is for per-agent settings, not per-recall arguments — and explains when the empty+beta<=0 form clears the override and returns to default. It doesn't explicitly name sibling tools like calibrate_threshold or get_recall_precision, but the context is enough for typical routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
storeAIdempotent
Store a message in agent memory for future recall. Every response carries result — the one field to branch on: 'stored' (a new row was written; {ok:true, result:'stored', id:, embedded:}, embedded true iff a local blob was persisted or the remote index push succeeded — false under EMBEDDING_MODE=none; the response also carries truncated:true when content exceeded the length cap and was shortened, and nodes:{status:'queued'} when the text runs past the embedding window and its overflow-tree nodes were queued for construction — absent when it fits, when the embedding server cannot report tokens, or with the task queue disabled), 'skipped' (nothing written and nothing wrong: {ok:true, result:'skipped', reason:...}; the msg_id / content dedup branches echo the pre-existing row's id, the OR IGNORE fallback reason='duplicate (unique index)' omits id by design — TOCTOU seam), or 'rejected' (nothing written because the request was refused: {ok:false, result:'rejected', reason:...} — empty content, content that sanitizes to empty, or an operating-context project_id refusal, which also carries error). Note for pre-2.5.2b1 callers: ok is no longer unconditionally true, and skipped:true is gone — a rejection used to look like a success. reason is human-readable, not a stable machine token. Under pause_persistence the write is skipped (result:'skipped') and the response carries persisted:false (id:'no-persist', embedded:false) — branch on persisted to tell a paused write apart from a dedup hit.
| Name | Required | Description | Default |
|---|---|---|---|
| channel | No | Memory channel for context separation (e.g. 'chat', 'discord'). Default: '' (shared). | |
| message | Yes | ClotoMessage to store. Legacy source shapes are normalized server-side where unambiguous (e.g. lowercase type words, Rust serde externally-tagged dicts, bare 'user'/'assistant' strings); unknown shapes are stored verbatim and surfaced by check_health(invalid_source_type). | |
| agent_id | Yes | Agent identifier | |
| project_id | No | v2.4.17 isolation axis. Optional — omit or pass '' to store in the global pool. Reads via γ semantics: a recall with project_id='X' returns 'X' rows + global pool. v2.5.1: pass '@auto' to resolve this agent's default from the server's operating context (the resolution is echoed as resolved_project_id; an unmapped agent yields operating_context_warning). bug-186: resolution requires a configured operating context. With none — the default, and equally the outcome of a sidecar that fails to parse — the sentinel is NOT resolved: it is stored and filtered as the literal project_id '@auto', resolved_project_id echoes '@auto', and no warning is raised. Read resolved_project_id before relying on the resolution. | |
| session_key | No | Opaque session identity you declare: a partition hint, not authentication and not a data filter. Selects which no-persist pause applies to this call. Omit to share one bucket with every caller that omits it. Full text on recall. | |
| associations | No | Associative memory to declare alongside this call: entities the text mentions, with their aliases, and subject–predicate–object relations. Stored verbatim; the server extracts nothing and infers nothing. On store, the stored memory is recorded as mentioning every entity named here and anchors every relation. Malformed items are reported in the response's associations.dropped and skipped; the memory is stored regardless. Optional. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description goes far beyond the annotations: it discloses dedup behavior, the TOCTOU seam, paused-persistence semantics, the breaking change in ok/skipped, the non-stable nature of reason, embedding-mode conditionality, overflow-tree queueing, and project_id '@auto' resolution failure behavior. This is exceptionally transparent for a mutation tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is information-dense and front-loaded with the most important branch ('stored'), but it is a long, nesting-heavy prose block. The pre-2.5.2b1 compatibility note and bug-186 digression add real value but also increase parsing load. It is thorough rather than concise; the structure sacrifices scannability for completeness.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With 6 parameters, nested objects, dedup semantics, rejection conditions, pausing, and version-dependent behavior, this is a high-complexity tool. The description covers response shapes, edge cases, fallback behavior, and migration notes comprehensively. Since there is no output schema, the description had to explain return semantics, and it does so exhaustively.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema already has 100% coverage, so the baseline is 3, but the description adds meaning beyond the schema: it explains how message.id participates in dedup, how content that sanitizes to empty is rejected, how project_id '@auto' behaves when the operating context is missing, how session_key only selects a pause bucket, and how associations are stored verbatim with no extraction. This is complementary, not redundant.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Store a message in agent memory for future recall.' It clearly identifies this as the write-path tool among read-side siblings like recall, traverse, and reconstruct, and explicitly names the response discriminator ('stored' / 'skipped' / 'rejected') so the agent knows what to branch on.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides rich guidance on when the write succeeds, when it is skipped, and when it is rejected, including exact conditions like dedup, empty content, pausing, and project_id refusal. It does not explicitly contrast with sibling tools such as declare_associations or update_memory, but the write-to-memory intent is clear enough that an agent will not confuse it with recall or traverse.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
traverseARead-only
The neighbourhood of a declared entity, as a graph: the entity named, its aliases, the entity -> entity relations declared on it and on what they reach, up to max_hops in either direction, and the refs of the records that mention each entity. Only what was declared (declare_associations, or associations on store); nothing is inferred. No record text: expand a ref with get_contents. entity is a name or an alias, compared after normalization; when it names more than one entity this call can read (a project's and the global pool's), all are starts. ORDER: entities by hops, then by the most recently declared relation that reached them, then by id; mentions by record id; relations most recently declared first. limit bounds both the entities returned and the refs listed per entity. Response: {entity, max_hops, limit, entities:[{id, name, hops, aliases?, mentions?, mentions_omitted?}], relations:[{id, subject, predicate, object, declared_by, declared_at, anchor_ref?}] (subject/object are entity ids from entities; a relation is listed when both ends are), entities_omitted?, bounds?:{omitted:[max_hops | limit]}, reason?:'no_such_entity'}. Relation ids are what declare_associations' retract takes. Records are listed only when this call could read them: the project, channel and source_id filters apply as in recall.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Maximum entities returned, and maximum mentioning refs listed per entity. | |
| entity | Yes | A declared entity name or alias. | |
| channel | No | Memory channel filter -- same semantics as in `recall`. | |
| agent_id | Yes | Agent identifier | |
| max_hops | No | Relations to follow from the entity, in either direction. 0 returns the entity alone. | |
| source_id | No | Per-user source filter on the mentioning records -- same semantics as in `recall`. | |
| project_id | No | γ filter -- same semantics as in `recall`. v2.5.1: pass '@auto' to resolve this agent's default from the server's operating context (the resolution is echoed as resolved_project_id; an unmapped agent yields operating_context_warning). bug-186: resolution requires a configured operating context. With none — the default, and equally the outcome of a sidecar that fails to parse — the sentinel is NOT resolved: it is stored and filtered as the literal project_id '@auto', resolved_project_id echoes '@auto', and no warning is raised. Read resolved_project_id before relying on the resolution. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only include readOnlyHint. The description adds extensive behavioral detail: only declared associations are returned, nothing is inferred, no record text, deterministic ordering rules, limit semantics, access-controlled record listing, and the no_such_entity error case. No contradiction exists with readOnlyHint.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long, but the first sentence front-loads the core purpose and every later sentence adds operational specifics (ordering, response shape, access rules, retract linkage). The only mild redundancy is restating `limit` semantics already present in the schema; overall the length is justified by the tool's complexity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description fully specifies the response shape, including optional fields, omitted flags, bounds, and the no_such_entity reason. It also covers ordering, limit behavior, access filtering, and cross-references to get_contents and recall, leaving nothing critical missing for a correct call.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema documents all parameters (100% coverage), so the baseline is 3. The description adds meaningful nuance beyond the schema, especially for `entity` (normalized comparison; multiple start entities when the name maps to more than one readable entity) and clarifies that filters behave as in recall. It does not fully detail every filter's edge cases, but the added value justifies a 4.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: it traverses the neighbourhood of a declared entity as a graph of entities, aliases, relations, and mentioning record refs. It distinguishes itself from siblings by pointing to get_contents for record text and by referencing recall for filter semantics.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Clear context is given about what the tool produces, and alternatives are referenced (get_contents for record text, recall for filter semantics). However, there is no explicit 'use this when ... and not when ...' statement, so exclusions and preference rules are left mostly implicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
unlock_memoryBIdempotent
Unlock a memory to allow deletion and editing. Ownership enforced when agent_id provided.
| Name | Required | Description | Default |
|---|---|---|---|
| agent_id | No | Agent ID for ownership verification | |
| memory_id | Yes | Memory ID to unlock | |
| session_key | No | Opaque session identity you declare: a partition hint, not authentication and not a data filter. Selects which no-persist pause applies to this call. Omit to share one bucket with every caller that omits it. Full text on recall. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, idempotentHint=true, and destructiveHint=false. The description adds one behavioral trait beyond them: ownership enforcement when agent_id is provided. It does not disclose failure behavior when ownership check fails, the locked→unlocked state transition semantics, or reversibility, which a mutation tool should surface.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences totaling 13 words with the primary action and effect front-loaded. Both sentences earn their place: the first states the operation and outcome, the second adds the ownership condition. No filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutating tool with no output schema, the description is thin: it omits the state transition semantics (what exactly 'unlocked' means), failure behavior when ownership fails, reversibility, and how session_key interacts with unlock. The schema's odd session_key note ('Full text on recall') is not reconciled in the description, leaving an agent with open questions for a state-changing call.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is effectively 100%, so all three parameters are already documented and the baseline is 3. The description's agent_id note is consistent with the schema's 'Agent ID for ownership verification' but adds no new meaning; memory_id and session_key are not elaborated in the description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb (unlock) and resource (memory) and explains the effect ('to allow deletion and editing'), which clearly differentiates it from the sibling lock_memory and from delete/update operations. It stops short of explicitly naming sibling contrast, so it misses the top score.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The purpose clause 'to allow deletion and editing' implies when the tool should be called (before editing or deleting a locked memory), and the ownership condition qualifies agent_id usage. However, no alternatives are named and there is no explicit when-not-to-use guidance or exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
update_memoryAIdempotent
Update memory content by ID. Rejects if memory is locked. Ownership enforced when agent_id provided. The new content passes through the same sanitizer as store: it is capped at the content length limit (the response carries truncated:true when the cap bit) and [Memory from ...] annotations are stripped, so content consisting only of those is refused rather than written as an empty row. A new text that runs past the embedding window gets nodes:{status:'queued'}, as on store.
| Name | Required | Description | Default |
|---|---|---|---|
| content | Yes | New content for the memory | |
| agent_id | No | Agent ID for ownership verification | |
| memory_id | Yes | Memory ID to update | |
| session_key | No | Opaque session identity you declare: a partition hint, not authentication and not a data filter. Selects which no-persist pause applies to this call. Omit to share one bucket with every caller that omits it. Full text on recall. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only establish that the tool writes, is idempotent, and is non-destructive. The description greatly expands on this with concrete behavioral details: locked-memory rejection, conditional ownership enforcement, sanitizer pass-through, truncation signaling, refusal of annotation-only content, and queued embed status. This is exactly the kind of context an agent needs beyond annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but every sentence carries operational information—no filler. The second sentence is long and contains minor jargon/phrasing issue ('when the cap bit'), which slightly hurts readability and prevents a perfect score.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers preconditions, sanitization, truncation, refusal behavior, and embedding queue side effects—all critical for correct invocation. There is no output schema, and while key response branches are mentioned, the normal success response and how to observe the updated content are not described, leaving a minor completeness gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3 even without description-level parameter help. The description adds real meaning by explaining that content is sanitizer-limited and can be refused, and that agent_id triggers ownership enforcement. This goes beyond the schema's generic descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a clear verb+resource: 'Update memory content by ID.' This immediately separates it from destructive siblings like delete_memory and creation siblings like store. It does not explicitly name the sibling alternatives it is not, so it stops just short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit guidance about when to use this tool versus store, delete_memory, lock_memory, or unlock_memory. Preconditions such as 'Rejects if memory is locked' and 'Ownership enforced when agent_id provided' are useful, but they do not tell an agent how to route among the many memory-related siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
update_profileA
Save a pre-computed agent profile to the database. The text passes through the same sanitizer as store, against the profile's own ceiling: it is capped at 2000 characters (CPERSONA_MAX_PROFILE_LENGTH) and the response carries truncated:true when the cap bit — branch on it, the discarded remainder is not stored anywhere else.
| Name | Required | Description | Default |
|---|---|---|---|
| profile | Yes | Profile text to save (pre-computed by caller). Capped at 2000 characters (CPERSONA_MAX_PROFILE_LENGTH); the response says truncated:true when the cap cut it. | |
| agent_id | Yes | Agent identifier | |
| session_key | No | Opaque session identity you declare: a partition hint, not authentication and not a data filter. Selects which no-persist pause applies to this call. Omit to share one bucket with every caller that omits it. Full text on recall. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description goes beyond the annotations by disclosing the 2000-character cap, the truncated:true response flag, the sanitizer behavior shared with store, and that discarded content is not stored elsewhere. This gives the agent actionable information about side effects and output, though it does not describe auth needs or full response behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is reasonably concise and front-loaded with the core purpose, then adds the critical truncation behavior. The phrasing 'when the cap bit — branch on it' is slightly awkward, but every sentence contributes meaningful information without redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple save operation with no output schema, the description covers the main behavioral caveat the agent must handle: truncation and the truncated:true flag. It does not specify the success response shape or whether an existing profile is overwritten, which are minor gaps but not blocking for invoking the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3. The description adds value beyond the schema by naming the constant CPERSONA_MAX_PROFILE_LENGTH, mentioning the shared sanitizer, and warning that the discarded remainder is not persisted. It does not add detail on session_key, but the schema already explains it adequately.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly says the tool saves a pre-computed agent profile to the database, with a specific verb, resource, and scope. It does not explicitly differentiate update_profile from its sibling store, though the 'pre-computed' qualifier and the reference to store's sanitizer imply a distinct role.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The intended use is implied: call this when you already have a pre-computed agent profile to save. However, it does not explicitly state when to prefer this over store or any other sibling, nor does it give exclusions or alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
8 tool updates
v2.5.12- Added
declare_associations - Changed
deep_check1 field changed- changed
Input schema / properties / checks / descriptionPrevious value: -"Checks to run (empty = all). Options: anonymous_source, short_content, stale_profile, orphaned_episodes, calibration_staleness, near_duplicate"New value: +"Checks to run (empty = all). Options: anonymous_source, short_content, stale_profile, orphaned_episodes, calibration_staleness, near_duplicate, unnormalized_content, embedding_norm"
- Changed
get_contents3 fields changed- changed
Input schema / properties / refs / descriptionPrevious value: -"Refs from recall messages, e.g. ['mem:123', 'ep:45'] (max 20 per call)"New value: +"Refs from recall messages ('mem:<id>' / 'ep:<id>'), or range objects such as {'ref': 'mem:<id>', 'node': [2, 3]} / {'ref': 'ep:<id>', 'span': [0, 800]} (max 20 per call)" - added
Input schema / properties / refs / items / anyOfAdded value: +[ + { + "type": "string" + }, + { + "properties": { + "node": { + "anyOf": [ + { + "minimum": 0, + "type": "integer" + }, + { + "items": { + "minimum": 0, + "type": "integer" + }, + "maxItems": 2, + "minItems": 2, + "type": "array" + } + ] + }, + "ref": { + "type": "string" + }, + "span": { + "items": { + "minimum": 0, + "type": "integer" + }, + "maxItems": 2, + "minItems": 2, + "type": "array" + } + }, + "required": [ + "ref" + ], + "type": "object" + } +] - removed
Input schema / properties / refs / items / typeRemoved value: -"string"
- Changed
recall1 field changed- changed
Input schema / properties / exclude_contents / descriptionPrevious value: -"Normalized content strings to exclude from results (starts-with match). Used to prevent duplication with conversation context already known to the caller."New value: +"Normalized content strings to exclude from results (starts-with match). Used to prevent duplication with conversation context already known to the caller. bug-399: the starts-with rule holds only at or above 32 characters. A shorter entry has to EQUAL the stored content (after the normalization this parameter already asks for: stripped and lower-cased) — a short prefix would otherwise suppress every memory beginning with it, inside the retrievers and with nothing in the response reporting the exclusion. Size entries at or above that length when you mean a prefix."
- Changed
recall_with_context8 fields changed- changed
Input schema / properties / external_context / descriptionPrevious value: -"Conversation history entries [{role, name?, user_id?, content, timestamp?}, ...]"New value: +"Conversation history entries [{role, content, name?, user_id?, timestamp?}, ...]. Every declared field is a string; one that is not is read as absent and reported in context_field_issues (CPERSONA_EXTERNAL_CONTEXT_MODE)." - added
Input schema / properties / external_context / items / properties / content / descriptionAdded value: +"The entry's text. A string — a value that is not one is read as absent and reported in context_field_issues." - removed
Input schema / properties / external_context / items / properties / content / typeRemoved value: -"string" - added
Input schema / properties / external_context / items / properties / nameAdded value: +{ + "description": "Display label for a role=user entry; becomes source.name, and source.id when user_id is absent. Default 'User'. A string — a value that is not one is read as absent and reported in context_field_issues." +} - added
Input schema / properties / external_context / items / properties / role / descriptionAdded value: +"'user' or 'assistant'; other roles filter the recall without being merged. A string — a value that is not one is read as absent and reported in context_field_issues." - removed
Input schema / properties / external_context / items / properties / role / typeRemoved value: -"string" - added
Input schema / properties / external_context / items / properties / timestampAdded value: +{ + "description": "ISO-8601 stamp deciding where this entry lands in the merged chronology. An entry without one — or with one that names no instant — sorts ahead of every dated message. A string — a value that is not one is read as absent and reported in context_field_issues." +} - added
Input schema / properties / external_context / items / properties / user_idAdded value: +{ + "description": "Stable id for a role=user entry; becomes source.id as 'discord:<user_id>'. A string — a value that is not one is read as absent and reported in context_field_issues." +}
- Added
reconstruct - Changed
store2 fields changed- added
Input schema / properties / associationsAdded value: +{ + "description": "Associative memory to declare alongside this call: entities the text mentions, with their aliases, and subject–predicate–object relations. Stored verbatim; the server extracts nothing and infers nothing. On store, the stored memory is recorded as mentioning every entity named here and anchors every relation. Malformed items are reported in the response's associations.dropped and skipped; the memory is stored regardless. Optional.", + "properties": { + "entities": { + "description": "Entities to register (if new) and mark as mentioned. Names are compared after normalization (NFKC, case-folded, whitespace collapsed).", + "items": { + "properties": { + "aliases": { + "description": "Other names for the same entity. An alias resolves to at most one entity per scope; a second claim on it is dropped.", + "items": { + "type": "string" + }, + "type": "array" + }, + "name": { + "description": "The canonical name, kept as written.", + "type": "string" + } + }, + "required": [ + "name" + ], + "type": "object" + }, + "type": "array" + }, + "relations": { + "description": "Declared relations. subject / object are entity names (registered if new) or record refs 'mem:<id>' / 'ep:<id>' of this agent; predicate is free text, normalized. A predicate from the role vocabulary (supports, supersedes, corrects, qualifies, contradicts, temporal_predecessor) on a record → record relation is read by reconstruct as that role.", + "items": { + "properties": { + "object": { + "type": "string" + }, + "predicate": { + "type": "string" + }, + "subject": { + "type": "string" + } + }, + "required": [ + "subject", + "predicate", + "object" + ], + "type": "object" + }, + "type": "array" + } + }, + "type": "object" +} - changed
Input schema / properties / message / properties / source / typePrevious value: -"object"New value: +[ + "object", + "string", + "null" +]
- Added
traverse
23 tool updates
v2.5.10- Changed
archive_episode1 field changed- added
Input schema / properties / session_keyAdded value: +{ + "default": "", + "description": "Opaque session identity you declare: a partition hint, not authentication and not a data filter. Selects which no-persist pause applies to this call. Omit to share one bucket with every caller that omits it. Full text on recall.", + "type": "string" +}
- Changed
calibrate_threshold1 field changed- added
Input schema / properties / session_keyAdded value: +{ + "default": "", + "description": "Opaque session identity you declare: a partition hint, not authentication and not a data filter. Selects which no-persist pause applies to this call. Omit to share one bucket with every caller that omits it. Full text on recall.", + "type": "string" +}
- Changed
check_health1 field changed- added
Input schema / properties / session_keyAdded value: +{ + "default": "", + "description": "Opaque session identity you declare: a partition hint, not authentication and not a data filter. Selects which no-persist pause applies to this call. Omit to share one bucket with every caller that omits it. Full text on recall.", + "type": "string" +}
- Added
check_update - Changed
deep_check1 field changed- added
Input schema / properties / session_keyAdded value: +{ + "default": "", + "description": "Opaque session identity you declare: a partition hint, not authentication and not a data filter. Selects which no-persist pause applies to this call. Omit to share one bucket with every caller that omits it. Full text on recall.", + "type": "string" +}
- Changed
delete_agent_data1 field changed- added
Input schema / properties / session_keyAdded value: +{ + "default": "", + "description": "Opaque session identity you declare: a partition hint, not authentication and not a data filter. Selects which no-persist pause applies to this call. Omit to share one bucket with every caller that omits it. Full text on recall.", + "type": "string" +}
- Changed
delete_episode1 field changed- added
Input schema / properties / session_keyAdded value: +{ + "default": "", + "description": "Opaque session identity you declare: a partition hint, not authentication and not a data filter. Selects which no-persist pause applies to this call. Omit to share one bucket with every caller that omits it. Full text on recall.", + "type": "string" +}
- Changed
delete_memory1 field changed- added
Input schema / properties / session_keyAdded value: +{ + "default": "", + "description": "Opaque session identity you declare: a partition hint, not authentication and not a data filter. Selects which no-persist pause applies to this call. Omit to share one bucket with every caller that omits it. Full text on recall.", + "type": "string" +}
- Added
get_session_findings - Changed
import_memories1 field changed- added
Input schema / properties / session_keyAdded value: +{ + "default": "", + "description": "Opaque session identity you declare: a partition hint, not authentication and not a data filter. Selects which no-persist pause applies to this call. Omit to share one bucket with every caller that omits it. Full text on recall.", + "type": "string" +}
- Changed
lock_memory1 field changed- added
Input schema / properties / session_keyAdded value: +{ + "default": "", + "description": "Opaque session identity you declare: a partition hint, not authentication and not a data filter. Selects which no-persist pause applies to this call. Omit to share one bucket with every caller that omits it. Full text on recall.", + "type": "string" +}
- Changed
merge_memories1 field changed- added
Input schema / properties / session_keyAdded value: +{ + "default": "", + "description": "Opaque session identity you declare: a partition hint, not authentication and not a data filter. Selects which no-persist pause applies to this call. Omit to share one bucket with every caller that omits it. Full text on recall.", + "type": "string" +}
- Changed
migrate_channel_axis1 field changed- added
Input schema / properties / session_keyAdded value: +{ + "default": "", + "description": "Opaque session identity you declare: a partition hint, not authentication and not a data filter. Selects which no-persist pause applies to this call. Omit to share one bucket with every caller that omits it. Full text on recall.", + "type": "string" +}
- Changed
pause_persistence1 field changed- added
Input schema / properties / session_keyAdded value: +{ + "default": "", + "description": "Opaque session identity you declare: a partition hint, not authentication and not a data filter. Selects which no-persist pause applies to this call. Omit to share one bucket with every caller that omits it. Full text on recall.", + "type": "string" +}
- Changed
persistence_status1 field changed- added
Input schema / properties / session_keyAdded value: +{ + "default": "", + "description": "Opaque session identity you declare: a partition hint, not authentication and not a data filter. Selects which no-persist pause applies to this call. Omit to share one bucket with every caller that omits it. Full text on recall.", + "type": "string" +}
- Changed
recall2 fields changed- changed
Input schema / properties / limit / descriptionPrevious value: -"Per-retriever search depth, not a pure response cap (CSC #716): the value is handed to each retrieval channel (vector / episode FTS / keyword) as its top-K, so lowering it shrinks the candidate pool itself — rows beyond the depth are unreachable at any gate value, and score normalization / autocut operate on the smaller pool, which can also reorder what remains. Fewer rows than this may be returned. (Agent-facing cap; the library layer accepts up to the scan window for direct callers.)"New value: +"Per-retriever search depth, not a pure response cap: the value is handed to each retrieval channel (vector / episode FTS / keyword) as its top-K, so lowering it shrinks the candidate pool itself — rows beyond the depth are unreachable at any gate value, and score normalization / autocut operate on the smaller pool, which can also reorder what remains. Fewer rows than this may be returned. (Agent-facing cap; the library layer accepts up to the scan window for direct callers.)" - added
Input schema / properties / session_keyAdded value: +{ + "default": "", + "description": "Opaque session identity you declare — a partition hint, NOT authentication. It scopes this process's per-session state: the degraded-recall advisory's \"already told you\" memory, and which no-persist pause applies to this call. It does NOT filter stored data (use agent_id / project_id / channel for that), and it never reaches the database. Omit it to share one bucket with every other caller that omits it, which is the behaviour that predates this parameter.", + "type": "string" +}
- Changed
recall_with_context2 fields changed- changed
Input schema / properties / limit / descriptionPrevious value: -"Per-retriever search depth for the underlying recall, not a pure response cap — same semantics as recall's limit (CSC #716): lowering it shrinks the candidate pool itself, not just the rows returned. (Agent-facing cap; the library layer accepts up to the scan window for direct callers.)"New value: +"Per-retriever search depth for the underlying recall, not a pure response cap — same semantics as recall's limit: lowering it shrinks the candidate pool itself, not just the rows returned. (Agent-facing cap; the library layer accepts up to the scan window for direct callers.)" - added
Input schema / properties / session_keyAdded value: +{ + "default": "", + "description": "Opaque session identity you declare — a partition hint, NOT authentication. It scopes this process's per-session state: the degraded-recall advisory's \"already told you\" memory, and which no-persist pause applies to this call. It does NOT filter stored data (use agent_id / project_id / channel for that), and it never reaches the database. Omit it to share one bucket with every other caller that omits it, which is the behaviour that predates this parameter.", + "type": "string" +}
- Changed
resume_persistence1 field changed- added
Input schema / properties / session_keyAdded value: +{ + "default": "", + "description": "Opaque session identity you declare: a partition hint, not authentication and not a data filter. Selects which no-persist pause applies to this call. Omit to share one bucket with every caller that omits it. Full text on recall.", + "type": "string" +}
- Changed
set_recall_precision1 field changed- added
Input schema / properties / session_keyAdded value: +{ + "default": "", + "description": "Opaque session identity you declare: a partition hint, not authentication and not a data filter. Selects which no-persist pause applies to this call. Omit to share one bucket with every caller that omits it. Full text on recall.", + "type": "string" +}
- Changed
store1 field changed- added
Input schema / properties / session_keyAdded value: +{ + "default": "", + "description": "Opaque session identity you declare: a partition hint, not authentication and not a data filter. Selects which no-persist pause applies to this call. Omit to share one bucket with every caller that omits it. Full text on recall.", + "type": "string" +}
- Changed
unlock_memory1 field changed- added
Input schema / properties / session_keyAdded value: +{ + "default": "", + "description": "Opaque session identity you declare: a partition hint, not authentication and not a data filter. Selects which no-persist pause applies to this call. Omit to share one bucket with every caller that omits it. Full text on recall.", + "type": "string" +}
- Changed
update_memory1 field changed- added
Input schema / properties / session_keyAdded value: +{ + "default": "", + "description": "Opaque session identity you declare: a partition hint, not authentication and not a data filter. Selects which no-persist pause applies to this call. Omit to share one bucket with every caller that omits it. Full text on recall.", + "type": "string" +}
- Changed
update_profile1 field changed- added
Input schema / properties / session_keyAdded value: +{ + "default": "", + "description": "Opaque session identity you declare: a partition hint, not authentication and not a data filter. Selects which no-persist pause applies to this call. Omit to share one bucket with every caller that omits it. Full text on recall.", + "type": "string" +}
2 tool updates
v2.5.6- Changed
recall1 field changed- changed
Input schema / properties / deep / descriptionPrevious value: -"Deep recall — disable time and completion decay for exhaustive search"New value: +"Deep recall — halves the quality gate (and the calibrated fused gate), so weaker matches are admitted. It also disables time and completion decay, which are inert unless CPERSONA_CONFIDENCE_ENABLED=true, and it does NOT widen the scan window (CPERSONA_MAX_MEMORIES) — deep is about how weak a match may be, not how far back the search reaches."
- Changed
recall_with_context1 field changed- changed
Input schema / properties / deep / descriptionPrevious value: -"Disable time decay"New value: +"Deep recall — same semantics as in `recall`: halves the quality gate so weaker matches are admitted."
5 tool updates
v2.5.4- Changed
calibrate_threshold1 field changed- added
Input schema / properties / method / enumAdded value: +[ + "separation", + "percentile", + "zscore" +]
- Changed
recall1 field changed- changed
Input schema / properties / limit / descriptionPrevious value: -"Max memories to return (agent-facing cap; the library layer accepts up to the scan window for direct callers)"New value: +"Per-retriever search depth, not a pure response cap (CSC #716): the value is handed to each retrieval channel (vector / episode FTS / keyword) as its top-K, so lowering it shrinks the candidate pool itself — rows beyond the depth are unreachable at any gate value, and score normalization / autocut operate on the smaller pool, which can also reorder what remains. Fewer rows than this may be returned. (Agent-facing cap; the library layer accepts up to the scan window for direct callers.)"
- Changed
recall_with_context1 field changed- changed
Input schema / properties / limit / descriptionPrevious value: -"Max recalled memories (agent-facing cap; the library layer accepts up to the scan window for direct callers)"New value: +"Per-retriever search depth for the underlying recall, not a pure response cap — same semantics as recall's limit (CSC #716): lowering it shrinks the candidate pool itself, not just the rows returned. (Agent-facing cap; the library layer accepts up to the scan window for direct callers.)"
- Changed
store2 fields changed- changed
Input schema / properties / message / properties / source / properties / type / descriptionPrevious value: -"Producer role — this enum IS the contract; send one of these. Legacy producers that cannot are folded server-side at the write seam ('ai' / 'assistant' are normalized to 'Agent'; 'session' is normalized to 'System' (type words are matched case-insensitively)), and shapes outside that table are stored verbatim for check_health(invalid_source_type) to surface."New value: +"Producer role — send one of 'User', 'Agent', 'System'. Legacy producers that cannot are folded server-side at the write seam ('ai' / 'assistant' are normalized to 'Agent'; 'session' is normalized to 'System' (type words are matched case-insensitively)), and shapes outside that table are stored verbatim for check_health(invalid_source_type) to surface." - removed
Input schema / properties / message / properties / source / properties / type / enumRemoved value: -[ - "User", - "Agent", - "System" -]
- Changed
update_profile1 field changed- changed
Input schema / properties / profile / descriptionPrevious value: -"Profile text to save (pre-computed by caller)"New value: +"Profile text to save (pre-computed by caller). Capped at 2000 characters (CPERSONA_MAX_PROFILE_LENGTH); the response says truncated:true when the cap cut it."
17 tool updates
v2.5.2- Changed
archive_episode1 field changed- changed
Input schema / properties / project_id / descriptionPrevious value: -"v2.4.17 isolation axis. Omit or pass '' for the global pool."New value: +"v2.4.17 isolation axis. Omit or pass '' for the global pool. v2.5.1: pass '@auto' to resolve this agent's default from the server's operating context (the resolution is echoed as resolved_project_id; an unmapped agent yields operating_context_warning). bug-186: resolution requires a configured operating context. With none — the default, and equally the outcome of a sidecar that fails to parse — the sentinel is NOT resolved: it is stored and filtered as the literal project_id '@auto', resolved_project_id echoes '@auto', and no warning is raised. Read resolved_project_id before relying on the resolution."
- Added
calibrate_threshold - Added
check_health - Added
delete_episode - Added
delete_memory - Added
export_memories - Added
get_operating_context - Added
get_recall_precision - Changed
list_episodes1 field changed- changed
Input schema / properties / project_id / descriptionPrevious value: -"v2.4.17 γ filter. Same semantics as list_memories."New value: +"v2.4.17 γ filter. Same semantics as list_memories. v2.5.1: pass '@auto' to resolve this agent's default from the server's operating context (the resolution is echoed as resolved_project_id; an unmapped agent yields operating_context_warning). bug-186: resolution requires a configured operating context. With none — the default, and equally the outcome of a sidecar that fails to parse — the sentinel is NOT resolved: it is stored and filtered as the literal project_id '@auto', resolved_project_id echoes '@auto', and no warning is raised. Read resolved_project_id before relying on the resolution."
- Changed
list_memories1 field changed- changed
Input schema / properties / project_id / descriptionPrevious value: -"v2.4.17 γ filter. Omit → no filter; '' → global pool only; 'X' → 'X' ∪ global pool."New value: +"v2.4.17 γ filter. Omit → no filter; '' → global pool only; 'X' → 'X' ∪ global pool. v2.5.1: pass '@auto' to resolve this agent's default from the server's operating context (the resolution is echoed as resolved_project_id; an unmapped agent yields operating_context_warning). bug-186: resolution requires a configured operating context. With none — the default, and equally the outcome of a sidecar that fails to parse — the sentinel is NOT resolved: it is stored and filtered as the literal project_id '@auto', resolved_project_id echoes '@auto', and no warning is raised. Read resolved_project_id before relying on the resolution."
- Added
pause_persistence - Added
recall - Added
recall_with_context - Added
resume_persistence - Changed
store4 fields changed- changed
Input schema / properties / message / properties / content / descriptionPrevious value: -"The text to store. Empty content is skipped."New value: +"The text to store. Content that is empty — or that sanitizes to empty — is refused with ok:false, result:'rejected'." - changed
Input schema / properties / message / properties / metadata / descriptionPrevious value: -"Free-form JSON object for producer-specific context. Empty when unused."New value: +"Free-form JSON object for producer-specific context. Empty when unused. Serialised size is capped at 8000 characters (same cap for source); an oversized field is refused with result='rejected' rather than truncated, because a truncated JSON document is not a JSON document." - changed
Input schema / properties / message / properties / source / properties / type / descriptionPrevious value: -"Producer role. 'Assistant' / 'ai' are normalized to 'Agent'; 'session' is normalized to 'System'."New value: +"Producer role — this enum IS the contract; send one of these. Legacy producers that cannot are folded server-side at the write seam ('ai' / 'assistant' are normalized to 'Agent'; 'session' is normalized to 'System' (type words are matched case-insensitively)), and shapes outside that table are stored verbatim for check_health(invalid_source_type) to surface." - changed
Input schema / properties / project_id / descriptionPrevious value: -"v2.4.17 isolation axis. Optional — omit or pass '' to store in the global pool. Reads via γ semantics: a recall with project_id='X' returns 'X' rows + global pool. v2.5.1: pass '@auto' to resolve this agent's default from the server's operating context (echoed as resolved_project_id)."New value: +"v2.4.17 isolation axis. Optional — omit or pass '' to store in the global pool. Reads via γ semantics: a recall with project_id='X' returns 'X' rows + global pool. v2.5.1: pass '@auto' to resolve this agent's default from the server's operating context (the resolution is echoed as resolved_project_id; an unmapped agent yields operating_context_warning). bug-186: resolution requires a configured operating context. With none — the default, and equally the outcome of a sidecar that fails to parse — the sentinel is NOT resolved: it is stored and filtered as the literal project_id '@auto', resolved_project_id echoes '@auto', and no warning is raised. Read resolved_project_id before relying on the resolution."
- Added
unlock_memory - Added
update_profile
15 tool updates
v2.5.1- Changed
archive_episode1 field changed- changed
Input schema / properties / history / descriptionPrevious value: -"Original conversation messages (used for timestamp extraction and embedding)"New value: +"Original conversation messages (used for start/end timestamp extraction; the episode embedding is computed from summary)"
- Removed
calibrate_threshold - Removed
check_health - Removed
delete_episode - Removed
delete_memory - Removed
export_memories - Added
get_contents - Removed
get_recall_precision - Removed
pause_persistence - Removed
recall - Removed
recall_with_context - Removed
resume_persistence - Changed
store3 fields changed- changed
Input schema / properties / message / descriptionPrevious value: -"ClotoMessage to store (id, content, source, timestamp, metadata)"New value: +"ClotoMessage to store. Legacy source shapes are normalized server-side where unambiguous (e.g. lowercase type words, Rust serde externally-tagged dicts, bare 'user'/'assistant' strings); unknown shapes are stored verbatim and surfaced by check_health(invalid_source_type)." - added
Input schema / properties / message / propertiesAdded value: +{ + "content": { + "description": "The text to store. Empty content is skipped.", + "type": "string" + }, + "id": { + "description": "Caller-supplied message id used for msg_id-based dedup (γ-project-scoped). Optional.", + "type": "string" + }, + "metadata": { + "description": "Free-form JSON object for producer-specific context. Empty when unused.", + "type": "object" + }, + "source": { + "description": "Attribution of who produced the content. Canonical shape is {type, id, name}. Type is the discriminator; id / name identify the concrete producer. Store null / empty {} only when the producer is genuinely unknown. A null source is normalized to {} at the write seam, so both persist (and recall) as the anonymous {}.", + "properties": { + "id": { + "description": "Stable producer id (e.g. discord user id, agent id). Empty when anonymous.", + "type": "string" + }, + "name": { + "description": "Human-readable label for display. Empty when unknown.", + "type": "string" + }, + "type": { + "description": "Producer role. 'Assistant' / 'ai' are normalized to 'Agent'; 'session' is normalized to 'System'.", + "enum": [ + "User", + "Agent", + "System" + ], + "type": "string" + } + }, + "type": "object" + }, + "timestamp": { + "description": "UTC ISO-8601 timestamp with offset (e.g. '2026-07-22T12:00:00+00:00'). Defaults to server-time UTC when omitted. Aware non-UTC offsets are accepted; naive strings are surfaced by check_health(timestamp_format_drift).", + "type": "string" + } +} - changed
Input schema / properties / project_id / descriptionPrevious value: -"v2.4.17 isolation axis. Optional — omit or pass '' to store in the global pool. Reads via γ semantics: a recall with project_id='X' returns 'X' rows + global pool."New value: +"v2.4.17 isolation axis. Optional — omit or pass '' to store in the global pool. Reads via γ semantics: a recall with project_id='X' returns 'X' rows + global pool. v2.5.1: pass '@auto' to resolve this agent's default from the server's operating context (echoed as resolved_project_id)."
- Removed
unlock_memory - Removed
update_profile
2 tool updates
v2.4.37- Changed
check_health1 field changed- added
Input schema / properties / checksAdded value: +{ + "description": "Registry check names to run (empty = all). See cpersona.checks.HEALTH_CHECK_NAMES.", + "items": { + "type": "string" + }, + "type": "array" +}
- Changed
deep_check1 field changed- changed
Input schema / properties / checks / descriptionPrevious value: -"Checks to run (empty = all). Options: anonymous_source, short_content, stale_profile, orphaned_episodes"New value: +"Checks to run (empty = all). Options: anonymous_source, short_content, stale_profile, orphaned_episodes, calibration_staleness, near_duplicate"
13 tool updates
v2.4.34- Changed
archive_episode2 fields changed- added
Input schema / properties / channelAdded value: +{ + "description": "v2.4.22 conversation-channel tag (e.g. a Discord channel id). Default '' (= unscoped). Channel-scoped recall returns episodes whose channel matches; this powers the per-channel episodic loop.", + "type": "string" +} - added
Input schema / properties / project_idAdded value: +{ + "description": "v2.4.17 isolation axis. Omit or pass '' for the global pool.", + "type": "string" +}
- Changed
calibrate_threshold3 fields changed- added
Input schema / properties / methodAdded value: +{ + "description": "'percentile' (default), 'zscore', or 'separation' (two-population, learns the operating point from null vs nearest-neighbour positives)", + "type": "string" +} - added
Input schema / properties / percentileAdded value: +{ + "description": "Null-distribution quantile for method='percentile' (default: 0.95, higher = stricter)", + "type": "number" +} - changed
Input schema / properties / z_factor / descriptionPrevious value: -"Z-score multiplier (default: 1.0, higher = stricter)"New value: +"Z-score multiplier for method='zscore' (default: 1.0, higher = stricter)"
- Added
get_recall_precision - Changed
list_episodes1 field changed- added
Input schema / properties / project_idAdded value: +{ + "description": "v2.4.17 γ filter. Same semantics as list_memories.", + "type": "string" +}
- Changed
list_memories1 field changed- added
Input schema / properties / project_idAdded value: +{ + "description": "v2.4.17 γ filter. Omit → no filter; '' → global pool only; 'X' → 'X' ∪ global pool.", + "type": "string" +}
- Added
migrate_channel_axis - Added
pause_persistence - Added
persistence_status - Changed
recall2 fields changed- added
Input schema / properties / project_idAdded value: +{ + "description": "v2.4.17 γ filter. Omit → no filter (all projects). '' → global pool only. 'X' → 'X' bucket ∪ global pool. Threaded through cascade / RRF / vector / FTS / keyword paths.", + "type": "string" +} - added
Input schema / properties / source_idAdded value: +{ + "default": "", + "description": "v2.4.20 per-user source filter. Empty (default) = no filter. Non-empty = prefix match against json_extract(source, '$.id'), e.g. 'discord:12345' to restrict to one Discord user, or 'discord:' to scope to all Discord-sourced memories. Episodes are skipped when set (no per-user source tagging).", + "type": "string" +}
- Changed
recall_with_context2 fields changed- added
Input schema / properties / project_idAdded value: +{ + "description": "v2.4.17 γ filter — passed through to recall. Same semantics as in `recall`.", + "type": "string" +} - added
Input schema / properties / source_idAdded value: +{ + "default": "", + "description": "v2.4.20 per-user source filter — passed through to recall. Same semantics as in `recall`.", + "type": "string" +}
- Added
resume_persistence - Added
set_recall_precision - Changed
store1 field changed- added
Input schema / properties / project_idAdded value: +{ + "description": "v2.4.17 isolation axis. Optional — omit or pass '' to store in the global pool. Reads via γ semantics: a recall with project_id='X' returns 'X' rows + global pool.", + "type": "string" +}
6 tool updates
v2.4.10- Added
deep_check - Added
lock_memory - Changed
recall1 field changed- added
Input schema / properties / exclude_contentsAdded value: +{ + "description": "Normalized content strings to exclude from results (starts-with match). Used to prevent duplication with conversation context already known to the caller.", + "items": { + "type": "string" + }, + "type": "array" +}
- Added
recall_with_context - Added
unlock_memory - Added
update_memory
16 tool updates
v0.1.0- First observed
archive_episode - First observed
calibrate_threshold - First observed
check_health - First observed
delete_agent_data - First observed
delete_episode - First observed
delete_memory - First observed
export_memories - First observed
get_profile - First observed
get_queue_status - First observed
import_memories - First observed
list_episodes - First observed
list_memories - First observed
merge_memories - First observed
recall - First observed
store - First observed
update_profile
TDQS
Scored across 34 tools
Most tools map to a distinct verb+resource pair, but the health/diagnostic cluster (check_health, get_session_findings, deep_check) overlaps heavily, and recall/reconstruct/traverse sit close together. The very detailed descriptions help an agent after reading them, but selection mistakes are plausible.
The dominant pattern is verb_noun snake_case (list_memories, update_memory, set_recall_precision), which is consistent and predictable. A few outliers — store, traverse, reconstruct, deep_check, persistence_status — deviate slightly but do not obscure meaning.
34 tools is well beyond the comfortable range and makes the surface hard to survey. The breadth reflects a genuinely feature-rich memory server, but several clusters (diagnostics, persistence, recall variants) could be consolidated.
The surface covers memory/episode/profile CRUD, recall and reconstruction, associations, precision tuning, import/export/merge, persistence controls, and deep diagnostics. There are no obvious dead ends; every operation has a read/write pairing and a maintenance path.
Maintenance
Related MCP Connectors
Persistent, outcome-grounded episodic memory for Claude. 14ms CPU retrieval, no GPU, no vector DB.
Persistent memory for AI agents. Semantic search, memory graph, W3C DID identity.
Persistent memory for AI agents across Claude, ChatGPT and any MCP client.
- AmberOAuthcom.ambermem
Long-term memory for AI assistants. Hybrid retrieval, query expansion, auto-topics.
Related MCP Servers
- AlicenseNot gradedqualityAmaintenancePersistent AI memory server with 3-layer hybrid search (vector + FTS5 + keyword), confidence scoring via Reciprocal Rank Fusion, episodic/profile memory, and 16 tools. Zero LLM dependency. Works standalone with Claude Desktop and Claude Code. MIT licensed.3Business Source 1.1
- AlicenseAqualityCmaintenanceLocal-first memory for Claude Code and any MCP client: hybrid vector + keyword search and a bi-temporal knowledge graph in one SQLite file. Local embeddings, no API key, $0/token.51168 npm2PolyForm Noncommercial 1.0.0
- FlicenseNot gradedqualityDmaintenanceA lightweight MCP memory server built on SQLite + FTS5, providing cross-session long-term memory for Claude Code.-
- FlicenseNot gradedqualityDmaintenancePersistent memory server for AI assistants with semantic search and three-layer context (global, project, personality). Works with MCP-compatible AI tools like Claude Code, Cursor, Continue, Cline, and more.1-