CPersona
OfficialCPersona is a local, LLM-free MCP memory server: it stores an agent's memories in a single SQLite file and recalls them later by hybrid (vector + FTS5 + keyword) search.
Store and manage memory —
storemessages (facts/session content) with optional entities, aliases and relations;update_memory,delete_memory,lock_memory/unlock_memory, andarchive_episodefor summarized conversations.Recall memory —
recall(hybrid vector/FTS5/keyword, with time cues, project/channel/source filters, excerpts and trace),recall_with_context(merge in external conversation history), andreconstruct— the recommended read path — which returns items quoting the stored rows behind them, traced viaget_contents.Trace and navigate evidence —
get_contentsexpands preview refs, exact character spans, overflow-tree nodes and revisions;traversewalks the declared entity/relation graph.Build associative memory —
declare_associationsrecords entities, aliases and subject–predicate–object relations verbatim (withretractto remove);reconstructreads them as evidence and roles.Agent profile —
get_profile/update_profilefor the accumulated, caller-computed profile.Browse and port data —
list_memories,list_episodes(budgeted, preview-tiered listings),export_memoriesto JSONL,import_memoriesfrom JSONL,merge_memoriesbetween agents.Tune retrieval —
calibrate_threshold(separation/percentile/zscore),set_recall_precisionandget_recall_precision(strict/balanced/lenient or raw beta), recalibrating and persisting the quality gate live.Operate and diagnose —
check_health(36-check registry, auto-fix, lossy repairs),deep_check(data-quality judgment),get_session_findings(whole-DB integrity findings),get_queue_status.Control persistence —
pause_persistence/resume_persistence/persistence_statusto no-op all writes for a TTL window (scoped to asession_keybucket) for benchmarking or AB testing; reads still work.Housekeeping and updates —
migrate_channel_axisto re-channel bridge memories, andcheck_updateto see whether a newer release exists and optionally install it (never automatic; restart required).Operator doctrine —
get_operating_contextreads the server-side operating context (preview tier or one section's full body); it is edited by the operator on disk, not via MCP.Isolation and honesty — everything is scoped by
agent_id,project_idandchannel; the server never calls a generative model, and it explicitly signals degraded recall (no embedding backend → FTS5/keyword only,gate_fallback) rather than quietly returning less.
CPersona
MCP Memory Server
Persistent memory for AI agents, over MCP.
One SQLite file you own. No LLM in the loop. Honest when recall degrades.
Ask your agent to save a project's design decisions and their reasons, then recall them in a later coding session (setup and verification).
Use reconstruct to read recalled items with quotes from the stored records (traceable recall).
Documentation · Getting Started · Architecture · Tools · PyPI · Zenn Book (JP)
Standalone repository — This is the standalone version for use with Claude Desktop, Claude Code, Codex CLI, Cursor, VS Code, and any other MCP client (registration table). If you are a ClotoCore user, install CPersona from the in-app marketplace (ClotoHub) instead — it distributes this same repository.
Project status — 2.4.x is Stable; 2.6.x is Current, where all fixes land, pending production-soak certification. Its first final, 2.6.0, migrates the schema from version 13 to 17 on its first start and changes some tool behaviour, so take a 2.5 store across with Upgrading from 2.5 to 2.6. 2.5.x is Candidate: no longer the newest line and not yet certified, reachable by exact version. Which version to run, and how long each line keeps receiving fixes: SUPPORT.md. Where the lines are heading: the roadmap.
Upgrading from 2.5.2 or earlier? Two things need a decision from you. v2.5.3 will not start the HTTP transport without
CPERSONA_AUTH_TOKEN, wherever it binds — set one, or opt out withCPERSONA_ALLOW_UNAUTHENTICATED_HTTP=true(why; stdio is unaffected). v2.5.2 changed tool response shapes — branch onok is false, and treat any response carryingerroras a failure whether or notokis present (contract §10).
The Problem
Claude forgets everything between sessions. Every conversation starts from zero — no context about your project, your preferences, or what you discussed yesterday.
cpersona fixes this. It's an MCP server that stores memories in a local SQLite file and retrieves them through hybrid search. Claude remembers you. It runs against any MCP-compatible host — Claude Desktop, Claude Code, ClotoCore (the AI agent platform where cpersona originated, and whose memory layer it is), or a client of your own.
Related MCP server: mcp-memory-graph
Quick Start
Setting this up as an AI agent? Read the Agent Skill —
skills/cpersona-memory/SKILL.mdin a clone of this repository — and follow it end to end. It is the complete procedure: install, embedding server, registration, and the two steps a quick start tends to skip — verifying from a new session, and persisting the memory policy so the next session knows whichagent_idholds the memories.Using Claude Code yourself? The same skill ships inside the wheel. Once
cpersonais installed, copy it in and say "Set up CPersona."python -c "import cpersona,pathlib,shutil; s=pathlib.Path(cpersona.__file__).parent/'skills'/'cpersona-memory'; shutil.copytree(s, pathlib.Path.home()/'.claude/skills/cpersona-memory', dirs_exist_ok=True)"
1. Install — Python 3.11+, and uv for the one-command path.
uvx cpersona # run directly, no install step
pip install cpersona # or install it2. Run an embedding server — strongly recommended; it powers the vector layer
uvx --from "cembedding[onnx]" cembedding-download-model --model jina-v5-nano
EMBEDDING_PROVIDER=onnx_jina_v5_nano uvx --from "cembedding[onnx]" cembedding # serves http://127.0.0.1:8401/embedThe reference server's lifetime is bound to its stdin. Started with stdin closed — by a service manager, by nohup … </dev/null, or from an agent's background shell — it binds the port and exits within the same second with status 0. Give it a stdin that stays open (sleep infinity | cembedding); Getting Started has the details.
Any endpoint implementing the embedding contract works and is equally recommended; CEmbedding is the reference implementation. The choice of backend is yours — the recommendation is to connect one, not to connect that one.
Without a backend, cpersona still runs — FTS5 + keyword search, and it says on every recall that it is degraded rather than quietly returning less. That is a supported fallback, not a recommended way to run: recall then matches on shared words, so a memory phrased differently from your question can be missed, and so can an older one.
3. Register it with your MCP client
For Codex, Cursor, VS Code, or Claude Desktop, follow the client registration table. The command below is for Claude Code.
claude mcp add-json cpersona '{"type":"stdio","command":"uvx","args":["cpersona"],"env":{"CPERSONA_DB_PATH":"/home/you/.claude/cpersona.db","EMBEDDING_MODE":"http","EMBEDDING_HTTP_URL":"http://127.0.0.1:8401/embed"}}' -s user4. Verify from a new session — ask the agent to store something, then recall it in a fresh session. Surviving the session boundary is the whole point.
5. Make it stick — registration gives the agent the tools. It does not tell the next session to use them, or which agent_id holds the memories: recall is scoped to an exact agent_id, so a session that guesses the wrong one gets nothing back. Persist the short policy block into the file your client loads every session (~/.claude/CLAUDE.md, AGENTS.md, …) — Getting Started §5.
At startup the server asks pypi.org whether a newer release exists and tells the
calling agent through recall; set CPERSONA_UPDATE_CHECK=false to turn that
off. Updating is never automatic.
Claude Desktop config, Windows paths, installing from source and the full walkthrough: Getting Started.
What You Get
Hybrid search — vector (the layer an embedding server powers), FTS5 (trigram, so it works on Japanese and other space-less scripts) and keyword, fused by rank or relative score. The FTS and keyword layers rescue what vectors miss: identifiers, error strings, exact names.
Evidence you can trace —
reconstruct, the recommended way to answer from memory, returns items that quote the stored rows behind them, within a character budget. Text past a long record's embedding window stays reachable (block reach, on by default).Three memory types — facts, session summaries and an accumulated profile.
Zero LLM dependency — cpersona never calls a generative model; your agent summarizes and hands over the result. Recall is deterministic given a calibrated gate, but the gate is sampled, so two installs on identical data can settle differently.
Single-file SQLite — no external database;
sqlite3 .backupcopies the corpus (the calibration sidecar beside it needs copying too).Operable — auto-calibrated thresholds, a health check with auto-repair, an advisory when the embedding layer dies, JSONL export/import, agent-to-agent merge.
Isolation —
agent_id,project_idandchannellet several agents and projects share one database without bleeding into each other.
How it fits together: Architecture · what the tools do: Tools · what you may rely on: Behavior Contracts.
Benchmarks
Measured on OmniMemEval's LongMemEval-S
pipeline: the 500 questions of the cleaned LongMemEval-S, answered by gpt-4.1-mini
and judged by gpt-4o-mini, through the same harness, prompts and models as the
memory backends OmniMemEval reproduces. Accuracy is the share of questions judged
correct; Context Tokens is the average number of tokens sent to the answer model per
question, its prompt included.
Backend | Deployment | SS-User | SS-Asst | SS-Pref | Temp. Reas | Multi-S | Know. Upd | Overall | Context Tokens |
CPersona 2.6.3a1 | local | 90.00 | 78.57 | 86.67 | 85.71 | 67.67 | 85.90 | 80.80 | 2,354.6 |
CPersona v1.2 (2.6.4a1) | local | 91.43 | 80.36 | 90.00 | 83.46 | 71.43 | 84.62 | 81.60 | 1,786.7 |
81.60% at 1,786.7 tokens a question. Of the 12 backends OmniMemEval reproduces through the same pipeline, MemOS answers more questions (89.20%, at 4,151 tokens) and Mem0's cloud service sends fewer tokens (856, at 56.00%); 2.6.4a1 scored higher with fewer tokens than each of the other ten, the closest of them within one run's sampling error. CPersona stores each conversation as it is and calls no model to store or recall.
2.6.3a1 is the code of 2.6.3, the release a plain install gets; 2.6.4a1 is a
pre-release. Each was run once, by this project. The comparison rows, the registered
rules and every departure from a clean run are in the
results.
Retrieval-only measurements on LMEB (22 tasks) and the harness behind them:
benchmarks/.
Documentation
cloto-dev.github.io/CPersona is canonical — when this README disagrees with it, the site wins.
Install, embedding server, client registration, verification | |
What you may rely on: recall ordering, dedup, scan window, response shapes | |
Every tool, grouped by what you reach for it for | |
Storage, the retrieval pipeline, isolation axes | |
What each release line is for and may break; planned retrieval features and the scale ladder | |
Backup, degradation detection, tuning, CJK guidance, corpus sync | |
Every environment variable and its default | |
How a release is gated: audits, the bug ledger, structural and mutation gates | |
Short answers to the questions operators actually ask |
Japanese translations are in the language selector (English is canonical) and
agents can read llms.txt.
Longer reads in Japanese: a book
on the design and setup, and an article
on the token economics of session-end → /clear → recall.
Quality Assurance
Every release is gated by a machine-verifiable process: multi-agent audit rounds with adversarial verification, a bug ledger that fails CI if a fix marker disappears or a removed defect returns, structural gates for invariants a plain test cannot express, a mutation proof that those gates go red when the invariant is broken, and gates holding the documented counts, defaults and version claims to the source that defines them.
Behind it: ~2,917 test functions across ~205 test modules (~3,730 cases parametrised, more test code than server code), on Schema v18 — how a release is gated.
Support
Three tiers — Stable (production-certified, critical fixes only), Current (newest line, all fixes land here) and Experimental (opt-in pre-releases). A superseded line keeps critical-fix support for 30 more days. Read SUPPORT.md § Known issues before pinning a version — some of them change what you should run.
Found a bug, or something the docs do not explain? Open a bug report or feature request, even when you are not certain — a configuration problem mistaken for a bug means the documentation was unclear, which is a defect of its own. Report security vulnerabilities privately via SECURITY.md.
Sponsorship
CPersona is MIT-licensed and stays fully usable whether or not anyone sponsors it. Sponsorship buys no feature, no release tier and no position in the issue queue — issues are triaged by impact, reproducibility and safety, and that does not change for anyone.
If CPersona has earned a place in your workflow and you would like the work to continue, you can sponsor Cloto-dev on GitHub. The same page covers CPersona, ClotoCore and the other projects published under that account; sponsorship goes toward development time, testing and infrastructure, documentation and maintenance.
Money is not the only thing that helps, and it is not the thing this project needs most. Starring the repository, saying which part of the setup was confusing, filing a reproducible issue, or correcting a sentence in the documentation all move it forward.
License
MIT — free to use from any MCP host without restriction.
Available Tools
34 toolsarchive_episodeA
Archive a conversation episode with pre-computed summary, keywords, and resolved status. All LLM processing is performed by the caller. A summary that runs past the embedding window adds nodes:{status:'queued'} to the response, as on store.
| Name | Required | Description | Default |
|---|---|---|---|
| channel | No | v2.4.22 conversation-channel tag (e.g. a Discord channel id). Default '' (= unscoped). Channel-scoped recall returns episodes whose channel matches; this powers the per-channel episodic loop. | |
| history | No | Original conversation messages (used for start/end timestamp extraction; the episode embedding is computed from summary) | |
| summary | Yes | Episode summary (pre-computed by caller) | |
| agent_id | Yes | Agent identifier | |
| keywords | No | Space-separated keywords (pre-computed by caller) | |
| resolved | No | Whether the topic was completed/concluded | |
| project_id | No | v2.4.17 isolation axis. Omit or pass '' for the global pool. v2.5.1: pass '@auto' to resolve this agent's default from the server's operating context (the resolution is echoed as resolved_project_id; an unmapped agent yields operating_context_warning). bug-186: resolution requires a configured operating context. With none — the default, and equally the outcome of a sidecar that fails to parse — the sentinel is NOT resolved: it is stored and filtered as the literal project_id '@auto', resolved_project_id echoes '@auto', and no warning is raised. Read resolved_project_id before relying on the resolution. | |
| session_key | No | Opaque session identity you declare: a partition hint, not authentication and not a data filter. Selects which no-persist pause applies to this call. Omit to share one bucket with every caller that omits it. Full text on recall. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnlyHint=false, idempotentHint=false, destructiveHint=false). The description adds genuine value beyond that: it discloses the caller-side LLM contract and the specific overflow behavior (nodes:{status:'queued'} when the summary exceeds the embedding window). It stops short of describing persistence semantics or auth requirements.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences with no waste; the core purpose is front-loaded and the caller-processing constraint and queue behavior follow immediately. Nothing is redundant.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description usefully covers the one notable return-value shape (the queued node) and the caller-processing contract. For an 8-parameter mutation tool this is largely sufficient, though it could say more about persistence/dedup behavior.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all 8 parameters in detail — including the complex project_id '@auto' resolution semantics. The description adds the summary/embedding-window interaction but no syntax or meaning for the remaining parameters, so the baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (archive) and resource (conversation episode) plus the key inputs (summary, keywords, resolved status). It implicitly separates itself from the sibling `store` by saying the queued behavior is 'as on store', which signals this is the episode-specific archive path rather than the generic store.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'All LLM processing is performed by the caller' does imply the usage contract (caller must pre-compute the summary/keywords), but it never states when to choose this over siblings like store, list_episodes, or delete_episode. Usage is implied rather than explicitly routed.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
calibrate_thresholdA
Auto-calibrate the vector search threshold from the null (random-pair) cosine distribution. Samples random memory pairs and places the threshold ABOVE the null mean so unrelated pairs are rejected. method='separation' (default) learns the operating point from two populations — null pairs vs temporally-adjacent same-session positives (nearest-neighbour fallback when too few exist); method='percentile' uses a quantile of the null distribution (robust to anisotropic models such as bge-m3); method='zscore' uses mean + z*std. No labels used, purely statistical. Adapts to both embedding model and corpus characteristics.
| Name | Required | Description | Default |
|---|---|---|---|
| method | No | 'separation' (default; two-population — learns the operating point from null pairs vs temporally-adjacent same-session positives, falling back to nearest-neighbour when too few exist), 'percentile', or 'zscore' | |
| agent_id | Yes | Agent ID whose memories to sample | |
| z_factor | No | Z-score multiplier for method='zscore' (default: 1.0, higher = stricter) | |
| percentile | No | Null-distribution quantile for method='percentile' (default: 0.95, higher = stricter) | |
| sample_size | No | Number of embeddings to sample (default: 200) | |
| session_key | No | Opaque session identity you declare: a partition hint, not authentication and not a data filter. Selects which no-persist pause applies to this call. Omit to share one bucket with every caller that omits it. Full text on recall. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=false and non-idempotent, and the description says the calibration is 'purely statistical' with no labels, adding algorithmic transparency. However, it does not disclose operational behavior: whether the computed threshold is persisted to the agent, whether it overwrites a prior threshold, the cost of sampling 200 embeddings, or any auth requirement.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The purpose is front-loaded in the first sentence, followed by method semantics. It is a bit dense with parenthetical detail, but nearly every clause (null mean placement, no-labels, model/corpus adaptation) carries information an agent needs.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 6-parameter, no-output-schema tool, the description thoroughly explains the calibration algorithm and method choices, which is the core complexity. It leaves minor gaps (persistence of the resulting threshold, response shape), which keeps it from a 5.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds genuine value beyond the schema by explaining the tradeoffs between the method enum values (separation's two-population learning, percentile's robustness to anisotropic models like bge-m3, zscore's mean+z*std), which helps an agent pick the right value.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'Auto-calibrate the vector search threshold from the null (random-pair) cosine distribution.' An agent immediately understands the operation. It does not, however, name or contrast with any sibling tool (e.g. set_recall_precision), so it stops short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives guidance on which of the three methods to choose (percentile for anisotropic models like bge-m3), which is useful, but it never states when to call this tool versus alternatives such as set_recall_precision, nor any precondition like requiring existing memories. Usage is implied rather than explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
check_healthA
Check memory database health (36-check registry, each issue tagged with severity critical/warn/info). Detects contamination, duplicates, oversized content, embedding issues, FTS integrity (count + content-level), schema version/object drift (missing UNIQUE indexes or FTS triggers), SQLite file integrity, project_id naming drift, invalid JSON/timestamps, timestamp format drift, stale tasks, missing profiles, empty content, invalid/anonymous sources. Returns storage stats incl. project_id/channel distributions. Set fix=true to auto-repair (agent-scoped, locked-safe); the one exception is dedup_msg_id_index, whose repair blanks colliding msg_id values under every agent because the UNIQUE index it restores is a global schema object — with an ACL configured that repair demands read-write on '*', so exclude it via checks to stay agent-scoped. critical file-integrity findings are report-only. Two repairs are lossy and irreversible, each against its own cap: oversized memories are cut to CPERSONA_MAX_CONTENT_LENGTH (default 16000 since 2.5.4a2) and the agent's profile row to CPERSONA_MAX_PROFILE_LENGTH (default 2000), keeping the start. Lower either cap and a fix run shortens rows that were within the old one. Some repairs are bounded per run (source canonicalisation classifies at most 10000 rows); a fix response carrying remaining > 0 with a re-run hint has NOT converged — run fix again until remaining stops decreasing. Use checks parameter to run a subset — an unknown name is rejected (ok=false) rather than silently running nothing, and every response echoes checks_run. The verdict is status: healthy / degraded / unhealthy, derived from severity counts (info never degrades). The pre-2.5.2b1 healthy boolean (len(issues) == 0) is gone — it reported False for an info-only database that status called healthy; read issues / severity_summary for the underlying counts. Read status as a verdict on what is IN the database, not on whether the pipeline that fills it is working: a corpus where every embedding is NULL is internally consistent, so it scores healthy while semantic recall is dead. Nothing here contacts the embedding backend unless fix=true — on a report-only run the liveness findings cannot appear at all, and their absence is not evidence the backend answered. The null_embedding finding carries the reason its repair cannot run; read that before reading status.
| Name | Required | Description | Default |
|---|---|---|---|
| fix | No | Auto-fix detected issues | |
| checks | No | Registry check names to run (empty = all). See cpersona.checks.HEALTH_CHECK_NAMES. | |
| agent_id | No | Agent ID to check (empty = all agents) | |
| session_key | No | Opaque session identity you declare: a partition hint, not authentication and not a data filter. Selects which no-persist pause applies to this call. Omit to share one bucket with every caller that omits it. Full text on recall. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Extraordinarily rich disclosure beyond the readOnlyHint=false annotation: which repairs are lossy/irreversible and against which caps, that dedup_msg_id_index repair is global and needs read-write on '*', that some repairs are bounded per run with a `remaining` counter, that critical file-integrity findings are report-only, and that a report-only run never contacts the embedding backend so liveness findings are absent. No contradiction with the annotation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with purpose and every caveat is substantive, but the single dense paragraph runs very long and is hard to parse; several clauses (pre-2.5.2b1 history, cap defaults) could be tightened without losing meaning.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite no output schema, the description explains the return contract thoroughly (status verdict, issues, severity_summary, checks_run, remaining) and warns how to read status as a verdict on database contents rather than pipeline liveness. An agent has everything needed to call and interpret it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3, but the description adds real meaning: checks rejects unknown names (ok=false) and every response echoes checks_run; fix carries the lossy-repair and per-run-bound semantics; the caps that govern fix behavior are named. Only session_key is left to the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Check memory database health') and enumerates the 36-check registry in detail, so the scope is unmistakable. It does not, however, differentiate itself from the sibling deep_check, which an agent could reasonably confuse it with.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives concrete invocation guidance: set fix=true to repair, use checks to run a subset, and re-run fix until `remaining` stops decreasing. It explains how to stay agent-scoped by excluding dedup_msg_id_index. What it lacks is an explicit statement of when to reach for this tool rather than a sibling like deep_check.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
check_updateA
Report whether a newer release of this server exists, and — only if you ask — install it. The check itself runs ONCE per process start, in a background task that nothing waits on, and its verdict is cached for 24h (CPERSONA_UPDATE_CHECK_INTERVAL_SECONDS) in a file beside the database; a bare call here reads that verdict and reaches neither the network nor the disk. state is one of: ok (running the newest release) / newer (a newer final release exists — pre-releases are never proposed) / yanked (every file of the RUNNING version has been withdrawn on PyPI; reason carries the publisher's text) / unlisted (this version is not on the index at all — a development checkout; not a defect) / unknown (no check has completed, e.g. no network) / disabled. install names how this process was installed (uvx / pip / checkout / unknown) and the exact command that would update it. refresh=true performs the fetch now (3s budget) and updates the cache. apply=true runs that command as an argv list (never a shell), returning exit_code and the last 40 lines of output — supported for pip and checkout installs only; under uvx the environment is a cache entry keyed by the launch arguments, so the update belongs in your client's config (uvx cpersona@latest), and an install here would be discarded on the next launch. A checkout parked at a tag (detached HEAD) is likewise refused before anything runs, and answers with the git fetch --tags && git checkout <tag> form to use instead. Updating is NEVER automatic and never a side effect of any other call. A RESTART IS ALWAYS REQUIRED afterwards: this process keeps serving the old code until it is replaced. Unaffected by pause_persistence — an install writes no memory row, so a no-persist session can still repair a withdrawn version. Set CPERSONA_UPDATE_CHECK=false to disable the feature entirely: no fetch, no cache, no notice on recall or check_health, and this tool answers state=disabled.
| Name | Required | Description | Default |
|---|---|---|---|
| apply | No | Run the detected update command (pip / checkout installs only). Off by default; a restart is required afterwards. | |
| refresh | No | Fetch the package index now instead of reading the cached verdict (3s budget; a failure answers state=unknown). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description extensively discloses behavior beyond the minimal readOnlyHint:false annotation: the once-per-process background check, 24h cache, no network/disk on bare calls, exact state meanings, restart requirement, no automatic updates, install-method constraints, and independence from pause_persistence. This is exemplary transparency for a tool with side-effecting capabilities.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The content is dense and every major point is relevant, but it is delivered as one long monolithic paragraph with no bullets or headings, making it harder to parse. It is somewhat longer than necessary, with minor repetitions around 'never automatic' and 'only if you ask'.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description fully covers the return contract: state values, reason, install command, exit_code, and output truncation. It also addresses edge cases such as detached HEAD, uvx cache behavior, and the CPERSONA_UPDATE_CHECK disable path, so an agent has everything needed to invoke the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, but the description adds substantial meaning: refresh has a 3s budget and failure yields state=unknown; apply returns exit_code and last 40 lines, never uses a shell, and is restricted to pip/checkout installs. The description enriches both boolean parameters well beyond their schema definitions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a precise verb and resource: 'Report whether a newer release of this server exists' and explicitly scopes the install capability as opt-in. It clearly separates check, refresh, and apply behaviors, making the tool's purpose unambiguous even among many siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context for when a bare call, refresh=true, and apply=true are appropriate, and explicitly states when apply is refused (checkout at tag, uvx). It does not explicitly name alternative sibling tools, but the usage boundaries are otherwise detailed enough for an agent to decide correctly.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
declare_associationsADestructiveIdempotent
Declare associative memory after the fact: entities with aliases, and subject–predicate–object relations, recorded verbatim and walked by reconstruct. The server extracts nothing and infers nothing — coverage is exactly what was declared. anchor_ref names the record (mem:<id> / ep:<id>) the declaration is evidenced by: every entity named is recorded as mentioned by it and every relation carries it. Malformed items are reported in dropped and skipped; nothing else in the call is refused for them. retract removes relations by id and mentions by {entity, ref} — the only way a declaration leaves the store. Response: {ok, result:'declared', entities:[{id, name, created}], mentions, relations:[ids], dropped:[{item, reason}], retracted?:{relations, mentions}}. Under pause_persistence nothing is written (result:'skipped', persisted:false).
| Name | Required | Description | Default |
|---|---|---|---|
| channel | No | Memory channel the declaration belongs to. Default: '' (shared). | |
| retract | No | Declarations to remove: {relations: [relation ids], mentions: [{entity: <entity id>, ref: 'mem:<id>'}]}. Only this agent's rows are touched. | |
| agent_id | Yes | Agent identifier | |
| anchor_ref | No | The record this declaration is evidenced by: 'mem:<id>' or 'ep:<id>' of this agent. Optional. | |
| project_id | No | Project the declaration belongs to. Optional — omit or pass '' for the global pool. v2.5.1: pass '@auto' to resolve this agent's default from the server's operating context (the resolution is echoed as resolved_project_id; an unmapped agent yields operating_context_warning). bug-186: resolution requires a configured operating context. With none — the default, and equally the outcome of a sidecar that fails to parse — the sentinel is NOT resolved: it is stored and filtered as the literal project_id '@auto', resolved_project_id echoes '@auto', and no warning is raised. Read resolved_project_id before relying on the resolution. | |
| session_key | No | Opaque session identity you declare: a partition hint, not authentication and not a data filter. Selects which no-persist pause applies to this call. Omit to share one bucket with every caller that omits it. Full text on recall. | |
| associations | No | The entities and relations to declare. Malformed items are reported in `dropped` and skipped. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Goes well beyond the annotations (destructiveHint=true, idempotentHint=true) by disclosing that nothing is inferred or extracted, that malformed items land in `dropped` while the rest of the call proceeds, that retract is the only path by which a declaration leaves the store, and that pause_persistence causes a skip with persisted:false.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core purpose and behavior, and the response shape is compressed into one line. It is dense and long, but nearly every clause conveys a distinct behavioral fact rather than restating the schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, yet the description supplies the full response envelope and the retraction/pause variants. Combined with the 100%-documented schema, an agent has everything needed to invoke and interpret the call.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds meaning the schema does not: anchor_ref's propagation to every named entity and relation, and the retract/response contract. It does not fully re-explain the deeply documented project_id/session_key fields, which the schema already carries.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a precise verb (declare) and resource (associative memory: entities with aliases plus subject-predicate-object relations) and distinguishes itself from siblings by noting coverage is verbatim and read back via reconstruct. An agent can tell this apart from store/recall/reconstruct without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Frames the use case ('declare after the fact') and points to reconstruct as the consumer of what is declared, and to retract as the removal path. It does not explicitly contrast with store or recall, so the when-not guidance is implied rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
deep_checkA
Deep heuristic analysis of memory data quality. Detects issues requiring recovery or judgment (anonymous sources, short/trivial content, stale profiles, orphaned episodes, stale threshold calibration, embedding-space near-duplicate pairs as merge candidates). fix=true applies repairs for: anonymous_source, short_content. Report-only (fix is accepted and ignored): stale_profile, orphaned_episodes, calibration_staleness, near_duplicate, unnormalized_content, embedding_norm — apply those decisions via merge_memories / delete_memory / calibrate_threshold / update_profile. Use checks parameter to select specific checks.
| Name | Required | Description | Default |
|---|---|---|---|
| fix | No | Apply repairs (default: dry-run preview only) | |
| checks | No | Checks to run (empty = all). Options: anonymous_source, short_content, stale_profile, orphaned_episodes, calibration_staleness, near_duplicate, unnormalized_content, embedding_norm | |
| agent_id | Yes | Agent ID to check (required) | |
| session_key | No | Opaque session identity you declare: a partition hint, not authentication and not a data filter. Selects which no-persist pause applies to this call. Omit to share one bucket with every caller that omits it. Full text on recall. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Goes well past the readOnlyHint=false annotation by disclosing the exact mutation semantics: fix defaults to dry-run preview, and fix=true only affects two of the eight checks while the rest ignore it. This partial-write behavior is the single most important thing an agent must know before setting fix=true and it is stated plainly.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core purpose, then a parenthetical inventory of issue types, then the fix/report-only split, then follow-up tools. Dense but every clause carries information; the parenthetical enumeration pushes it toward a run-on sentence.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a four-parameter tool with no output schema, the description covers what it detects, what it mutates, what it ignores, and where to route unresolved issues. The main omission is any indication of the report shape or findings structure an agent should expect back.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3 and the schema already documents fix, checks options, and session_key. The description still adds real meaning by mapping checks to their fix behavior and clarifying that an empty checks list runs everything. It does not, however, expand on the unusual session_key partitioning hint beyond what the schema says.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb+resource: 'Deep heuristic analysis of memory data quality' with an enumerated set of detected issue classes (anonymous sources, stale profiles, orphaned episodes, near-duplicate pairs). An agent immediately knows this is a diagnostics/repair pass. It never explicitly contrasts itself with the sibling check_health, so sibling differentiation is only implicit via the check list.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Strong routing guidance: it states which checks are auto-repairable (anonymous_source, short_content) and, for each report-only check, names the sibling to apply the decision (merge_memories / delete_memory / calibrate_threshold / update_profile). No explicit 'when to call this vs check_health' statement, so it stops short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
delete_agent_dataADestructiveIdempotent
Delete ALL data (memories, profiles, episodes) for a specific agent. Used by kernel during agent deletion.
| Name | Required | Description | Default |
|---|---|---|---|
| agent_id | Yes | Agent ID whose data should be purged | |
| session_key | No | Opaque session identity you declare: a partition hint, not authentication and not a data filter. Selects which no-persist pause applies to this call. Omit to share one bucket with every caller that omits it. Full text on recall. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true and idempotentHint=true, so the safety profile is covered. The description adds genuine value by enumerating exactly what is destroyed (memories, profiles, episodes) and the scope (single agent), which the annotations do not convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences with zero padding; the destructive scope is front-loaded before the caller-context note. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an irreversible bulk-delete tool, the description pairs the destructive scope with annotations that flag destructiveness, and no output schema exists to explain. It could state irreversibility or permission requirements, but nothing critical for correct invocation is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and both parameters carry detailed schema descriptions, including a lengthy note on session_key semantics, so the schema does the heavy lifting. The description adds no parameter-level meaning, making the baseline 3 correct.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Delete) plus resource scope (ALL data for a specific agent) and enumerates what is covered: memories, profiles, episodes. This distinguishes it from the finer-grained delete_memory and delete_episode siblings, though it never names them explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
"Used by kernel during agent deletion" gives context about who calls it and when, but provides no guidance for an agent choosing between this and delete_memory/delete_episode/archive_episode, and no when-not-to-use condition or prerequisites.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
delete_episodeADestructiveIdempotent
Delete a single episode by ID. Ownership is enforced when agent_id is provided.
| Name | Required | Description | Default |
|---|---|---|---|
| agent_id | No | Agent ID for ownership verification (injected by kernel) | |
| episode_id | Yes | Episode ID to delete | |
| session_key | No | Opaque session identity you declare: a partition hint, not authentication and not a data filter. Selects which no-persist pause applies to this call. Omit to share one bucket with every caller that omits it. Full text on recall. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotations already declare destructiveHint=true, idempotentHint=true, and readOnlyHint=false, covering the safety profile. The description adds meaningful context by stating that ownership is enforced only when agent_id is provided, which affects authorization behavior. It does not explain cascading effects, but this is a valuable addition beyond structured fields.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with zero waste. The core action and ownership condition are front-loaded and concise.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive single-resource deletion with rich annotations and full schema coverage, the description provides the key ownership condition. Minor gaps exist, such as no guidance on choosing archive_episode over deletion or on data cascade, but these are not critical given the tool's complexity and the available structured data.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all three parameters thoroughly. The description restates the episode_id and agent_id ownership concept but adds no syntax, format, or constraint details beyond the schema. The session_key parameter is not mentioned in the description at all, though it is fully documented in the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Delete') and resource ('a single episode by ID'), making the action clear. It implies scoping to one episode, which helps distinguish it from sibling operations like delete_agent_data or archive_episode, but it does not explicitly name any sibling for differentiation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit guidance on when to use this tool versus alternatives such as archive_episode or delete_agent_data. The word 'single' weakly implies a non-bulk operation, but no routing conditions or exclusions are stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
delete_memoryBDestructiveIdempotent
Delete a single memory by ID. Ownership is enforced when agent_id is provided.
| Name | Required | Description | Default |
|---|---|---|---|
| agent_id | No | Agent ID for ownership verification (injected by kernel) | |
| memory_id | Yes | Memory ID to delete | |
| session_key | No | Opaque session identity you declare: a partition hint, not authentication and not a data filter. Selects which no-persist pause applies to this call. Omit to share one bucket with every caller that omits it. Full text on recall. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, idempotentHint=true, and readOnlyHint=false, so the safety profile is covered structurally. The description adds useful authorization context ('Ownership is enforced when agent_id is provided'), but does not state that the deletion is permanent/irreversible or what happens to associated data.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, front-loaded with the core action and followed by the ownership condition. Efficient with no padding, though minimally terse.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive mutation with strong annotations and full schema coverage, the description covers the core action and one behavioral constraint. It omits irreversibility, cascade effects, and failure/error behavior when ownership check fails, leaving meaningful gaps for a delete tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so all three parameters (agent_id, memory_id, session_key) are already documented in the schema, including the ownership-verification note on agent_id. The description restates the ownership rule but adds no syntax or format detail beyond the schema. Baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb+resource (delete a single memory by ID), which distinguishes it from siblings like delete_agent_data and delete_episode. However, it does not explicitly name those siblings or scope the difference beyond the singular 'single memory'.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use or when-not-to-use guidance, and no alternatives named. The only conditional is about ownership enforcement, which is behavioral rather than a usage rule.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
export_memoriesADestructive
Export memories, episodes, and profiles to a JSONL file for backup or portability.
| Name | Required | Description | Default |
|---|---|---|---|
| agent_id | Yes | Agent identifier (empty string to export all agents) | |
| output_path | Yes | File path for the JSONL output | |
| include_embeddings | No | Include embedding BLOBs as base64 (default false, usually not needed) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description describes an export operation, which is typically non-destructive. However, annotations set destructiveHint to true, implying the tool may have destructive side effects (e.g., file overwrite). The description does not disclose this, contradicting the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, straightforward sentence with no wasted words. It is front-loaded with the action and purpose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description mentions the data types exported (memories, episodes, profiles) and the format (JSONL), but does not address potential side effects like file overwriting despite the destructiveHint annotation. Given no output schema, more detail on the return value or behavior would improve completeness.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema covers all three parameters (agent_id, output_path, include_embeddings) with descriptions. The tool description does not add additional meaning beyond what the schema provides, so baseline score of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool exports memories, episodes, and profiles to a JSONL file for backup or portability. It uses a specific verb and resource, and distinguishes from siblings like import_memories.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides context for when to use the tool (backup or portability) but does not explicitly state when not to use it or mention alternatives. Sibling list makes the purpose clear enough.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_contentsARead-only
Fetch full, untrimmed content for recall preview refs. Use after a preview-tier recall to expand only the rows that matter instead of opting the whole recall out with full_content=true. Bounded twice: at most 20 refs per call, and a 40,000-character budget across the batch (2.5.4a2) that does not move when CPERSONA_MAX_CONTENT_LENGTH does. Rows are never cut to fit — when the budget is spent the remaining refs come back in deferred (absent otherwise) alongside budget_chars; re-fetch them in a second call. A single row larger than the budget is still returned in full, because this tool is the only path back to a row's complete text. RANGES: a ref may instead be an object that names part of its record -- {ref, node: i} or {ref, node: [first, last]} (inclusive) for overflow-tree nodes, e.g. the node.index of a reconstruct quote and its neighbours, or {ref, span: [start, end]} for characters. Offsets are in the stored text (a memory's content, an episode's summary without the '[Episode] ' label). The item then carries that slice as content and range = {span, content_len, and node + of when nodes were named, or block + of when blocks were}; a span end past the text is clamped and range.span says what was served. A ref may also carry revision, the digest a reconstruct quote's expand hands out for the text its offsets were measured in: a record rewritten since refuses rather than serving different characters under the same numbers. A range that cannot be served exactly is never widened to the whole row: it comes back in unresolved (absent otherwise) as {ref, reason}, reason one of invalid_range, no_current_nodes (the record has no complete node set -- short records have none, and a new long one gets them shortly after store), node_out_of_range, span_out_of_range, no_current_blocks, block_out_of_range, stale_revision. Only the slice counts against the budget.
| Name | Required | Description | Default |
|---|---|---|---|
| refs | Yes | Refs from recall messages ('mem:<id>' / 'ep:<id>'), or range objects such as {'ref': 'mem:<id>', 'node': [2, 3]} / {'ref': 'ep:<id>', 'span': [0, 800]} / {'ref': 'mem:<id>', 'block': 4, 'revision': '<from an expand>'} (max 20 per call). At most one of node / span / block | |
| agent_id | Yes | Agent identifier (ownership check — another agent's refs come back in `missing`) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations carry only readOnlyHint=true; the description shoulders the rest and does so richly: two hard bounds (20 refs, 40,000-char budget), the deliberate non-truncation policy with `deferred`/`budget_chars`, the oversized-single-row exception, staleness refusal via `revision`, and enumerated `unresolved` reasons. Nothing here contradicts the read-only annotation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
It is long, but the tool's surface genuinely is large (ranges, node vs block vs span, revisions, deferral), and the purpose is front-loaded before the mechanics. The version tag and env-var mention are minor noise, keeping this just under a 5.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description must explain returns — and it does: `deferred`, `budget_chars`, `unresolved` with enumerated reasons, `range` contents, `missing` for foreign refs, and the oversized-row pass-through. An agent has everything needed to call and interpret results correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Although schema coverage is 100%, the description adds substantially more than the schema: concrete ref formats ('mem:<id>' / 'ep:<id>'), example range objects, offset semantics (offsets are in stored text; episode summary excludes the '[Episode] ' label), the clamp behavior for spans past the end, and the meaning of `revision`. This far exceeds the schema's 'At most one of node / span / block' note.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Fetch full, untrimmed content for recall preview refs') and positions itself relative to sibling behaviors: it expands preview-tier rows rather than opting the whole recall out with full_content=true. An agent can distinguish this from recall and reconstruct without inspecting schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly names the trigger ('after a preview-tier recall') and the anti-pattern it replaces (setting full_content=true on the whole recall), steering the agent toward 'expand only the rows that matter.' The alternative path and the condition selecting it are stated, not implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_operating_contextARead-only
Read the server-served operating context (v2.5.1): the operator-owned doctrine distributed to every connected client. Without arguments returns the preview tier — context_revision, instructions_summary, project_id registry (+ enforce mode), @auto defaults, and doctrine section names. Pass section to fetch one section's full body. Read-only: the context is edited by the operator on the filesystem (~/.cpersona/operating-context.toml), never via MCP.
| Name | Required | Description | Default |
|---|---|---|---|
| section | No | Doctrine section name to fetch in full (from doctrine_sections). Empty = preview tier. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description not only aligns with the readOnlyHint annotation but adds significant context: the context is edited on the filesystem (~/.cpersona/operating-context.toml), never via MCP. This discloses the source of truth and mutation path, which annotations alone do not provide.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three concise sentences, front-loaded with the core purpose, followed by mode details and behavioral note. No wasted words. Structure is logical: what, how, important note.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given one optional parameter, no output schema, and full annotations, the description covers all necessary aspects: return types (preview tier components, full section body), usage modes, and behavioral constraints (read-only, filesystem editing). It is sufficient for an AI agent to correctly invoke this tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% with a description for the 'section' parameter. The description adds value by explaining the default behavior (preview tier) and that the section is from 'doctrine_sections'. It clarifies the parameter's effect beyond the schema's own description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool reads the 'server-served operating context' with a specific version (v2.5.1). It identifies the resource and its nature as 'operator-owned doctrine'. This is specific and distinct from sibling tools like 'get_profile' or 'get_contents'.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly states the two modes: without arguments returns the preview tier, and with a 'section' argument returns the full body. It mentions read-only and that editing is done via filesystem, not MCP. While it doesn't contrast with siblings, it provides clear usage context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_profileCRead-only
Get the current profile for an agent.
| Name | Required | Description | Default |
|---|---|---|---|
| agent_id | Yes | Agent identifier |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Description aligns with readOnlyHint by stating 'Get', but adds no further behavioral details such as error handling for missing agents, return format, or scope of the profile.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Single sentence, concise and front-loaded. Could include more detail without being overly long, but the brevity aids readability.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Adequate for a simple 1-parameter tool with no output schema. However, lack of return value description may leave the agent uncertain about the output format.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% with a clear description for 'agent_id'. The tool description does not add any additional meaning beyond the schema, so baseline score applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description 'Get the current profile for an agent' clearly states the action (get) and resource (profile) with specifier 'for an agent'. It distinguishes from sibling 'update_profile' but is slightly redundant with the tool name.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this tool vs alternatives, no prerequisites or limitations mentioned. The only implicit guidance is from the readOnlyHint annotation.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_queue_statusARead-only
Get the status of the background task queue (pending tasks, retry config).
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, so the description's addition of 'pending tasks, retry config' provides some context. However, it could be more transparent about the return format or any other behavioral traits.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence of 13 words. Every word is purposeful and concise.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple, parameterless tool with readOnlyHint annotation, the description is fairly complete. It could benefit from specifying the output structure, but given no output schema, it's adequate.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has zero parameters, so baseline is 4. The description does not add parameter details, which is acceptable since there are none.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states a specific verb 'Get' and resource 'background task queue status', and mentions what is included (pending tasks, retry config). This distinguishes it from sibling tools like check_health or list_episodes.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for checking queue status but provides no explicit guidance on when to use it versus alternatives, nor when not to use it. No sibling comparisons mentioned.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_recall_precisionARead-only
Read an agent's effective recall precision (knob 3) — the read-back companion to set_recall_precision. Returns the resolved specificity weight (beta) and its named precision level (strict / balanced / lenient, or 'custom' for a raw beta), and flags whether the value is a per-agent override or the global CPERSONA_RECALL_PRECISION default (overridden + global_precision / global_beta). Read-only: it never recalibrates and never persists, so a UI can load the current setting, let the user edit it, and write it back instead of the control being write-only.
| Name | Required | Description | Default |
|---|---|---|---|
| agent_id | Yes | Agent whose precision to read |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true. The description goes further by stating it never recalibrates or persists, and details the returned fields (beta, precision level, override flags). This adds behavioral context beyond the annotation, though it doesn't cover all edge cases.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single paragraph that front-loads the main purpose and then provides additional details. It is reasonably concise, though some sentences could be tightened. Every sentence adds value.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the simple input (1 param, no output schema), the description thoroughly explains the return value and its relationship to the global default and override behavior. It is complete for a read-only tool of this complexity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema covers the single required parameter (agent_id) with a description. The tool description does not add meaning beyond that, but since schema coverage is 100%, a baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool reads an agent's effective recall precision, identifies it as the read-back companion to set_recall_precision, and specifies it is read-only. This distinguishes it from its sibling and provides a specific verb-resource combination.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explains when to use this tool (as a read-back companion to set_recall_precision, for UI loading before editing) and implies it should be used before writing. However, it does not explicitly mention alternatives or when not to use it, though the sibling set tool is clearly the counterpart.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_session_findingsARead-only
Pull the storage-integrity findings on demand (SuperAuditor v1 pull contract, docs/SUPERAUDITOR_STANDARD.md) instead of reading them off check_health. Same detector as check_health(fix=false) over the WHOLE database, delivered as findings: each carries kind (the finding's name: a check registry name, or an escalation tier this seam mints for a runner that grades its own severity, e.g. null_embedding_pipeline_down — a tier is NOT a registry name), check (the registry name that produced it, so check_health(checks=[finding['check']]) re-runs exactly that probe) and a static per-kind severity (critical = the read contract is broken now / warn = two stored facts contradict / info = an observation). check_health's own instance verdict rides along as health_severity; a probe that raised is reported as kind check_crashed instead of failing the pull, so a partial result says which probe is missing. Read-only, never repairs. NOT free, though: the registry runs unfiltered, which includes two whole-database reads (the FTS5 integrity-check over both indexes, and PRAGMA quick_check over the file), so every pull is O(database) on a channel meant to be pulled once a session — budget it by call frequency. There is deliberately no cheap subset: choosing which probes run would be choosing which forgotten state stays forgotten. Findings are NOT filtered by agent_id or project_id — the channel surfaces forgotten state, and slicing it by the caller's bucket would hide exactly the rows that were forgotten (scope a repair with check_health(agent_id=...)). Honest caps: findings holds at most per_kind_limit rows per kind, capped_kinds names every kind that had more (observed, not inferred from count == limit), total and the counts describe the RETURNED set only, and per_kind_limit echoes the limit applied. summary restates the same trimmed set in prose (pass include_summary=false to skip paying for it). On a shared remote transport with no session_key declared the response carries identity_shared: true — this server has no session-scoped probes, so the key is a partition hint, not a filter. _meta.server_version identifies the running instance.
| Name | Required | Description | Default |
|---|---|---|---|
| session_key | No | Opaque client-declared session identity (partition hint, not authentication). Empty on a non-stdio transport marks the response identity_shared. | |
| per_kind_limit | No | Maximum findings returned per kind (default 5, minimum 1). Kinds that hit it are listed in capped_kinds. | |
| include_summary | No | Include the human-readable `summary` rendering (default true). It restates `findings` in prose — set false when machine-reading. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Adds far more than readOnlyHint provides: the O(database) cost with two named whole-database reads, the crash-as-finding behavior (kind `check_crashed`), cap semantics for per_kind_limit/capped_kinds/total describing only the returned set, identity_shared on shared transports, and the severity taxonomy. Read-only/no-repair claim is consistent with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded and informative, but heavily overstuffed — the kind/check/severity semantics are restated and parenthetically re-explained multiple times, and clauses like the tier-vs-registry-name aside add length without changing agent behavior. Structure is sound; economy is not.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the full return-value burden and does so thoroughly: fields of each finding, health_severity, capped_kinds, total, per_kind_limit, summary, identity_shared, and _meta.server_version. An agent has enough to interpret a partial or capped response correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is already 100%, but the description adds real meaning beyond it: session_key is framed as a partition hint rather than a filter and tied to identity_shared, per_kind_limit is linked to capped_kinds, and include_summary is given a cost rationale. It stops short of adding format/syntax detail the schema lacks, so it sits just above the baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Pull the storage-integrity findings on demand') and immediately positions itself against the sibling check_health, specifying it is the 'same detector as check_health(fix=false) over the WHOLE database'. An agent can distinguish it from check_health and deep_check without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit when-to-use ('instead of reading them off check_health'), when-not (no cheap subset exists, deliberately), frequency budget ('a channel meant to be pulled once a session'), and the alternative for scoped repair (check_health(agent_id=...)). This is unusually complete routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
import_memoriesADestructiveIdempotent
Import memories, episodes, and profiles from a JSONL file. Idempotent: memories deduplicate on msg_id (and on content within a project/channel), episodes on their summary within a project/channel.
| Name | Required | Description | Default |
|---|---|---|---|
| dry_run | No | Count records without writing to DB (preview mode) | |
| input_path | Yes | Path to the JSONL file to import | |
| session_key | No | Opaque session identity you declare: a partition hint, not authentication and not a data filter. Selects which no-persist pause applies to this call. Omit to share one bucket with every caller that omits it. Full text on recall. | |
| target_agent_id | No | Remap all records to this agent ID (empty to use original agent_id from file) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare write (readOnlyHint=false), idempotentHint=true, and destructiveHint=true, so the safety profile is covered. The description adds the actual deduplication mechanism (msg_id, and content within project/channel for memories; summary for episodes), which is genuinely useful behavioral detail beyond the annotations. It does not disclose what the destructive aspect overwrites, but the idempotency disclosure is the valuable part.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with purpose then idempotency detail, with zero filler. Each sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, but the description plus 100%-covered parameters and safety annotations give an agent enough to call this correctly. Minor gaps remain around conflict/overwrite behavior and auth, but the core operation and idempotency contract are covered.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the parameters (input_path, dry_run, session_key, target_agent_id) are documented in detail in the schema itself, including the opaque session_key semantics. The description adds no parameter-level meaning, so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Import) plus the exact resources (memories, episodes, profiles) and the source format (JSONL file), which clearly distinguishes it from export_memories by direction. It never names a sibling explicitly, but the operation is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no when-to-use, prerequisites, or alternative routing (e.g., vs store, migrate_channel_axis, or export_memories). The idempotency note is behavioral, not usage guidance, so an agent gets no help deciding when this tool is the right choice.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_episodesARead-only
List archived episodes for an agent (for dashboard display). bug-385: limit is clamped to 200 rows. When the caller asked for more and rows past the cap exist, the response carries budget_rows (the cap), so a capped listing can be told from one that reached the end of the data — reach the rest through export_memories or a narrower filter, not a larger limit. bug-255: within that cap the response holds an 800,000-character budget across summary and keywords together, with the same degradation and ceiling semantics as list_memories — rows past the budget that exceed the preview cap carry pure prefixes plus summary_truncated/summary_len and keywords_truncated/keywords_len; budget_chars appears iff at least one row was degraded. Their ref expands the summary via get_contents (under the row's own agent_id); a full keywords string is only available through export_data.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Max episodes to return | |
| agent_id | No | Agent identifier (empty for all agents) | |
| project_id | No | v2.4.17 γ filter. Same semantics as list_memories. v2.5.1: pass '@auto' to resolve this agent's default from the server's operating context (the resolution is echoed as resolved_project_id; an unmapped agent yields operating_context_warning). bug-186: resolution requires a configured operating context. With none — the default, and equally the outcome of a sidecar that fails to parse — the sentinel is NOT resolved: it is stored and filtered as the literal project_id '@auto', resolved_project_id echoes '@auto', and no warning is raised. Read resolved_project_id before relying on the resolution. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnlyHint annotation, the description discloses the 200-row limit clamp, the budget_rows cap signal, an 800,000-character budget across summary+keywords, conditional budget_chars emission, and per-field truncation markers (summary_truncated/summary_len, keywords_truncated/keywords_len). These are non-obvious behaviors an agent could not infer from structured fields.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
It is front-loaded with purpose before the behavioral detail, and every sentence carries actionable information. The single dense paragraph and inline bug-ticket tags make it heavier to parse than ideal, but there is little waste.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the full burden of explaining return values, and it does so thoroughly: budget_rows, budget_chars, the truncation flags and their paired length fields, plus how to recover the truncated content. Nothing needed to call or interpret the result is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds meaningful parameter-relevant meaning: limit is silently clamped to 200 and a bigger value does not help, and project_id's '@auto' resolution caveat (bug-186) is reinforced. This is genuine added value beyond the schema text.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ("List archived episodes for an agent") plus intended scope ("for dashboard display"), so the agent immediately knows the operation and its context. It is distinguishable from the neighboring list_memories, whose truncation semantics it references as a parallel case.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It routes the agent to alternatives for specific goals: reach truncated rows via export_memories or a narrower filter rather than a larger limit, expand a ref via get_contents, and get a full keywords string only through export_data. It stops short of an explicit when-to-use-this-vs-list_memories rule, so it earns a 4.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_memoriesARead-only
List recent memories for an agent (for dashboard display). bug-385: limit is clamped to 500 rows. When the caller asked for more and rows past the cap exist, the response carries budget_rows (the cap), so a capped listing can be told from one that reached the end of the data — reach the rest through export_memories or a narrower filter, not a larger limit. bug-255: within that cap the response holds a 1,000,000-character content budget. Rows are returned newest-first and none is dropped by the budget; once it is spent, later rows LONGER than the preview cap (CPERSONA_RECALL_PREVIEW_CHARS, default 500) degrade to a pure prefix with content_truncated/content_len and a ref that get_contents expands under the row's own agent_id (in an all-agents listing, pair the ref with the row's agent_id field). budget_chars appears iff at least one row was degraded. The effective ceiling is the budget plus one whole row plus the degraded rows' prefixes, so it scales with the preview cap; preview cap 0 disables trimming and the budget with it.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Max memories to return | |
| agent_id | No | Agent identifier (empty for all agents) | |
| project_id | No | v2.4.17 γ filter. Omit → no filter; '' → global pool only; 'X' → 'X' ∪ global pool. v2.5.1: pass '@auto' to resolve this agent's default from the server's operating context (the resolution is echoed as resolved_project_id; an unmapped agent yields operating_context_warning). bug-186: resolution requires a configured operating context. With none — the default, and equally the outcome of a sidecar that fails to parse — the sentinel is NOT resolved: it is stored and filtered as the literal project_id '@auto', resolved_project_id echoes '@auto', and no warning is raised. Read resolved_project_id before relying on the resolution. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With only readOnlyHint available, the description carries the full behavioral burden and does so in depth: a 500-row clamp, a 1,000,000-character content budget, newest-first ordering, no row dropped, prefix degradation of long rows past the preview cap, and the conditions that populate budget_rows/budget_chars. This is far beyond what annotations convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose is front-loaded, but the body is a dense wall of bug-tagged clauses (bug-385, bug-255, resolved_project_id, preview cap) that is hard to scan. Content is largely non-redundant, yet the run-on sentences and inline parentheticals hurt readability.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description fully enumerates the response fields that matter (budget_rows, budget_chars, content_truncated, content_len, ref) and their triggering conditions. An agent has enough to interpret a capped or degraded listing correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3, but the description adds real meaning: `limit` is clamped to 500, and the `ref` must be paired with the row's own `agent_id` in all-agents listings. `project_id` sentinel/resolution behavior is covered by the schema itself, so the description's added value is concentrated on `limit`.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('List recent memories for an agent') plus scope hint (dashboard display). It partially differentiates from siblings by routing extra data to export_memories, but does not clearly separate itself from recall/recall_with_context, which also return memories.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly warns against raising `limit` past the cap and names alternatives ('reach the rest through export_memories or a narrower filter'). This gives clear when-not and alternative guidance, though it never states when list_memories beats recall for a plain retrieval.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
lock_memoryAIdempotent
Lock a memory to prevent deletion and editing. Ownership enforced when agent_id provided.
| Name | Required | Description | Default |
|---|---|---|---|
| agent_id | No | Agent ID for ownership verification | |
| memory_id | Yes | Memory ID to lock | |
| session_key | No | Opaque session identity you declare: a partition hint, not authentication and not a data filter. Selects which no-persist pause applies to this call. Omit to share one bucket with every caller that omits it. Full text on recall. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, idempotentHint=true and destructiveHint=false, so the safety profile is covered. The description adds genuine context beyond them: the ownership verification behavior triggered when agent_id is supplied, which is a conditional authorization rule not expressed in the schema. It omits failure modes (already locked, non-owner caller), keeping it from a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, zero waste, with the primary effect front-loaded ahead of the conditional ownership note. Nothing could be trimmed without losing information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple 3-parameter, 1-required mutation tool with annotations carrying the safety profile and no output schema, the description covers the essential state change and the conditional authorization rule. What is missing is error/edge behavior, such as what happens if the memory is already locked or the caller is not the owner.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description adds a small amount of meaning for agent_id by tying it to enforced ownership, but says nothing extra about memory_id or the unusual session_key partition semantics, so it does not rise above baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb and resource ('Lock a memory') and states the concrete effect ('prevent deletion and editing'), so the agent knows exactly what state change results. It does not, however, name or contrast with the obvious sibling unlock_memory, so differentiation is left implicit.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The intended use is implied by the effect clause — call this when you want a memory protected from modification — but there is no explicit when-to-use framing, no prerequisite conditions, and no mention of unlock_memory as the inverse operation. Adequate but with clear gaps.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
merge_memoriesADestructiveIdempotent
Merge memories, episodes, and profiles from one agent into another. Atomic one-shot equivalent of export→import without intermediate files. Strategy 'skip' deduplicates by msg_id (memories) and summary (episodes).
| Name | Required | Description | Default |
|---|---|---|---|
| mode | No | Merge mode: 'copy' (preserve source) or 'move' (delete source after merge) | copy |
| dry_run | No | Preview merge without writing to DB | |
| strategy | No | Merge strategy: 'skip' (default) — skip duplicates, keep target's version | skip |
| session_key | No | Opaque session identity you declare: a partition hint, not authentication and not a data filter. Selects which no-persist pause applies to this call. Omit to share one bucket with every caller that omits it. Full text on recall. | |
| source_agent_id | Yes | Agent ID to merge FROM | |
| target_agent_id | Yes | Agent ID to merge INTO |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare write (readOnlyHint=false), idempotent, and destructive behavior, so the bar is lower. The description adds meaningful non-annotation context: atomicity, no intermediate files, and the deduplication key (msg_id for memories, summary for episodes). It does not detail what 'move' destroys beyond what the schema states, but it meaningfully supplements the safety profile.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences, front-loaded with the core action, then the differentiating equivalence, then the dedup semantics. Zero filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive, multi-entity merge with a fully documented schema and no output schema, the description covers purpose, atomicity, and dedup rules adequately. Minor gaps remain around recommended usage context (dataset size, ordering, permission prerequisites), which keeps it just short of complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3; the description goes beyond the schema's terse 'skip duplicates' by specifying that dedup is by msg_id (memories) and summary (episodes). This adds real semantic meaning for the strategy parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource (merge memories, episodes, and profiles from one agent into another) and explicitly distinguishes itself from siblings by framing it as the 'atomic one-shot equivalent of export→import'. An agent can tell it apart from export_memories/import_memories without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides clear context for when to prefer this over the export→import alternative ('without intermediate files'), which routes the agent between siblings. However it gives no explicit when-not conditions or prerequisites (e.g. permissions, same-namespace constraints).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
migrate_channel_axisA
Re-channel bridge-type memories to their concrete channel (knob2 v2 default flip prep). Memories the kernel filed under the bridge type ('discord') are rewritten to the concrete channel recovered from the stored session_id ('{channel_id}:{user_id}:{chunk}' | '{channel_id}:shared' → channel_id), so per-channel recall can match them. Non-destructive (only the channel column changes) and idempotent (re-running is a no-op once moved). dry_run=true (default) reports the recoverable count, the channels that would be recovered, and an unrecoverable bucket (channel='discord' rows with no snowflake session_id) without mutating. globalize_unrecoverable=true moves the unrecoverable bucket to channel='' (global, matched by every channel-scoped recall) so the flip orphans nothing; default false (report only).
| Name | Required | Description | Default |
|---|---|---|---|
| dry_run | No | Preview counts only, no mutation (default: true) | |
| agent_id | No | Agent ID to migrate (empty = all agents) | |
| session_key | No | Opaque session identity you declare: a partition hint, not authentication and not a data filter. Selects which no-persist pause applies to this call. Omit to share one bucket with every caller that omits it. Full text on recall. | |
| globalize_unrecoverable | No | Also move channel='discord' rows with no snowflake session_id to channel='' (global). Default false. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only give readOnlyHint=false; the description carries the rest and does it well, disclosing non-destructiveness ('only the channel column changes'), idempotence ('re-running is a no-op'), the default dry_run preview, and precisely what globalize_unrecoverable alters. This is exactly the mutation context an agent needs.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action, then progressively the mechanics, safety properties, and the two flag behaviors. Sentences are dense and jargon-heavy ('snowflake session_id', '{channel_id}:shared'), but every clause conveys information about behavior or defaults rather than filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description compensates by describing what dry_run returns (recoverable count, recovered channels, unrecoverable bucket) and what the two flags do. For a 4-param mutation tool with safety annotations, nothing an agent needs to invoke it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3, but the description adds real meaning: it explains that dry_run=true (default) reports recoverable count, channels, and an unrecoverable bucket without mutating, and that globalize_unrecoverable=true moves the unrecoverable bucket to channel='' (global). That clarifies the effect of both booleans beyond their terse schema text.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('re-channel bridge-type memories to their concrete channel') and explains the underlying mechanism (session_id '{channel_id}:{user_id}:{chunk}' parsing). No sibling tool shares this migration purpose, so the agent can immediately tell it apart from recall/store/profile tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Frames the when ('knob2 v2 default flip prep' so per-channel recall can match) and explains the default dry_run=true preview mode as the safe entry point. It doesn't name an explicit alternative tool or state when-not-to-use, but no plausible sibling overlaps, so the guidance is effectively complete.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
pause_persistenceAIdempotent
Pause write operations on this MCP server for an opt-in TTL window. While paused, every write tool — store, declare_associations, archive_episode, update_memory, delete_memory, delete_episode, delete_agent_data, lock_memory, unlock_memory, update_profile, import_memories, merge_memories, calibrate_threshold, set_recall_precision — returns a no-op response carrying persisted: false, dry_run: true and a reason (with the TTL remaining) instead of writing to the database. persisted: false is the authoritative signal: branch on it, not on an id. Where the success shape has an id, it reads "no-persist" (store, archive_episode); action-specific id keys (deleted_id / updated_id / locked_id / unlocked_id / episode_id) are blanked to null so a truthy echo cannot read as success. migrate_channel_axis is gated differently — it is forced to dry_run and reports repairs_skipped rather than returning a skipped-response, so it carries no persisted key. check_health and deep_check are not blocked but downgrade to fix=false (they answer with repairs_skipped: true). Read tools (recall, list_*, get_profile, etc.) still answer normally, except that recall suppresses its recall_count / last_recalled_at bump — a write that would otherwise move ranking state during a paused session. Blast radius follows session_key (response scope). Pass the same session_key here and on your write calls and the pause covers that key alone (scope: "session"): a session that sends a different key is neither silenced by it nor able to clear it. The key is a partition hint, not a credential — it is compared, never verified — so anyone who sends the same string shares the pause. Omit it and you arm the bucket every keyless caller shares (scope: "process") — on a streamable-HTTP deployment a single process serves every connected client, so a keyless pause silences writes for every other keyless session until resume or TTL elapse, and those sessions get no signal. Under stdio (one process per client) that bucket is the session. This affects only this MCP server (cpersona); call cscheduler's pause_persistence too if you want both paused. Use for benchmarking, AB testing, or ephemeral exploration where memory contamination must be avoided. Default TTL: 1800 seconds (30 minutes); upper bound: 86400 seconds (1 day).
| Name | Required | Description | Default |
|---|---|---|---|
| session_key | No | Opaque session identity you declare: a partition hint, not authentication and not a data filter. Selects which no-persist pause applies to this call. Omit to share one bucket with every caller that omits it. Full text on recall. | |
| ttl_seconds | No | TTL until automatic resume. Min 1, max 86400 (clamped). Default 1800. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Discloses far more than the annotations: no-op response shapes, persisted:false as the authoritative signal, id key blanking, the migrate_channel_axis special case, check_health/deep_check downgrades, and recall's suppressed ranking bump. It never states whether the pause is itself reversible/clearable beyond resume/TTL, but the safety profile is thoroughly covered.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The core action, the blast-radius warning, and the TTL bounds are ordered well and the critical session_key caveat is bolded. It is long, but nearly every clause carries operational information; only minor tightening of the response-shape enumeration is possible.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and a non-trivial side-effecting tool, the description fully compensates: it describes the return payloads (persisted, dry_run, reason, no-persist id), the read-tool carve-outs, and the scope semantics an agent needs to call it safely.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, but the description still adds real meaning: session_key is a comparison-only partition hint (not a credential), blast radius follows it, omitting it arms a shared process bucket, and TTL bounds/defaults are restated operationally.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource with scope: 'Pause write operations on this MCP server for an opt-in TTL window.' It enumerates the affected write tools by name and is unmistakably distinct from siblings like resume_persistence and persistence_status.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit use cases (benchmarking, AB testing, ephemeral exploration to avoid memory contamination), names the counter-action (resume or TTL elapse), warns that cscheduler's pause_persistence must be called separately, and spells out the when-not consequence of omitting session_key.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
persistence_statusARead-only
Report whether persistence is currently paused and the TTL remaining (in seconds). It reports the bucket session_key selects (response scope), not the server as a whole: with a session_key it answers for your session only, so paused: false here does not mean no other session is paused. Without one it reflects the bucket every keyless caller shares, which on a streamable-HTTP deployment means paused: true may have been armed by a different keyless session.
| Name | Required | Description | Default |
|---|---|---|---|
| session_key | No | Opaque session identity you declare: a partition hint, not authentication and not a data filter. Selects which no-persist pause applies to this call. Omit to share one bucket with every caller that omits it. Full text on recall. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, so the safety profile is covered. Beyond that, the description adds valuable behavioral context: what the `scope` response reflects, and the non-obvious caveat that `paused: false` for your session does not imply no other session is paused, including the keyless-HTTP scenario.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the core purpose in the first sentence, then spends the remaining text on the genuinely useful scoping caveats. It is dense but every sentence carries weight; no filler or repetition of the schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description correctly enumerates what is returned (paused flag, TTL, scope). For a single-parameter read tool with nuanced scoping, the description supplies everything an agent needs to interpret the result correctly, including the keyless-bucket edge case.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the schema already documents session_key in depth (opaque partition hint, default, maxLength). The description reinforces which bucket it selects but adds no syntax or format detail beyond the schema, so the baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (report) and resource (persistence paused state) plus the returned value (TTL in seconds). It also distinguishes itself from the broader system by clarifying it reports the session-scoped bucket rather than the server as a whole.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description thoroughly explains the two usage modes (with vs. without session_key) and their differing semantics, which is genuine guidance. However, it never names the natural alternatives (pause_persistence, resume_persistence) or states explicitly when to call this tool versus them, leaving routing to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
recallARead-only
Recall relevant memories using multi-strategy search (vector + FTS5 + keyword). To answer a question from memory, prefer reconstruct, the recommended way to read it: it returns items that quote the rows supporting them, within a character budget. Message content is returned as a preview tier by default — expand selected rows with get_contents(refs), or opt out wholesale with full_content=true. full_content is itself budgeted (200k chars per response, bug-211): rows past the budget degrade to the preview tier and the response carries full_content_budget_chars (absent when the budget never bites). 2.6 additive: a message whose content the preview cut also carries excerpt — the part of the record that matched the query, at most 800 characters (CPERSONA_RECALL_EXCERPT_CHARS), separate passages joined by ' … ' in text order — and excerpt_basis (blocks: the record's block set; lexical: divided at read time, ranked by shared words; start: the record is one block, so its start). Read the excerpt before deciding to expand a row; content stays the record's start. Absent under full_content and on rows shown whole. v2.5.2 additive: each scored message carries match_reason={signal, score, ...} where signal is the branch the quality gate keyed on (rsf > cosine > rrf; confidence only under CPERSONA_CONFIDENCE_ORDERING=legacy — from 2.6.0a7 an enabled confidence score is returned beside each row but neither orders nor gates) and the remaining keys (cosine / rrf / rsf) surface the internal per-retriever contributions present on that row; prior, when present, is the age weight that ordered the row (CPERSONA_PRIOR_AGE_RATE). Unscored rows (cascade FTS/keyword) omit match_reason. A response carrying gate_fallback=true (absent otherwise) means every candidate fell below the quality gate and the below-gate lexical matches were returned instead of an empty result — treat them as low-confidence. A response may carry suggestion (absent otherwise, at most once per session): something the server noticed that only the user can decide — that this scope has grown past the scan window with no coarse index for a time cue to search the rest. Relay its message, and run its fix only if the user agrees.
| Name | Required | Description | Default |
|---|---|---|---|
| deep | No | Deep recall — halves the quality gate (and the calibrated fused gate), so weaker matches are admitted. It also disables time and completion decay, which are inert unless CPERSONA_CONFIDENCE_ENABLED=true, and it does NOT widen the scan window (CPERSONA_MAX_MEMORIES) — deep is about how weak a match may be, not how far back the search reaches. | |
| limit | No | Per-retriever search depth, not a pure response cap: the value is handed to each retrieval channel (vector / episode FTS / keyword) as its top-K, so lowering it shrinks the candidate pool itself — rows beyond the depth are unreachable at any gate value, and score normalization / autocut operate on the smaller pool, which can also reorder what remains. Fewer rows than this may be returned. (Agent-facing cap; the library layer accepts up to the scan window for direct callers.) | |
| query | Yes | Search query (empty returns recent memories) | |
| trace | No | 2.6 recall trace: true adds `trace` to the response — which rows each stage (retrieval arms, fusion, quality gate, autocut, final order, count cut, reserved seats) kept, dropped or reordered, and why, with ranks and scores. It carries references only, never stored text, and is not stored on the server. trace_version identifies its shape. False (the default) returns the response unchanged. | |
| channel | No | Filter memories by channel (e.g. 'chat', 'discord'). Default: '' (all channels). | |
| agent_id | Yes | Agent identifier | |
| time_cue | No | 2.6: when the answer was stored, as far as you remember. Give after and/or before (a date YYYY-MM-DD, which includes that whole day, or an ISO-8601 timestamp), or ago, plus confidence. The server also searches that period. Among the rows it returns anyway, one found there moves up by at most 3 / 2 / 1 places (sure / likely / vague), and as many extra seats (3 / 2 / 1) hold the best records that search found which the answer does not already hold: one no other search reached, or one the count cut. The returned rows are those of a recall without the cue, reordered, plus at most those seats: no row is removed or re-scored, and which rows pass the quality gate does not change. If the period holds nothing it is widened once, one confidence step. likely widens the period by half its length on each side, vague by its whole length. The response then carries time_cue: the period searched, whether it was widened, how many rows moved and how many seats were used, and remainder when the period held more records than the vector search reads and no coarse index could search the rest. Pass it only when the request itself says when (a date, a month, "last week", "in the spring"); omit it when the request names no time, and never fill it with today's date or a guess. A cue whose own period starts within the last 24 hours (today, or later) is not used: the rows are those of a recall without a cue, and time_cue in the response says ignored. A cue that cannot be read returns an error and no messages. | |
| source_id | No | v2.4.20 per-user source filter. Empty (default) = no filter. Non-empty = prefix match against json_extract(source, '$.id'), e.g. 'discord:12345' to restrict to one Discord user, or 'discord:' to scope to all Discord-sourced memories. Episodes carry no per-user source tagging, so they are skipped when this is set — UNLESS channel is also set, which scopes episodes to one conversation and re-admits them. | |
| project_id | No | v2.4.17 γ filter. Omit → no filter (all projects). '' → global pool only. 'X' → 'X' bucket ∪ global pool. Threaded through cascade / RRF / vector / FTS / keyword paths. v2.5.1: pass '@auto' to resolve this agent's default from the server's operating context (the resolution is echoed as resolved_project_id; an unmapped agent yields operating_context_warning). bug-186: resolution requires a configured operating context. With none — the default, and equally the outcome of a sidecar that fails to parse — the sentinel is NOT resolved: it is stored and filtered as the literal project_id '@auto', resolved_project_id echoes '@auto', and no warning is raised. Read resolved_project_id before relying on the resolution. | |
| session_key | No | Opaque session identity you declare — a partition hint, NOT authentication. It scopes this process's per-session state: the degraded-recall advisory's "already told you" memory, and which no-persist pause applies to this call. It does NOT filter stored data (use agent_id / project_id / channel for that), and it never reaches the database. Omit it to share one bucket with every other caller that omits it, which is the behaviour that predates this parameter. | |
| full_content | No | v2.5.0 preview tier opt-out. By default message content longer than the preview cap (CPERSONA_RECALL_PREVIEW_CHARS, default 500) is returned as a pure prefix with content_truncated/content_len markers; each message's `ref` expands via get_contents. true returns full text. | |
| exclude_contents | No | Normalized content strings to exclude from results (starts-with match). Used to prevent duplication with conversation context already known to the caller. bug-399: the starts-with rule holds only at or above 32 characters. A shorter entry has to EQUAL the stored content (after the normalization this parameter already asks for: stripped and lower-cased) — a short prefix would otherwise suppress every memory beginning with it, inside the retrievers and with nothing in the response reporting the exclusion. Size entries at or above that length when you mean a prefix. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare readOnlyHint=true, so the description carries the safety/behavioral burden and discharges it richly: preview-tier defaults, the 200k full_content budget with degradation to preview, the meaning of gate_fallback, excerpt/excerpt_basis, match_reason signals, and the once-per-session suggestion semantics. This is far more than the annotation provides.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The purpose is front-loaded, but the body degenerates into versioned changelog prose ('2.6 additive', 'v2.5.2 additive', 'v2.4.20') with internal bug IDs and config-flag asides that an agent does not need to select or invoke the tool. Far longer than the decision requires.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With 12 parameters, nested objects, and no output schema, the description does explain response-side fields (full_content_budget_chars, excerpt, match_reason, gate_fallback, suggestion) that the schema cannot. It is complete enough to call correctly, though the changelog framing makes the relevant behavior harder to extract than it needs to be.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the per-parameter schema text is already dense (deep, limit, trace, time_cue, source_id, project_id, session_key). The description mostly restates or lightly extends that, adding marginal value; baseline 3 is appropriate when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening states a specific verb and resource ('Recall relevant memories') plus the retrieval mechanism (vector + FTS5 + keyword), and explicitly distinguishes this from the sibling `reconstruct`. An agent can tell the two apart without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives explicit routing ('To answer a question from memory, prefer `reconstruct`'), tells the caller how to expand previews via get_contents(refs), when to set full_content, when to pass time_cue ('only when the request itself says when... never fill it with today's date'), and how to treat gate_fallback rows. When-to-use and when-not-to-use are both present.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
recall_with_contextARead-only
Recall memories and merge with external conversation context. Automatically deduplicates, sorts chronologically, and returns a unified list. Replaces separate recall + manual merge in the caller. Content is preview-tiered by default — see recall's full_content / get_contents (full_content shares recall's 200k-char response budget, bug-211); a recalled row the preview cut carries excerpt / excerpt_basis as recall's do. Every external_context entry's content filters the recall (the caller already holds that text), but only role=user / role=assistant entries are merged into messages. When entries of other roles are present the response carries context_filter_only={roles:[...]} — those entries filtered the recall without appearing in the output, whether or not they dropped a memory this time. context_field_issues={entries:[{index, fields}]} (absent otherwise) names entries whose declared field was not a string: it was read as absent and the entry merged without it. CPERSONA_EXTERNAL_CONTEXT_MODE=reject refuses such a call instead. gate_fallback=true (absent otherwise) is forwarded from the underlying recall: every candidate fell below the quality gate and the below-gate lexical matches were returned instead of an empty result — treat them as low-confidence.
| Name | Required | Description | Default |
|---|---|---|---|
| deep | No | Deep recall — same semantics as in `recall`: halves the quality gate so weaker matches are admitted. | |
| limit | No | Per-retriever search depth for the underlying recall, not a pure response cap — same semantics as recall's limit: lowering it shrinks the candidate pool itself, not just the rows returned. (Agent-facing cap; the library layer accepts up to the scan window for direct callers.) | |
| query | Yes | Search query | |
| channel | No | Memory channel filter | |
| agent_id | Yes | Agent ID | |
| source_id | No | v2.4.20 per-user source filter — passed through to recall. Same semantics as in `recall`. | |
| project_id | No | v2.4.17 γ filter — passed through to recall. Same semantics as in `recall`. v2.5.1: pass '@auto' to resolve this agent's default from the server's operating context (the resolution is echoed as resolved_project_id; an unmapped agent yields operating_context_warning). bug-186: resolution requires a configured operating context. With none — the default, and equally the outcome of a sidecar that fails to parse — the sentinel is NOT resolved: it is stored and filtered as the literal project_id '@auto', resolved_project_id echoes '@auto', and no warning is raised. Read resolved_project_id before relying on the resolution. | |
| session_key | No | Opaque session identity you declare — a partition hint, NOT authentication. It scopes this process's per-session state: the degraded-recall advisory's "already told you" memory, and which no-persist pause applies to this call. It does NOT filter stored data (use agent_id / project_id / channel for that), and it never reaches the database. Omit it to share one bucket with every other caller that omits it, which is the behaviour that predates this parameter. | |
| full_content | No | v2.5.0 preview tier opt-out — same semantics as in `recall`. | |
| external_context | No | Conversation history entries [{role, content, name?, user_id?, timestamp?}, ...]. Every declared field is a string; one that is not is read as absent and reported in context_field_issues (CPERSONA_EXTERNAL_CONTEXT_MODE). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare readOnlyHint=true; the description adds substantial behavior beyond that: automatic dedup and chronological sort, the fact that every external_context entry filters the recall while only user/assistant roles merge, context_filter_only, context_field_issues, the CPERSONA_EXTERNAL_CONTEXT_MODE=reject refusal, and gate_fallback's low-confidence semantics. This is exactly the extra context annotations can't carry.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The purpose is front-loaded in the first sentence, but the body is a dense wall of parenthetical version tags and bug references (v2.4.20, bug-211, v2.5.1, bug-186) that add noise rather than decision value for an agent. Length is partly justified by the tool's complexity, but the versioning asides do not earn their place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries return-value burden and does so well: unified list, context_filter_only, context_field_issues, gate_fallback, and preview-tiered rows with excerpt/excerpt_basis. It defers the detailed row shape to recall, which is a minor gap rather than a blocker.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds interaction semantics the schema doesn't: full_content shares recall's 200k-char response budget, and the role-based filtering vs. merging split. Some of it duplicates the nested item descriptions (context_field_issues, role handling), keeping it from a 5.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ("Recall memories and merge with external conversation context") and immediately differentiates from the sibling: "Replaces separate recall + manual merge in the caller." An agent can distinguish this from recall without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Clear context for when to reach for it — you already hold external conversation text and want it merged with recall results — and it names the alternative it supersedes (separate recall + manual merge). It stops short of an explicit when-not-to-use or a direct routing rule against recall.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
reconstructARead-only
The recommended way to answer a question from memory (10 items unless count says otherwise). Assemble recall ITEMS from the candidate rows a recall produces: units of memory, each traceable to the canonical rows that support it. Reconstruction means select, order and assign roles -- never compose. No model is called and nothing is summarised: content quotes the item's head claim verbatim -- the parts of its record that matched, filled up to a fixed size -- and expands through head_ref via get_contents. HEAD CLAIM: the most relevant row in the item; if newer versions of that record (same message id in the same stored project) are present, their latest version. When max_evidence cuts an item, the head is kept and the most relevant remaining rows fill the rest. Stored rows are never modified. COUNT IS A CEILING, NOT A FILL TARGET AND NOT A SEARCH DEPTH: base = forced ?? requested ?? server default, effective = min(base, maximum), and 0 <= returned <= effective. Every response states effective_count and returned_count. ONE EXCEPTION, WITH THE BLOCK ARM ON: records only that arm reached are held beside the window, as in recall -- up to the block reservation, after the window's items, each marked admission='reservation' and counted in reserved_count (a held item the budget left out is counted in reserved_omitted). They never take or displace a place in the window, so returned_count may exceed effective_count by reserved_count. A RESPONSE SAYS MORE ONLY WHEN THE SERVER DID SOMETHING OTHER THAN WHAT WAS ASKED: requested_count + count_policy {source, clamped, reason} when the count was clamped or operator-forced; requested_budget + budget_policy when the budget was clamped, raised or forced; effective_budget + used_budget when the budget withheld an item or an excerpt; bounds when a bound dropped rows, was reached, or was lowered by the library ceiling; reconstruction.excluded_without_provenance when rows were excluded. A response without them was served as asked. trace=true returns the full audit every time. Fewer items than the window is a NORMAL result and carries shortfall_reason (no_relevant_evidence / below_quality_threshold / exhausted_candidates); a shortfall is never padded with duplicates, fragments, or a cluster split in two. BREADTH IS SEPARATE FROM COUNT: top_k (candidate depth), max_hops (relation hops) and max_evidence are declared independently and none is derived from count -- changing count alone does not move the candidate id set. WHAT THE RESPONSE ADMITS: bounds.omitted names a bound that DROPPED rows the tool held (max_evidence -- each cut item also counts them in claims_omitted -- or max_hops: a declared relation was left unfollowed); bounds.reached names a bound that was only MET (top_k: retrieval returned as many rows as it was allowed; max_evidence: an entity the walk reached is mentioned by more records than were read -- whether more lay beyond is not known). Both are absent when empty. quote_selection: lexical_only appears when no query embedding was available and nodes were ranked by shared trigrams alone; an item whose cut quote is merely the start of its record carries node_unavailable (no_nodes, or not_current when nodes exist but are partial or another model's). ABSENCE IS NOT A VERDICT: a response without these fields does not say its items suffice to answer, that the whole store was searched, or that the rows were checked for contradiction -- conflicts detects one narrow case only. BREADTH BEFORE DEPTH: budget bounds the characters of quoted text -- each item's content and its excerpts -- where count bounds how many items. The quoted text is one fixed sequence: every head in item order, then each item's most relevant remaining excerpt, then the next, and the response is its longest prefix that fits. An excerpt the budget cannot carry is omitted (counted in excerpts_omitted, absent when zero; its claim and ref stay); an item is dropped only when its head does not fit, with shortfall_reason budget_exhausted. Raising the budget alone never removes an item or an excerpt. When budget is omitted the default is the configured default or the window's quote sizes summed, whichever is more, so a count you name is not cut by a budget you did not set; a budget you do name is taken as given. QUOTES: content quotes the head claim and each excerpts[] entry quotes another retained claim, most relevant first; all are verbatim. The head quote is the record's passages that matched the query, taken in ranking order while they fit the item's quote size -- CPERSONA_RECONSTRUCT_QUOTE_CHARS (800) for the first CPERSONA_RECONSTRUCT_FULL_QUOTES (5) items, CPERSONA_RECONSTRUCT_TAIL_QUOTE_CHARS (400) after them -- and shown in text order, passages with text between them joined by ' … ' -- the recall excerpt's filling; ranges gives their character spans in the record, and content_len appears when the record is longer than its quote (trace=true adds quote_basis and content_truncated). A best passage longer than the size is cut and carries context_incomplete and expand ({ref, span}) for get_contents. Excerpts of the other claims -- and the head when CPERSONA_RECONSTRUCT_QUOTE_CHARS=0 -- are quoted as before 2.6 and cut as the preview tier cuts: a long record with overflow-tree nodes is quoted from the node that best matches the query (rank by embedding similarity and by shared character trigrams, fused), and node gives its index, node count and character span in the stored text; a record without nodes is quoted from its start. READ FURTHER IN STEPS, SMALLEST FIRST: a node quote is the start of a node several times its length, and an item whose quote was cut carries expand -- pass it to get_contents as it is to read the rest of that node. If that is not enough, read its neighbours with {ref, node: [index - 1, index + 1]}. Pass the bare ref, the whole record, only when the parts did not answer: a record can be tens of times a node. Nodes are read after items are chosen, so they never change which items come back or their order. No relevance score is returned. ITEM SHAPE (the same for every item): claims carries one entry per retained row, newest first, each with ref, as_of, why (the key that admitted the row; relation:<predicate> when a declared relation did; absent when the search returned the row itself, seed), hops when the relation walk reached the row, and roles when it has any -- sort by as_of for a chronological view. excerpts and excerpts_omitted are absent when empty. trace=true adds reconstruction (policy, candidate / cluster / selected counts), candidate refs, clusters and, for each record quoted by node, node_order -- its best few node indices, best first, as places to read next (an order, not a confidence). Gate fallback remains visible even when count is filled; zero count states count_zero. Retrieval degradation and update notices are delivered unchanged. If the library ceiling clamps top_k, bounds.effective_top_k reports the applied bound, including when the candidate pool is empty. independence_reason says why this is a separate item (absent when nothing joined it to another: singleton); conflicts appears only when two rows cannot be ordered. ROLE DIRECTION: roles[].role names what the REFERENCED row is to this claim (the ref is the subject, the claim is the object): the referenced episode SUPPORTS this claim, the referenced newer row SUPERSEDES it. The vocabulary is fixed at supports / supersedes / corrects / qualifies / contradicts / temporal_predecessor. The server derives supersedes (same message id, time order) and supports (episode span containment); any role word can also be DECLARED as a record -> record relation (declare_associations), whose subject is the ref. Ignore a role you do not know. BUNDLING KEYS are deterministic and never semantic: same message id within the same stored project (unknown project context cannot establish identity), containment in a candidate episode's time span, and adjacent timestamps FROM THE SAME SOURCE within the same project and channel, with the entire burst bounded by the time window (source alone is not a key -- in a single-agent store it is constant and would fold the whole pool into one item), and a declared record -> record relation between two candidates. Sharing a declared entity does not bundle. DECLARED ASSOCIATIONS (declare_associations, or associations on store) are read here and nowhere else, and change nothing when none apply: the names and aliases of entities the query mentions are added to the LEXICAL search only (the query's meaning, and so the vector search, is unchanged; the extra match is a vote, not a pass through the quality gate); and from each item's candidates the relation walk follows declared entity -> entity relations, either direction, up to max_hops, adding records that mention an entity it reached as evidence inside that item -- never as an item, never twice in one response, kept fewest hops first, then most recently declared relation, then lowest record id. This tool is additive: the recall contract is untouched.
| Name | Required | Description | Default |
|---|---|---|---|
| deep | No | Deep recall for the candidate stage -- same semantics as in `recall`. | |
| count | No | Ceiling on recall items returned -- not a fill target, not a search depth. Declare per call; omit to take the server default (10 unless configured). An operator-forced value overrides both. Requesting 5 with only 2 valid items returns 2; neither setting requires filling the window. Clamped to the server maximum, and the clamp is reported in count_policy rather than applied silently. | |
| query | Yes | Search query (empty returns recent memories) | |
| top_k | No | Candidate depth: how many rows the retrieval hands to bundling. This is the breadth knob; it is independent of `count` and is what to raise when items are missing evidence. | |
| trace | No | Include candidate refs and cluster membership for local diagnosis, and the trace of the recall it made as trace.recall; no full text is added. | |
| budget | No | Payload budget: characters of quoted text (item `content` plus `excerpts`) the response may carry. Bounds depth, where `count` bounds breadth, and breadth wins: excerpts are omitted before any item is. Omit for the server default, which is never less than one quote per item of the window; an operator-forced value overrides both; clamped to the server maximum and raised to one head quote, and budget_policy says which. | |
| channel | No | Memory channel filter | |
| agent_id | Yes | Agent identifier | |
| max_hops | No | Relation hops the walk may follow from an item's candidates through declared entity -> entity relations. 0 adds no walked evidence. A relation left unfollowed at the bound is named in bounds.omitted. At most 5, as for traverse: a larger value is lowered to it, and bounds.max_hops then states the value applied. | |
| time_cue | No | 2.6: when the answer was stored, as far as you remember — the same object recall takes (see recall's time_cue), applied to the candidate recall this reconstruction reads. The response carries time_cue when one was applied. | |
| source_id | No | Per-user source filter -- same semantics as in `recall`. | |
| project_id | No | γ filter -- same semantics as in `recall`, including the '@auto' sentinel, which resolves this agent's default from the server's operating context and echoes the resolution as resolved_project_id. With no configured operating context the sentinel is NOT resolved: it is filtered as the literal project_id '@auto'. Read resolved_project_id before relying on the resolution. | |
| session_key | No | Opaque session identity you declare: a partition hint, not authentication and not a data filter. Forwarded to the candidate recall. | |
| max_evidence | No | Maximum retained rows per item, bounding its claims and role targets. A cut is named in bounds.omitted and counted in the item's claims_omitted. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations provide only readOnlyHint=true, so the description carries the behavioral burden and does so exhaustively: no model is called, nothing is summarised, stored rows are never modified, count is a ceiling with clamp reporting (count_policy), budget precedence and excerpt-before-item dropping, shortfall_reason values, reserved_count under the block arm, and node quoting with expand handles. This is far beyond what the single annotation conveys.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The purpose statement is correctly front-loaded, but the body is a sprawling multi-thousand-token essay mixing response-field inventories, historical asides ('quoted as before 2.6'), environment-variable constants, and shouting capitalisation ('COUNT IS A CEILING', 'HEAD CLAIM', 'ABSENCE IS NOT A VERDICT'). Much is necessary spec, but the same content could be a fraction of the length with structure; sentences do not each earn their place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With 14 parameters, nested objects, and no output schema, the description must cover the response surface itself and does: effective_count/returned_count, count_policy/budget_policy, bounds.omitted vs bounds.reached, reconstruction.excluded_without_provenance, node/expand shape, and role vocabulary. Nothing an agent needs to interpret the return or call correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the per-parameter descriptions already restate much of what appears here (count as ceiling, top_k as breadth knob, budget as depth). The description's added value is largely cross-parameter: the independence of top_k/max_hops/max_evidence from count, the budget ordering rule (breadth wins), and the max_hops/window interaction. That is real added meaning, but a significant portion is duplication of the schema text.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Opens with a specific verb+resource+scope: 'The recommended way to answer a question from memory... Assemble recall ITEMS from the candidate rows a recall produces.' It explicitly delimits reconstruction ('select, order and assign roles -- never compose') and distinguishes itself from siblings recall (candidate production) and get_contents (expansion). An agent can tell exactly what it produces versus recall.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Names alternatives and the conditions selecting them: 'READ FURTHER IN STEPS, SMALLEST FIRST' routes to get_contents with the exact expand payload; the recall contract is described as producing candidates that this tool assembles. It also states when to raise top_k ('what to raise when items are missing evidence') versus count versus budget, giving explicit when/when-not guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
resume_persistenceAIdempotent
Re-enable persistence immediately, clearing this caller's active no-persist TTL. Returns was_active=true if THIS bucket was paused before the call. It clears only the bucket session_key selects (response scope): with a session_key, your own pause and no other session's; without one, the shared keyless bucket, which on a streamable-HTTP deployment re-enables writes for every other keyless session too.
| Name | Required | Description | Default |
|---|---|---|---|
| session_key | No | Opaque session identity you declare: a partition hint, not authentication and not a data filter. Selects which no-persist pause applies to this call. Omit to share one bucket with every caller that omits it. Full text on recall. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations cover idempotency (idempotentHint=true) and non-destructiveness, and the description goes well beyond them: it discloses the return value (was_active), the response `scope` field, and critically warns that omitting session_key on a streamable-HTTP deployment re-enables writes for every other keyless session. That shared-side-effect warning is exactly the kind of behavioral context annotations cannot carry.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The action is front-loaded in the first sentence, followed by return semantics and then scope caveats. It is dense but every sentence earns its place; the bolded scope clause is long but carries the highest-stakes information, so it is justified rather than padding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description must carry return-value meaning, and it does (was_active, response `scope`). Combined with the side-effect disclosure, an agent has everything needed to invoke it correctly and understand the consequences of the session_key choice.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds real semantic value beyond the schema: it explains that session_key selects which pause bucket is cleared, that with a key only your own pause is affected, and that omitting it targets a shared keyless bucket. This meaningfully clarifies the parameter's blast radius.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Re-enable persistence immediately, clearing this caller's active no-persist TTL'), which precisely names the action and its mechanism. It is unambiguously the inverse of the sibling pause_persistence, so an agent can select it without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear usage context: re-enable persistence, and the scope of what gets re-enabled depending on session_key. However, it never explicitly names pause_persistence as the counterpart or states conditions for choosing between them, so routing relies on inference from the obvious inverse relationship.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
set_recall_precisionAIdempotent
Set an agent's recall precision (knob 3) and recalibrate its quality gate. precision = strict | balanced | lenient maps to a specificity weight beta of 2.0 / 1.0 / 0.5 in the gate separation objective (sensitivity + beta*specificity): strict sits the gate higher (fewer contaminants, more misses), lenient lower (fewer misses, more contaminants). A raw beta > 0 overrides the named level; an empty precision with beta <= 0 clears the per-agent override and returns the agent to the global CPERSONA_RECALL_PRECISION default. The gate is recalibrated at the new beta immediately and persisted, so the change is live without a restart. Precision is a per-agent setting, not a per-recall argument: the gate threshold is precomputed on the separation curve at a fixed beta, so this tool recalibrates once instead.
| Name | Required | Description | Default |
|---|---|---|---|
| beta | No | Raw specificity weight; overrides the named precision when > 0. | |
| agent_id | Yes | Agent whose precision to set | |
| precision | No | strict / balanced / lenient. Empty (with beta <= 0) clears the override. | |
| session_key | No | Opaque session identity you declare: a partition hint, not authentication and not a data filter. Selects which no-persist pause applies to this call. Omit to share one bucket with every caller that omits it. Full text on recall. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare the safety profile (readOnlyHint=false, idempotentHint=true, destructiveHint=false), and the description adds genuinely new behavior: recalibration happens immediately and is persisted, the change is live without a restart, and clearing the override reverts to the global CPERSONA_RECALL_PRECISION default. It does not mention permissions, failure modes, or the returned payload, which keeps it at 4 rather than 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the action and the recalibration consequence, and the level/beta mapping earns its space given the non-obvious semantics. The final sentence restates the recalibration that was already stated up front, and the level trade-offs are repeated in both the mapping and the parenthetical, so a little trimming is possible.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with no output schema, the description covers what changes, when it takes effect, and how the override is cleared or reset to the global default, which is what an agent needs to call it correctly. It omits authorization requirements and error/edge-case behavior (e.g. invalid beta or unknown agent_id), leaving a small gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3, but the description adds real meaning beyond the schema: the numeric beta values behind each named level (2.0/1.0/0.5), the precedence rule for beta over precision, and the effect of each level on gate placement. It still does not explain session_key beyond what the schema already says.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Set an agent's recall precision') plus the side effect ('recalibrate its quality gate'), which cleanly separates it from the read-side sibling get_recall_precision and from calibrate_threshold. The '(knob 3)' label and the beta mapping make the target unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear selection rules for the levels (strict/balanced/lenient) and for the override path (raw beta > 0 beats the named level; empty precision with beta <= 0 clears the override). It also explains the per-agent vs per-recall distinction, which steers correct use. It never names a sibling alternative or an explicit when-not-to-use, so it stops short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
storeAIdempotent
Store a message in agent memory for future recall. Every response carries result — the one field to branch on: 'stored' (a new row was written; {ok:true, result:'stored', id:, embedded:}, embedded true iff a local blob was persisted or the remote index push succeeded — false under EMBEDDING_MODE=none; the response also carries truncated:true when content exceeded the length cap and was shortened, and nodes:{status:'queued'} when the text runs past the embedding window and its overflow-tree nodes were queued for construction — absent when it fits, when the embedding server cannot report tokens, or with the task queue disabled), 'skipped' (nothing written and nothing wrong: {ok:true, result:'skipped', reason:...}; the msg_id / content dedup branches echo the pre-existing row's id, the OR IGNORE fallback reason='duplicate (unique index)' omits id by design — TOCTOU seam), or 'rejected' (nothing written because the request was refused: {ok:false, result:'rejected', reason:...} — empty content, content that sanitizes to empty, or an operating-context project_id refusal, which also carries error). Note for pre-2.5.2b1 callers: ok is no longer unconditionally true, and skipped:true is gone — a rejection used to look like a success. reason is human-readable, not a stable machine token. Under pause_persistence the write is skipped (result:'skipped') and the response carries persisted:false (id:'no-persist', embedded:false) — branch on persisted to tell a paused write apart from a dedup hit.
| Name | Required | Description | Default |
|---|---|---|---|
| channel | No | Memory channel for context separation (e.g. 'chat', 'discord'). Default: '' (shared). | |
| message | Yes | ClotoMessage to store. Legacy source shapes are normalized server-side where unambiguous (e.g. lowercase type words, Rust serde externally-tagged dicts, bare 'user'/'assistant' strings); unknown shapes are stored verbatim and surfaced by check_health(invalid_source_type). | |
| agent_id | Yes | Agent identifier | |
| project_id | No | v2.4.17 isolation axis. Optional — omit or pass '' to store in the global pool. Reads via γ semantics: a recall with project_id='X' returns 'X' rows + global pool. v2.5.1: pass '@auto' to resolve this agent's default from the server's operating context (the resolution is echoed as resolved_project_id; an unmapped agent yields operating_context_warning). bug-186: resolution requires a configured operating context. With none — the default, and equally the outcome of a sidecar that fails to parse — the sentinel is NOT resolved: it is stored and filtered as the literal project_id '@auto', resolved_project_id echoes '@auto', and no warning is raised. Read resolved_project_id before relying on the resolution. | |
| session_key | No | Opaque session identity you declare: a partition hint, not authentication and not a data filter. Selects which no-persist pause applies to this call. Omit to share one bucket with every caller that omits it. Full text on recall. | |
| associations | No | Associative memory to declare alongside this call: entities the text mentions, with their aliases, and subject–predicate–object relations. Stored verbatim; the server extracts nothing and infers nothing. On store, the stored memory is recorded as mentioning every entity named here and anchors every relation. Malformed items are reported in the response's associations.dropped and skipped; the memory is stored regardless. Optional. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare readOnlyHint=false, idempotentHint=true, destructiveHint=false; the description goes far beyond, disclosing dedup behavior, exact rejection conditions (empty/sanitized-empty content, operating-context project_id refusal), truncation under the length cap, embedding-mode effects, and pause_persistence semantics. Nothing it claims conflicts with the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with a clean one-line purpose, and most content is substantive rather than padding. However the response contract is delivered as a single dense run-on with deeply nested parentheticals, which impedes scanning despite the useful information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description must carry the return contract itself — and it does exhaustively, enumerating 'stored'/'skipped'/'rejected' with their exact fields, edge cases, and a migration note. For a 6-param, nested-object tool this leaves nothing an agent needs missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents every parameter at length (project_id '@auto' resolution, session_key partition semantics, message fields). The description adds little parameter-level detail beyond that, so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening sentence gives a precise verb+resource ('Store a message in agent memory for future recall'), and 'for future recall' implicitly delineates it from retrieval siblings like recall and reconstruct. An agent can distinguish it from declare_associations, list_memories, and archive_episode without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied rather than stated: it explains that under pause_persistence the write is skipped and to 'branch on persisted', and to read resolved_project_id before relying on resolution. However, there is no explicit when-to-use/when-not guidance, and the overlap with the sibling declare_associations (the tool itself accepts associations) is never addressed.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
traverseARead-only
The neighbourhood of a declared entity, as a graph: the entity named, its aliases, the entity -> entity relations declared on it and on what they reach, up to max_hops in either direction, and the refs of the records that mention each entity. Only what was declared (declare_associations, or associations on store); nothing is inferred. No record text: expand a ref with get_contents. entity is a name or an alias, compared after normalization; when it names more than one entity this call can read (a project's and the global pool's), all are starts. ORDER: entities by hops, then by the most recently declared relation that reached them, then by id; mentions by record id; relations most recently declared first. limit bounds both the entities returned and the refs listed per entity. Response: {entity, max_hops, limit, entities:[{id, name, hops, aliases?, mentions?, mentions_omitted?}], relations:[{id, subject, predicate, object, declared_by, declared_at, anchor_ref?}] (subject/object are entity ids from entities; a relation is listed when both ends are), entities_omitted?, bounds?:{omitted:[max_hops | limit]}, reason?:'no_such_entity'}. Relation ids are what declare_associations' retract takes. Records are listed only when this call could read them: the project, channel and source_id filters apply as in recall.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Maximum entities returned, and maximum mentioning refs listed per entity. | |
| entity | Yes | A declared entity name or alias. | |
| channel | No | Memory channel filter -- same semantics as in `recall`. | |
| agent_id | Yes | Agent identifier | |
| max_hops | No | Relations to follow from the entity, in either direction. 0 returns the entity alone. | |
| source_id | No | Per-user source filter on the mentioning records -- same semantics as in `recall`. | |
| project_id | No | γ filter -- same semantics as in `recall`. v2.5.1: pass '@auto' to resolve this agent's default from the server's operating context (the resolution is echoed as resolved_project_id; an unmapped agent yields operating_context_warning). bug-186: resolution requires a configured operating context. With none — the default, and equally the outcome of a sidecar that fails to parse — the sentinel is NOT resolved: it is stored and filtered as the literal project_id '@auto', resolved_project_id echoes '@auto', and no warning is raised. Read resolved_project_id before relying on the resolution. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only include readOnlyHint. The description adds extensive behavioral detail: only declared associations are returned, nothing is inferred, no record text, deterministic ordering rules, limit semantics, access-controlled record listing, and the no_such_entity error case. No contradiction exists with readOnlyHint.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long, but the first sentence front-loads the core purpose and every later sentence adds operational specifics (ordering, response shape, access rules, retract linkage). The only mild redundancy is restating `limit` semantics already present in the schema; overall the length is justified by the tool's complexity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description fully specifies the response shape, including optional fields, omitted flags, bounds, and the no_such_entity reason. It also covers ordering, limit behavior, access filtering, and cross-references to get_contents and recall, leaving nothing critical missing for a correct call.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema documents all parameters (100% coverage), so the baseline is 3. The description adds meaningful nuance beyond the schema, especially for `entity` (normalized comparison; multiple start entities when the name maps to more than one readable entity) and clarifies that filters behave as in recall. It does not fully detail every filter's edge cases, but the added value justifies a 4.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: it traverses the neighbourhood of a declared entity as a graph of entities, aliases, relations, and mentioning record refs. It distinguishes itself from siblings by pointing to get_contents for record text and by referencing recall for filter semantics.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Clear context is given about what the tool produces, and alternatives are referenced (get_contents for record text, recall for filter semantics). However, there is no explicit 'use this when ... and not when ...' statement, so exclusions and preference rules are left mostly implicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
unlock_memoryAIdempotent
Unlock a memory to allow deletion and editing. Ownership enforced when agent_id provided.
| Name | Required | Description | Default |
|---|---|---|---|
| agent_id | No | Agent ID for ownership verification | |
| memory_id | Yes | Memory ID to unlock | |
| session_key | No | Opaque session identity you declare: a partition hint, not authentication and not a data filter. Selects which no-persist pause applies to this call. Omit to share one bucket with every caller that omits it. Full text on recall. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations cover the safety profile (idempotent, non-destructive, not read-only), and the description adds genuinely new behavior: the unlock gate that enables subsequent edits/deletes, plus the conditional 'Ownership enforced when agent_id provided' authorization rule. It omits failure behavior for non-owners and whether the unlock state persists, so not a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences with zero filler; the purpose and its consequence are front-loaded before the authorization caveat.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 3-param state-changing gate tool with no output schema, the definition covers purpose, downstream effect, and the ownership rule. Minor gaps remain around non-owner error behavior and the scope/duration of the unlock, but the essential call-time information is present.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the schema already describes agent_id as 'Agent ID for ownership verification', memory_id as the target, and session_key at length. The description's 'Ownership enforced when agent_id provided' adds only a marginal conditional nuance (what happens when it is omitted), so baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Unlock a memory') and the concrete effect ('to allow deletion and editing'), which is enough to tell it apart from the inverse sibling lock_memory without opening a schema. It stops short of naming that sibling explicitly, but the intent is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied: by saying unlocking enables deletion and editing, an agent can infer it precedes delete_memory/update_memory. There is no explicit when-to-use statement, no prerequisites, and no named alternative, so it lands at minimum-viable guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
update_memoryAIdempotent
Update memory content by ID. Rejects if memory is locked. Ownership enforced when agent_id provided. The new content passes through the same sanitizer as store: it is capped at the content length limit (the response carries truncated:true when the cap bit) and [Memory from ...] annotations are stripped, so content consisting only of those is refused rather than written as an empty row. A new text that runs past the embedding window gets nodes:{status:'queued'}, as on store.
| Name | Required | Description | Default |
|---|---|---|---|
| content | Yes | New content for the memory | |
| agent_id | No | Agent ID for ownership verification | |
| memory_id | Yes | Memory ID to update | |
| session_key | No | Opaque session identity you declare: a partition hint, not authentication and not a data filter. Selects which no-persist pause applies to this call. Omit to share one bucket with every caller that omits it. Full text on recall. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Adds substantial behavior beyond the annotations (readOnlyHint=false, idempotentHint=true, destructiveHint=false): the sanitizer passthrough, the content-length cap and the truncated:true response signal, stripping of [Memory from ...] annotations with refusal of annotation-only content, and nodes:{status:'queued'} when text exceeds the embedding window. These are non-obvious side effects an agent could not infer from structured fields.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the core action ('Update memory content by ID') before listing caveats, and every sentence carries functional detail. The later sentences are dense and somewhat run-on, but no content is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description compensates by naming the response signals (truncated:true, nodes:{status:'queued'}) and the refusal case, giving an agent everything needed to interpret results. The mutation semantics, constraints, and failure modes are all covered.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds real meaning: content is sanitized/capped rather than written verbatim, and agent_id activates ownership verification rather than being a passive identifier. memory_id's lock requirement is also surfaced, which the schema does not note.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource with an identifier ('Update memory content by ID') and anchors it against the sibling 'store' by describing shared sanitization behavior. It doesn't explicitly distinguish itself from update_profile, but the memory-vs-profile resource split is clear from the name and text.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives concrete operating conditions: rejection when the memory is locked, and ownership enforcement when agent_id is supplied. This tells the agent when the call will fail and what agent_id buys, though it doesn't state an explicit when-to-use or a named alternative to use instead.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
update_profileA
Save a pre-computed agent profile to the database. The text passes through the same sanitizer as store, against the profile's own ceiling: it is capped at 2000 characters (CPERSONA_MAX_PROFILE_LENGTH) and the response carries truncated:true when the cap bit — branch on it, the discarded remainder is not stored anywhere else.
| Name | Required | Description | Default |
|---|---|---|---|
| profile | Yes | Profile text to save (pre-computed by caller). Capped at 2000 characters (CPERSONA_MAX_PROFILE_LENGTH); the response says truncated:true when the cap cut it. | |
| agent_id | Yes | Agent identifier | |
| session_key | No | Opaque session identity you declare: a partition hint, not authentication and not a data filter. Selects which no-persist pause applies to this call. Omit to share one bucket with every caller that omits it. Full text on recall. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false and destructiveHint=false, and the description adds real behavioral value beyond them: it names the sanitizer, discloses the 2000-char ceiling, reveals the truncated:true response flag, and warns that the discarded remainder is not stored elsewhere. It omits whether an existing profile is overwritten on save, which is the one meaningful gap for a save operation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the core action, then a dense but relevant clause on cap behavior. Every clause carries information; the parenthetical constant name is the only mild filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, and the description compensates by explaining the response's truncated:true flag and its consequence. For a non-destructive write tool with full annotation and schema coverage, this is nearly complete; only overwrite/replace semantics are left unstated.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all three parameters, including the same 2000-char cap, making 3 the baseline. The description's only added nuance is that discarded text is not persisted anywhere, which is behavioral rather than parameter-meaning clarification.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ("Save a pre-computed agent profile to the database") and the 'pre-computed' qualifier tells the agent the caller supplies the text. It references the sibling 'store' as the sanitizer source but never explicitly contrasts itself with get_profile or store, so sibling differentiation is only implicit.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is a weak hint that the profile is 'pre-computed by caller', implying this is a write-only sink rather than a generator. But there is no statement of when to use this versus get_profile, store, or update_memory, and no preconditions are named. The agent must infer routing from the name alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
1 tool update
v2.6.3- Changed
recall1 field changed- changed
Input schema / properties / time_cue / descriptionPrevious value: -"2.6: when the answer was stored, as far as you remember. Give after and/or before (a date YYYY-MM-DD, which includes that whole day, or an ISO-8601 timestamp), or ago, plus confidence. The server also searches that period. Among the rows it returns anyway, one found there moves up by at most 3 / 2 / 1 places (sure / likely / vague), and as many extra seats (3 / 2 / 1) hold the best records that search found which the answer does not already hold: one no other search reached, or one the count cut. The returned rows are those of a recall without the cue, reordered, plus at most those seats: no row is removed or re-scored, and which rows pass the quality gate does not change. If the period holds nothing it is widened once, one confidence step. likely widens the period by half its length on each side, vague by its whole length. The response then carries time_cue: the period searched, whether it was widened, how many rows moved and how many seats were used. Pass it only when the request itself says when (a date, a month, \"last week\", \"in the spring\"); omit it when the request names no time, and never fill it with today's date or a guess. A cue whose own period starts within the last 24 hours (today, or later) is not used: the rows are those of a recall without a cue, and time_cue in the response says ignored. A cue that cannot be read returns an error and no messages."New value: +"2.6: when the answer was stored, as far as you remember. Give after and/or before (a date YYYY-MM-DD, which includes that whole day, or an ISO-8601 timestamp), or ago, plus confidence. The server also searches that period. Among the rows it returns anyway, one found there moves up by at most 3 / 2 / 1 places (sure / likely / vague), and as many extra seats (3 / 2 / 1) hold the best records that search found which the answer does not already hold: one no other search reached, or one the count cut. The returned rows are those of a recall without the cue, reordered, plus at most those seats: no row is removed or re-scored, and which rows pass the quality gate does not change. If the period holds nothing it is widened once, one confidence step. likely widens the period by half its length on each side, vague by its whole length. The response then carries time_cue: the period searched, whether it was widened, how many rows moved and how many seats were used, and remainder when the period held more records than the vector search reads and no coarse index could search the rest. Pass it only when the request itself says when (a date, a month, \"last week\", \"in the spring\"); omit it when the request names no time, and never fill it with today's date or a guess. A cue whose own period starts within the last 24 hours (today, or later) is not used: the rows are those of a recall without a cue, and time_cue in the response says ignored. A cue that cannot be read returns an error and no messages."
25 tool updates
v2.6.1- Changed
archive_episode1 field changed- added
Input schema / properties / session_key / maxLengthAdded value: +256
- Changed
calibrate_threshold1 field changed- added
Input schema / properties / session_key / maxLengthAdded value: +256
- Changed
check_health1 field changed- added
Input schema / properties / session_key / maxLengthAdded value: +256
- Changed
declare_associations3 fields changed- changed
Input schema / properties / associations / descriptionPrevious value: -"Associative memory to declare alongside this call: entities the text mentions, with their aliases, and subject–predicate–object relations. Stored verbatim; the server extracts nothing and infers nothing. On store, the stored memory is recorded as mentioning every entity named here and anchors every relation. Malformed items are reported in the response's associations.dropped and skipped; the memory is stored regardless. Optional."New value: +"The entities and relations to declare. Malformed items are reported in `dropped` and skipped." - changed
Input schema / properties / associations / properties / entities / descriptionPrevious value: -"Entities to register (if new) and mark as mentioned. Names are compared after normalization (NFKC, case-folded, whitespace collapsed)."New value: +"Entities to register (if new) and mark as mentioned. Names and aliases are compared after normalization (NFKC, case-folded, whitespace collapsed), so two that normalize alike name one entity." - added
Input schema / properties / session_key / maxLengthAdded value: +256
- Changed
deep_check1 field changed- added
Input schema / properties / session_key / maxLengthAdded value: +256
- Changed
delete_agent_data1 field changed- added
Input schema / properties / session_key / maxLengthAdded value: +256
- Changed
delete_episode1 field changed- added
Input schema / properties / session_key / maxLengthAdded value: +256
- Changed
delete_memory1 field changed- added
Input schema / properties / session_key / maxLengthAdded value: +256
- Changed
get_contents2 fields changed- changed
Input schema / properties / refs / descriptionPrevious value: -"Refs from recall messages ('mem:<id>' / 'ep:<id>'), or range objects such as {'ref': 'mem:<id>', 'node': [2, 3]} / {'ref': 'ep:<id>', 'span': [0, 800]} (max 20 per call)"New value: +"Refs from recall messages ('mem:<id>' / 'ep:<id>'), or range objects such as {'ref': 'mem:<id>', 'node': [2, 3]} / {'ref': 'ep:<id>', 'span': [0, 800]} / {'ref': 'mem:<id>', 'block': 4, 'revision': '<from an expand>'} (max 20 per call). At most one of node / span / block" - changed
Input schema / properties / refs / items / anyOfPrevious value: -[ - { - "type": "string" - }, - { - "properties": { - "node": { - "anyOf": [ - { - "minimum": 0, - "type": "integer" - }, - { - "items": { - "minimum": 0, - "type": "integer" - }, - "maxItems": 2, - "minItems": 2, - "type": "array" - } - ] - }, - "ref": { - "type": "string" - }, - "span": { - "items": { - "minimum": 0, - "type": "integer" - }, - "maxItems": 2, - "minItems": 2, - "type": "array" - } - }, - "required": [ - "ref" - ], - "type": "object" - } -]New value: +[ + { + "type": "string" + }, + { + "properties": { + "block": { + "anyOf": [ + { + "minimum": 0, + "type": "integer" + }, + { + "items": { + "minimum": 0, + "type": "integer" + }, + "maxItems": 2, + "minItems": 2, + "type": "array" + } + ] + }, + "node": { + "anyOf": [ + { + "minimum": 0, + "type": "integer" + }, + { + "items": { + "minimum": 0, + "type": "integer" + }, + "maxItems": 2, + "minItems": 2, + "type": "array" + } + ] + }, + "ref": { + "type": "string" + }, + "revision": { + "type": "string" + }, + "span": { + "items": { + "minimum": 0, + "type": "integer" + }, + "maxItems": 2, + "minItems": 2, + "type": "array" + } + }, + "required": [ + "ref" + ], + "type": "object" + } +]
- Changed
get_session_findings1 field changed- added
Input schema / properties / session_key / maxLengthAdded value: +256
- Changed
import_memories1 field changed- added
Input schema / properties / session_key / maxLengthAdded value: +256
- Changed
lock_memory1 field changed- added
Input schema / properties / session_key / maxLengthAdded value: +256
- Changed
merge_memories1 field changed- added
Input schema / properties / session_key / maxLengthAdded value: +256
- Changed
migrate_channel_axis1 field changed- added
Input schema / properties / session_key / maxLengthAdded value: +256
- Changed
pause_persistence1 field changed- added
Input schema / properties / session_key / maxLengthAdded value: +256
- Changed
persistence_status1 field changed- added
Input schema / properties / session_key / maxLengthAdded value: +256
- Changed
recall3 fields changed- added
Input schema / properties / session_key / maxLengthAdded value: +256 - added
Input schema / properties / time_cueAdded value: +{ + "additionalProperties": false, + "description": "2.6: when the answer was stored, as far as you remember. Give after and/or before (a date YYYY-MM-DD, which includes that whole day, or an ISO-8601 timestamp), or ago, plus confidence. The server also searches that period. Among the rows it returns anyway, one found there moves up by at most 3 / 2 / 1 places (sure / likely / vague), and as many extra seats (3 / 2 / 1) hold the best records that search found which the answer does not already hold: one no other search reached, or one the count cut. The returned rows are those of a recall without the cue, reordered, plus at most those seats: no row is removed or re-scored, and which rows pass the quality gate does not change. If the period holds nothing it is widened once, one confidence step. likely widens the period by half its length on each side, vague by its whole length. The response then carries time_cue: the period searched, whether it was widened, how many rows moved and how many seats were used. Pass it only when the request itself says when (a date, a month, \"last week\", \"in the spring\"); omit it when the request names no time, and never fill it with today's date or a guess. A cue whose own period starts within the last 24 hours (today, or later) is not used: the rows are those of a recall without a cue, and time_cue in the response says ignored. A cue that cannot be read returns an error and no messages.", + "properties": { + "after": { + "type": "string" + }, + "ago": { + "description": "Relative to now: {\"unit\": \"days\" | \"weeks\" | \"months\", \"value\": N} names the unit-long period centred N units ago (a month is 30 days); \"long_ago\" names the oldest third of what this scope holds.", + "oneOf": [ + { + "additionalProperties": false, + "properties": { + "unit": { + "enum": [ + "days", + "weeks", + "months" + ], + "type": "string" + }, + "value": { + "minimum": 0, + "type": "integer" + } + }, + "required": [ + "unit", + "value" + ], + "type": "object" + }, + { + "enum": [ + "long_ago" + ], + "type": "string" + } + ] + }, + "before": { + "type": "string" + }, + "confidence": { + "enum": [ + "sure", + "likely", + "vague" + ], + "type": "string" + } + }, + "required": [ + "confidence" + ], + "type": "object" +} - added
Input schema / properties / traceAdded value: +{ + "default": false, + "description": "2.6 recall trace: true adds `trace` to the response — which rows each stage (retrieval arms, fusion, quality gate, autocut, final order, count cut, reserved seats) kept, dropped or reordered, and why, with ranks and scores. It carries references only, never stored text, and is not stored on the server. trace_version identifies its shape. False (the default) returns the response unchanged.", + "type": "boolean" +}
- Changed
recall_with_context1 field changed- added
Input schema / properties / session_key / maxLengthAdded value: +256
- Changed
reconstruct7 fields changed- changed
Input schema / properties / budget / descriptionPrevious value: -"Payload budget: characters of quoted text (item `content` plus `excerpts`) the response may carry. Bounds depth, where `count` bounds breadth, and breadth wins: excerpts are omitted before any item is. Omit for the server default, which is never less than one quote per item of the window; an operator-forced value overrides both; clamped to the server maximum and raised to one preview-tier excerpt, and budget_policy says which."New value: +"Payload budget: characters of quoted text (item `content` plus `excerpts`) the response may carry. Bounds depth, where `count` bounds breadth, and breadth wins: excerpts are omitted before any item is. Omit for the server default, which is never less than one quote per item of the window; an operator-forced value overrides both; clamped to the server maximum and raised to one head quote, and budget_policy says which." - changed
Input schema / properties / count / descriptionPrevious value: -"Ceiling on recall items returned -- not a fill target, not a search depth. Declare per call; omit to take the server default (1 unless configured). An operator-forced value overrides both. Requesting 5 with only 2 valid items returns 2; neither setting requires filling the window. Clamped to the server maximum, and the clamp is reported in count_policy rather than applied silently."New value: +"Ceiling on recall items returned -- not a fill target, not a search depth. Declare per call; omit to take the server default (10 unless configured). An operator-forced value overrides both. Requesting 5 with only 2 valid items returns 2; neither setting requires filling the window. Clamped to the server maximum, and the clamp is reported in count_policy rather than applied silently." - changed
Input schema / properties / max_hops / descriptionPrevious value: -"Relation hops the walk may follow from an item's candidates through declared entity -> entity relations. 0 adds no walked evidence. A relation left unfollowed at the bound is named in bounds.omitted."New value: +"Relation hops the walk may follow from an item's candidates through declared entity -> entity relations. 0 adds no walked evidence. A relation left unfollowed at the bound is named in bounds.omitted. At most 5, as for traverse: a larger value is lowered to it, and bounds.max_hops then states the value applied." - added
Input schema / properties / max_hops / maximumAdded value: +5 - added
Input schema / properties / session_key / maxLengthAdded value: +256 - added
Input schema / properties / time_cueAdded value: +{ + "description": "2.6: when the answer was stored, as far as you remember — the same object recall takes (see recall's time_cue), applied to the candidate recall this reconstruction reads. The response carries time_cue when one was applied.", + "type": "object" +} - changed
Input schema / properties / trace / descriptionPrevious value: -"Include candidate refs and cluster membership for local diagnosis; no full text is added."New value: +"Include candidate refs and cluster membership for local diagnosis, and the trace of the recall it made as trace.recall; no full text is added."
- Changed
resume_persistence1 field changed- added
Input schema / properties / session_key / maxLengthAdded value: +256
- Changed
set_recall_precision1 field changed- added
Input schema / properties / session_key / maxLengthAdded value: +256
- Changed
store2 fields changed- changed
Input schema / properties / associations / properties / entities / descriptionPrevious value: -"Entities to register (if new) and mark as mentioned. Names are compared after normalization (NFKC, case-folded, whitespace collapsed)."New value: +"Entities to register (if new) and mark as mentioned. Names and aliases are compared after normalization (NFKC, case-folded, whitespace collapsed), so two that normalize alike name one entity." - added
Input schema / properties / session_key / maxLengthAdded value: +256
- Changed
unlock_memory1 field changed- added
Input schema / properties / session_key / maxLengthAdded value: +256
- Changed
update_memory1 field changed- added
Input schema / properties / session_key / maxLengthAdded value: +256
- Changed
update_profile1 field changed- added
Input schema / properties / session_key / maxLengthAdded value: +256
8 tool updates
v2.5.12- Added
declare_associations - Changed
deep_check1 field changed- changed
Input schema / properties / checks / descriptionPrevious value: -"Checks to run (empty = all). Options: anonymous_source, short_content, stale_profile, orphaned_episodes, calibration_staleness, near_duplicate"New value: +"Checks to run (empty = all). Options: anonymous_source, short_content, stale_profile, orphaned_episodes, calibration_staleness, near_duplicate, unnormalized_content, embedding_norm"
- Changed
get_contents3 fields changed- changed
Input schema / properties / refs / descriptionPrevious value: -"Refs from recall messages, e.g. ['mem:123', 'ep:45'] (max 20 per call)"New value: +"Refs from recall messages ('mem:<id>' / 'ep:<id>'), or range objects such as {'ref': 'mem:<id>', 'node': [2, 3]} / {'ref': 'ep:<id>', 'span': [0, 800]} (max 20 per call)" - added
Input schema / properties / refs / items / anyOfAdded value: +[ + { + "type": "string" + }, + { + "properties": { + "node": { + "anyOf": [ + { + "minimum": 0, + "type": "integer" + }, + { + "items": { + "minimum": 0, + "type": "integer" + }, + "maxItems": 2, + "minItems": 2, + "type": "array" + } + ] + }, + "ref": { + "type": "string" + }, + "span": { + "items": { + "minimum": 0, + "type": "integer" + }, + "maxItems": 2, + "minItems": 2, + "type": "array" + } + }, + "required": [ + "ref" + ], + "type": "object" + } +] - removed
Input schema / properties / refs / items / typeRemoved value: -"string"
- Changed
recall1 field changed- changed
Input schema / properties / exclude_contents / descriptionPrevious value: -"Normalized content strings to exclude from results (starts-with match). Used to prevent duplication with conversation context already known to the caller."New value: +"Normalized content strings to exclude from results (starts-with match). Used to prevent duplication with conversation context already known to the caller. bug-399: the starts-with rule holds only at or above 32 characters. A shorter entry has to EQUAL the stored content (after the normalization this parameter already asks for: stripped and lower-cased) — a short prefix would otherwise suppress every memory beginning with it, inside the retrievers and with nothing in the response reporting the exclusion. Size entries at or above that length when you mean a prefix."
- Changed
recall_with_context8 fields changed- changed
Input schema / properties / external_context / descriptionPrevious value: -"Conversation history entries [{role, name?, user_id?, content, timestamp?}, ...]"New value: +"Conversation history entries [{role, content, name?, user_id?, timestamp?}, ...]. Every declared field is a string; one that is not is read as absent and reported in context_field_issues (CPERSONA_EXTERNAL_CONTEXT_MODE)." - added
Input schema / properties / external_context / items / properties / content / descriptionAdded value: +"The entry's text. A string — a value that is not one is read as absent and reported in context_field_issues." - removed
Input schema / properties / external_context / items / properties / content / typeRemoved value: -"string" - added
Input schema / properties / external_context / items / properties / nameAdded value: +{ + "description": "Display label for a role=user entry; becomes source.name, and source.id when user_id is absent. Default 'User'. A string — a value that is not one is read as absent and reported in context_field_issues." +} - added
Input schema / properties / external_context / items / properties / role / descriptionAdded value: +"'user' or 'assistant'; other roles filter the recall without being merged. A string — a value that is not one is read as absent and reported in context_field_issues." - removed
Input schema / properties / external_context / items / properties / role / typeRemoved value: -"string" - added
Input schema / properties / external_context / items / properties / timestampAdded value: +{ + "description": "ISO-8601 stamp deciding where this entry lands in the merged chronology. An entry without one — or with one that names no instant — sorts ahead of every dated message. A string — a value that is not one is read as absent and reported in context_field_issues." +} - added
Input schema / properties / external_context / items / properties / user_idAdded value: +{ + "description": "Stable id for a role=user entry; becomes source.id as 'discord:<user_id>'. A string — a value that is not one is read as absent and reported in context_field_issues." +}
- Added
reconstruct - Changed
store2 fields changed- added
Input schema / properties / associationsAdded value: +{ + "description": "Associative memory to declare alongside this call: entities the text mentions, with their aliases, and subject–predicate–object relations. Stored verbatim; the server extracts nothing and infers nothing. On store, the stored memory is recorded as mentioning every entity named here and anchors every relation. Malformed items are reported in the response's associations.dropped and skipped; the memory is stored regardless. Optional.", + "properties": { + "entities": { + "description": "Entities to register (if new) and mark as mentioned. Names are compared after normalization (NFKC, case-folded, whitespace collapsed).", + "items": { + "properties": { + "aliases": { + "description": "Other names for the same entity. An alias resolves to at most one entity per scope; a second claim on it is dropped.", + "items": { + "type": "string" + }, + "type": "array" + }, + "name": { + "description": "The canonical name, kept as written.", + "type": "string" + } + }, + "required": [ + "name" + ], + "type": "object" + }, + "type": "array" + }, + "relations": { + "description": "Declared relations. subject / object are entity names (registered if new) or record refs 'mem:<id>' / 'ep:<id>' of this agent; predicate is free text, normalized. A predicate from the role vocabulary (supports, supersedes, corrects, qualifies, contradicts, temporal_predecessor) on a record → record relation is read by reconstruct as that role.", + "items": { + "properties": { + "object": { + "type": "string" + }, + "predicate": { + "type": "string" + }, + "subject": { + "type": "string" + } + }, + "required": [ + "subject", + "predicate", + "object" + ], + "type": "object" + }, + "type": "array" + } + }, + "type": "object" +} - changed
Input schema / properties / message / properties / source / typePrevious value: -"object"New value: +[ + "object", + "string", + "null" +]
- Added
traverse
23 tool updates
v2.5.10- Changed
archive_episode1 field changed- added
Input schema / properties / session_keyAdded value: +{ + "default": "", + "description": "Opaque session identity you declare: a partition hint, not authentication and not a data filter. Selects which no-persist pause applies to this call. Omit to share one bucket with every caller that omits it. Full text on recall.", + "type": "string" +}
- Changed
calibrate_threshold1 field changed- added
Input schema / properties / session_keyAdded value: +{ + "default": "", + "description": "Opaque session identity you declare: a partition hint, not authentication and not a data filter. Selects which no-persist pause applies to this call. Omit to share one bucket with every caller that omits it. Full text on recall.", + "type": "string" +}
- Changed
check_health1 field changed- added
Input schema / properties / session_keyAdded value: +{ + "default": "", + "description": "Opaque session identity you declare: a partition hint, not authentication and not a data filter. Selects which no-persist pause applies to this call. Omit to share one bucket with every caller that omits it. Full text on recall.", + "type": "string" +}
- Added
check_update - Changed
deep_check1 field changed- added
Input schema / properties / session_keyAdded value: +{ + "default": "", + "description": "Opaque session identity you declare: a partition hint, not authentication and not a data filter. Selects which no-persist pause applies to this call. Omit to share one bucket with every caller that omits it. Full text on recall.", + "type": "string" +}
- Changed
delete_agent_data1 field changed- added
Input schema / properties / session_keyAdded value: +{ + "default": "", + "description": "Opaque session identity you declare: a partition hint, not authentication and not a data filter. Selects which no-persist pause applies to this call. Omit to share one bucket with every caller that omits it. Full text on recall.", + "type": "string" +}
- Changed
delete_episode1 field changed- added
Input schema / properties / session_keyAdded value: +{ + "default": "", + "description": "Opaque session identity you declare: a partition hint, not authentication and not a data filter. Selects which no-persist pause applies to this call. Omit to share one bucket with every caller that omits it. Full text on recall.", + "type": "string" +}
- Changed
delete_memory1 field changed- added
Input schema / properties / session_keyAdded value: +{ + "default": "", + "description": "Opaque session identity you declare: a partition hint, not authentication and not a data filter. Selects which no-persist pause applies to this call. Omit to share one bucket with every caller that omits it. Full text on recall.", + "type": "string" +}
- Added
get_session_findings - Changed
import_memories1 field changed- added
Input schema / properties / session_keyAdded value: +{ + "default": "", + "description": "Opaque session identity you declare: a partition hint, not authentication and not a data filter. Selects which no-persist pause applies to this call. Omit to share one bucket with every caller that omits it. Full text on recall.", + "type": "string" +}
- Changed
lock_memory1 field changed- added
Input schema / properties / session_keyAdded value: +{ + "default": "", + "description": "Opaque session identity you declare: a partition hint, not authentication and not a data filter. Selects which no-persist pause applies to this call. Omit to share one bucket with every caller that omits it. Full text on recall.", + "type": "string" +}
- Changed
merge_memories1 field changed- added
Input schema / properties / session_keyAdded value: +{ + "default": "", + "description": "Opaque session identity you declare: a partition hint, not authentication and not a data filter. Selects which no-persist pause applies to this call. Omit to share one bucket with every caller that omits it. Full text on recall.", + "type": "string" +}
- Changed
migrate_channel_axis1 field changed- added
Input schema / properties / session_keyAdded value: +{ + "default": "", + "description": "Opaque session identity you declare: a partition hint, not authentication and not a data filter. Selects which no-persist pause applies to this call. Omit to share one bucket with every caller that omits it. Full text on recall.", + "type": "string" +}
- Changed
pause_persistence1 field changed- added
Input schema / properties / session_keyAdded value: +{ + "default": "", + "description": "Opaque session identity you declare: a partition hint, not authentication and not a data filter. Selects which no-persist pause applies to this call. Omit to share one bucket with every caller that omits it. Full text on recall.", + "type": "string" +}
- Changed
persistence_status1 field changed- added
Input schema / properties / session_keyAdded value: +{ + "default": "", + "description": "Opaque session identity you declare: a partition hint, not authentication and not a data filter. Selects which no-persist pause applies to this call. Omit to share one bucket with every caller that omits it. Full text on recall.", + "type": "string" +}
- Changed
recall2 fields changed- changed
Input schema / properties / limit / descriptionPrevious value: -"Per-retriever search depth, not a pure response cap (CSC #716): the value is handed to each retrieval channel (vector / episode FTS / keyword) as its top-K, so lowering it shrinks the candidate pool itself — rows beyond the depth are unreachable at any gate value, and score normalization / autocut operate on the smaller pool, which can also reorder what remains. Fewer rows than this may be returned. (Agent-facing cap; the library layer accepts up to the scan window for direct callers.)"New value: +"Per-retriever search depth, not a pure response cap: the value is handed to each retrieval channel (vector / episode FTS / keyword) as its top-K, so lowering it shrinks the candidate pool itself — rows beyond the depth are unreachable at any gate value, and score normalization / autocut operate on the smaller pool, which can also reorder what remains. Fewer rows than this may be returned. (Agent-facing cap; the library layer accepts up to the scan window for direct callers.)" - added
Input schema / properties / session_keyAdded value: +{ + "default": "", + "description": "Opaque session identity you declare — a partition hint, NOT authentication. It scopes this process's per-session state: the degraded-recall advisory's \"already told you\" memory, and which no-persist pause applies to this call. It does NOT filter stored data (use agent_id / project_id / channel for that), and it never reaches the database. Omit it to share one bucket with every other caller that omits it, which is the behaviour that predates this parameter.", + "type": "string" +}
- Changed
recall_with_context2 fields changed- changed
Input schema / properties / limit / descriptionPrevious value: -"Per-retriever search depth for the underlying recall, not a pure response cap — same semantics as recall's limit (CSC #716): lowering it shrinks the candidate pool itself, not just the rows returned. (Agent-facing cap; the library layer accepts up to the scan window for direct callers.)"New value: +"Per-retriever search depth for the underlying recall, not a pure response cap — same semantics as recall's limit: lowering it shrinks the candidate pool itself, not just the rows returned. (Agent-facing cap; the library layer accepts up to the scan window for direct callers.)" - added
Input schema / properties / session_keyAdded value: +{ + "default": "", + "description": "Opaque session identity you declare — a partition hint, NOT authentication. It scopes this process's per-session state: the degraded-recall advisory's \"already told you\" memory, and which no-persist pause applies to this call. It does NOT filter stored data (use agent_id / project_id / channel for that), and it never reaches the database. Omit it to share one bucket with every other caller that omits it, which is the behaviour that predates this parameter.", + "type": "string" +}
- Changed
resume_persistence1 field changed- added
Input schema / properties / session_keyAdded value: +{ + "default": "", + "description": "Opaque session identity you declare: a partition hint, not authentication and not a data filter. Selects which no-persist pause applies to this call. Omit to share one bucket with every caller that omits it. Full text on recall.", + "type": "string" +}
- Changed
set_recall_precision1 field changed- added
Input schema / properties / session_keyAdded value: +{ + "default": "", + "description": "Opaque session identity you declare: a partition hint, not authentication and not a data filter. Selects which no-persist pause applies to this call. Omit to share one bucket with every caller that omits it. Full text on recall.", + "type": "string" +}
- Changed
store1 field changed- added
Input schema / properties / session_keyAdded value: +{ + "default": "", + "description": "Opaque session identity you declare: a partition hint, not authentication and not a data filter. Selects which no-persist pause applies to this call. Omit to share one bucket with every caller that omits it. Full text on recall.", + "type": "string" +}
- Changed
unlock_memory1 field changed- added
Input schema / properties / session_keyAdded value: +{ + "default": "", + "description": "Opaque session identity you declare: a partition hint, not authentication and not a data filter. Selects which no-persist pause applies to this call. Omit to share one bucket with every caller that omits it. Full text on recall.", + "type": "string" +}
- Changed
update_memory1 field changed- added
Input schema / properties / session_keyAdded value: +{ + "default": "", + "description": "Opaque session identity you declare: a partition hint, not authentication and not a data filter. Selects which no-persist pause applies to this call. Omit to share one bucket with every caller that omits it. Full text on recall.", + "type": "string" +}
- Changed
update_profile1 field changed- added
Input schema / properties / session_keyAdded value: +{ + "default": "", + "description": "Opaque session identity you declare: a partition hint, not authentication and not a data filter. Selects which no-persist pause applies to this call. Omit to share one bucket with every caller that omits it. Full text on recall.", + "type": "string" +}
2 tool updates
v2.5.6- Changed
recall1 field changed- changed
Input schema / properties / deep / descriptionPrevious value: -"Deep recall — disable time and completion decay for exhaustive search"New value: +"Deep recall — halves the quality gate (and the calibrated fused gate), so weaker matches are admitted. It also disables time and completion decay, which are inert unless CPERSONA_CONFIDENCE_ENABLED=true, and it does NOT widen the scan window (CPERSONA_MAX_MEMORIES) — deep is about how weak a match may be, not how far back the search reaches."
- Changed
recall_with_context1 field changed- changed
Input schema / properties / deep / descriptionPrevious value: -"Disable time decay"New value: +"Deep recall — same semantics as in `recall`: halves the quality gate so weaker matches are admitted."
5 tool updates
v2.5.4- Changed
calibrate_threshold1 field changed- added
Input schema / properties / method / enumAdded value: +[ + "separation", + "percentile", + "zscore" +]
- Changed
recall1 field changed- changed
Input schema / properties / limit / descriptionPrevious value: -"Max memories to return (agent-facing cap; the library layer accepts up to the scan window for direct callers)"New value: +"Per-retriever search depth, not a pure response cap (CSC #716): the value is handed to each retrieval channel (vector / episode FTS / keyword) as its top-K, so lowering it shrinks the candidate pool itself — rows beyond the depth are unreachable at any gate value, and score normalization / autocut operate on the smaller pool, which can also reorder what remains. Fewer rows than this may be returned. (Agent-facing cap; the library layer accepts up to the scan window for direct callers.)"
- Changed
recall_with_context1 field changed- changed
Input schema / properties / limit / descriptionPrevious value: -"Max recalled memories (agent-facing cap; the library layer accepts up to the scan window for direct callers)"New value: +"Per-retriever search depth for the underlying recall, not a pure response cap — same semantics as recall's limit (CSC #716): lowering it shrinks the candidate pool itself, not just the rows returned. (Agent-facing cap; the library layer accepts up to the scan window for direct callers.)"
- Changed
store2 fields changed- changed
Input schema / properties / message / properties / source / properties / type / descriptionPrevious value: -"Producer role — this enum IS the contract; send one of these. Legacy producers that cannot are folded server-side at the write seam ('ai' / 'assistant' are normalized to 'Agent'; 'session' is normalized to 'System' (type words are matched case-insensitively)), and shapes outside that table are stored verbatim for check_health(invalid_source_type) to surface."New value: +"Producer role — send one of 'User', 'Agent', 'System'. Legacy producers that cannot are folded server-side at the write seam ('ai' / 'assistant' are normalized to 'Agent'; 'session' is normalized to 'System' (type words are matched case-insensitively)), and shapes outside that table are stored verbatim for check_health(invalid_source_type) to surface." - removed
Input schema / properties / message / properties / source / properties / type / enumRemoved value: -[ - "User", - "Agent", - "System" -]
- Changed
update_profile1 field changed- changed
Input schema / properties / profile / descriptionPrevious value: -"Profile text to save (pre-computed by caller)"New value: +"Profile text to save (pre-computed by caller). Capped at 2000 characters (CPERSONA_MAX_PROFILE_LENGTH); the response says truncated:true when the cap cut it."
17 tool updates
v2.5.2- Changed
archive_episode1 field changed- changed
Input schema / properties / project_id / descriptionPrevious value: -"v2.4.17 isolation axis. Omit or pass '' for the global pool."New value: +"v2.4.17 isolation axis. Omit or pass '' for the global pool. v2.5.1: pass '@auto' to resolve this agent's default from the server's operating context (the resolution is echoed as resolved_project_id; an unmapped agent yields operating_context_warning). bug-186: resolution requires a configured operating context. With none — the default, and equally the outcome of a sidecar that fails to parse — the sentinel is NOT resolved: it is stored and filtered as the literal project_id '@auto', resolved_project_id echoes '@auto', and no warning is raised. Read resolved_project_id before relying on the resolution."
- Added
calibrate_threshold - Added
check_health - Added
delete_episode - Added
delete_memory - Added
export_memories - Added
get_operating_context - Added
get_recall_precision - Changed
list_episodes1 field changed- changed
Input schema / properties / project_id / descriptionPrevious value: -"v2.4.17 γ filter. Same semantics as list_memories."New value: +"v2.4.17 γ filter. Same semantics as list_memories. v2.5.1: pass '@auto' to resolve this agent's default from the server's operating context (the resolution is echoed as resolved_project_id; an unmapped agent yields operating_context_warning). bug-186: resolution requires a configured operating context. With none — the default, and equally the outcome of a sidecar that fails to parse — the sentinel is NOT resolved: it is stored and filtered as the literal project_id '@auto', resolved_project_id echoes '@auto', and no warning is raised. Read resolved_project_id before relying on the resolution."
- Changed
list_memories1 field changed- changed
Input schema / properties / project_id / descriptionPrevious value: -"v2.4.17 γ filter. Omit → no filter; '' → global pool only; 'X' → 'X' ∪ global pool."New value: +"v2.4.17 γ filter. Omit → no filter; '' → global pool only; 'X' → 'X' ∪ global pool. v2.5.1: pass '@auto' to resolve this agent's default from the server's operating context (the resolution is echoed as resolved_project_id; an unmapped agent yields operating_context_warning). bug-186: resolution requires a configured operating context. With none — the default, and equally the outcome of a sidecar that fails to parse — the sentinel is NOT resolved: it is stored and filtered as the literal project_id '@auto', resolved_project_id echoes '@auto', and no warning is raised. Read resolved_project_id before relying on the resolution."
- Added
pause_persistence - Added
recall - Added
recall_with_context - Added
resume_persistence - Changed
store4 fields changed- changed
Input schema / properties / message / properties / content / descriptionPrevious value: -"The text to store. Empty content is skipped."New value: +"The text to store. Content that is empty — or that sanitizes to empty — is refused with ok:false, result:'rejected'." - changed
Input schema / properties / message / properties / metadata / descriptionPrevious value: -"Free-form JSON object for producer-specific context. Empty when unused."New value: +"Free-form JSON object for producer-specific context. Empty when unused. Serialised size is capped at 8000 characters (same cap for source); an oversized field is refused with result='rejected' rather than truncated, because a truncated JSON document is not a JSON document." - changed
Input schema / properties / message / properties / source / properties / type / descriptionPrevious value: -"Producer role. 'Assistant' / 'ai' are normalized to 'Agent'; 'session' is normalized to 'System'."New value: +"Producer role — this enum IS the contract; send one of these. Legacy producers that cannot are folded server-side at the write seam ('ai' / 'assistant' are normalized to 'Agent'; 'session' is normalized to 'System' (type words are matched case-insensitively)), and shapes outside that table are stored verbatim for check_health(invalid_source_type) to surface." - changed
Input schema / properties / project_id / descriptionPrevious value: -"v2.4.17 isolation axis. Optional — omit or pass '' to store in the global pool. Reads via γ semantics: a recall with project_id='X' returns 'X' rows + global pool. v2.5.1: pass '@auto' to resolve this agent's default from the server's operating context (echoed as resolved_project_id)."New value: +"v2.4.17 isolation axis. Optional — omit or pass '' to store in the global pool. Reads via γ semantics: a recall with project_id='X' returns 'X' rows + global pool. v2.5.1: pass '@auto' to resolve this agent's default from the server's operating context (the resolution is echoed as resolved_project_id; an unmapped agent yields operating_context_warning). bug-186: resolution requires a configured operating context. With none — the default, and equally the outcome of a sidecar that fails to parse — the sentinel is NOT resolved: it is stored and filtered as the literal project_id '@auto', resolved_project_id echoes '@auto', and no warning is raised. Read resolved_project_id before relying on the resolution."
- Added
unlock_memory - Added
update_profile
15 tool updates
v2.5.1- Changed
archive_episode1 field changed- changed
Input schema / properties / history / descriptionPrevious value: -"Original conversation messages (used for timestamp extraction and embedding)"New value: +"Original conversation messages (used for start/end timestamp extraction; the episode embedding is computed from summary)"
- Removed
calibrate_threshold - Removed
check_health - Removed
delete_episode - Removed
delete_memory - Removed
export_memories - Added
get_contents - Removed
get_recall_precision - Removed
pause_persistence - Removed
recall - Removed
recall_with_context - Removed
resume_persistence - Changed
store3 fields changed- changed
Input schema / properties / message / descriptionPrevious value: -"ClotoMessage to store (id, content, source, timestamp, metadata)"New value: +"ClotoMessage to store. Legacy source shapes are normalized server-side where unambiguous (e.g. lowercase type words, Rust serde externally-tagged dicts, bare 'user'/'assistant' strings); unknown shapes are stored verbatim and surfaced by check_health(invalid_source_type)." - added
Input schema / properties / message / propertiesAdded value: +{ + "content": { + "description": "The text to store. Empty content is skipped.", + "type": "string" + }, + "id": { + "description": "Caller-supplied message id used for msg_id-based dedup (γ-project-scoped). Optional.", + "type": "string" + }, + "metadata": { + "description": "Free-form JSON object for producer-specific context. Empty when unused.", + "type": "object" + }, + "source": { + "description": "Attribution of who produced the content. Canonical shape is {type, id, name}. Type is the discriminator; id / name identify the concrete producer. Store null / empty {} only when the producer is genuinely unknown. A null source is normalized to {} at the write seam, so both persist (and recall) as the anonymous {}.", + "properties": { + "id": { + "description": "Stable producer id (e.g. discord user id, agent id). Empty when anonymous.", + "type": "string" + }, + "name": { + "description": "Human-readable label for display. Empty when unknown.", + "type": "string" + }, + "type": { + "description": "Producer role. 'Assistant' / 'ai' are normalized to 'Agent'; 'session' is normalized to 'System'.", + "enum": [ + "User", + "Agent", + "System" + ], + "type": "string" + } + }, + "type": "object" + }, + "timestamp": { + "description": "UTC ISO-8601 timestamp with offset (e.g. '2026-07-22T12:00:00+00:00'). Defaults to server-time UTC when omitted. Aware non-UTC offsets are accepted; naive strings are surfaced by check_health(timestamp_format_drift).", + "type": "string" + } +} - changed
Input schema / properties / project_id / descriptionPrevious value: -"v2.4.17 isolation axis. Optional — omit or pass '' to store in the global pool. Reads via γ semantics: a recall with project_id='X' returns 'X' rows + global pool."New value: +"v2.4.17 isolation axis. Optional — omit or pass '' to store in the global pool. Reads via γ semantics: a recall with project_id='X' returns 'X' rows + global pool. v2.5.1: pass '@auto' to resolve this agent's default from the server's operating context (echoed as resolved_project_id)."
- Removed
unlock_memory - Removed
update_profile
2 tool updates
v2.4.37- Changed
check_health1 field changed- added
Input schema / properties / checksAdded value: +{ + "description": "Registry check names to run (empty = all). See cpersona.checks.HEALTH_CHECK_NAMES.", + "items": { + "type": "string" + }, + "type": "array" +}
- Changed
deep_check1 field changed- changed
Input schema / properties / checks / descriptionPrevious value: -"Checks to run (empty = all). Options: anonymous_source, short_content, stale_profile, orphaned_episodes"New value: +"Checks to run (empty = all). Options: anonymous_source, short_content, stale_profile, orphaned_episodes, calibration_staleness, near_duplicate"
13 tool updates
v2.4.34- Changed
archive_episode2 fields changed- added
Input schema / properties / channelAdded value: +{ + "description": "v2.4.22 conversation-channel tag (e.g. a Discord channel id). Default '' (= unscoped). Channel-scoped recall returns episodes whose channel matches; this powers the per-channel episodic loop.", + "type": "string" +} - added
Input schema / properties / project_idAdded value: +{ + "description": "v2.4.17 isolation axis. Omit or pass '' for the global pool.", + "type": "string" +}
- Changed
calibrate_threshold3 fields changed- added
Input schema / properties / methodAdded value: +{ + "description": "'percentile' (default), 'zscore', or 'separation' (two-population, learns the operating point from null vs nearest-neighbour positives)", + "type": "string" +} - added
Input schema / properties / percentileAdded value: +{ + "description": "Null-distribution quantile for method='percentile' (default: 0.95, higher = stricter)", + "type": "number" +} - changed
Input schema / properties / z_factor / descriptionPrevious value: -"Z-score multiplier (default: 1.0, higher = stricter)"New value: +"Z-score multiplier for method='zscore' (default: 1.0, higher = stricter)"
- Added
get_recall_precision - Changed
list_episodes1 field changed- added
Input schema / properties / project_idAdded value: +{ + "description": "v2.4.17 γ filter. Same semantics as list_memories.", + "type": "string" +}
- Changed
list_memories1 field changed- added
Input schema / properties / project_idAdded value: +{ + "description": "v2.4.17 γ filter. Omit → no filter; '' → global pool only; 'X' → 'X' ∪ global pool.", + "type": "string" +}
- Added
migrate_channel_axis - Added
pause_persistence - Added
persistence_status - Changed
recall2 fields changed- added
Input schema / properties / project_idAdded value: +{ + "description": "v2.4.17 γ filter. Omit → no filter (all projects). '' → global pool only. 'X' → 'X' bucket ∪ global pool. Threaded through cascade / RRF / vector / FTS / keyword paths.", + "type": "string" +} - added
Input schema / properties / source_idAdded value: +{ + "default": "", + "description": "v2.4.20 per-user source filter. Empty (default) = no filter. Non-empty = prefix match against json_extract(source, '$.id'), e.g. 'discord:12345' to restrict to one Discord user, or 'discord:' to scope to all Discord-sourced memories. Episodes are skipped when set (no per-user source tagging).", + "type": "string" +}
- Changed
recall_with_context2 fields changed- added
Input schema / properties / project_idAdded value: +{ + "description": "v2.4.17 γ filter — passed through to recall. Same semantics as in `recall`.", + "type": "string" +} - added
Input schema / properties / source_idAdded value: +{ + "default": "", + "description": "v2.4.20 per-user source filter — passed through to recall. Same semantics as in `recall`.", + "type": "string" +}
- Added
resume_persistence - Added
set_recall_precision - Changed
store1 field changed- added
Input schema / properties / project_idAdded value: +{ + "description": "v2.4.17 isolation axis. Optional — omit or pass '' to store in the global pool. Reads via γ semantics: a recall with project_id='X' returns 'X' rows + global pool.", + "type": "string" +}
6 tool updates
v2.4.10- Added
deep_check - Added
lock_memory - Changed
recall1 field changed- added
Input schema / properties / exclude_contentsAdded value: +{ + "description": "Normalized content strings to exclude from results (starts-with match). Used to prevent duplication with conversation context already known to the caller.", + "items": { + "type": "string" + }, + "type": "array" +}
- Added
recall_with_context - Added
unlock_memory - Added
update_memory
16 tool updates
v0.1.0- First observed
archive_episode - First observed
calibrate_threshold - First observed
check_health - First observed
delete_agent_data - First observed
delete_episode - First observed
delete_memory - First observed
export_memories - First observed
get_profile - First observed
get_queue_status - First observed
import_memories - First observed
list_episodes - First observed
list_memories - First observed
merge_memories - First observed
recall - First observed
store - First observed
update_profile
TDQS
Scored across 34 tools
Most tools target clearly distinct operations, and the very detailed descriptions carefully separate recall/reconstruct/recall_with_context and check_health/get_session_findings/deep_check. However, the recall variants and the two overlapping health-reporting seams mean an agent must read carefully to choose correctly.
All tool names use consistent snake_case, overwhelmingly in a verb_noun pattern (store, recall, get_profile, delete_memory, pause_persistence, set_recall_precision). There is no camelCase or mixed-convention drift, and the few bare verbs (recall, store) remain readable and conventional.
34 tools is well above the 3–15 sweet spot and past the 25-tool threshold, making the surface heavy for an agent to navigate even if most tools have distinct purposes. Several capabilities are split into read/write/status trios (pause/resume/persistence_status, set/get_recall_precision), which adds further bulk.
The surface covers memory CRUD, episodes, profiles, declared associations, recall/reconstruction, health checks, persistence control, precision tuning, import/export/merge, migration, and update checks — a broad lifecycle. Minor gaps remain, such as no standalone delete_profile, but core workflows are well covered.
Maintenance
Related MCP Connectors
Persistent memory, hybrid search and a goal graph for AI agents, over stdio or remote HTTP.
Persistent, outcome-grounded episodic memory for Claude. 14ms CPU retrieval, no GPU, no vector DB.
Persistent memory for AI agents. Semantic search, memory graph, W3C DID identity.
Persistent memory for AI agents across Claude, ChatGPT and any MCP client.
Related MCP Servers
- AlicenseNot gradedqualityAmaintenancePersistent AI memory server with 3-layer hybrid search (vector + FTS5 + keyword), confidence scoring via Reciprocal Rank Fusion, episodic/profile memory, and 16 tools. Zero LLM dependency. Works standalone with Claude Desktop and Claude Code. MIT licensed.3Business Source 1.1
- AlicenseAqualityCmaintenanceLocal-first memory for Claude Code and any MCP client: hybrid vector + keyword search and a bi-temporal knowledge graph in one SQLite file. Local embeddings, no API key, $0/token.51168 npm2PolyForm Noncommercial 1.0.0
- FlicenseNot gradedqualityDmaintenanceA lightweight MCP memory server built on SQLite + FTS5, providing cross-session long-term memory for Claude Code.-

Loreofficial
AlicenseBqualityAmaintenanceCross-agent memory server: hybrid (vector + full-text + knowledge-graph) recall, per-user private/shared visibility, write-side secret/PII redaction, and as-of-date bi-temporal fact queries. Auto-injection hooks for Claude Code / Cursor / Codex; self-hostable on PostgreSQL or SQLite.4657 PyPI8MIT