recall
Search stored agent memories with vector, FTS5, and keyword strategies to retrieve relevant context for answering questions.
Instructions
Recall relevant memories using multi-strategy search (vector + FTS5 + keyword). To answer a question from memory, prefer reconstruct, the recommended way to read it: it returns items that quote the rows supporting them, within a character budget. Message content is returned as a preview tier by default — expand selected rows with get_contents(refs), or opt out wholesale with full_content=true. full_content is itself budgeted (200k chars per response, bug-211): rows past the budget degrade to the preview tier and the response carries full_content_budget_chars (absent when the budget never bites). 2.6 additive: a message whose content the preview cut also carries excerpt — the part of the record that matched the query, at most 800 characters (CPERSONA_RECALL_EXCERPT_CHARS), separate passages joined by ' … ' in text order — and excerpt_basis (blocks: the record's block set; lexical: divided at read time, ranked by shared words; start: the record is one block, so its start). Read the excerpt before deciding to expand a row; content stays the record's start. Absent under full_content and on rows shown whole. v2.5.2 additive: each scored message carries match_reason={signal, score, ...} where signal is the branch the quality gate keyed on (rsf > cosine > rrf; confidence only under CPERSONA_CONFIDENCE_ORDERING=legacy — from 2.6.0a7 an enabled confidence score is returned beside each row but neither orders nor gates) and the remaining keys (cosine / rrf / rsf) surface the internal per-retriever contributions present on that row; prior, when present, is the age weight that ordered the row (CPERSONA_PRIOR_AGE_RATE). Unscored rows (cascade FTS/keyword) omit match_reason. A response carrying gate_fallback=true (absent otherwise) means every candidate fell below the quality gate and the below-gate lexical matches were returned instead of an empty result — treat them as low-confidence. A response may carry suggestion (absent otherwise, at most once per session): something the server noticed that only the user can decide — that this scope has grown past the scan window with no coarse index for a time cue to search the rest. Relay its message, and run its fix only if the user agrees.
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| deep | No | Deep recall — halves the quality gate (and the calibrated fused gate), so weaker matches are admitted. It also disables time and completion decay, which are inert unless CPERSONA_CONFIDENCE_ENABLED=true, and it does NOT widen the scan window (CPERSONA_MAX_MEMORIES) — deep is about how weak a match may be, not how far back the search reaches. | |
| limit | No | Per-retriever search depth, not a pure response cap: the value is handed to each retrieval channel (vector / episode FTS / keyword) as its top-K, so lowering it shrinks the candidate pool itself — rows beyond the depth are unreachable at any gate value, and score normalization / autocut operate on the smaller pool, which can also reorder what remains. Fewer rows than this may be returned. (Agent-facing cap; the library layer accepts up to the scan window for direct callers.) | |
| query | Yes | Search query (empty returns recent memories) | |
| trace | No | 2.6 recall trace: true adds `trace` to the response — which rows each stage (retrieval arms, fusion, quality gate, autocut, final order, count cut, reserved seats) kept, dropped or reordered, and why, with ranks and scores. It carries references only, never stored text, and is not stored on the server. trace_version identifies its shape. False (the default) returns the response unchanged. | |
| channel | No | Filter memories by channel (e.g. 'chat', 'discord'). Default: '' (all channels). | |
| agent_id | Yes | Agent identifier | |
| time_cue | No | 2.6: when the answer was stored, as far as you remember. Give after and/or before (a date YYYY-MM-DD, which includes that whole day, or an ISO-8601 timestamp), or ago, plus confidence. The server also searches that period. Among the rows it returns anyway, one found there moves up by at most 3 / 2 / 1 places (sure / likely / vague), and as many extra seats (3 / 2 / 1) hold the best records that search found which the answer does not already hold: one no other search reached, or one the count cut. The returned rows are those of a recall without the cue, reordered, plus at most those seats: no row is removed or re-scored, and which rows pass the quality gate does not change. If the period holds nothing it is widened once, one confidence step. likely widens the period by half its length on each side, vague by its whole length. The response then carries time_cue: the period searched, whether it was widened, how many rows moved and how many seats were used, and remainder when the period held more records than the vector search reads and no coarse index could search the rest. Pass it only when the request itself says when (a date, a month, "last week", "in the spring"); omit it when the request names no time, and never fill it with today's date or a guess. A cue whose own period starts within the last 24 hours (today, or later) is not used: the rows are those of a recall without a cue, and time_cue in the response says ignored. A cue that cannot be read returns an error and no messages. | |
| source_id | No | v2.4.20 per-user source filter. Empty (default) = no filter. Non-empty = prefix match against json_extract(source, '$.id'), e.g. 'discord:12345' to restrict to one Discord user, or 'discord:' to scope to all Discord-sourced memories. Episodes carry no per-user source tagging, so they are skipped when this is set — UNLESS channel is also set, which scopes episodes to one conversation and re-admits them. | |
| project_id | No | v2.4.17 γ filter. Omit → no filter (all projects). '' → global pool only. 'X' → 'X' bucket ∪ global pool. Threaded through cascade / RRF / vector / FTS / keyword paths. v2.5.1: pass '@auto' to resolve this agent's default from the server's operating context (the resolution is echoed as resolved_project_id; an unmapped agent yields operating_context_warning). bug-186: resolution requires a configured operating context. With none — the default, and equally the outcome of a sidecar that fails to parse — the sentinel is NOT resolved: it is stored and filtered as the literal project_id '@auto', resolved_project_id echoes '@auto', and no warning is raised. Read resolved_project_id before relying on the resolution. | |
| session_key | No | Opaque session identity you declare — a partition hint, NOT authentication. It scopes this process's per-session state: the degraded-recall advisory's "already told you" memory, and which no-persist pause applies to this call. It does NOT filter stored data (use agent_id / project_id / channel for that), and it never reaches the database. Omit it to share one bucket with every other caller that omits it, which is the behaviour that predates this parameter. | |
| full_content | No | v2.5.0 preview tier opt-out. By default message content longer than the preview cap (CPERSONA_RECALL_PREVIEW_CHARS, default 500) is returned as a pure prefix with content_truncated/content_len markers; each message's `ref` expands via get_contents. true returns full text. | |
| exclude_contents | No | Normalized content strings to exclude from results (starts-with match). Used to prevent duplication with conversation context already known to the caller. bug-399: the starts-with rule holds only at or above 32 characters. A shorter entry has to EQUAL the stored content (after the normalization this parameter already asks for: stripped and lower-cased) — a short prefix would otherwise suppress every memory beginning with it, inside the retrievers and with nothing in the response reporting the exclusion. Size entries at or above that length when you mean a prefix. |