Skip to main content
Glama
Cloto-dev

CPersona

Official
by Cloto-dev

recall

Read-only

Search stored agent memories with vector, FTS5, and keyword strategies to retrieve relevant context for answering questions.

Instructions

Recall relevant memories using multi-strategy search (vector + FTS5 + keyword). To answer a question from memory, prefer reconstruct, the recommended way to read it: it returns items that quote the rows supporting them, within a character budget. Message content is returned as a preview tier by default — expand selected rows with get_contents(refs), or opt out wholesale with full_content=true. full_content is itself budgeted (200k chars per response, bug-211): rows past the budget degrade to the preview tier and the response carries full_content_budget_chars (absent when the budget never bites). 2.6 additive: a message whose content the preview cut also carries excerpt — the part of the record that matched the query, at most 800 characters (CPERSONA_RECALL_EXCERPT_CHARS), separate passages joined by ' … ' in text order — and excerpt_basis (blocks: the record's block set; lexical: divided at read time, ranked by shared words; start: the record is one block, so its start). Read the excerpt before deciding to expand a row; content stays the record's start. Absent under full_content and on rows shown whole. v2.5.2 additive: each scored message carries match_reason={signal, score, ...} where signal is the branch the quality gate keyed on (rsf > cosine > rrf; confidence only under CPERSONA_CONFIDENCE_ORDERING=legacy — from 2.6.0a7 an enabled confidence score is returned beside each row but neither orders nor gates) and the remaining keys (cosine / rrf / rsf) surface the internal per-retriever contributions present on that row; prior, when present, is the age weight that ordered the row (CPERSONA_PRIOR_AGE_RATE). Unscored rows (cascade FTS/keyword) omit match_reason. A response carrying gate_fallback=true (absent otherwise) means every candidate fell below the quality gate and the below-gate lexical matches were returned instead of an empty result — treat them as low-confidence. A response may carry suggestion (absent otherwise, at most once per session): something the server noticed that only the user can decide — that this scope has grown past the scan window with no coarse index for a time cue to search the rest. Relay its message, and run its fix only if the user agrees.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
deepNoDeep recall — halves the quality gate (and the calibrated fused gate), so weaker matches are admitted. It also disables time and completion decay, which are inert unless CPERSONA_CONFIDENCE_ENABLED=true, and it does NOT widen the scan window (CPERSONA_MAX_MEMORIES) — deep is about how weak a match may be, not how far back the search reaches.
limitNoPer-retriever search depth, not a pure response cap: the value is handed to each retrieval channel (vector / episode FTS / keyword) as its top-K, so lowering it shrinks the candidate pool itself — rows beyond the depth are unreachable at any gate value, and score normalization / autocut operate on the smaller pool, which can also reorder what remains. Fewer rows than this may be returned. (Agent-facing cap; the library layer accepts up to the scan window for direct callers.)
queryYesSearch query (empty returns recent memories)
traceNo2.6 recall trace: true adds `trace` to the response — which rows each stage (retrieval arms, fusion, quality gate, autocut, final order, count cut, reserved seats) kept, dropped or reordered, and why, with ranks and scores. It carries references only, never stored text, and is not stored on the server. trace_version identifies its shape. False (the default) returns the response unchanged.
channelNoFilter memories by channel (e.g. 'chat', 'discord'). Default: '' (all channels).
agent_idYesAgent identifier
time_cueNo2.6: when the answer was stored, as far as you remember. Give after and/or before (a date YYYY-MM-DD, which includes that whole day, or an ISO-8601 timestamp), or ago, plus confidence. The server also searches that period. Among the rows it returns anyway, one found there moves up by at most 3 / 2 / 1 places (sure / likely / vague), and as many extra seats (3 / 2 / 1) hold the best records that search found which the answer does not already hold: one no other search reached, or one the count cut. The returned rows are those of a recall without the cue, reordered, plus at most those seats: no row is removed or re-scored, and which rows pass the quality gate does not change. If the period holds nothing it is widened once, one confidence step. likely widens the period by half its length on each side, vague by its whole length. The response then carries time_cue: the period searched, whether it was widened, how many rows moved and how many seats were used, and remainder when the period held more records than the vector search reads and no coarse index could search the rest. Pass it only when the request itself says when (a date, a month, "last week", "in the spring"); omit it when the request names no time, and never fill it with today's date or a guess. A cue whose own period starts within the last 24 hours (today, or later) is not used: the rows are those of a recall without a cue, and time_cue in the response says ignored. A cue that cannot be read returns an error and no messages.
source_idNov2.4.20 per-user source filter. Empty (default) = no filter. Non-empty = prefix match against json_extract(source, '$.id'), e.g. 'discord:12345' to restrict to one Discord user, or 'discord:' to scope to all Discord-sourced memories. Episodes carry no per-user source tagging, so they are skipped when this is set — UNLESS channel is also set, which scopes episodes to one conversation and re-admits them.
project_idNov2.4.17 γ filter. Omit → no filter (all projects). '' → global pool only. 'X' → 'X' bucket ∪ global pool. Threaded through cascade / RRF / vector / FTS / keyword paths. v2.5.1: pass '@auto' to resolve this agent's default from the server's operating context (the resolution is echoed as resolved_project_id; an unmapped agent yields operating_context_warning). bug-186: resolution requires a configured operating context. With none — the default, and equally the outcome of a sidecar that fails to parse — the sentinel is NOT resolved: it is stored and filtered as the literal project_id '@auto', resolved_project_id echoes '@auto', and no warning is raised. Read resolved_project_id before relying on the resolution.
session_keyNoOpaque session identity you declare — a partition hint, NOT authentication. It scopes this process's per-session state: the degraded-recall advisory's "already told you" memory, and which no-persist pause applies to this call. It does NOT filter stored data (use agent_id / project_id / channel for that), and it never reaches the database. Omit it to share one bucket with every other caller that omits it, which is the behaviour that predates this parameter.
full_contentNov2.5.0 preview tier opt-out. By default message content longer than the preview cap (CPERSONA_RECALL_PREVIEW_CHARS, default 500) is returned as a pure prefix with content_truncated/content_len markers; each message's `ref` expands via get_contents. true returns full text.
exclude_contentsNoNormalized content strings to exclude from results (starts-with match). Used to prevent duplication with conversation context already known to the caller. bug-399: the starts-with rule holds only at or above 32 characters. A shorter entry has to EQUAL the stored content (after the normalization this parameter already asks for: stripped and lower-cased) — a short prefix would otherwise suppress every memory beginning with it, inside the retrievers and with nothing in the response reporting the exclusion. Size entries at or above that length when you mean a prefix.

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed1 schema field changedv2.6.3
    • changedInput schema / properties / time_cue / description
      Previous value: -"2.6: when the answer was stored, as far as you remember. Give after and/or before (a date YYYY-MM-DD, which includes that whole day, or an ISO-8601 timestamp), or ago, plus confidence. The server also searches that period. Among the rows it returns anyway, one found there moves up by at most 3 / 2 / 1 places (sure / likely / vague), and as many extra seats (3 / 2 / 1) hold the best records that search found which the answer does not already hold: one no other search reached, or one the count cut. The returned rows are those of a recall without the cue, reordered, plus at most those seats: no row is removed or re-scored, and which rows pass the quality gate does not change. If the period holds nothing it is widened once, one confidence step. likely widens the period by half its length on each side, vague by its whole length. The response then carries time_cue: the period searched, whether it was widened, how many rows moved and how many seats were used. Pass it only when the request itself says when (a date, a month, \"last week\", \"in the spring\"); omit it when the request names no time, and never fill it with today's date or a guess. A cue whose own period starts within the last 24 hours (today, or later) is not used: the rows are those of a recall without a cue, and time_cue in the response says ignored. A cue that cannot be read returns an error and no messages."New value: +"2.6: when the answer was stored, as far as you remember. Give after and/or before (a date YYYY-MM-DD, which includes that whole day, or an ISO-8601 timestamp), or ago, plus confidence. The server also searches that period. Among the rows it returns anyway, one found there moves up by at most 3 / 2 / 1 places (sure / likely / vague), and as many extra seats (3 / 2 / 1) hold the best records that search found which the answer does not already hold: one no other search reached, or one the count cut. The returned rows are those of a recall without the cue, reordered, plus at most those seats: no row is removed or re-scored, and which rows pass the quality gate does not change. If the period holds nothing it is widened once, one confidence step. likely widens the period by half its length on each side, vague by its whole length. The response then carries time_cue: the period searched, whether it was widened, how many rows moved and how many seats were used, and remainder when the period held more records than the vector search reads and no coarse index could search the rest. Pass it only when the request itself says when (a date, a month, \"last week\", \"in the spring\"); omit it when the request names no time, and never fill it with today's date or a guess. A cue whose own period starts within the last 24 hours (today, or later) is not used: the rows are those of a recall without a cue, and time_cue in the response says ignored. A cue that cannot be read returns an error and no messages."
  2. Changed3 schema fields changedv2.6.1
    • addedInput schema / properties / session_key / maxLength
      Added value: +256
    • addedInput schema / properties / time_cue
      Added value: +{
      +  "additionalProperties": false,
      +  "description": "2.6: when the answer was stored, as far as you remember. Give after and/or before (a date YYYY-MM-DD, which includes that whole day, or an ISO-8601 timestamp), or ago, plus confidence. The server also searches that period. Among the rows it returns anyway, one found there moves up by at most 3 / 2 / 1 places (sure / likely / vague), and as many extra seats (3 / 2 / 1) hold the best records that search found which the answer does not already hold: one no other search reached, or one the count cut. The returned rows are those of a recall without the cue, reordered, plus at most those seats: no row is removed or re-scored, and which rows pass the quality gate does not change. If the period holds nothing it is widened once, one confidence step. likely widens the period by half its length on each side, vague by its whole length. The response then carries time_cue: the period searched, whether it was widened, how many rows moved and how many seats were used. Pass it only when the request itself says when (a date, a month, \"last week\", \"in the spring\"); omit it when the request names no time, and never fill it with today's date or a guess. A cue whose own period starts within the last 24 hours (today, or later) is not used: the rows are those of a recall without a cue, and time_cue in the response says ignored. A cue that cannot be read returns an error and no messages.",
      +  "properties": {
      +    "after": {
      +      "type": "string"
      +    },
      +    "ago": {
      +      "description": "Relative to now: {\"unit\": \"days\" | \"weeks\" | \"months\", \"value\": N} names the unit-long period centred N units ago (a month is 30 days); \"long_ago\" names the oldest third of what this scope holds.",
      +      "oneOf": [
      +        {
      +          "additionalProperties": false,
      +          "properties": {
      +            "unit": {
      +              "enum": [
      +                "days",
      +                "weeks",
      +                "months"
      +              ],
      +              "type": "string"
      +            },
      +            "value": {
      +              "minimum": 0,
      +              "type": "integer"
      +            }
      +          },
      +          "required": [
      +            "unit",
      +            "value"
      +          ],
      +          "type": "object"
      +        },
      +        {
      +          "enum": [
      +            "long_ago"
      +          ],
      +          "type": "string"
      +        }
      +      ]
      +    },
      +    "before": {
      +      "type": "string"
      +    },
      +    "confidence": {
      +      "enum": [
      +        "sure",
      +        "likely",
      +        "vague"
      +      ],
      +      "type": "string"
      +    }
      +  },
      +  "required": [
      +    "confidence"
      +  ],
      +  "type": "object"
      +}
    • addedInput schema / properties / trace
      Added value: +{
      +  "default": false,
      +  "description": "2.6 recall trace: true adds `trace` to the response — which rows each stage (retrieval arms, fusion, quality gate, autocut, final order, count cut, reserved seats) kept, dropped or reordered, and why, with ranks and scores. It carries references only, never stored text, and is not stored on the server. trace_version identifies its shape. False (the default) returns the response unchanged.",
      +  "type": "boolean"
      +}
  3. Changed1 schema field changedv2.5.12
    • changedInput schema / properties / exclude_contents / description
      Previous value: -"Normalized content strings to exclude from results (starts-with match). Used to prevent duplication with conversation context already known to the caller."New value: +"Normalized content strings to exclude from results (starts-with match). Used to prevent duplication with conversation context already known to the caller. bug-399: the starts-with rule holds only at or above 32 characters. A shorter entry has to EQUAL the stored content (after the normalization this parameter already asks for: stripped and lower-cased) — a short prefix would otherwise suppress every memory beginning with it, inside the retrievers and with nothing in the response reporting the exclusion. Size entries at or above that length when you mean a prefix."
  4. Changed2 schema fields changedv2.5.10
    • changedInput schema / properties / limit / description
      Previous value: -"Per-retriever search depth, not a pure response cap (CSC #716): the value is handed to each retrieval channel (vector / episode FTS / keyword) as its top-K, so lowering it shrinks the candidate pool itself — rows beyond the depth are unreachable at any gate value, and score normalization / autocut operate on the smaller pool, which can also reorder what remains. Fewer rows than this may be returned. (Agent-facing cap; the library layer accepts up to the scan window for direct callers.)"New value: +"Per-retriever search depth, not a pure response cap: the value is handed to each retrieval channel (vector / episode FTS / keyword) as its top-K, so lowering it shrinks the candidate pool itself — rows beyond the depth are unreachable at any gate value, and score normalization / autocut operate on the smaller pool, which can also reorder what remains. Fewer rows than this may be returned. (Agent-facing cap; the library layer accepts up to the scan window for direct callers.)"
    • addedInput schema / properties / session_key
      Added value: +{
      +  "default": "",
      +  "description": "Opaque session identity you declare — a partition hint, NOT authentication. It scopes this process's per-session state: the degraded-recall advisory's \"already told you\" memory, and which no-persist pause applies to this call. It does NOT filter stored data (use agent_id / project_id / channel for that), and it never reaches the database. Omit it to share one bucket with every other caller that omits it, which is the behaviour that predates this parameter.",
      +  "type": "string"
      +}
  5. Changed1 schema field changedv2.5.6
    • changedInput schema / properties / deep / description
      Previous value: -"Deep recall — disable time and completion decay for exhaustive search"New value: +"Deep recall — halves the quality gate (and the calibrated fused gate), so weaker matches are admitted. It also disables time and completion decay, which are inert unless CPERSONA_CONFIDENCE_ENABLED=true, and it does NOT widen the scan window (CPERSONA_MAX_MEMORIES) — deep is about how weak a match may be, not how far back the search reaches."
  6. Changed1 schema field changedv2.5.4
    • changedInput schema / properties / limit / description
      Previous value: -"Max memories to return (agent-facing cap; the library layer accepts up to the scan window for direct callers)"New value: +"Per-retriever search depth, not a pure response cap (CSC #716): the value is handed to each retrieval channel (vector / episode FTS / keyword) as its top-K, so lowering it shrinks the candidate pool itself — rows beyond the depth are unreachable at any gate value, and score normalization / autocut operate on the smaller pool, which can also reorder what remains. Fewer rows than this may be returned. (Agent-facing cap; the library layer accepts up to the scan window for direct callers.)"
  7. Addedv2.5.2
  8. Removedv2.5.1
  9. Changed2 schema fields changedv2.4.34
    • addedInput schema / properties / project_id
      Added value: +{
      +  "description": "v2.4.17 γ filter. Omit → no filter (all projects). '' → global pool only. 'X' → 'X' bucket ∪ global pool. Threaded through cascade / RRF / vector / FTS / keyword paths.",
      +  "type": "string"
      +}
    • addedInput schema / properties / source_id
      Added value: +{
      +  "default": "",
      +  "description": "v2.4.20 per-user source filter. Empty (default) = no filter. Non-empty = prefix match against json_extract(source, '$.id'), e.g. 'discord:12345' to restrict to one Discord user, or 'discord:' to scope to all Discord-sourced memories. Episodes are skipped when set (no per-user source tagging).",
      +  "type": "string"
      +}
  10. Changed1 schema field changedv2.4.10
    • addedInput schema / properties / exclude_contents
      Added value: +{
      +  "description": "Normalized content strings to exclude from results (starts-with match). Used to prevent duplication with conversation context already known to the caller.",
      +  "items": {
      +    "type": "string"
      +  },
      +  "type": "array"
      +}
  11. First observedv0.1.0

TDQS

A4.3/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare readOnlyHint=true, so the description carries the safety/behavioral burden and discharges it richly: preview-tier defaults, the 200k full_content budget with degradation to preview, the meaning of gate_fallback, excerpt/excerpt_basis, match_reason signals, and the once-per-session suggestion semantics. This is far more than the annotation provides.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The purpose is front-loaded, but the body degenerates into versioned changelog prose ('2.6 additive', 'v2.5.2 additive', 'v2.4.20') with internal bug IDs and config-flag asides that an agent does not need to select or invoke the tool. Far longer than the decision requires.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With 12 parameters, nested objects, and no output schema, the description does explain response-side fields (full_content_budget_chars, excerpt, match_reason, gate_fallback, suggestion) that the schema cannot. It is complete enough to call correctly, though the changelog framing makes the relevant behavior harder to extract than it needs to be.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the per-parameter schema text is already dense (deep, limit, trace, time_cue, source_id, project_id, session_key). The description mostly restates or lightly extends that, adding marginal value; baseline 3 is appropriate when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The opening states a specific verb and resource ('Recall relevant memories') plus the retrieval mechanism (vector + FTS5 + keyword), and explicitly distinguishes this from the sibling `reconstruct`. An agent can tell the two apart without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives explicit routing ('To answer a question from memory, prefer `reconstruct`'), tells the caller how to expand previews via get_contents(refs), when to set full_content, when to pass time_cue ('only when the request itself says when... never fill it with today's date'), and how to treat gate_fallback rows. When-to-use and when-not-to-use are both present.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.