Skip to main content
Glama
Cloto-dev

CPersona

Official
by Cloto-dev

check_health

Diagnose memory database health: detects contamination, duplicates, embedding failures, and schema drift, returning a severity-based verdict. Run report-only or set fix=true to auto-repair issues.

Instructions

Check memory database health (36-check registry, each issue tagged with severity critical/warn/info). Detects contamination, duplicates, oversized content, embedding issues, FTS integrity (count + content-level), schema version/object drift (missing UNIQUE indexes or FTS triggers), SQLite file integrity, project_id naming drift, invalid JSON/timestamps, timestamp format drift, stale tasks, missing profiles, empty content, invalid/anonymous sources. Returns storage stats incl. project_id/channel distributions. Set fix=true to auto-repair (agent-scoped, locked-safe); the one exception is dedup_msg_id_index, whose repair blanks colliding msg_id values under every agent because the UNIQUE index it restores is a global schema object — with an ACL configured that repair demands read-write on '*', so exclude it via checks to stay agent-scoped. critical file-integrity findings are report-only. Two repairs are lossy and irreversible, each against its own cap: oversized memories are cut to CPERSONA_MAX_CONTENT_LENGTH (default 16000 since 2.5.4a2) and the agent's profile row to CPERSONA_MAX_PROFILE_LENGTH (default 2000), keeping the start. Lower either cap and a fix run shortens rows that were within the old one. Some repairs are bounded per run (source canonicalisation classifies at most 10000 rows); a fix response carrying remaining > 0 with a re-run hint has NOT converged — run fix again until remaining stops decreasing. Use checks parameter to run a subset — an unknown name is rejected (ok=false) rather than silently running nothing, and every response echoes checks_run. The verdict is status: healthy / degraded / unhealthy, derived from severity counts (info never degrades). The pre-2.5.2b1 healthy boolean (len(issues) == 0) is gone — it reported False for an info-only database that status called healthy; read issues / severity_summary for the underlying counts. Read status as a verdict on what is IN the database, not on whether the pipeline that fills it is working: a corpus where every embedding is NULL is internally consistent, so it scores healthy while semantic recall is dead. Nothing here contacts the embedding backend unless fix=true — on a report-only run the liveness findings cannot appear at all, and their absence is not evidence the backend answered. The null_embedding finding carries the reason its repair cannot run; read that before reading status.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
fixNoAuto-fix detected issues
checksNoRegistry check names to run (empty = all). See cpersona.checks.HEALTH_CHECK_NAMES.
agent_idNoAgent ID to check (empty = all agents)
session_keyNoOpaque session identity you declare: a partition hint, not authentication and not a data filter. Selects which no-persist pause applies to this call. Omit to share one bucket with every caller that omits it. Full text on recall.

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed1 schema field changedv2.6.1
    • addedInput schema / properties / session_key / maxLength
      Added value: +256
  2. Changed1 schema field changedv2.5.10
    • addedInput schema / properties / session_key
      Added value: +{
      +  "default": "",
      +  "description": "Opaque session identity you declare: a partition hint, not authentication and not a data filter. Selects which no-persist pause applies to this call. Omit to share one bucket with every caller that omits it. Full text on recall.",
      +  "type": "string"
      +}
  3. Addedv2.5.2
  4. Removedv2.5.1
  5. Changed1 schema field changedv2.4.37
    • addedInput schema / properties / checks
      Added value: +{
      +  "description": "Registry check names to run (empty = all). See cpersona.checks.HEALTH_CHECK_NAMES.",
      +  "items": {
      +    "type": "string"
      +  },
      +  "type": "array"
      +}
  6. First observedv0.1.0

TDQS

A4.2/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Extraordinarily rich disclosure beyond the readOnlyHint=false annotation: which repairs are lossy/irreversible and against which caps, that dedup_msg_id_index repair is global and needs read-write on '*', that some repairs are bounded per run with a `remaining` counter, that critical file-integrity findings are report-only, and that a report-only run never contacts the embedding backend so liveness findings are absent. No contradiction with the annotation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with purpose and every caveat is substantive, but the single dense paragraph runs very long and is hard to parse; several clauses (pre-2.5.2b1 history, cap defaults) could be tightened without losing meaning.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Despite no output schema, the description explains the return contract thoroughly (status verdict, issues, severity_summary, checks_run, remaining) and warns how to read status as a verdict on database contents rather than pipeline liveness. An agent has everything needed to call and interpret it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so baseline is 3, but the description adds real meaning: checks rejects unknown names (ok=false) and every response echoes checks_run; fix carries the lossy-repair and per-run-bound semantics; the caps that govern fix behavior are named. Only session_key is left to the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Check memory database health') and enumerates the 36-check registry in detail, so the scope is unmistakable. It does not, however, differentiate itself from the sibling deep_check, which an agent could reasonably confuse it with.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives concrete invocation guidance: set fix=true to repair, use checks to run a subset, and re-run fix until `remaining` stops decreasing. It explains how to stay agent-scoped by excluding dedup_msg_id_index. What it lacks is an explicit statement of when to reach for this tool rather than a sibling like deep_check.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.