ats serve
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@ats serverecall my past sessions about the auth refactor"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Agent Trace Signals
Memory analytics layer for coding agent sessions. Ingests traces from Claude Code (~/.claude/projects/), Gemini CLI (~/.gemini/tmp/), and Opencode (~/.local/share/opencode/opencode.db), extracts entities and memories, stores everything in a local SQLite database with vector embeddings, and serves memories back to Claude Code via MCP. Runs entirely offline via Ollama.
Architecture
Every extraction output is kept as a distinct, separately-typed row (not collapsed into a single blob) and consolidated into cross-session structure before it's ever retrieved:
Related MCP server: AI Agent History RAG MCP Server
Prerequisites
Get started
git clone https://github.com/LanGuo/agent-trace-signals.git
cd agent-trace-signals
# Install dependencies
uv sync
uv pip install -e '.[ner]' # GLiNER NER model (~400 MB, downloaded on first use)
# Pull required Ollama models
ollama pull gemma4:e4b # combined per-chunk extraction (summary + entities + memories + patterns)
ollama pull gemma3:12b # session summaries, explicit memory extraction
ollama pull nomic-embed-text # embeddings (ats embed + ats recall/chunk-search)
# Start Ollama (if not already running)
ollama serve &Web UI
uv pip install -e '.[ui]' # one-time: adds streamlit + plotly
uv run ats ui # opens http://localhost:8501Five pages:
Page | What it shows |
Home / Sessions | Sessions table — click a row to drive session detail below. Metrics strip, clickable memory-type breakdown, memories-per-chunk chart. Tabs: Entities, Related Sessions (same workspace / shared file-commit-PR entity / shared cluster-memory pattern — see |
Query |
|
Explore Sessions | Corpus-wide session EDA (formerly "Agent Compare"/"Explore", now also absorbing "Session Compare") — session scatter with freely selectable x/y metrics, memory/entity-type composition by structural session type, token efficiency timeline, cache hit % distribution, recovery leaderboard, structural embedding UMAP, group summary, and a session connection graph (community detection over |
Explore Memories | Corpus-wide memory EDA (formerly "Memory Browser", now also absorbing the memory-yield-by-session-type chart) — overview totals, clickable memory-type breakdown, memories-over-time, memory yield per chunk by structural session type, cluster-themes chart, memory-level embedding UMAP colorable by type/cluster-status/harness/workspace, memories-by-workspace. Filterable table (type, method, session, keyword) — click a row for full content, source sessions, and (for |
Evaluation | Static results from |
Pipeline order
ats ingest → ats embed → ats cluster-sessions → ats cluster-memories → ats recall / ats serve
(embeddings) (HDBSCAN session tags) (cross-session memory consolidation)Step | What it does | Needs Ollama? |
| Parse agent traces (Claude Code JSONL, Gemini CLI JSON/JSONL, Opencode SQLite) → a combined extraction prompt extracts | Yes (extraction model) |
| Embed memories, chunks, entities, and session summaries | Yes (nomic-embed-text) |
| HDBSCAN over structural embeddings → | Yes (labeling step; |
| Embedding-similarity clustering of memory+pattern rows → consolidated | Yes ( |
| Hybrid BM25 + semantic retrieval over memories and chunks | Yes (nomic-embed-text for query embedding) |
ats cluster-sessions and ats cluster-memories are never run automatically — ats ingest
does not trigger either one. Both are safe to rerun anytime (each does a fresh HDBSCAN pass and
fully rewrites session_type/cluster memories, not an incremental patch), but if you skip rerunning
them after ingesting a meaningful batch of new sessions, the new sessions sit with session_type IS NULL and no cluster memories reference them — nothing errors, the UI just silently shows
"unlabeled"/empty sections for anything untouched by the last clustering run. Rerun both after any
ingest that adds or meaningfully changes the corpus, not just once at initial setup.
Usage
# See what's ingested
uv run ats sessions
# Preview pending files
uv run ats ingest --dry-run
# Step 1: Ingest — combined extraction prompt extracts entities, memories, and patterns per chunk
uv run ats ingest
uv run ats ingest --no-summaries # skip chunk/session summaries (faster)
# Step 2: Embed everything (memories, chunks, entities, session summaries)
uv run ats embed
# Step 3: Session clustering — HDBSCAN over structural embeddings → session_type tags,
# then LLM-labels each cluster's shared interaction shape (e.g. "Iterative troubleshooting")
uv run ats cluster-sessions
uv run ats cluster-sessions --no-label # tag only, skip the LLM labeling step
# Step 4: Memory clustering — cross-session consolidation into cluster memories.
# Newly minted cluster memories are embedded automatically at the end (--no-embed to skip
# and embed later by hand instead).
uv run ats cluster-memories
uv run ats cluster-memories --min-cluster-size 3 --min-recurrence 1 # defaults shown
uv run ats cluster-memories --no-embed # skip auto-embed of newly minted memories
# Query memories and chunks
uv run ats recall "entity extraction pipeline"
uv run ats recall "error recovery pattern" --memory-type pattern_recovery --top-k 5
uv run ats recall "auth bug" --method semantic # semantic-only (no BM25)
uv run ats recall "auth bug" --workspace agent-trace-signals # restrict to one project
uv run ats chunk-search "extraction patterns" --top-k 10
uv run ats chunk-search "auth bug" --session <session-prefix>
uv run ats chunk-search "auth bug" --method lexical # BM25-only
uv run ats chunk-search "auth bug" --workspace coral-ai # restrict to one project
# Find relevant sessions by query
uv run ats session-search "Symphony multi-repo orchestration"
uv run ats session-search "coral ai" # matches workspace name
uv run ats session-search "PDF download" --method semantic
uv run ats session-search "PDF download" --workspace coral-ai # restrict to one project
# Find sessions sharing a workspace, structural entities (files/commits/PRs), or a
# consolidated memory pattern (requires ats cluster-memories to have been run), with a given session
uv run ats graph-walk <session-prefix>
# Start the MCP server (see Claude Code integration below)
uv run ats serve
# DB stats
uv run ats stats
# Maintenance: remove orphaned vec0 embedding rows from past re-ingestions
uv run ats clean-orphansPlanned / In Progress
Area | Status | Notes |
| ✅ done | HDBSCAN over structural embeddings → |
| ✅ done | Embedding-similarity clustering of memory+pattern rows → consolidated |
Recall measurement | 🔄 partial | Level 1a eval harness exists; needs annotated ground-truth sessions |
Raw-trace session embeddings | ✅ done | Stored in |
Known limitations
Clustering is a manual step.
ats ingestdoes not triggerats cluster-sessionsorats cluster-memories; rerun both after any ingest that changes the corpus, orsession_typeand cluster memories go stale.Conflict detection has no resolution workflow. Contradictions between memories in the same cluster are flagged in the
conflictstable and shown in Explore Memories, but nothing resolves them,conflict_typeis free text rather than a taxonomy, and memories that never cluster together are never compared.ModelConfig.conflict_verifieris reserved for a future re-check pass and currently unused.Retrieval ignores recency. Ranking does not use memory
status(open vs. resolved) or age.No in-place schema migration for taxonomy v2. Older databases must be re-ingested rather than upgraded.
Cluster-to-cluster relations (clusters that share sessions) are not yet exposed as a query.
session_searchis CLI-only; the MCP server exposesrecall,chunk_search, andgraph_walk.The
revertedmemory status has not yet been observed firing on real data.
Evaluate
# --- Per-session entity annotation (Level 0 eval) ---
uv run ats annotate-init <session_prefix> # pre-fill from extracted entities
uv run ats annotate-view <session_prefix> # HTML span viewer — open in browser
uv run ats eval0 --annotations annotations/ # precision/recall vs annotations (annotations/ is created by you; gitignored)
# --- Cross-session memory annotation (Level 1a eval) ---
uv run ats annotate-memories # generate HTML viewer, open in browser, save labels
uv run ats eval1a # structural quality + recall/FPR if annotated
Claude Code integration (MCP)
ats serve starts an MCP server on stdio exposing three tools to Claude Code: recall,
chunk_search, and graph_walk.
Register it with the claude mcp add CLI (not by hand-editing settings.json — recent Claude
Code versions don't read an mcpServers key there):
claude mcp add ats-memory --scope user \
-e ATS_DB=/path/to/agent-trace-signals/traces.db \
-- uv run --directory /path/to/agent-trace-signals ats serve--scope user registers it once for every project, not just this repo. Verify it connected:
claude mcp listThen in any Claude Code session:
Use the ats-memory recall tool to find memories about entity extractionMCP tool | What it does |
| Hybrid BM25 + semantic search over memories, fused via RRF |
| Hybrid search over raw conversation chunks |
| BFS over |
include_graph=true on recall/chunk_search (and graph_walk itself) expands the top hits with graph-connected memories/chunks from other sessions — exploratory cross-session context, not precision retrieval. Measured on the Evaluation page: it never improves fact-lookup hit rate over the base call, even now that both tools are capped (recall ≤30 extra memories since 2026-07-30, chunk_search ≤66 extra chunks since 2026-08-04) — see design_decisions.md's entries on those two dates for the fan-out bugs this fixed.
Search methods and scoring
All three CLI search commands (ats recall, ats chunk-search, ats session-search) support --method hybrid|semantic|lexical. All methods return RRF scores on a consistent 0–0.033 scale regardless of which legs are active — a single-leg method (e.g. --method lexical) feeds only that leg into RRF, producing comparable scores.
Method | BM25 leg | Semantic leg | When to use |
| FTS5 porter-stemmed BM25 | vec0 cosine ANN | Best overall after |
| FTS5 only | — | Works without embeddings; good for exact terms |
| — | vec0 only | Intent-based queries after |
ats session-search lexical leg searches both session_summary and workspace_id — so ats session-search "coral ai" matches the -Users-you-src-coral-ai workspace even without embeddings. Semantic leg uses avg-chunk embeddings per session, so it works even for sessions ingested with --no-summaries.
ats chunk-search lexical leg searches both chunk_text and chunk_summary — raw conversation text plus the LLM-generated summary for each chunk.
--workspace/-w (on ats recall, ats chunk-search, ats session-search, and the recall/chunk_search MCP tools) restricts results to one project via a substring (LIKE '%...%') match against workspace_id, not exact equality — the same real project can be ingested under different workspace_id strings depending on the source plugin and code path (e.g. -Users-you-src-coral-ai vs. bare coral-ai), so an exact match would silently miss some of a project's sessions.
Project structure
src/agent_trace_signals/
cli.py # ats CLI entry point
config.py # all knobs: ScannerConfig, PipelineConfig, ModelConfig, LightAnalyticsConfig
models.py # Pydantic data models
embedder.py # Ollama batch embedding with cache
db/ # SQLite schema + SQLiteStore (sqlite-vec embeddings)
pipeline/ # IngestionPipeline, ClusteringAnalyticsPipeline, RetrievalEngine
mcp/ # FastMCP server exposing recall, chunk_search, graph_walk
annotation/ # HTML viewers for entity and memory annotation
eval/ # Level 0 and Level 1a evaluation harnesses
plugins/ # ClaudeCodeSource (JSONL), GeminiSource (JSON/JSONL), OpencodeSource (SQLite)
providers/ # OllamaProvider, AnthropicProvider
ui/ # Streamlit app (Query, Explore Sessions, Explore Memories, Evaluation)
design/
exploration.md # overarching thesis and system architecture
design_decisions.md # chronological evolution of design choices (primary reference)
deprecated_*.md # superseded designs, kept for history
experiments/ # experiment scripts + write-ups (raw trace data not included)
notes/ # research notes (related papers, landscape, early findings)
demo/ # Remotion source for the demo videoConfiguration
All defaults live in src/agent_trace_signals/config.py. Edit that file to change any setting — there is no separate config file. Changes take effect on the next run.
Switching models
Every LLM call goes through a named field in ModelConfig. You can change any of them independently:
Field | Default | Used for |
|
| Per-chunk entity extraction (short prompts) and multi-chunk batch verification |
|
| Combined per-chunk D-prompt: summary + entities + memories + patterns + preferences in one call. Was |
|
| Session-level summaries during ingest |
|
| Explicit memory extraction during ingest |
|
| Conflict verification in full analytics |
|
| All embeddings via |
|
| "Generate Answer" RAG synthesis on the Query page — separate from |
Note: ats cluster-memories's consolidation LLM call uses its own --model CLI flag (default gemma4:31b), not a ModelConfig field — cluster_labeler in config.py is currently unused dead config (a cleanup candidate, along with Memory.entity_id and Memory.action_orientation — see design_decisions.md 2026-07-13 entry).
gemma4:e4b vs gemma3:12b: gemma4:e4b is a thinking model (Gemma 4 architecture with chain-of-thought reasoning). It is faster than gemma3:12b for short prompts due to architectural improvements, but it consumes its generation budget on reasoning tokens before producing output. For short per-chunk extraction prompts this is fine; for longer prompts (e.g. analytics fact-compression with many occurrences) use gemma3:12b. The defaults reflect this split.
Note: gemma4:e4b uses entity_verifier_max_tokens (default 2048) rather than the standard 512-token budget. This is because its chain-of-thought reasoning tokens count against num_predict, so multi-chunk verification prompts need extra headroom. This does not affect KV cache memory — that is controlled solely by ollama_num_ctx.
Ollama memory / performance
Field | Default | Effect |
|
| KV cache token window per call. Ollama's model default is 131072 (~32 GB GPU RAM); 8192 drops that to ~2 GB. Increase to |
|
|
|
|
|
|
|
| Ollama server URL — change if running Ollama on a different host or port |
Ingestion
Field | Default | Effect |
|
| Target chunk size. Exchanges are greedily packed until this limit; a single exchange that exceeds it alone is split into overlapping sub-chunks. |
|
| Character overlap between sub-chunks when splitting an oversized single exchange (10% of the split window size). |
|
| Sentences either side of entity mention captured as context |
Entity types (combined extraction pipeline)
The default pipeline uses a combined extraction analyzer — one LLM call per chunk extracts summary, entities, memories, and patterns together. Entity types are configurable via PipelineConfig.d_entity_types (the d_ prefix is a naming remnant from early experimentation, not meaningful on its own):
# config.py — PipelineConfig
d_entity_types: dict = {
"technology": "libraries, tools, CLIs, APIs, languages, platforms, services",
"framework": "orchestration and workflow libraries (→ stored as technology)",
"algorithm": "specific named algorithms or methods (→ stored as concept)",
"concept": "abstractions, patterns, architectural ideas, methodologies",
"model": "ML/AI models and model families",
"dataset": "named benchmarks, corpora, or datasets",
"person": "named individuals",
"project": "named repos, products, or systems being built (→ stored as technology)",
}Both the type label and its description are sent to the LLM so it classifies consistently. Types are normalized to the shared entity vocabulary (technology, concept, model, dataset, person, org) before storage so D-path and legacy entities use the same type system. Any unrecognized type falls back to concept.
To add a domain-specific type (e.g. "protocol": "communication or data exchange protocols"): add it to d_entity_types and, if it maps to an existing shared type, add an entry to D_TYPE_NORMALIZATION in pipeline/chunk_analyzer.py.
Legacy pipeline entity types (--legacy-pipeline)
When --legacy-pipeline is passed, GLiNER is used instead. Entity types are controlled by two separate fields:
ner_entity_labels— natural-language labels sent to GLiNER (e.g."software library or framework")GLiNER maps these to canonical types via
_GLINER_LABEL_TO_TYPEinentity_extractor.py
Session similarity (embedding distance)
Session-level similarity can be approximated by averaging chunk embeddings per session (requires ats embed to have run):
import sqlite3, struct, numpy as np
import sqlite_vec
conn = sqlite3.connect("traces.db")
conn.enable_load_extension(True); sqlite_vec.load(conn); conn.enable_load_extension(False)
cur = conn.cursor()
def decode(blob): return np.array(struct.unpack(f"{len(blob)//4}f", blob))
def cosine(a, b): return float(np.dot(a,b) / (np.linalg.norm(a) * np.linalg.norm(b)))
cur.execute("SELECT r.session_id, re.embedding FROM record_embeddings re JOIN records r ON r.id = re.record_id")
from collections import defaultdict
sess_vecs = defaultdict(list)
for sid, blob in cur.fetchall():
if blob: sess_vecs[sid].append(decode(blob))
avg_vecs = {sid: np.mean(vs, axis=0) for sid, vs in sess_vecs.items()}
sids = list(avg_vecs)
for i, a in enumerate(sids):
for b in sids[i+1:]:
print(f"{a[:8]} vs {b[:8]}: {cosine(avg_vecs[a], avg_vecs[b]):.4f}")Note: session_embeddings (the per-session embedding stored during ats embed) requires embedding_text to be populated on the session row — this is only set when --no-summaries is NOT used. For sessions ingested with --no-summaries, use the chunk-average approach above.
Session metadata
Per-session metadata is extracted during ats ingest and stored in the session_metadata table (FK → sessions). Run ats backfill-metadata to populate it for sessions ingested before this feature landed. See design/session_metadata.md for the full schema.
Available from Claude Code sessions:
Field | Source | Example |
|
|
|
|
|
|
|
|
|
|
|
|
| sum of | |
| sum of | |
| sum of | |
| last − first message timestamp | |
| count of | |
|
|
|
Available from Gemini sessions:
Field | Source | Example |
|
|
|
| sum of | |
| sum of | |
| sum of | |
|
|
Available from Opencode sessions:
Field | Source | Example |
|
|
|
|
|
|
|
| |
|
| |
|
| |
|
| |
|
| |
|
| |
| count of | |
|
|
|
|
|
|
|
|
Backfilling existing sessions: Once implemented, run ats backfill-metadata (planned) to extract metadata from already-ingested files without re-ingesting. This re-parses source files only — no LLM calls, no changes to records/memories.
Available Tools
3 toolschunk_searchA
Search raw conversation chunks from agent session history using hybrid BM25 + semantic search.
Args: query: Natural language query session_id: Optional — restrict to a specific session (full ID) top_k: Number of chunks to return (default 20) include_graph: Expand top hits with chunks (from any session) that mention the same entity — a cross-session hop via shared file/commit/PR/etc. workspace: Optional — restrict to one project/workspace (substring match)
Returns: JSON list of chunks with chunk_text, session_id, chunk_index, and rrf_score
| Name | Required | Description | Default |
|---|---|---|---|
| query | Yes | ||
| top_k | No | ||
| workspace | No | ||
| session_id | No | ||
| include_graph | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden, and it does reasonably well: it explains that include_graph performs a cross-session hop via shared entities, that session_id must be a full ID, and that workspace matches by substring. It does not state read-only/rate-limit characteristics, but for a search tool that is a minor gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded purpose sentence followed by Args/Returns blocks; each line is informative and none is filler. Slightly more verbose than strictly needed but well structured.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, the description need not enumerate return fields, and it correctly gives only a brief Returns summary. It is complete for calling the tool, though the missing routing guidance against siblings leaves a small gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and there are 5 params, so the description must compensate entirely, and it does: it documents query, session_id (full ID, optional), top_k (default 20), workspace (substring match), and include_graph with its cross-session expansion semantics.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource (search raw conversation chunks from agent session history) plus the retrieval mechanism (hybrid BM25 + semantic). Clear enough to separate from graph_walk and recall, though it never states what those siblings do.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No explicit when-to-use or when-not-to-use guidance relative to recall or graph_walk. The include_graph note hints at graph-style behavior but the description never names an alternative tool or the condition that should select it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
graph_walkA
Walk the session graph from a seed memory to find related memories.
Args: seed_memory_id: Full memory ID to start from depth: BFS depth (1 = direct neighbours, 2 = neighbours of neighbours)
Returns: JSON list of related memories
| Name | Required | Description | Default |
|---|---|---|---|
| depth | No | ||
| seed_memory_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It discloses the traversal semantics and that the result is a JSON list of related memories, but says nothing about read-only safety, performance or cost characteristics, or traversal limits — meaningfully incomplete for a tool with zero annotation coverage.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The Args/Returns structure is front-loaded and each line carries information; there is no padding. It is slightly verbose in restating the seed parameter already visible in the schema, but overall efficient.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so the Returns line is somewhat redundant, and the core traversal mechanics are explained. What is missing is sibling differentiation and any behavioral caveats (annotations are absent), leaving the agent without guidance on when this tool beats recall.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must carry the parameter burden, and it does: seed_memory_id is described as the full memory ID to start from, and depth is explained as BFS depth with '1 = direct neighbours, 2 = neighbours of neighbours.' This adds real meaning beyond the bare titles. It omits the default value (2) listed in the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'Walk the session graph from a seed memory to find related memories.' An agent can immediately tell this is a graph-traversal tool, but the description never differentiates it from siblings recall and chunk_search, which also surface related content.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied by the phrase 'to find related memories' — there is no explicit statement of when to prefer graph traversal over recall or chunk_search, nor any exclusion criteria. An agent must infer the selection conditions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
recallB
Retrieve relevant memories from agent session history using hybrid BM25 + semantic search.
Args: query: Natural language query (e.g. "PDF rendering error", "auth bug pattern") memory_type: Optional filter — 'episodic', 'procedural', or 'preference' top_k: Number of memories to return (default 10) include_graph: Whether to expand results with graph-connected memories workspace: Optional — restrict to one project/workspace (substring match)
Returns: JSON list of memories with content, type, extraction_method, and rrf_score
| Name | Required | Description | Default |
|---|---|---|---|
| query | Yes | ||
| top_k | No | ||
| workspace | No | ||
| memory_type | No | ||
| include_graph | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It usefully discloses the hybrid BM25 + semantic ranking approach and the return fields (content, type, extraction_method, rrf_score), but says nothing about read-only safety, permissions, rate limits, result size limits, or what include_graph expansion costs.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the one-line purpose, then cleanly separates Args and Returns. The structured layout is easy to scan; the only minor waste is restating return fields that the output schema may already cover.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With five parameters at 0% schema coverage, the description compensates well by documenting every argument, and it goes beyond the existing output schema with a Returns summary. The remaining gap is the absence of any sibling routing or behavioral/permission context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must document all parameters, and it largely does: query with concrete examples, memory_type with its three valid values, top_k default, workspace as a substring match, and include_graph's expansion behavior. It stops short of clarifying interaction between top_k and include_graph, but the semantic coverage is strong.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Retrieve relevant memories from agent session history') plus the retrieval mechanism (hybrid BM25 + semantic search). An agent can tell this is memory retrieval rather than chunk/graph traversal, but the description never names or contrasts the siblings chunk_search and graph_walk, so sibling differentiation is left to inference.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit when-to-use guidance, no when-not-to-use, and no mention of the alternatives chunk_search or graph_walk even though include_graph overlaps conceptually with graph_walk. Usage must be inferred entirely from the tool's name and the query examples.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
3 tool updates
v0.1.0- First observed
chunk_search - First observed
graph_walk - First observed
recall
TDQS
Scored across 3 tools
recall and chunk_search both perform hybrid search over agent session history, but they target distinct granularities (extracted memories vs raw conversation chunks). graph_walk is clearly separate. The overlap is manageable given descriptions, but an agent could still confuse the two search tools without careful reading.
All names are snake_case, but 'recall' is a bare verb while 'chunk_search' and 'graph_walk' follow a noun_verb pattern. This mixed verb style is readable but not fully consistent.
Three tools is a focused set for a memory retrieval service, and each tool has a clear role. However, it sits at the lower boundary of typical well-scoped sets and could benefit from one or two more retrieval utilities.
The surface covers search over memories and chunks plus graph expansion, but lacks a direct get-by-ID operation (problematic since graph_walk needs a seed memory ID) and session/workspace listing. These are notable gaps for a memory service, though core search workflows are covered.
Maintenance
Related MCP Connectors
Persistent memory for AI agents. Search, store, and recall across sessions.
Shared memory for coding agents. Stop re-explaining your codebase every session.
Persistent memory and knowledge graphs for AI agents. Hybrid search, context checkpoints, and more.
Project memory, semantic code search, and grounded agent context.
Related MCP Servers
- AlicenseAqualityDmaintenanceProvides persistent memory for AI agents using hybrid search (vector embeddings + BM25) with neural reranking, enabling storage and retrieval of insights, debugging solutions, and patterns across coding sessions.8MIT
- AlicenseAqualityBmaintenanceProvides persistent, searchable memory across AI coding agent and chat history (Claude Code, Codex, Gemini CLI, ChatGPT, and more) via retrieval-augmented generation, enabling semantic and hybrid search to retain context across sessions.5MIT
- AlicenseBqualityCmaintenanceProvides local, agentic semantic recall over Claude Code session history, enabling the agent to search past discussions semantically, expand turns, and grep transcripts.512MIT
- AlicenseAqualityBmaintenanceEnables coding agents to search a local, offline archive of past sessions and repository Markdown docs using hybrid lexical and semantic retrieval, and read matching records to ground new work in prior knowledge without external calls.2MIT