Skip to main content
Glama

The Question

When you close a deal, ship a feature, or lose a client — how did it happen? Which decisions, in what order, against which context, on which hypotheses? Today's AI agents can't answer that. They follow skills and instructions perfectly, but they don't accumulate grounded knowledge about how outcomes actually arrived.

OpenExp captures every human-AI decision as a step in a trajectory, links those steps into coherent journeys, and grades each journey retroactively when reality returns its verdict — a deal closes, a sprint ships, a payment lands. The result is a continuously growing labeled dataset of decisions tied to outcomes, ready to train domain-specific intuition.

Related MCP server: Longhand

What It Is Not

  • Not a Q-learning memory system. We tried Q-values for 8 months. Mean Q-value across 27,000 memories was 0.006; 90% of memories never received any reward signal. Removed on 2026-04-26.

  • Not Mem0 / Zep / Letta. Those are storage layers. Storage is the easy part — semantic search alone doesn't tell you which memory actually led to a result.

  • Not a replacement for skills or CLAUDE.md. Those say how to do something. OpenExp captures what happened and how it ended.

The Methodological Core: No Pre-Labeling

We do not hand-craft features at step level (tone: urgent, signal: positive, hypothesis: probable). Pre-labeling injects the labeler's biases and corrupts the eventual training signal. Same hygiene as credit scoring: collect rich features per applicant, label only the terminal outcome (paid / didn't), let the model learn what predicts repayment from data alone.

Only terminal outcomes get labels:

  • outcome — closed_won / closed_lost / failed / abandoned

  • grade — 0.0 to 1.0, school-style

Steps are stored raw. Authors annotate their own intent, hypotheses, and decisions ("I believed X at this point", "I chose Y because Z"). They do not label the signal quality of individual events — that's what the eventual model learns.

Casual analogy: kids in school don't get annotations on every homework problem. They turn in work, get a grade at the end of the term, and develop intuition over hundreds of grades.

Quick Start

git clone https://github.com/anthroos/openexp.git
cd openexp
./setup.sh

That installs the four hooks into Claude Code, makes sure Qdrant is up, and registers the MCP server.

If something is already serving Qdrant on localhost:6333, the script uses it and leaves it alone. Otherwise it starts Qdrant in Docker for you.

Prerequisites: Python 3.11+, jq, and Qdrant — either Docker (the script handles the container) or a native Qdrant binary you run yourself.

No API key required for core functionality. Embeddings run locally via FastEmbed. An Anthropic API key is optional and only powers the two-prompt pipeline (anonymize + extract experience) when you publish.

How It Works

Four hooks run automatically inside Claude Code:

Hook

When

What

SessionStart

Session opens

Searches Qdrant for relevant memories, injects top results as context

UserPromptSubmit

Every message

Lightweight per-prompt recall

PostToolUse

After Write / Edit / Bash

Captures observations as JSONL

SessionEnd

Session closes

Ingests transcript into Qdrant; extracts decisions via Opus 4.x (async)

Retrieval ranks via semantic similarity + BM25 + recency. No magic numbers. No Q-value scoring component.

The Pipeline

When you decide to publish an experience — turn a real, terminal trajectory into a shareable artifact — two prompts do the work:

  1. prompts/anonymize.md — takes raw trajectory data (transcripts, emails, decisions) and produces an anonymized YAML trajectory. PII is replaced by category tokens (<counterparty_cto>, <regulated_industry>, <value:10k-100k>, <local_currency>, day_+5) while structural features are preserved. The prompt enforces a reverse-identification rule: tokens narrow enough to identify a counterparty in jurisdiction must be generalized one level up before publication.

  2. prompts/extract_experience.md — reads the anonymized trajectory plus the terminal outcome label and produces a facts-only meta.yaml (id, outcome label, duration, step count, category tokens, license). It deliberately refuses to write applies_when, searchable_summary, or a grade reason — those are interpretations and belong to the reader's Claude at use time, not to the publisher at publish time.

You run both prompts inside your own Claude Code, against your own Qdrant. Nothing is sent to a central server.

Publishing an Experience

A published experience is four files in a UUID-named directory (schema v3, 2026-04-27):

experiences/<uuid>/
├── meta.yaml                    # facts only: id, outcome label, duration, category tokens, license
├── trajectory.anonymized.yaml   # raw ordered timeline of N steps, anonymized
├── README.md                    # human-readable face for the marketplace
└── SKILL.md                     # Claude entry point — read first when skill is invoked

meta.yaml shape (abridged from seed d49e0997):

pack:
  id: d49e0997-8455-4d3c-90ca-d6cf54d0f662
  author: ivan-pasichnyk
  license: MIT
  schema_version: 3

  outcome:
    label: closed_won            # fact, not interpretation
    closed_at: day_+57

  duration_days: 57
  step_count: 26

  category_tokens:               # what appears in the trajectory
    - <counterparty_cto>
    - <counterparty_pm>
    - <regulated_industry>
    - <e_signing_platform_local>
    # ...

No applies_when, no searchable_summary, no grade_reason. Earlier schemas (v2) baked the publisher's read of the timeline into the artifact — one Claude's interpretation, frozen. Schema v3 inverts that: the pack ships raw, and the reader's Claude derives match on the fly against the reader's actual situation. Different readers, different contexts, different inferences from the same trajectory. See CHANGELOG.md for the full v2 → v3 transition rationale.

Install as a Claude Code skill

A published experience is a namespaced Claude Code skill:

openexp:<author-handle>:<experience-slug>

Drop the pack into ~/.claude/skills/openexp:<author>:<slug>/ (rename the directory to the skill-namespaced form on copy). Claude Code auto-discovers it on the next session.

# Install the seed pack as a skill
cp -r ~/openexp/experiences/d49e0997 \
  ~/.claude/skills/openexp:ivan-pasichnyk:inbound-acquisition-with-free-pilot

Two layers of identity:

  • Author identity is public — it signs the pack, like authorship on a research paper.

  • Counterparty identity stays anonymized — the skill name reveals who created the pack, never who they were dealing with.

SKILL.md inside the pack is the entry point — it tells the user's Claude when to invoke, how to use the trajectory, and what not to do (no fabrication, no de-anonymization, attribution required).

See docs/skill-architecture.md for the full naming convention, install flow, and design rationale.

The experiences/ directory in this repo is the seed of an eventual marketplace. Published packs are listed in CATALOG.md. The publication format works; seeds will accumulate. A directory of installable experiences is the eventual surface, not a built product today.

MCP Tools

Five focused tools (hippocampus model — write everything, retrieve selectively):

Tool

Description

search_memory

Hybrid search: semantic similarity + BM25 + recency

add_memory

Store a memory. Supports client_id for entity tagging

log_prediction

Log a pack-grounded prediction. Required when an installed experience pack cites a specific relative_day as the basis for an action recommendation.

log_outcome

Resolve a prediction with the observed signal — interpretation-free record.

memory_stats

Collection stats: point counts by source/type, session count

Prediction / outcome instrumentation

Pack-grounded predictions are how the system learns whether a published experience pack actually moves real-world outcomes. Without prediction/outcome pairs, pack value cannot be measured against any baseline, and any future experiment (cross-pack voting, embedding retrieval, new packs from new authors) is unfalsifiable.

Trigger criterion is sharp. Logging fires only when the assistant cites a pack's specific relative_day as the reason for an action recommendation. No day-citation → no log. Description of a situation without a recommendation → no log. This keeps the dataset honest and the cost low.

log_prediction (new path, schema_version 2)

Field

Required

Purpose

pack_id

yes

The pack's slug

pack_author

yes

Author handle

cited_step

yes

The exact day +N cited

case_id

yes

External reference (CRM lead_id, ticket ID, deal ID — opaque string)

applied_action

yes

What was recommended TO do

expected_signal

yes

Observable resolution

expected_window_days

yes

Deadline in days for log_outcome

prevented_action

optional

Negative-space prediction — what was recommended NOT to do (often the higher-value half)

notes

optional

Free-text context

log_outcome (new path, schema_version 2)

Field

Required

Purpose

prediction_id

yes

ID returned from log_prediction

actual_signal

yes

What was observed — raw fact, no interpretation

days_to_resolve

yes

How many days from prediction to resolution

notes

optional

Free-text, e.g. unexpected events

What's deliberately NOT in the schema: confidence (Claude-side confidence is uncalibrated until ≥30 outcome datapoints), alternative_action_if_no_pack and predicted_outcome_alternative (the same Claude that writes the prediction would invent the counterfactual, biased toward "the pack helped" — real ablation needs a pack-blind run, separate track).

Backward compatibility. The legacy schema (prediction, confidence, strategic_value, memory_ids_used) is still accepted by both tools and recorded as data. As of 2026-07-03 nothing updates Q-values — the Q engine is fully deleted. New-path entries are marked schema_version: 2 in the JSONL row.

CLI

openexp search -q "stalled enterprise procurement" -n 5
openexp ingest          # ingest pending transcripts into Qdrant
openexp stats           # collection + prediction stats

Configuration

Environment variables (.env):

Variable

Default

Description

QDRANT_HOST

localhost

Qdrant server host

QDRANT_PORT

6333

Qdrant server port

OPENEXP_COLLECTION

openexp_memories

Qdrant collection name

OPENEXP_DATA_DIR

~/.openexp/data

Predictions, retrieval logs

OPENEXP_OBSERVATIONS_DIR

~/.openexp/observations

Hook output

OPENEXP_SESSIONS_DIR

~/.openexp/sessions

Session summaries

OPENEXP_EMBEDDING_MODEL

BAAI/bge-small-en-v1.5

Embedding model (local, free)

ANTHROPIC_API_KEY

(optional)

Required only for the publishing pipeline

Status

Pilot. Architecture freeze landed 2026-04-26. First experience seed published: experiences/d49e0997/ — a 57-day inbound acquisition that closed at grade 1.0 (author's own assessment), anonymized to category tokens.

Honest about what isn't done:

  • The marketplace UI is just a directory in this repo. No web surface yet.

  • Anonymization is conservative but not bulletproof for readers with deep domain knowledge.

  • Schema may iterate — author-annotation fields (author_intent, author_hypothesis, author_decision) are a likely near-term addition.

  • The eventual ML model trained on this corpus does not exist yet. ≥30 graded trajectories first.

See docs/redesign-2026-04-26.md for the full architecture freeze and docs/claude-design-brief.md for the v2 product framing.

Contributing

This project is in early stages. See CONTRIBUTING.md for setup and workflow.

The most useful contribution right now is publishing a real experience. Take one of your own closed trajectories, run it through prompts/anonymize.md and prompts/extract_experience.md, and open a PR adding a new directory under experiences/.

License

MIT &copy; Ivan Pasichnyk

Available Tools

5 tools
add_memoryC

Store a new memory with FastEmbed embedding and LLM enrichment

ParametersJSON Schema
NameRequiredDescriptionDefault
typeNofact
agentNomain
contentYes
client_idNoAssociated client/entity ID

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It only states that it 'stores' a memory, implying a write operation, but fails to mention potential side effects (e.g., whether it overwrites existing memories, any limits, or required permissions). The technical detail about FastEmbed and LLM enrichment does not address behavioral consequences for the caller.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence, highly efficient, and front-loaded with the core action. It avoids redundancy and maintains a logical structure, though it sacrifices necessary detail for brevity. It is appropriately sized for the simplicity of the tool but could have included more context without becoming verbose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has 4 parameters with very low schema coverage, no output schema, and no annotations, the description is severely incomplete. It does not explain how parameters should be used, what the return value looks like, or any edge cases. An agent would have almost no guidance on how to invoke this tool correctly beyond guessing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 25% (only client_id has a description), and the tool description adds no parameter information at all. It does not explain the meaning or required format of 'content', 'type', or 'agent', nor does it clarify how they interact. The description completely fails to compensate for the sparse schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('Store a new memory') with a specific resource, and the tool name 'add_memory' aligns perfectly. The mention of FastEmbed embedding and LLM enrichment provides specific technical context that distinguishes it from siblings like search_memory or log_prediction, making its purpose unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage for storing memories but does not explicitly state when to use it versus alternatives. Siblings like search_memory and memory_stats have distinct purposes, and an agent could infer the correct choice from the name alone, but no explicit guidance on conditions or exclusions is provided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

log_outcomeA

Resolve a prediction with observed facts: provide actual_signal and days_to_resolve — an interpretation-free record of what happened. Legacy outcome/reward fields are accepted and recorded as data.

ParametersJSON Schema
NameRequiredDescriptionDefault
notesNoOptional free-text context (e.g. unexpected events)
rewardNo[deprecated — recorded as data, updates nothing]
outcomeNo[deprecated alias for actual_signal — accepted for backward compat]
actual_signalNoWhat was observed — raw fact, no interpretation. Required on the new path.
prediction_idYesID from log_prediction
cause_categoryNo[deprecated, accepted for backward compat]
days_to_resolveNoHow many days from prediction to resolution. Required on the new path.

TDQS

A3.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It discloses that legacy fields are 'accepted and recorded as data' and that the record is 'interpretation-free', which explains what happens to those fields. However, it does not state whether the action is reversible, whether the prediction is permanently marked as resolved, what happens if the prediction_id does not exist, or what the tool returns. These are meaningful gaps for a mutation tool with no annotation support.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is one concise, front-loaded sentence. It states the core purpose first ('Resolve a prediction with observed facts'), then the required fields, then the legacy behavior. No filler or redundancy; every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description provides enough to make a basic call (prediction_id plus actual_signal/days_to_resolve) and explains legacy field handling, but with no output schema and no annotations, it omits important context like return values, error conditions, idempotency, and whether resolution is final. For a tool with 7 parameters and a state-changing action, these gaps make it only partially complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description adds 'interpretation-free record' and the new-path vs legacy-path framing, but the schema already contains these details (e.g., 'Required on the new path' and 'deprecated — recorded as data, updates nothing'). The description does not significantly extend the semantic understanding beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Resolve') and resource ('a prediction'), and clarifies the tool's role by naming the core fields (actual_signal, days_to_resolve) and framing it as an 'interpretation-free record'. This clearly differentiates it from sibling tools like log_prediction, which creates predictions, and the memory tools, which manage memory.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear context: use this when you have observed facts to resolve a prediction, and instructs the caller to provide actual_signal and days_to_resolve. It does not explicitly mention alternatives or exclusion cases (e.g., 'use log_prediction instead'), but the purpose is specific enough that an agent can infer when to call it. No exclusions are stated, so it falls short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

log_predictionA

Log a pack-grounded prediction. REQUIRED whenever the assistant cites a specific relative_day of an installed experience pack as the basis for a real-world action recommendation. Captures: which step was cited, which case it applies to, what was recommended (and what was explicitly NOT recommended), the observable signal that resolves the prediction, and the window in days. Returns prediction_id for later log_outcome.

ParametersJSON Schema
NameRequiredDescriptionDefault
notesNoOptional free-text context
case_idNoExternal reference for this case (CRM lead_id, ticket ID, etc.) — opaque string
pack_idNoThe pack's slug (e.g. 'inbound-acquisition-with-free-pilot')
client_idNo[deprecated alias for case_id, accepted for backward compat]
cited_stepNoThe exact relative_day cited (e.g. 'day +57')
confidenceNo[deprecated, removed from required schema 2026-04-27 — Claude confidence is uncalibrated until ≥30 outcome datapoints. Accepted for backward compat.]
predictionNo[deprecated] Free-text prediction. Use applied_action + expected_signal instead. Accepted for backward compat.
pack_authorNoThe pack's author handle (e.g. 'ivan-pasichnyk')
applied_actionNoWhat the assistant recommended TO do, derived from the cited step
expected_signalNoObservable signal that resolves this prediction (e.g. 'counterparty signs both sides')
memory_ids_usedNoMemory IDs that were retrieved for this prediction (recorded as data)
strategic_valueNo[deprecated, accepted for backward compat]
prevented_actionNoWhat the assistant recommended NOT to do (negative-space prediction). Optional but encouraged — often the higher-value half.
expected_window_daysNoDeadline in days for log_outcome to be called

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are present, so the description carries the burden. It discloses that this is a logging operation, captures the prediction details, and returns a prediction_id. It also transparently explains deprecated fields (confidence, prediction, strategic_value) and the rationale for confidence deprecation, which adds behavioral context beyond a simple 'logs prediction.'

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, well-structured paragraph that front-loads the purpose and requirement, then lists captured fields without redundancy. Each sentence contributes value, and the deprecated-field notes are appropriately placed in the schema rather than bloating the description.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with 14 optional parameters and no output schema, the description covers the when, what, and return value (prediction_id). It doesn't detail error handling or edge cases, but given the logging nature and the explicit trigger, it provides sufficient context for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the schema already documents every parameter. The description adds semantic nuance by explaining the trigger (relative_day citation), emphasizing that prevented_action is 'often the higher-value half,' and clarifying the deprecated confidence field's uncalibrated nature. This goes beyond the schema's descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and object: 'Log a pack-grounded prediction.' It goes further to state the exact trigger condition ('REQUIRED whenever the assistant cites a specific relative_day...') and enumerates the data captured, which distinguishes it from the sibling log_outcome by clarifying it records the prediction itself, not the outcome.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly states when the tool is required (when citing a relative_day as basis for a recommendation) and implicitly routes the next step via 'Returns prediction_id for later log_outcome.' It does not enumerate exclusions or alternative tools, but the requirement framing is unambiguous.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

memory_statsA

Get memory system health: point counts by source/role, pending predictions, date range, Q-cache size

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. The verb 'Get' and the term 'health' strongly imply a read-only operation, but the description does not explicitly state 'no side effects', auth needs, or output size. It does add value by listing what stats are returned.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single sentence that front-loads the action and immediately defines the return scope. Every phrase adds useful information: source/role, pending predictions, date range, and Q-cache size.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a parameterless health-check tool with no output schema, the description provides enough context about what is returned. It could be more explicit about the output shape or any naming conventions of the metrics, but the listed categories give an agent a solid starting point.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, so the description cannot add parameter-level semantics. Per the rubric, 0 params earns a baseline of 4; no further parameter information is necessary.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Get') and resource ('memory system health') and enumerates the returned categories: point counts by source/role, pending predictions, date range, Q-cache size. This clearly differentiates it from sibling tools like search_memory and add_memory.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies a monitoring/health-check use case and the sibling names suggest alternatives, but it never explicitly states when to use it instead of search_memory, add_memory, log_prediction, or log_outcome. Agents must infer the distinction from the term 'health'.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_memoryA

Search memories with FastEmbed + Qdrant: hybrid semantic + BM25 + recency + importance scoring with lifecycle filtering

ParametersJSON Schema
NameRequiredDescriptionDefault
roleNoFilter by role: user or assistant
typeNoFilter by memory type
agentNoFilter by agent name
limitNo
queryYesSearch query
sourceNoFilter by source: transcript, decision, etc.
date_toNoEnd date (ISO format, e.g. 2026-04-08)
client_idNoFilter by client ID
date_fromNoStart date (ISO format, e.g. 2026-04-01)
session_idNoFilter by session ID

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the behavioral disclosure burden. It does disclose a read-only search with scoring and lifecycle filtering, but leaves unexplained the result ordering, pagination, lifecycle semantics, and any access or rate limits.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single dense sentence that front-loads the action and packs useful distinguishing details without filler. The main drawback is jargon like 'FastEmbed', 'Qdrant', and 'lifecycle filtering' is not unpacked, slightly reducing clarity for an agent.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 10-parameter search tool with no output schema and no annotations, the description is too thin: it does not describe the result shape, pagination, how filters combine, or what lifecycle filtering excludes. An agent can infer high-level intent but not safely predict call outcomes.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is high at 90%, so the baseline is 3; the description adds no per-parameter meaning beyond what the schema already provides. It does not explain how query, date filters, or role/source filters interact with the scoring behavior.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource ('Search memories') and names the retrieval/scoring strategy (hybrid semantic + BM25 + recency + importance), which clearly distinguishes it from siblings like memory_stats and add_memory. The purpose is unambiguous and not a tautology.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The search intent is clear from the name and description, but there is no explicit guidance on when to use this tool over sibling alternatives, nor any stated exclusions or conditions. Usage is implied rather than instructed.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 2 tool updatesv2.0.0
    • Changedlog_outcome1 field changed
      • changedInput schema / properties / reward / description
        Previous value: -"[deprecated — only used on the legacy Q-update path. Omit on the new path.]"New value: +"[deprecated — recorded as data, updates nothing]"
    • Changedlog_prediction1 field changed
      • changedInput schema / properties / memory_ids_used / description
        Previous value: -"Memory IDs that were retrieved for this prediction (for legacy Q-value updates on log_outcome)"New value: +"Memory IDs that were retrieved for this prediction (recorded as data)"
  2. 5 tool updatesv0.1.0
    • First observedadd_memory
    • First observedlog_outcome
    • First observedlog_prediction
    • First observedmemory_stats
    • First observedsearch_memory

TDQS

A3.7/5.0

Scored across 5 tools

Disambiguation5/5

Each tool maps to a distinct concern: memory system health, memory search, memory creation, prediction logging, and outcome resolution. The prediction/outcome pair is clearly separated by their lifecycle roles.

Naming Consistency4/5

search_memory, add_memory, log_prediction, and log_outcome all follow a verb_noun pattern. memory_stats deviates by starting with a noun, making get_memory_stats the more consistent equivalent.

Tool Count5/5

Five tools is appropriately scoped for a memory plus prediction-logging server. Each tool covers a necessary function without redundancy.

Completeness4/5

The memory surface covers add, search, and health stats, while predictions cover both logging and outcome resolution. Missing prediction retrieval/listing and memory mutation/deletion are minor gaps rather than critical dead ends.

Maintenance

ActivityMaintained
ResponsivenessUnresponsive

Related MCP Connectors

Related MCP Servers

  • A
    license
    A
    quality
    C
    maintenance
    Persistent memory for Claude Code. Automatically indexes every conversation and provides production-grade hybrid search (BM25 + vectors + reranker) via MCP tools. 100% local, zero config, zero API keys, zero invoice.
    16
    50 npm
    7
    MIT
  • A
    license
    A
    quality
    A
    maintenance
    Persistent local memory for Claude Code that indexes every session's JSONL file verbatim into SQLite + ChromaDB. Exposes 17 MCP tools for semantic recall, deterministic file replay, and fuzzy "do you remember when..." queries across your entire session history — no API calls, nothing leaves the machine.
    17
    367 PyPI
    14
    MIT
  • A
    license
    Not graded
    quality
    B
    maintenance
    Persistent memory for AI coding agents that stores and recalls preferences, decisions, and conventions via semantic similarity, with zero cloud dependencies and plug-and-play MCP integration for Claude Code.
    Apache 2.0