Skip to main content
Glama

License: MIT Python 3.10+ MCP PyPI MCP Registry

Get started · Why · Compared · How it works · Benchmark · Install · Quickstart · MCP · Fleets

Get started in 30 seconds

Claude Code plugin — MCP server plus a skill that tells the agent when to recall and when to record:

/plugin install cogvault --marketplace NBibikov/cogvault

Claude Code before 2.1.275: /plugin marketplace add NBibikov/cogvault, then /plugin install cogvault@cogvault.

Memory lives in ~/.cogvault/memory (set COGVAULT_TENANT to change it). The plugin adds /cogvault:remember and /cogvault:doctor.

Any MCP client, one line (needs uv):

claude mcp add cogvault -- uvx cogvault mcp --tenant ~/agent/memory

Add to claude_desktop_config.json, ~/.cursor/mcp.json, or your client's MCP config:

{
  "mcpServers": {
    "cogvault": {
      "command": "uvx",
      "args": ["cogvault", "mcp", "--tenant", "~/agent/memory"]
    }
  }
}

Already have Markdown memory? Point the tenant at it — nothing to import. Claude Code's auto-memory works as-is (frontmatter type, [[links]] and all):

M=~/.claude/projects/<project>/memory
uvx cogvault index  --tenant $M --ignore MEMORY.md   # first pass embeds, later passes are incremental
uvx cogvault search --tenant $M "how do we deploy"

The MCP server indexes on start by itself; the CLI search reads the existing index. --ignore MEMORY.md keeps the index file from competing with the cards it points to.


Related MCP server: auxly-memory-cli

Why

Most "AI agent memory" tools want to be an autonomous LLM daemon that summarizes your work into an opaque database or a graph you can't read. For a fleet of coding agents that just need to reliably recall a decision, a bug fix, or an infra detail, that's the wrong trade.

cogvault makes the opposite bet:

  • Your Markdown files are the source of truth. Open them, edit them, git diff them. The SQLite index is a derived cache — delete it and it rebuilds from the files.

  • One library, many tenants. Each agent gets an isolated memory namespace via its own directory. A process loads each embedding model once and shares it across every tenant it touches; with the stdio MCP server that means one small process per agent session, not a central daemon.

  • No LLM in the loop. Ingest and retrieval are deterministic. Your agent is the LLM — it doesn't need a second one to remember.

  • Local, private, offline. FastEmbed runs on-device. Nothing leaves your machine.

Compared to

Checked against each project's README and docs on 2026-10-05.

cogvault

basic-memory

mem0

Graphiti

Letta Code

Source of truth

Markdown files

Markdown files

Vector DB (Qdrant / pgvector)

Graph DB (Neo4j, FalkorDB, …)

Markdown in a git repo per agent

LLM needed to store or recall

No

No (optional reranker)

Yes by default (add() extracts facts)

Yes to ingest

Yes — the agent edits its memory

Retrieval

Vector + BM25, RRF, decay, MMR

Full-text + vector, optional rerank

Semantic + BM25 + entities

Semantic + BM25 + graph

File search; hybrid optional

Infra

One process, SQLite file

One process, SQLite (Postgres optional)

Library, or Docker + Postgres server

A graph database

Letta backend

Pick something else when: you want an LLM to distil and merge facts for you (mem0), relationships between entities are the point (Graphiti), you want the agent to manage its own memory (Letta), or you want a richer notes app around the same Markdown idea, with Obsidian sync and a hosted option (basic-memory — the closest to cogvault).

Pick cogvault when: you run several agents and want each one's memory isolated in its own directory behind one process; you want recall to be deterministic and offline; and you want to measure it — cogvault eval scores recall on your agents' real queries and cogvault analyze lists what they tried to recall and couldn't.

How it works

Anatomy of a recall

Hybrid retrieval fuses semantic (vector) and keyword (BM25/FTS5) ranking with Reciprocal Rank Fusion, then applies optional temporal decay (recent memory outranks stale) and MMR (diverse top results, not five near-duplicates). Each card contributes only its best chunk, so one long file can't fill the whole result list.

Benchmark

Measured on real recall traffic, not synthetic questions: 66 queries sampled from the query logs of 6 live agent tenants (58% Ukrainian, the rest English), each judged against the actual cards — including answers that no configuration returned. One query has no answer in memory and counts as a gap, so 65 are scored. Every configuration was re-indexed from scratch on copies of the same tenants with cogvault 0.11.0.

Configuration

hit@1

hit@5

MRR@10

multilingual-e5-small, chunk_chars = 700, summary chunk

0.57

0.91

0.70

multilingual-e5-small, chunk_chars = 700, no summary chunk

0.55

0.86

0.69

paraphrase-multilingual-MiniLM-L12-v2 (built-in default)

0.54

0.80

0.66

bge-small-en-v1.5 (English-only)

0.57

0.77

0.65

What the numbers do and don't say:

  • hit@1 is a tie. All four land within 0.54–0.57, and the 95% bootstrap intervals overlap almost completely. Real agent queries read like card titles, so the right card usually wins on its name alone.

  • The gap is in the top 5. e5-small with the summary chunk puts the answer in the top 5 for 91% of queries vs. 80% for the default MiniLM and 77% for English-only bge on this mixed-language memory. That's what an agent reading 5 results actually feels.

  • 65 queries is still a small sample. Treat differences under ~0.1 as noise. The aggregate numbers are in assets/benchmark.json; the queries are private and stay in each tenant.

Run the same check on your own memory: put judged queries in <tenant>/.cogvault-golden.jsonl ({"query": "...", "relevant": ["file.md"]}, empty relevant = a known gap) and run cogvault eval --tenant DIR.

Choosing an embedding model

Agent memory is often not English-only. The default is multilingual so nothing is broken out of the box — but pick the model that matches your fleet's language mix (set COGVAULT_MODEL, or Config(model=...)). Switching models auto-rebuilds the index.

Model (COGVAULT_MODEL)

Dim

Size

Real-query hit@5*

Cyrillic / multilingual

When

paraphrase-multilingual-MiniLM-L12-v2 (default)

384

0.22 GB

0.80

✅ works

Mixed-language fleets; safe default

BAAI/bge-small-en-v1.5

384

0.13 GB

0.77

❌ Cyrillic vectors break

English-only memory

intfloat/multilingual-e5-small

384

0.47 GB

0.91

✅ best per GB (512-token window)

Mixed-language fleets; use chunk_chars = 700

intfloat/multilingual-e5-large

1024

2.24 GB

not measured

✅ best

Max quality, RAM to spare

*From the benchmark above: 65 real queries over mixed EN/UK memory, cogvault 0.11.0. On English-only memory bge-small-en is a fine choice; on Ukrainian content it returns a negative relevance margin (a distractor outranks the answer), so it is unsafe for non-English memory. Run cogvault eval on your own vault to decide.

COGVAULT_MODEL=BAAI/bge-small-en-v1.5 cogvault index --tenant ~/agent/memory

Pin the model per tenant so it travels with the data instead of relying on every command exporting COGVAULT_MODEL (forget it once and a model mismatch silently re-embeds the whole index). Drop a .cogvault.toml at the tenant root:

# ~/agent/memory/.cogvault.toml
model = "BAAI/bge-small-en-v1.5"
# optional: recursive = true, strip_frontmatter = true, ignore_globs = ["Templates/*"]

Now cogvault search --tenant ~/agent/memory "…" uses the right model with no env var. Precedence: explicit --model / $COGVAULT_MODEL > .cogvault.toml > built-in default.

Install

uv tool install cogvault        # CLI on PATH
# or
pip install cogvault

Or skip installing and run it on demand with uvx cogvault …. The first run downloads the embedding model (~0.2–0.5 GB, once per machine).

Add it to Claude Code as an MCP server in one line:

claude mcp add cogvault -- uvx cogvault mcp --tenant ~/agent/memory

Also listed in the official MCP Registry as io.github.NBibikov/cogvault. Wheels are attached to each GitHub release.

Quickstart

# index a tenant's markdown memory
cogvault index --tenant ~/agent/memory

# search (hybrid semantic + keyword)
cogvault search --tenant ~/agent/memory "how do I restart the worker service"

# enable temporal decay (recent wins) and tune diversity
cogvault search --tenant ~/agent/memory "deployment steps" --half-life 30 --mmr 0.5

# only cards of one frontmatter type (user / feedback / project / reference / …)
cogvault search --tenant ~/agent/memory "hard rules for deploys" --type feedback

Memory cards

cogvault understands two lightweight Markdown conventions (both optional — plain files index fine):

  • Frontmatter type — either flat (type: reference) or nested (metadata: → type: reference). Parsed at index time and filterable at search time (--type, MCP type param, search(card_type=...)). Every hit carries a type field; cards without frontmatter get null.

  • [[wiki-links]] — link targets are indexed, and the top search result includes a related list of linked cards that exist in the index (ghost links are dropped; matching is by exact filename stem).

As an MCP server (Claude Code, Cursor, any MCP client)

claude mcp add cogvault -- uvx cogvault mcp --tenant ~/agent/memory

Exposes two tools:

  • cogvault_recall — natural-language hybrid search over this agent's memory. Optional type param filters to one frontmatter card type; the top result includes a Related: line built from its [[wiki-links]].

  • cogvault_record — save a fact; it's written as a Markdown card and indexed

Indexing a folder tree (Obsidian vaults, knowledge bases)

By default a tenant is one flat directory of .md files. For a nested vault (e.g. Obsidian, with 01-Projects/…, frontmatter, and folders to skip), opt in:

cogvault index --tenant ~/vault \
  --recursive \
  --strip-frontmatter \
  --ignore ".obsidian/*" --ignore ".trash/*" --ignore "Templates/*"
  • --recursive walks subdirectories; files keep their path relative to the tenant, so two notes named Tasks.md in different folders never collide.

  • --strip-frontmatter drops a leading YAML --- … --- block so its keys don't pollute the embedding.

  • --ignore GLOB (repeatable) skips paths relative to the tenant root.

Same flags exist on search and mcp, and as Config(recursive=True, strip_frontmatter=True, ignore_globs=(...)) for the library. Indexing is incremental: the first pass embeds everything, later passes only re-embed changed files. (Reference: a ~3,500-note vault → ~9,500 chunks, first index ≈ 3–4 min, then warm recall in single-digit milliseconds.)

As a library

from cogvault import Vault, Config

vault = Vault("~/agent/memory", Config(half_life_days=30))
vault.reindex()
for hit in vault.search("where are credentials stored"):
    print(hit["score"], hit["file"], hit["snippet"])

Keeping a tenant healthy

cogvault doctor --tenant DIR reports what silently degrades recall: cards with no frontmatter or type, legacy timestamp filenames, frontmatter wrapped inside frontmatter, duplicate name: slugs, and [[links]] that resolve to nothing. Links resolve by frontmatter name:, filename stem, either separator style, and with or without the card-type prefix; links inside code and paths to files outside the tenant are not counted.

cogvault repair --tenant DIR fixes the mechanical half (dry run by default, --apply to write): unwraps nested frontmatter, infers a missing type, renames card-<timestamp>-….md to <type>_<slug>.md and rewrites every reference to it (MEMORY.md included), and adds minimal frontmatter to <type>_*.md cards that lack it. Healthy cards are left alone, and repaired cards keep their mtime so temporal decay is not reset.

Effectiveness logging

Every recall is logged (one JSONL line) so you can measure whether the memory is actually helping. cogvault analyze turns the log into a report:

cogvault analyze            # recalls, no-hit rate, latency p50/p95, avg top score
cogvault analyze --json     # machine-readable

The no-hit rate and recent no-hit queries are the signal that matters: they tell you what your agents tried to recall and couldn't — i.e. the memory gaps to fill. Set COGVAULT_LOG=off to disable, or COGVAULT_LOG=/path.jsonl to relocate.

Multi-tenant fleets

Point one process at many tenants — each directory is an isolated namespace, proven by the test suite (test_multi_tenant_isolation). Agent B can never recall Agent A's memory unless you point B at A's directory.

Design notes

Decision

Why

Markdown = source of truth

Human-readable, git-versionable, editable, never locked in a DB

SQLite + sqlite-vec + FTS5

Zero-infra hybrid search; one portable .db file; rebuildable

FastEmbed (multilingual MiniLM default, 384-d)

In-process ONNX, no server, no API key, ~220 MB

Content-hash cache

Re-indexing only embeds changed chunks

RRF + decay + MMR

Precision, recency, and diversity without a graph DB

One process, many tenants

Fleet infra, not a single-user desktop sidecar

WAL + incremental reindex

Concurrent agents read while one writes; only changed files re-embed

Roadmap

  • Importers (migrate existing memory from other stores)

  • Pluggable embedders (Ollama, OpenAI-compatible endpoint)

  • valid_until per-card temporal validity

  • Optional FastMCP transport

License

MIT — your memory, your files, your infrastructure. Forever.

Figures are hand-built SVG from assets/make_graphics.py — python assets/make_graphics.py --png regenerates them.

Available Tools

2 tools
cogvault_recallA

Search this agent's persistent memory. Pass a natural-language query; returns the most relevant memory snippets (hybrid semantic + keyword).

ParametersJSON Schema
NameRequiredDescriptionDefault
typeNoOnly recall cards of this frontmatter type (e.g. user, feedback, project, reference)
limitNo
queryYesWhat to recall

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations exist, so the description carries the full burden. It usefully discloses the retrieval mechanism ('hybrid semantic + keyword') and the return shape (snippets), but omits safety-relevant and operational traits like pagination, ordering, or whether results are scoped/searched exhaustively.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences, front-loaded with the purpose and then the return/mechanism. Every clause earns its place with no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple 3-parameter read tool with no output schema, the description covers purpose and return format adequately, but leaves the `type` filter and `limit` behavior to the schema and gives no guidance on result volume or refinement.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 67% (query and type documented; limit not). The description adds the useful semantic that the query is natural-language, but says nothing about the `type` filter or how `limit` bounds results, so it does not fully compensate for the gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource ('Search this agent's persistent memory') and even describes the result ('most relevant memory snippets'). It is clearly distinct in intent from the sibling cogvault_record, but it never explicitly contrasts itself with that write-oriented sibling, so it stops short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives one concrete usage cue ('Pass a natural-language query'), which implies how to call it, but offers no when-to-use/when-not framing and never names the alternative cogvault_record for storing memories. Usage is only implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

cogvault_recordA

Save a fact to persistent memory as a Markdown card. It becomes searchable on the next recall. Pass type so the card can be filtered on recall, and title so it gets a meaningful filename.

ParametersJSON Schema
NameRequiredDescriptionDefault
typeNoCard type: project (ongoing work/decision), feedback (a rule or correction), reference (pointer to a resource), user (who the user is)
titleNoShort title — becomes the card's name and filename
contentYesThe fact to remember

TDQS

A3.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden, and it does disclose one non-obvious trait: the card is only searchable on the next recall, implying indexing latency. However, it omits what happens on duplicate content (overwrite vs. append), permission requirements, persistence guarantees, and any return value — meaningful gaps for a write tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with the core action, then the guidance for the optional parameters. No filler and nothing duplicated from the schema descriptions.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple three-parameter write tool with no output schema, the description covers the action, the persistence model, and the purpose of both optional fields. It lacks only edge-case semantics such as duplicate handling, which is a minor gap rather than an omission an agent would trip over.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds rationale the schema lacks: `type` exists so the card can be filtered on recall, and `title` becomes the filename. That explains why the optional params matter rather than just restating them.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a concrete verb ("Save a fact") and resource ("persistent memory as a Markdown card"), plus the downstream consequence that it becomes searchable on recall. It contrasts functionally with cogvault_recall (write vs. retrieve) but never names the sibling, so it falls short of the explicit differentiation a 5 requires.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is only implied: the agent can infer it should call this when it wants a fact retained. There is no statement of when not to use it (e.g., duplicates, updates to existing cards) and no explicit alternative routing despite having exactly one sibling.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 2 tool updatesv0.11.2
    • First observedcogvault_recall
    • First observedcogvault_record

TDQS

A3.7/5.0

Scored across 2 tools

Disambiguation5/5

The two tools are a clean read/write pair: recall searches memory while record writes to it. There is no plausible way to confuse them, and their descriptions reinforce the distinction.

Naming Consistency5/5

Both tools use the identical `cogvault_` prefix followed by a single clear verb (recall, record). The pattern is perfectly predictable and readable.

Tool Count3/5

Two tools is a thin surface even for a focused memory store; the domain of persistent memory naturally invites a few more operations. It is defensible as a minimal core, but sits below the well-scoped 3-15 range.

Completeness3/5

The server covers create and read but has no update, delete/forget, or list operations, leaving common memory-lifecycle tasks (correcting a stale fact, removing sensitive data) as dead ends. Core recall/record workflows do work.

Maintenance

ActivityNo data
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    A
    maintenance
    Local-first, file-based memory layer for AI agents — one shared Markdown vault across Claude, Codex, Gemini, Cursor and any MCP client. Provides read/write memory tools with an audit trail, per-agent trust levels, and Git sync; no cloud and no lock-in.
    2
    MIT
  • A
    license
    Not graded
    quality
    A
    maintenance
    Enables AI agents to capture, structure, remember, and retrieve source-backed memory as local Markdown files, with reviewable writes and no cloud dependency.
    190
    MIT