Skip to main content
Glama

🧠 M3 Memory

A memory layer that outlives your agents. You switch from Claude Code to Cursor, upgrade your model, start fresh next week — and everything your tools learned about your project is gone. You re-explain the same decisions, the same preferences, the same hard-won context, over and over.

M3 fixes that. It's a private, local-first memory your agents share and build on — so your project's knowledge accumulates instead of resetting every time the agent does. One memory store, on your machine, that your tools and agents read from and write to — whether that's Claude Code, Cursor, Gemini CLI, or any MCP-compatible agent.

Under the hood, M3 treats agent memory as a distributed-systems infrastructure problem, not a simple retrieval feature — a shared, evolving, bitemporal, contradiction-aware knowledge base that multiple heterogeneous agents and machines read and write, built to stay consistent over months and years.

It runs where your data has to stay. A single pip install with no account, no API key, and no outbound calls — at home in a homelab, on a corporate or government network, or fully air-gapped. The embedder runs in-process and local, the store is a file you own, and installation works with no internet at all. On the metric that isolates the memory layer — retrieval accuracy, no answer model or judge involved — M3 reaches 99.2% session-hit-rate @ k=10 and 100% @ k=20 on LongMemEval-S.


šŸŽ¬ Quick video overview

One decision saved from a conversation, then recalled by a different agent in a new session, on a different machine. Captioned throughout, so it reads fine muted.

https://github.com/user-attachments/assets/09ab194a-d2a0-4fe5-a7db-69ae8225e39b

Player not loading? Download the video to play locally.


Related MCP server: memento

⚔ Quickstart

pip install m3-memory
m3 setup            # detects your agents, wires the MCP server, provisions the local embedder
m3 doctor           # verify: health, memory count, embedder, and which agents got wired

That's the whole install. No cloud account, no API key, no external embedding service.

What it does, in four lines

Save a decision — from any agent, or straight from the shell:

$ m3 memory memory_write --type decision --title "auth-jwt-algorithm" \
    --content "The auth service uses RS256 JWTs. HS256 was rejected because we need asymmetric verification at the edge."
"Created: 84a944fb-ef3e-403b-9240-f53ab3c015f7"

Next week, in a different agent, on a different model — ask in your own words:

$ m3 memory memory_search --query "which signing algorithm did we pick for tokens?" --k 3
Top 1 results:
----------------------------------------
1. [84a944fb-ef3e-403b-9240-f53ab3c015f7] score=0.7501  type: decision  title: auth-jwt-algorithm
Content:
The auth service uses RS256 JWTs. HS256 was rejected because we need asymmetric verification at the edge.
----------------------------------------

The query shares no keywords with the stored text — no "RS256", no "JWT" — and still finds it. That's the hybrid engine: BM25 for exact terms, local BGE-M3 vectors for meaning, MMR for diversity. Your agent calls the same tools over MCP, so it recalls this automatically instead of asking you again.

New here? The 5-Minute Getting Started Guide walks the same path with more context, and Core Tools lists the five you'll use most.


🧩 Beyond the core

The Quickstart above is the whole product for most people: shared memory, wired into your agents, working offline. Everything below is optional surface you can ignore until you want it — each row says what it costs to turn on.

Also a drop-in memory backend for LangChain / LangGraph, CrewAI, and PydanticAI — see the framework guides.

Every path gains automatic contradiction supersession, bitemporal historical queries, local sovereign embedding, and the full 100+ MCP tool set.


āš–ļø How M3 Compares

A full, feature-by-feature comparison table — M3 vs Mem0, Letta, Zep, Graphiti, LangChain Memory / LangMem, agentmemory, Chronos, Hindsight, Mastra OM, Memento, and more — with sourced benchmarks and honest "when to choose the other tool" guidance, lives in COMPARISON.md.

Short version: M3 is the local-first, MCP-native option that stays yours and works across every agent — where cloud services (Mem0), full agent runtimes (Letta), and graph-database systems (Zep, Graphiti) each ask you to adopt their infrastructure. See the comparison guide for the row-by-row detail.


šŸ’” Get Started Quickly:


šŸ“‘ Table of Contents


⚔ M3 at a Glance

Feature

Details

Works With

Claude Code Ā· Cursor Ā· Cline Ā· Gemini CLI Ā· Aider Ā· Google Antigravity Ā· OpenCode Ā· OpenClaw Ā· Hermes Ā· LangChain/LangGraph Ā· CrewAI Ā· PydanticAI Ā· Any MCP Agent

M3 Is

A persistent memory layer Ā· An MCP server Ā· A hybrid retrieval engine Ā· A bitemporal knowledge base

M3 Is Not

An LLM Ā· A chatbot Ā· A plain vector database Ā· A RAG framework Ā· An IDE

Core Promise

Private, offline-capable, locally owned memory shared securely across all your developer tools — with FIPS 140-3-ready crypto and atomic multi-agent writes for regulated and multi-agent environments.

Deploys In

Homelabs and self-hosted stacks Ā· corporate and government networks Ā· air-gapped and classified environments Ā· regulated industries (FIPS 140-3-ready, GDPR tooling, audit logs). No account, no API key, no outbound calls. See Sovereign & Air-Gapped Deployments.

Retrieval Accuracy

State-of-the-art for a local-first substrate — 99.2% session-hit-rate @ k=10, 100% @ k=20 on LongMemEval-S (no oracle routing), with a gold session as the #1 result for 91.8% of questions. SHR measures the memory layer alone — no answer model, no judge — which is why it, not end-to-end QA, is the like-for-like comparison between memory systems. See Benchmarks.

Context Efficiency

Exposes 100+ tools but occupies just ~1.8% of a 200K context window at startup — lazy domain-gating loads the rest on demand.

Maturity

Stable, battle-tested core engine (2,400+ tests) that's safe to build on today; new features and integrations are added actively. SQLite by default; PostgreSQL as a first-class primary backend (M3_DB_BACKEND=postgres) via a pluggable SQL storage seam. (See features.json)


🧠 Memory Model at a Glance

M3 is a typed, bitemporal, confidence-scored, self-maintaining knowledge base. Every feature listed below is implemented natively (see Memory Model Details):

  • Structured Metadata: Every memory contains a type, source, confidence, scope, provenance (change_agent), and salience (importance, decay_rate).

  • Verbatim, Non-Destructive Storage: Memory content is stored exactly as written and never altered in place — the raw text is always retrievable byte-for-byte. Corrections don't overwrite: a superseded fact is closed (its validity interval ends) and the new fact is linked to it, so both the original wording and its full edit history stay queryable. You get true verbatim recall and an audit trail, not one or the other.

  • Bitemporal History: Distinguishes valid-time from transaction-time. Because superseded facts are closed rather than deleted, you can query what the agent believed at any specific point in time.

  • Contradiction Management: Conflicting facts are resolved automatically on write. The stale fact is marked as superseded, and confidence values are updated dynamically via Bayesian confidence posteriors. Supersession fires above a deliberately conservative cosine bar (CONTRADICTION_THRESHOLD, default 0.92), so near-restatements of a claim close the old fact while genuinely different-but-related facts are both kept — use memory_supersede to close one explicitly. (See Technical Details.)

  • Self-Maintaining Lifecycle: Implements memory decay, deduplication, automatic consolidation into higher-order beliefs, TTL expiry, and GDPR erasure.

  • Procedural Memory: A first-class procedure type (skill / runbook / how-to / checklist) that is auto-distilled from successful task runs — the background loop rolls up a completed task and its step/result memories into a reusable, step-by-step procedure, preserved with distills_from provenance back to its sources. A "how do I…" query surfaces it via a procedural retrieval boost.

  • Write-Gating & Content Safety: Filters out low-signal noise via an enrichment queue and content safety guardrails before storage.

  • Explainable Retrieval: Hybrid engine combining vector similarity, BM25 (FTS5), MMR diversity, and reranking. memory_suggest returns the exact score breakdown per result. (See Confidence and Trust Guide).

  • Proven Accuracy: On LongMemEval-S, M3 delivers state-of-the-art retrieval for a local-first substrate — 99.2% session-hit-rate @ k=10 and 100% @ k=20 (no oracle routing), with a gold session as the #1 result for 91.8% of questions. End-to-end QA accuracy is 92.0% with no oracle metadata (see Benchmarking Report).


šŸ“¦ Installation

The Quickstart above covers the common path (pip install m3-memory → m3 setup). This section adds the alternatives: the shell installer, per-agent wiring, and manual MCP configuration.

The One-Liner (macOS & Linux)

curl -fsSL https://raw.githubusercontent.com/skynetcmd/m3-memory/main/install.sh | bash

Developer Setup Wizard

If you are developing inside python environments:

pip install m3-memory
m3 setup

The m3 setup wizard automatically detects your installed agents — Claude Code, Cursor, Cline, Gemini CLI, OpenCode, Antigravity, OpenClaw, Hermes — and wires the m3 memory MCP server into each, installs settings files/hooks, provisions the sovereign CPU embedder, and performs a system diagnostic. Detection and wiring re-run on every m3 update/m3 setup, and m3 doctor --fix repoints any config whose paths have moved — so an agent you install later gets picked up automatically the next time you run setup or update.

Integrating with AI Coding Tools

šŸ¤– Claude Code

Install as a plugin to unlock /m3:* slash commands, curation subagents, and automatic hooks:

/plugin marketplace add skynetcmd/m3-memory
/plugin install m3@skynetcmd

See Claude Code Plugin Reference and Claude.ai Connector Guide.

ā–· Cursor

Auto-detected and wired by the setup wizard — it writes the m3 memory MCP server into ~/.cursor/mcp.json:

m3 setup

Re-run after installing Cursor and it's picked up automatically; m3 doctor --fix repoints the entry if paths move. See MCP Client Install Guide.

ā—§ Cline (VS Code)

Auto-detected and wired by the setup wizard — it writes the m3 memory MCP server into Cline's cline_mcp_settings.json:

m3 setup

Also available from Cline's MCP marketplace (see llms-install.md). See MCP Client Install Guide.

🪐 Google Antigravity

Install the plugin directly:

agy plugin install https://github.com/skynetcmd/m3-memory

See Antigravity Plugin Reference.

🦊 Hermes Agent

Run the wizard to automatically wire up optimal memory providers:

m3 setup

See Hermes Plugin Integration Guide.

šŸ Python / LangChain & LangGraph

Use M3 as a drop-in Mem0 replacement or LangMem backend:

pip install m3-memory[langchain]

See LangChain Integration Guide.

šŸ‘„ CrewAI (v1.x)

A drop-in StorageBackend for CrewAI's unified memory:

pip install m3-memory[crewai]   # crewai>=1.10,<2 Ā· Python 3.10–3.13 (a 3.14 escape hatch is documented)

See CrewAI Integration Guide.

🧩 PydanticAI

m3 tools + auto-recall, or a formal M3MemoryToolset. Built on Pydantic v2 — runs natively on Python 3.14:

pip install m3-memory[pydantic-ai]   # pydantic-ai-slim>=2,<3

See PydanticAI Integration Guide.


Manual MCP Server Configuration

To expose M3 to any Model Context Protocol host, add it to your configuration file:

{
  "mcpServers": {
    "memory": {
      "command": "m3"
    }
  }
}

šŸŽšļø Domain Gating: the Full Catalog Without the Context Cost

M3 gives you the full 100+ tool surface while occupying just 1.8% of a 200K context window at startup — most MCP servers make you pay for every tool in every prompt. Tools are grouped into 9 domains (memory, chatlog, files, entity, agent, tasks, conversations, diagnostics, admin) and loaded lazily.

Only the essential core set (~18, ~3,540 tokens) registers at startup. When your agent needs advanced functionality, it calls tools_load_domain(domain="...") to fetch the rest on demand — so a large catalog costs near-zero context until you actually use a domain.

Gating Mode

Registered Tools

Tokens in Schema

% of 200K Window

Lazy (Default)

~18

~3,540

1.8%

Typical Active Session

64

~17,975

9.0%

Eager Mode (M3_TOOLS_LAZY=0)

110

~24,918

12.5%

šŸ› ļø Note: If your client does not support dynamic tool registration, set the environment variable M3_TOOLS_LAZY=0 to register all tools eagerly.


šŸ›”ļø Sovereign & Air-Gapped Deployments

M3 operates completely offline by default.

Sovereign Local Embedder

A high-performance BGE-M3 embedder runs locally after installation.

  • Default: in-process via the m3-core-rs native module (llama.cpp linked in-process, zero IPC — not a separate service you have to run or monitor). CPU execution using GGUF format (_assets/models/bge-m3-Q4_K_M.gguf). A local HTTP embed server on 127.0.0.1:8082 exists only as an automatic fallback if the in-process path can't load.

  • Hardware Acceleration (GPU): Execute m3 embedder install-gpu to compile with CUDA, Vulkan, or Metal.

  • External Provider Fallback: Set EMBED_BASE_URL to route requests to Ollama, LM Studio, or vLLM.

Rust-Oxidized Performance Core

M3 includes an optional Rust performance module (m3_core_rs) that speeds up MMR re-ranking, batch cosine distance calculations, and FTS compilations by 90Ɨ to 800Ɨ. If absent, M3 falls back to pure Python execution automatically. Disable with M3_CORE_RS_DISABLE=1. (See Oxidation Benchmarks).

Enterprise Security & Compliance

  • FIPS 140-3 Ready: Standardized encryption pathways allow routing through validated cryptographic modules (e.g., wolfSSL via M3_FIPS_MODE=1).

  • Air-Gapped Install: Supports installation without internet access via pre-compiled python wheels. (See Sovereign Deployment Guide & FIPS Boundary Reference).

  • Storage Location: State lives under three roots, so databases and configuration can be relocated and secured independently:

    Root

    Default

    Holds

    M3_ENGINE_ROOT

    ~/.m3/engine

    Databases + runtime state (agent_memory.db, agent_chatlog.db, files_database.db)

    M3_CONFIG_ROOT

    ~/.m3/config

    Configuration (chatlog config, salt)

    M3_MEMORY_ROOT

    ~/.m3-memory

    Payload / repo clone

    All three are overridable. Set any of them to relocate that root. M3_MEMORY_ROOT also acts as a master override — if set and the other two are unset, engine and config derive from it as <root>/engine and <root>/config. Precedence is M3_ENGINE_ROOT / M3_CONFIG_ROOT → M3_MEMORY_ROOT/… → the ~/.m3/… default, so a specific root always wins over the master. (See Architecture.)


šŸ”® What M3 Does

  • Memory Persistence: Saves system architecture, project decisions, and preferences across tool boundaries using a local SQLite database.

  • Autonomous Cognitive Loop: Background worker (m3_cognitive_loop.py) that periodically sweeps chat logs to extract facts, reconcile contradictions, and construct an entity relationship graph.

  • Hybrid Vector & Keyword Search: Seamlessly merges vector space, Full-Text Search (FTS5 BM25), and MMR diversity.

  • Hierarchical File Ingestion: A dedicated 26-tool files domain reads directories, chunks files, extracts facts, and reviews staleness — with ~4Ɨ faster incremental re-ingest (unchanged sections reuse cached embeddings).

  • Verbatim Chatlog Capture: A dedicated 10-tool chatlog domain records conversation turns before compaction, so prior Claude/Gemini sessions stay searchable and nothing is lost to context-window truncation.

  • Pluggable Storage Backend: SQLite by default; select PostgreSQL as a first-class primary store with M3_DB_BACKEND=postgres. Same semantics on either backend — the choice doesn't change behavior.

  • Cross-Device Sync: Optionally sync/federate to a PostgreSQL warehouse tier. Access the same memories on your laptop, desktop, or cloud environments.


šŸ“š Documentation Index

Start here, in this order: Getting Started → Memory Model (what a memory is, and how supersession works) → Agent Instructions (how to make your agent use it well). Everything else below is reference — reach for it when you hit the specific thing it covers.

Quick & Core

Advanced & Architecture

Integrations & Compliance

šŸš€ Getting Started Guide

šŸ—ļø System Architecture

🧩 LangChain/LangGraph

✨ Core Features

šŸ”§ Technical Implementation

🧩 Hermes Agent

āš™ļø Environment Variables

🧠 Memory Model Guide

šŸ›”ļø Compliance Guide (GDPR, FISMA)

šŸ› ļø Operations Playbook

⚔ Rust Oxidation benchmarks

šŸ›”ļø FIPS Cryptographic Boundary

šŸ¤– Agent Instructions & Rules

šŸ” Myths & Facts Guide

šŸ  Homelab Patterns

🧩 Tool Capability Matrix

šŸ¤– AI Context Injection Profile

šŸ”¢ Machine-Readable Features

More Documentation

Guide

Guide

Guide

šŸ—ŗļø Roadmap

šŸ”„ Cross-Device Sync

šŸ‘„ Multi-Agent Orchestration

āš–ļø Comparison vs Alternatives

ā“ FAQ

šŸ” Security Policy

🩹 Troubleshooting

āŒØļø CLI Reference

šŸ“– API Reference

šŸ“ Files Memory

šŸ’¬ Chat Log Subsystem

✨ Enrichment Guide

ā¬†ļø Upgrade Guide

🩺 Health FAQ

🧬 Dual Embedding

šŸ“œ Changelog

šŸ¤ Code of Conduct

šŸ—ļø Build Wheels


šŸŽÆ Who This Is For

M3 is a great fit if...

  • You run a homelab or self-hosted stack: M3 is a single pip install with no account, no API key, and no outbound calls — it runs on the hardware you already own, alongside your other self-hosted services. SQLite by default (zero infrastructure); PostgreSQL when you want a shared store across machines.

  • You operate under sovereignty or data-residency requirements — corporate, government, defence, healthcare, or any regulated environment: memory and embeddings never leave your boundary. The embedder is in-process and local, the store is a file you control, and installation works fully air-gapped from pre-compiled wheels. FIPS 140-3-ready crypto (M3_FIPS_MODE=1), GDPR gdpr_forget / gdpr_export, audit logs, and relocatable storage roots so databases and configuration can be secured independently.

  • You want the freedom to switch or add agents without losing what they know: change tools on the fly or down the road — Claude Code, Gemini, OpenClaw, Hermes, whatever comes next — and your project's knowledge carries over instead of disappearing with the switch.

  • You build with LangChain/LangGraph: An advanced replacement for standard memory models, adding bitemporal queries, contradiction management, and local embeddings.

  • You build with CrewAI (v1.10–1.x): A drop-in StorageBackend (Memory(storage=M3StorageBackend(user_id="crew-alpha"))) that gives CrewAI bitemporal recall, contradiction-aware supersession, and local embeddings — plus the thing single-vector stores can't do: a CrewAI-written memory can also be searchable by every other m3 agent (Claude Code, Gemini, LangChain) if you want. pip install m3-memory[crewai]. See the CrewAI integration guide.

  • You build with PydanticAI: m3-backed memory as either drop-in tools + auto-recall (register_m3_tools, m3_recall_processor) or a formal M3MemoryToolset (a real PydanticAI AbstractToolset). Built on Pydantic v2, so it runs on Python 3.14 with a plain pip install m3-memory[pydantic-ai]. See the PydanticAI integration guide.

  • You need security and compliance: Built-in gdpr_forget and gdpr_export tools, air-gapped support, and audit logs.

  • You value privacy: Zero external cloud requests or subscriptions required.

M3 is NOT a fit if...

  • You need a hosted SaaS dashboard with managed infrastructure (use Letta).

  • You don't want persistent memory: you want each session to start fresh, with no ability to retrieve prior sessions' knowledge — M3 exists to do the opposite, so your agent's built-in defaults are the simpler fit.


šŸ›”ļø Why Trust This

  • Benchmarked Retrieval: State-of-the-art for a local-first substrate — 99.2% session-hit-rate @ k=10, 100% @ k=20 on LongMemEval-S — with a published, reproducible methodology and no oracle routing. See Benchmarks.

  • Robust Coverage: Over 2,400 tests guarding correct behavior across search, sync, GDPR lifecycle, and files ingestion — run with warnings-as-errors, so a new warning fails the suite.

  • Audit Reports: Regular vulnerability reports (Bandit, secrets scans, pip-audit) published directly under docs/audits/.

  • Explainable Retrieval: No black-box queries; retrieval math is open, readable, and scoring parameters are outputted directly.

  • Open Source: Apache 2.0 licensed, free, with no SaaS walls or usage limits.


šŸ“Š Benchmarks

Read retrieval accuracy first — it is the only number that measures the memory layer.

Session Hit-Rate (SHR) asks one question: did the system surface the evidence that answers the query? No answer model is involved, so the score reflects the memory layer and nothing else. It is the like-for-like metric across memory systems.

End-to-end QA accuracy runs that retrieved context through an LLM and has a judge model grade the answer. Both choices move the score independently of retrieval: a stronger answerer lifts a weaker memory layer, a lenient judge lifts everyone, and neither is held constant across published comparisons. Two systems quoting QA numbers are usually not measuring the same thing.

Both are reported below. SHR is the headline; QA is context.

Retrieval Accuracy — Session Hit-Rate @ k (the memory-layer metric)

Evaluated on the 500-question LongMemEval-S dataset under default server configurations:

Retrieve Depth (k)

Session Hit-Rate (SHR) ⁂

Success Count

vs. Prior Version

1

91.8%

459 / 500

First Report †

5

98.2%

491 / 500

+2.0pp

10 (Default)

99.2%

496 / 500

+2.4pp

20

100.0%

500 / 500

First Report —

† SHR@1 is the strictest cut — a gold session as the single top-ranked result. M3 operates at k=10 (its default), where a gold session is present for 99.2% of questions; k=1 is reported here for completeness, not as the headline. Cross-system SHR/recall figures are usually quoted at k=5, k=10, k=20, or k=50, so comparing another system's k=10+ number against this k=1 figure is not a like-for-like comparison.

⁂ Which aggregation. These are binary per-question recall_any@k values — the convention adjacent LongMemEval submissions report. The benchmarking report's per-question-type table aggregates slightly differently and reads marginally higher at shallow depth (98.8% at k=5, 99.4% at k=10); k=20 is 100.0% either way. The table above quotes the more conservative figures.

— v3 improvement — the v3 engine reaches 100% SHR at k=20, exceeding the prior version's 97.8% measured at the deeper k=30 (LongMemEval issue #43) — higher recall at shallower depth. Both figures are retrieval-only SHR (no answerer). The "vs. Prior Version" deltas at k=5/k=10 compare v3 against the prior version's 96.2% / 96.8% at the same k.

End-to-End QA Accuracy (answer-model and judge dependent — not a memory-layer comparison)

92.0% accuracy (460/500 correct responses) with zero oracle metadata routing. Reported for completeness; see the note above on why this number is not comparable across systems the way SHR is:

Question Domain

Count (n)

Accuracy

single-session-user

70

94.3%

single-session-assistant

56

96.4%

single-session-preference

30

80.0%

multi-session

133

87.2%

temporal-reasoning

133

95.5%

knowledge-update

78

93.6%

Overall Summary

500

92.0%

Methodology and reproducibility details are located in the LongMemEval-S Benchmarking Report.


🧰 Core Tools

While M3 features 100+ tools, these five serve as your primary interface:

Tool Name

Operation Description

memory_write

Save a specific fact, project preference, or technical configuration.

memory_search

Run hybrid keyword (BM25) and semantic vector search.

memory_update

Edit existing facts to keep memory accurate.

memory_suggest

Query memories alongside a mathematically explicit score breakdown.

memory_get

Fetch details of a single memory using its unique ID.

Refer to the Agent Instructions Guide and Full MCP Tool Catalog for complete parameter definitions.


šŸ¤– For AI Agents

You can drop the agent ruleset file examples/AGENT_RULES.md into your workspace to teach your agent best practices (e.g., query before writing, update existing records instead of duplicating).

Command Installation Prompts

Copy and paste these prompts into your terminal client to let your agent set up M3 for you:

Claude Code Prompt

Install m3-memory for persistent memory. Run: pip install m3-memory
Then run: m3 setup
That wires the m3 "memory" MCP server into my agents and provisions the
local BGE-M3 embedder — no external embedding service is needed. If it
doesn't detect Claude Code, add {"mcpServers":{"memory":{"command":"m3"}}}
to my ~/.claude/settings.json under "mcpServers". Then use /mcp to verify
the memory server loaded.

Gemini CLI Prompt

Install m3-memory for persistent memory. Run: pip install m3-memory
Then run: m3 setup
That wires the m3 "memory" MCP server into my agents and provisions the
local BGE-M3 embedder — no external embedding service is needed. If it
doesn't detect Gemini CLI, add {"mcpServers":{"memory":{"command":"m3"}}}
to my ~/.gemini/settings.json under "mcpServers".

Active Chatlog Capture Plugin

To configure instant conversation logging and backup, tell your active coding agent:

Install the m3-memory chat log subsystem.

The agent executes bin/chatlog_init.py and configures execution triggers (see Chat Log Architecture Guide).


šŸŽ¬ See it in action

Contradiction Detection & Automatic Resolution

Hybrid Search Scoring Details

Multi-Device Database Sync


šŸ’¬ Community

Discord Badge &nbsp; GitHub Issues Badge

How to Contribute Ā· FAQ for Developers Ā· Good First Issues


šŸ“œ License & Attributions

This project is licensed under the Apache License 2.0. See LICENSE for details.

Built with

M3 Memory is authored and maintained by skynetCMD. It was built with the help of AI coding assistants — Gemini CLI, Claude Code, and Google Antigravity — which contributed code under the author's direction. (They are tools that assisted; they are not maintainers, sponsors, or co-owners of the project.)

Asset & Icon Credits

The provider badges under docs/badges/ embed small logo glyphs:

  • OpenClaw & OpenCode icons are from the MIT-licensed LobeHub icon set (lobe-icons).

  • The Hermes badge uses a generic caduceus glyph.

See NOTICE for the full third-party attribution list.

⭐ Star History

Chart regenerated on a schedule by .github/workflows/star-history.yml using the repo's own token — no third-party embed. Click through for the live interactive version.

Available Tools

20 tools
agent_listC

List registered agents, optionally filtered by status and/or role.

ParametersJSON Schema
NameRequiredDescriptionDefault
roleNo
statusNo
timeoutNo
databaseNo

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries full burden. It only says 'list' without disclosing side effects, authorization needs, result format, pagination, or behavior when no agents match. Minimal behavioral information.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single concise sentence that conveys the core functionality without extra words. It is front-loaded with the action and resource.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no output schema, no annotations, and 4 parameters with 0% schema coverage, the description is insufficiently complete. Missing details on result behavior, pagination, defaults, and error conditions reduce completeness for an agent to invoke reliably.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must explain parameters. It only mentions filtering by status and role, but omits timeout and database parameters entirely. 2 out of 4 parameters are unexplained.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states it lists registered agents with optional filtering by status and role. The verb 'list' and resource 'agents' are specific. However, it does not explicitly differentiate from sibling tools like task_list, though the different resource type implies distinction.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives, nor when not to use it. The description only states optional filters without context on selecting this tool over others.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

chatlog_statusB

One-call health summary of the chat log subsystem: mode, DB paths, row counts, queue depth, spill files, embed backlog, hook timestamps, redaction state, warnings.

ParametersJSON Schema
NameRequiredDescriptionDefault
timeoutNo
databaseNo

TDQS

B3.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the burden. It states what is returned but does not clarify if the tool is read-only, requires permissions, or has side effects.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence that front-loads the core purpose, though it lists many items which reduces readability slightly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no output schema and no annotations, the description lists the summary contents but fails to describe parameter behavior or return format, which is insufficient for a health tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% and the description does not mention the timeout or database parameters at all, leaving the agent to infer their meaning from the schema alone.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states it provides a health summary of the chat log subsystem and lists specific metrics included. However, it does not explicitly differentiate from sibling health tools like files_health.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage for checking chat log health but provides no guidance on when to use this tool vs alternatives or when not to use it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

chatlog_writeB

Append one chat turn to the chat log DB. Provenance (host_agent, provider, model_id, conversation_id) is required. Writes are async-queued — returns the row id immediately.

ParametersJSON Schema
NameRequiredDescriptionDefault
roleYes
contentYes
timeoutNo
user_idNo
agent_idNo
cost_usdNo
databaseNo
metadataNo{}
model_idYes
providerYes
tokens_inNo
host_agentYes
latency_msNo
tokens_outNo
turn_indexNo
conversation_idYes

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Discloses that writes are async-queued and returns the row id immediately, but lacks further behavioral details (e.g., error handling, idempotency, rate limits). With no annotations, the description carries the burden but is only partially transparent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with the core action and resource. No wasted words; every sentence adds value.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given 16 parameters and no output schema or annotations, the description is insufficient. It does not describe key parameters like role, content, or optional fields, and provides minimal return value information beyond 'row id'.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must add meaning. It only names required provenance fields (host_agent, provider, model_id, conversation_id) but does not explain other required or optional parameters, leaving many parameters undefined.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('append') and resource ('chat log DB'), with a specific verb-object pair. It distinguishes itself from sibling tools like chatlog_search and chatlog_status by being a write operation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No explicit guidance on when to use this tool over alternatives or when not to use it. The description only states what it does, not the context or prerequisites for use.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

files_corpus_listC

Enumerate corpora with row counts.

ParametersJSON Schema
NameRequiredDescriptionDefault
timeoutNo
databaseNo

TDQS

C2.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description must convey behavioral traits. It only says 'Enumerate corpora with row counts,' lacking disclosure on whether it is read-only, whether authentication is needed, or any side effects. This is insufficient for safe agent usage.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

At 4 words, the description is too short to be useful. While concise, it sacrifices necessary detail; it does not earn its place as it fails to inform adequately.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's simplicity (listing corpora with counts), the description covers the basic output. However, it omits context about what 'corpora' are, which database is used, and how results are structured, making it only minimally complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description should compensate by explaining the parameters. It does not mention 'timeout' or 'database' at all, leaving agents without context for these fields beyond the schema defaults.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description states 'Enumerate corpora with row counts,' which gives a basic verb and resource but fails to specify what 'corpora' refers to in context. It does not clearly distinguish from sibling tools like files_stats or files_search, leaving ambiguity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives. There is no mention of prerequisites, typical use cases, or when not to use it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

files_getA

Fetch one record by UUID. Tries file_nodes then leaves.

ParametersJSON Schema
NameRequiredDescriptionDefault
uuidYes
timeoutNo
databaseNo

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description discloses the lookup strategy (file_nodes then leaves), which is a key behavioral trait, but omits details on error handling, return format, or idempotency, especially important since no annotations exist.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely concise at one sentence, front-loading the core action and key behavioral detail without superfluous words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no annotations, no output schema, and three parameters, the description is too minimal—it lacks info on return values, error states, parameter constraints, and practical usage context, leaving the agent underinformed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 0% schema description coverage, the description only semantically enriches the required 'uuid' parameter by indicating it identifies the record; the 'timeout' and 'database' parameters are entirely unspecified.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool fetches one record by UUID and specifies the lookup order across file_nodes and leaves, distinguishing it from sibling tools like files_search which search or files_index which index.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage for fetching a single record by UUID but does not explicitly guide when to use this tool over alternatives like files_search or when the fallback behavior matters.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

files_healthC

DB integrity + FTS5 sync check. Set rebuild=True to fix drift.

ParametersJSON Schema
NameRequiredDescriptionDefault
rebuildNo
timeoutNo
databaseNo

TDQS

C2.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description must disclose behaviors. It mentions a check and a rebuild action, but does not describe potential side effects (e.g., time consumption, data modification during rebuild), failure modes, or the effect of timeout and database parameters.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely concise with one sentence, which is efficient, but it omits essential information about parameters and usage, making it under-specified rather than effectively concise.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the lack of output schema and annotations, the description is insufficient. It does not explain return values, how rebuild works, the effect of database selection, or timeout behavior. It fails to provide a complete understanding of the tool's capabilities and constraints.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must explain all parameters. Only 'rebuild' is mentioned. 'timeout' and 'database' have no description, leaving their purpose and acceptable values unclear.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states it performs a DB integrity and FTS5 sync check, and mentions the rebuild capability. It distinguishes itself from sibling tools like files_search and files_index by focusing on health rather than data retrieval or indexing, though it does not explicitly contrast them.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool over alternatives, or any prerequisites or conditions. The description implies using rebuild to fix drift but does not explain when that is appropriate or what the default behavior is.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

files_indexA

Return file-level summaries for triage (wiki-index primitive). Cheap-first retrieval -- no leaf content. Use BEFORE files_search to decide which files are worth deep-reading.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
corpusNo
corporaNo
timeoutNo
databaseNo
filetypeNo
directoryNo
filename_globNo
include_historyNo

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description notes 'no leaf content' and 'cheap-first retrieval', indicating the tool returns summaries only and is low cost. Since no annotations are provided, the description carries the full burden; it does not contradict any annotations. However, it omits details about rate limits, authentication, or side effects, though these are not critical for a read-only index tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is three short sentences that are front-loaded with the core purpose and usage recommendation. Every sentence adds value, with no redundancy or unnecessary information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has 9 parameters and no output schema, the description is minimalist. It captures the high-level intent but lacks details on output format, parameter usage, and behavioral boundaries. It is adequate for a simple index tool but incomplete for fully autonomous use.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 9 parameters and zero schema description coverage, the description provides no explanation of parameters like limit, corpus, corpora, database, filetype, etc. This is a significant gap, as the agent would need to infer parameter meanings from names alone.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool returns file-level summaries for triage as a cheap-first retrieval primitive, distinguishing it from file_search and file_get. It explicitly says 'Use BEFORE files_search to decide which files are worth deep-reading', differentiating among sibling tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It provides explicit guidance on when to use: 'Use BEFORE files_search to decide which files are worth deep-reading'. It also characterizes the tool as 'cheap-first retrieval', implying it should be used before more expensive operations.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

files_statsC

Corpus-level counters: file_nodes, leaves, embed coverage, by-filetype.

ParametersJSON Schema
NameRequiredDescriptionDefault
corpusNo
timeoutNo
databaseNo

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries full burden. It does not disclose whether the tool is read-only, destructive, requires authentication, or has any limitations. Only the output type is hinted but not detailed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is very concise (one phrase) but lacks structure and completeness. While it has no wasted words, it is too minimal to be fully useful.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no output schema and no annotations, the description should provide more context about what the counters represent, how to use parameters, and expected behavior. It only gives a high-level list of outputs, leaving significant gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, and the tool description does not explain any of the parameters (corpus, timeout, database). The description adds no meaning beyond the bare schema names.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool provides corpus-level counters such as file_nodes, leaves, embed coverage, and by-filetype. This specific verb and resource combination distinguishes it from sibling tools like files_health (health check) or files_search (search).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool versus alternatives. The description does not mention prerequisites, context, or when not to use it. Usage is only implied by its purpose.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

m3_callA

Invoke ANY m3 catalog tool by name without loading its domain — the low-token path to the full tool surface. Single call: pass tool (e.g. 'files_stats') and args (an object). Batch: pass batch, a list of {tool, args} (each isolated — one failure won't abort the rest; capped at 100). Set dry_run to validate args + check the destructive gate WITHOUT executing. Returns JSON. Call m3_index first if you don't know a tool's args. Destructive tools require MCP_PROXY_ALLOW_DESTRUCTIVE=1.

ParametersJSON Schema
NameRequiredDescriptionDefault
argsNo
toolNo
batchNo
dry_runNo
timeoutNo
databaseNo

TDQS

A4.3/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, description fully discloses batch isolation, cap at 100, dry_run validation, destructive gate requirement, and JSON return. This is comprehensive behavioral context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Description is well-structured and front-loaded with purpose. Every sentence adds value, though could be slightly tighter for the batch and dry_run explanations.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given tool complexity and no output schema, description covers invocation modes, error isolation, and destructive gate. Could expand on error handling or response format, but sufficient for typical use.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 0% schema coverage, description explains tool, args, batch, and dry_run but omits timeout and database parameters. Provides meaningful context for key parameters but not complete.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool invokes any m3 catalog tool by name without loading its domain. It distinguishes itself from siblings like m3_index by advising to call it first if args are unknown.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides guidance on when to use m3_index for unknown args, batch vs single call, and dry_run behavior. Missing explicit when-not-to-use scenarios but sufficiently covers common use cases.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

m3_help_capabilitiesA

Discover m3-memory tool capabilities, parameters, and availability. Allows filtering by a logical domain (memory, chatlog, files, entity, agent, tasks, conversations, admin, diagnostics) or searching by keywords.

ParametersJSON Schema
NameRequiredDescriptionDefault
queryNo
domainNo
timeoutNo
databaseNo

TDQS

A3.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It states the tool is for 'discovering' capabilities, implying a safe read-only operation. It does not disclose authentication needs, rate limits, or return format, but the non-destructive nature is implied.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise and front-loaded, effectively communicating the purpose in one sentence. However, it could be more structured by explicitly listing all parameters, which would improve usability.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has 0% schema coverage, no output schema, and no annotations, the description is only partially complete. It covers the core purpose and two parameters but omits timeout and database explanations. A more comprehensive description would benefit agent selection.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate. It explains 'domain' (logical domain filtering) and 'query' (search by keywords), but does not mention 'timeout' (default 30) or 'database' (default empty). These parameters are left unexplained, leaving gaps for an AI agent.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb ('discover') and resource ('m3-memory tool capabilities, parameters, and availability'). It distinguishes from sibling tools by being a meta-help tool for the entire system, while siblings like tools_list_domains or memory_write are domain-specific or operational.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear context on when to use the tool (to discover capabilities) and the available filters (domain list and keywords). However, it does not mention when not to use it or suggest alternative tools for specific needs, but the context is sufficient for basic guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

m3_indexA

List m3 catalog tools (optionally one domain) as structured rows: name, domain, one-line summary, destructive flag, and arg specs (name/type/required). Use this to discover the exact args for any tool before calling it via m3_call — cheaper than a failed call. Read-only catalog metadata; never returns tool output. Domains: memory, chatlog, files, entity, agent, tasks, conversations, diagnostics, admin.

ParametersJSON Schema
NameRequiredDescriptionDefault
domainNo
timeoutNo
databaseNo

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden. It states 'Read-only catalog metadata; never returns tool output', clearly indicating safety and no side effects. This is sufficient for a catalog tool, though it lacks details on rate limits or auth, which are less critical here.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise at 4 sentences, each serving a clear purpose: stating output, providing usage advice, declaring read-only behavior, and listing domains. It is front-loaded with the core function and highly efficient.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no output schema and 0% schema description coverage, the description is mostly adequate but misses details on the timeout and database parameters. It covers purpose, usage, behavior, and domains well, but the parameter gap reduces completeness.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema description coverage is 0%, so the description must explain all parameters. It partially explains the domain parameter with 'optionally one domain', but completely omits the timeout and database parameters, leaving them unexplained. This is a significant gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: 'List m3 catalog tools (optionally one domain) as structured rows'. It specifies the output fields and distinguishes itself from the sibling m3_call by advising to use this tool before calling m3_call.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides explicit usage guidance: 'Use this to discover the exact args for any tool before calling it via m3_call — cheaper than a failed call'. It also lists the available domains for filtering. However, it does not specify when not to use it or mention alternatives beyond m3_call.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

memory_getA

Retrieves a full MemoryItem; accepts full UUID or 8-char prefix; ambiguous prefixes return an error.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYes
timeoutNo
databaseNo

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Discloses error behavior for ambiguous prefixes, which adds value beyond a simple 'retrieve'. However, no annotations exist, so the description carries full burden, but it omits details like idempotency or side effects, though implied by 'get'.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Single sentence with no filler, all essential information is front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With 3 parameters and no output schema, the description is incomplete: it fails to explain 'timeout' and 'database' parameters, and does not describe the return value structure beyond 'full MemoryItem'.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so description must compensate. Only the 'id' parameter is partially described (accepts UUID or prefix), while 'timeout' and 'database' are entirely unmentioned.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Clearly states it retrieves a full MemoryItem by full UUID or 8-char prefix, distinguishing it from sibling tools like memory_search and memory_write.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explains the acceptable ID formats and error on ambiguous prefixes, but does not explicitly mention when not to use or provide alternatives like memory_search.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

memory_supersedeA

Explicitly supersede an existing memory with a new one. Use this to record an intentional update — 'this fact replaces that specific memory' — when you know the old memory's id. Unlike memory_write's automatic contradiction detection (a cosine + title heuristic that may link the wrong prior memory or none at all), this targets the given old_id deterministically. Non-destructive: the old memory is retained, its validity interval is closed (is_deleted=1, valid_to set), and a 'supersedes' edge is recorded new -> old. The old memory stays retrievable by id and via memory_history, and as_of-filtered search still sees it valid before the supersession point — it is only dropped from default search. Fields you omit (type, title, importance, scope) are inherited from the old memory, so pass only what changed. To hard-delete instead, that is a separate gated tool (memory_delete). old_id MUST be the full UUID — a prefix is rejected (full UUID required for mutation safety; memory_get accepts a prefix, this does not). Note: each supersede creates a NEW successor memory; call it once with the full id, do not chain supersedes.

ParametersJSON Schema
NameRequiredDescriptionDefault
typeNo
embedNo
scopeNo
titleNo
old_idYes
sourceNoagent
contentYes
timeoutNo
user_idNo
variantNo
agent_idNo
databaseNo
metadataNo{}
model_idNo
embed_textNo
importanceNo
valid_fromNo
change_agentNo

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description bears full burden. It discloses key behaviors: non-destructive operation, retention of old memory, closure of validity interval, recording of 'supersedes' edge, inheritance of omitted fields, requirement for full UUID, and prohibition of chaining. However, it does not explain the many other parameters (e.g., embed, timeout, user_id, etc.) and their effects, leaving some gaps.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is detailed but somewhat lengthy, covering multiple aspects. It is front-loaded with purpose and usage, then behavioral details and parameter notes. While informative, it could be more concise by reducing redundancy (e.g., repeating 'non-destructive' and 'full UUID' points). Still, it is well-structured and readable.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (18 parameters, no output schema, no annotations), the description explains the core functionality well but leaves many parameters (embed, timeout, source, etc.) unexplained. An agent may struggle to use these parameters correctly without additional information. The description is incomplete for fully autonomous use.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It explains only a few parameters: old_id (full UUID required), content (required), and mentions type, title, importance, scope as inheritable. The other 13 parameters (embed, source, timeout, user_id, etc.) are not described, leaving the agent without guidance for their purpose or usage. This is insufficient for a schema with 18 parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: to explicitly supersede an existing memory with a new one. It contrasts with memory_write's automatic detection, provides specifics on the operation (non-destructive, validity interval closed, supersedes edge), and distinguishes from a hard-delete tool. This is specific and helps differentiate from siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly explains when to use this tool: when the old memory's ID is known and an intentional update is needed. It contrasts with memory_write's automatic detection, advising against using this tool when that automatic linking is sufficient. It also warns not to chain supersedes. This provides clear guidance on usage vs. alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

memory_writeA

Creates a MemoryItem and optionally embeds it for semantic search. Contradiction detection is automatic — if new content conflicts with an existing memory of the same type/title, the old one is superseded. Use type='auto' to let the LLM decide the best category.

ParametersJSON Schema
NameRequiredDescriptionDefault
typeYes
embedNo
scopeNoagent
titleNo
sourceNoagent
contentYes
timeoutNo
user_idNo
variantNo
agent_idNo
databaseNo
metadataNo{}
model_idNo
valid_toNo
embed_textNo
importanceNo
refresh_onNo
valid_fromNo
change_agentNo
auto_classifyNo
refresh_reasonNo
conversation_idNo

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations exist, so description carries full burden. It discloses automatic contradiction detection and superseding behavior, but does not detail side effects or return values for the 22 parameters.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences efficiently convey the primary purpose and a key usage hint. No unnecessary words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (22 parameters, no output schema, no annotations), the description is too brief. It fails to cover optional parameters and their effects, making it incomplete for reliable agent invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%. The description only explains 'type' and 'content' implicitly; the other 20 parameters (e.g., scope, metadata, importance) are undocumented, leaving the agent uninformed.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states 'Creates a MemoryItem' with specific verb and resource. It distinguishes from siblings like memory_search and memory_get by indicating creation and embedding for semantic search.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description suggests using type='auto' but does not explicitly guide when to use this tool versus alternatives like memory_supersede. Usage context is implied rather than stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

task_listB

List tasks with optional filters. Newest updated first.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
stateNo
timeoutNo
databaseNo
owner_agentNo
parent_task_idNo

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations provided, so description is the sole source. It indicates a read operation and ordering, but does not mention pagination, behavior on empty results, or any rate limits. Adequate but not comprehensive.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Very concise at 7 words. However, it could include brief parameter hints without losing conciseness. Currently it is minimal but not wasteful.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema, no annotations, and 6 parameters with no descriptions. The description is too short to cover essential details like return format, pagination, or filter semantics.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% (no parameter descriptions). The description only says 'optional filters' without explaining each parameter. While names like 'limit' and 'state' are self-explanatory, 'timeout' and 'database' need clarification, which is missing.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description clearly states the verb 'List', the resource 'tasks', and the ordering 'Newest updated first'. It distinguishes from siblings as no other task-listing tool exists among the siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No explicit guidance on when to use or avoid this tool. However, it is the only task-listing tool, so usage is implied. Lacks alternative references.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

tools_list_domainsA

List m3 tool domains (memory, chatlog, files, entity, agent, tasks, conversations, diagnostics, admin) and their tool counts. Call tools_load_domain to expose a domain's full tool surface.

ParametersJSON Schema
NameRequiredDescriptionDefault
timeoutNo
databaseNo

TDQS

A3.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations provided, so the description carries the full burden. It implies a read-only list operation but does not explicitly state safety, side effects, or permission requirements. Adequate for a simple read tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two concise sentences, front-loaded with purpose. No wasted words; each sentence adds value.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a straightforward list tool with no output schema and simple optional parameters, the description provides the needed domain names and links to the follow-up tool. It does not detail output format or parameter usage, but given the simplicity, it is fairly complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the description does not mention or explain any of the two parameters (timeout, database). It adds no meaning beyond the schema, which is insufficient.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description clearly states the tool lists m3 tool domains and their tool counts, providing specific domain names. It also distinguishes the sibling 'tools_load_domain' by explaining what that tool does.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear context for when to use this tool (to list domains with counts) and points to an alternative (tools_load_domain) for exposing full tool surface. However, it does not explicitly state when not to use it or list other alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

tools_load_domainC

Register a tool domain's full surface for the current MCP session. Use when you need tools beyond the essentials (memory_search, memory_write, memory_get, chatlog_search, chatlog_write, files_search). Valid domains: memory, chatlog, files, entity, agent, tasks, conversations, diagnostics, admin.

ParametersJSON Schema
NameRequiredDescriptionDefault
domainYes
timeoutNo
databaseNo

TDQS

C2.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations provided, so description must fully disclose behavioral traits. It does not mention side effects, permissions, session-level impact, or whether loading is additive vs. replacement. Only states 'for the current MCP session' but no further details.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences are concise and front-loaded with purpose, but missing parameter info and behavioral details means it is incomplete rather than efficiently written. Could be restructured to include parameter hints.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given 3 parameters, no output schema, and no annotations, the description should fully explain usage. It does not mention parameters, return values, or implications of loading a domain. Incomplete for effective tool invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% and description adds no information about any of the three parameters (domain, timeout, database). The required 'domain' parameter is not mentioned, making it impossible for an agent to understand how to invoke correctly from description alone.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool registers a tool domain's full surface, and the use case distinguishes it from siblings like tools_list_domains. The verb 'register' and noun 'tool domain' are specific, though 'full surface' is slightly jargon.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly says 'Use when you need tools beyond the essentials' and lists those essentials, providing clear context. Lacks explicit when-not-to-use or alternatives but adequate.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

TDQS

B3/5.0
Disambiguation3/5

Tools are generally distinct, but the meta-tools (m3_help_capabilities, tools_list_domains, m3_index) overlap in purpose, and m3_call could confuse agents about when to use it versus directly calling a tool. Descriptions help, but some ambiguity remains.

Naming Consistency4/5

Naming follows a consistent snake_case pattern with clear prefixes (m3_, memory_, chatlog_, files_, agent_, task_, tools_), but deviations like files_stats (noun instead of verb_noun) and files_corpus_list (compound noun) break the pattern slightly.

Tool Count3/5

With 20 tools, the count is borderline high for a typical MCP server (3-15 is standard). The set includes 5 meta-tools for discovery and invocation, which adds overhead but is justified by the dynamic domain system.

Completeness3/5

Memory has create, read, update, but delete is missing from the default set. Files lack create/delete, and agent/task tools are read-only. These gaps may force workarounds, but core memory and chatlog operations are covered.

Maintenance

ActivityActive
ResponsivenessWithin a week

Resources

Unclaimed servers have limited discoverability.

Looking for Admin?

If you are the server author, to access and configure the admin panel.

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    D
    maintenance
    Local-first persistent memory for AI agents via MCP, enabling semantic search and memory sharing across agents with zero cloud cost and full privacy.
    16
    1
    MIT
  • A
    license
    Not graded
    quality
    A
    maintenance
    Provides persistent memory for AI coding agents via MCP, enabling agents to store and semantically recall facts, events, and lessons across sessions, all running locally without cloud dependencies.
    Apache 2.0
  • A
    license
    Not graded
    quality
    D
    maintenance
    Local-first AI memory layer with hybrid retrieval and brain-inspired namespaces. Enables agents to save, search, and manage memories directly via MCP tools.
    5
    MIT
  • A
    license
    A
    quality
    A
    maintenance
    Embedded memory and retrieval engine for AI agents, providing local-first memory with MCP support for multi-agent access control.
    3
    2
    MIT

Latest Blog Posts

MCP directory API

We provide all the information about MCP servers via our MCP API.

curl -X GET 'https://glama.ai/api/mcp/v1/servers/skynetcmd/m3-memory'

If you have feedback or need assistance with the MCP directory API, please join our Discord server