Skip to main content
Glama

@mnemosyne_os/mcp: MCP server for Mnemosyne OS. Gives AI agents access to vault memory, resonances, git context, and to what the other coding agents on this machine are doing.

Product mnemosyne-os.io · Company, press and labs mnemosyne-os.com · Documentation docs.mnemosyne-os.io

@mnemosyne_os/mcp

Give Claude, Cursor, Hermes Agent, Copilot, and any MCP-compatible agent access to your local Mnemosyne OS memory vault. Code, decisions, architecture notes, git history, semantically queryable. The vaults stay on your machine, and this server opens exactly one socket: 127.0.0.1:7799.

npm version License: MIT Node.js ≥18 Mnemosyne OS MCP server – quality and maintenance score on Glama


🍳 In a hurry? RECIPES.md gives your coding agent a persistent memory in one copy-paste block, for Claude Code, Cursor, Claude Desktop and the TypeScript SDK.

What this is

@mnemosyne_os/mcp is a Model Context Protocol server that turns your local Mnemosyne OS install into a queryable memory layer for any AI agent that speaks MCP.

Once configured, your agent can:

  • Query code, architecture, decisions and git history with real semantic ranking (Vertex, e5-base or nomic).

  • Persist new decisions, sessions or insights, so future agents can recover them.

  • Resume a project exactly where you left off, through Resonance positions.

  • Filter results by spineType: ARCHITECTURE, GIT, SOURCE_CODE, BUGFIX and the rest.

The MCP itself opens exactly one socket: 127.0.0.1:7799. It sends nothing anywhere else and keeps no state. Your claude.ai conversation sees only the chronicles you allow. What Mnemosyne OS does behind that socket follows the route you configured. mnemosyne_memory_ask runs whichever model you picked, local or cloud.


Related MCP server: auxly-memory-cli

Requirements

Node.js ≥ 18 is the only hard requirement.

The memory tools additionally need Mnemosyne OS Infinity Edition running. It owns your vaults and exposes the WebSocket gateway on ws://127.0.0.1:7799. It is a desktop application, and it is where your content lives.

The three agent-awareness tools need neither: mnemosyne_agent_list, mnemosyne_agent_collisions and mnemosyne_agent_files. They read transcript files your coding-agent harness already writes to disk, so they answer with the app closed, with no vault, and without spending a token. They read every harness they find, so a Claude Code session can see an Antigravity session running in the same repository.

The MCP is a thin bridge. It does not store anything itself. All data lives in Mnemosyne OS Infinity (%APPDATA%\@mnemosyne-workspace\infinity-edition\vaults\*.db on Windows, ~/Library/Application Support/... on macOS).

Getting your code, commits and docs in

Three different routes, and only one of them ingests anything. Knowing which is which saves you looking for a feature that is not where you expect it.

Your files: source, architecture notes, decision records. You declare a folder, the app watches it, and what lands there is ingested into the vault you chose. Nothing is uploaded and nothing is scanned that you did not name. This is the step people miss: installing the app gives you empty vaults, and declaring the folder is what fills them. See Getting started and DocWatch. The full walkthrough, including why there is an application behind this server at all, is Get your repository into memory.

Your commits. mnemosyne_git_log reads the repository the app is configured to read, at the moment you call it. Nothing is ingested and nothing is stored, so history is never stale and never doubled.

What your agents did. mnemosyne_agent_list, mnemosyne_agent_collisions and mnemosyne_agent_files read the transcript files your harness already writes. No app, no vault, no token. Those three work the minute this server is installed, which is why they are the ones to try first.

⚠️ Without the app running, everything else refuses and says so. The refusal names what it checked: whether ~/.mnemosyne exists tells it whether the app has ever run on this machine, and it says "install it" and "start it" as two different sentences, because they are two different problems.


Install, in 30 seconds

Claude Desktop, one click

Download Mnemosyne-OS-MCP-2.1.0.mcpb (4.1 MB), then open Claude Desktop → Settings → Extensions and drop the file into that panel. That is the whole install. The 25 tools appear straight away, and the same panel offers the three optional settings: default vault, other vaults, and the port the desktop application listens on.

The bundle rides on the application's current release, v1.5.1-infinity, because this package has no release of its own. The file is named after this package version, not after the application one in the tag. Earlier bundles stay attached to v1.4.5-infinity, so a link someone saved keeps working.

Two things worth knowing before you do it:

  • Claude Desktop shows a red "unverified developer" notice first. Every unsigned bundle does. This one is built from the repository linked at the top of this file, by packages/mcp/scripts/build-mcpb.mjs.

  • Double-clicking the file does nothing if your Claude Desktop came from the Microsoft Store: a Store app does not register the .mcpb extension with Windows. Drop it into the Extensions panel instead.

Claude Desktop, config file

If you would rather not install an extension, or you are on a build that has no Extensions panel:

Open Claude Desktop → Settings → Developer → Edit config (or edit claude_desktop_config.json directly):

  • Windows: %APPDATA%\Claude\claude_desktop_config.json

  • macOS: ~/Library/Application Support/Claude/claude_desktop_config.json

Add:

{
  "mcpServers": {
    "mnemosyne": {
      "command": "npx",
      "args": ["-y", "@mnemosyne_os/mcp"],
      "env": {
        "MNEMO_DEFAULT_VAULT": "DEV",
        "MNEMO_VAULTS": "DEV,PERSONAL,SOCIAL"
      }
    }
  }
}

Fully quit and relaunch Claude Desktop (close from the tray icon, not just the window). The mnemosyne server should show up under Settings → Developer → Local MCP Servers with the running badge.

Claude Code

Add to .mcp.json at the root of your project:

{
  "mcpServers": {
    "mnemosyne": {
      "command": "npx",
      "args": ["-y", "@mnemosyne_os/mcp"],
      "env": {
        "MNEMO_DEFAULT_VAULT": "DEV",
        "MNEMO_VAULTS": "DEV,PERSONAL,SOCIAL"
      }
    }
  }
}

Reload the Claude Code session.

Cursor

Add to .cursor/mcp.json:

{
  "mcpServers": {
    "mnemosyne": {
      "command": "npx",
      "args": ["-y", "@mnemosyne_os/mcp"],
      "env": {
        "MNEMO_DEFAULT_VAULT": "DEV",
        "MNEMO_VAULTS": "DEV,PERSONAL,SOCIAL"
      }
    }
  }
}

Hermes Agent

Hermes Agent from Nous Research ships with MCP support, so there is no extra install step. Add to ~/.hermes/config.yaml:

mcp_servers:
  mnemosyne:
    command: "npx"
    args: ["-y", "@mnemosyne_os/mcp"]
    env:
      MNEMO_DEFAULT_VAULT: "DEV"
      MNEMO_VAULTS: "DEV,NOTES"

Restart Hermes. Your agent now has a sovereign long-term memory it can query semantically, and everything stays on your machine, which is exactly the deal Hermes promises you.

Recommended for autonomous agents: grant read scopes narrowly, to the vaults the task needs and no others. Point MNEMO_DEFAULT_VAULT at a vault dedicated to agent work rather than at your personal notes.

The mnemosyne-memory skill is a portable agentskills.io-standard skill that teaches any agent the governance rules: discover vaults first, respect protection levels, ingest with provenance, read before you write.

Any other MCP client

npx -y @mnemosyne_os/mcp

The MCP speaks standard JSON-RPC over stdio.


Configure your vaults

Mnemosyne OS Infinity exposes one vault per tracked folder (the folder name uppercased), plus three reserved names: DEV, PERSONAL, SOCIAL. Tell the MCP which ones you want your agent to reach via two env vars:

Env var

Purpose

Default

MNEMO_DEFAULT_VAULT

Vault used when the agent does not specify one.

DEV

MNEMO_VAULTS

Comma-separated list of vaults the MCP declares scopes for.

DEV,PERSONAL,SOCIAL

Examples

A developer whose Infinity tracks ~/Documents/INFINITY/code-projects/:

"env": {
  "MNEMO_DEFAULT_VAULT": "CODE_PROJECTS",
  "MNEMO_VAULTS": "CODE_PROJECTS,NOTES,RESEARCH"
}

A researcher who keeps everything in ~/Documents/INFINITY/papers/:

"env": {
  "MNEMO_DEFAULT_VAULT": "PAPERS",
  "MNEMO_VAULTS": "PAPERS,REFS,IDEAS"
}

If your agent queries a vault that is not in MNEMO_VAULTS, the server returns SCOPE_DENIED. Add the vault name to the list and restart your MCP client.


Optional: let your agent render a voice

Mnemosyne OS ships local, offline text-to-speech engines that can clone a voice from a short reference clip. With one env var, your agent gets three extra tools that turn a written script into a WAV file on disk. They are made for voice-overs: TikTok, YouTube, podcast, narration.

"env": {
  "MNEMO_VOICE": "1"
}

Tool

What it does

mnemosyne_voice_list

List the local engines, installed or not, and the reference voices available for cloning. Call it first.

mnemosyne_voice_speak

Render a script to a WAV. Long scripts are split at sentence boundaries and re-assembled into one file, with nothing truncated. It returns a job: the tool waits, then hands back a job id if the render is still going.

mnemosyne_voice_status

Poll or cancel a render. It returns the file path once the render is done.

Off by default, on purpose. Turning it on makes Mnemosyne OS ask you to authorize voice:speak. That permission is never auto-granted, not even to first-party apps like this one, because its subject is your identity rather than your data. You approve it once, in a dialog that says what it means.

Two limits. The agent never creates or records a voice; you do that in the app, under Settings → Voice. A clone name that does not exist is refused rather than quietly replaced, because a voice-over in the wrong voice sounds perfect and is worthless.

Requirements: the Mnemosyne OS app must be running, since the engines are Python sidecars inside it and the headless daemon cannot speak. A local voice must be installed, and local neural TTS is a licensed feature.


Verify it works

Open a new conversation with your agent and ask, for example:

Use mnemosyne_memory_query to search my vault for "authentication flow", spine_type_filter ARCHITECTURE only.

You should see a structured response with 5–10 chronicles, each tagged with its spineType, score, source, and a content snippet. If the agent says it cannot connect, see Troubleshooting.


The 25 tools your agent gets

Twenty-two are in the table below. The three that read the other agents on this machine have their own section further down. Setting MNEMO_VOICE=1 adds the three voice tools documented further up, and MNEMO_FORGET=1 adds the erasure tool, for 29 in all.

Six of them read or write the To-do backlog and the calendar. All six carry the same requirement:

Works with the app closed on a dev install (the headless daemon reads the file); an npm install has no daemon and needs the app running.

Tool

What it does

mnemosyne_about

Re-read the briefing the agent was handed on connect: the governance tenet, the vault protection model, the spine model, and the rules for an agent working on someone else's memory. Call it if your client did not surface the server instructions.

mnemosyne_memory_query

Semantic search. It returns raw chronicles, ranked by vector similarity fused with a local BM25 channel and weighted by spineType. Supports spine_type_filter, max_content_chars and limit (50 max).

mnemosyne_memory_ask

Ask a question, get a synthesized prose answer grounded in the vault, plus its source chronicles. Use it for why, who and how questions that span many memories. It runs the full RAG pipeline, which also makes it the better retriever, and it is slower than query.

mnemosyne_vault_list

List the vaults Mnemosyne OS exposes, with each one's token, name and chronicle count. Call it to discover valid vault targets.

mnemosyne_memory_forget

Erase one chronicle for good, by the id mnemosyne_memory_query returned. Absent unless you set MNEMO_FORGET=1, and the app wants the FORGET intent on top of that. You arm erasure by hand, because writing a new chronicle cannot undo it.

mnemosyne_memory_ingest

Persist a memory. Pick a spine_type: ARCHITECTURE, DECISION, BUGFIX, FEATURE, NOTE, SESSION, RESONANCE or CUSTOM.

mnemosyne_todo_add

Put tasks into the user's To-do backlog, in order and optionally under named steps. Name the list, or pass create_list: true. An unknown name is refused with the lists that exist. With a window open the write goes through the widget's own store; on macOS with the window closed the app writes the file directly. Scope todo:write.

mnemosyne_cockpit_update

Your own status card on the user's canvas: state working, waiting, done, blocked or closed, a title, one status line, up to 4 detail lines. You declare the state, and the app prints it next to the time since your last call, so call at real milestones. "waiting" and "blocked" make the card pulse and the taskbar flash. The answer carries the messages the user left on your card. Send desktop once, with the name of a board the user gives you, and this session’s cards go there for the rest of the conversation; leave it out and they go to whichever board answers for this project. Needs the app window open. Scope cockpit:write.

mnemosyne_pheme_watch

Add a subreddit, a Hacker News query or a topic to the user's Pheme radar, or take one off. The lists are theirs. An operation that would empty one is refused, each one reports its own outcome, and the user gets a receipt in Pheme naming the agent and what changed. It posts nowhere. Needs the app running, though Pheme itself can be closed. Scope pheme:profile.

mnemosyne_pheme_radar

Read what the Pheme radar last found: fresh threads in the watched subreddits and Hacker News queries, with a topic score and, after the user's Mnemosyne OS pass, a tier. The answer leads with when the scan ran, because it is only as fresh as the last scan the user ran. "No radar yet" is said in words rather than returned as an empty list. Draft the reply; the user posts it. Needs the app running. Scope pheme:read.

mnemosyne_agenda_add

Put appointments or deadlines into the user's calendar. start and the optional end are ISO 8601, and a time with no offset is read as the user's own machine local time. Supports all_day, recurrence (daily, weekly, monthly, yearly) and alarm_minutes_before. To change or remove one, see mnemosyne_agenda_update and mnemosyne_agenda_remove. Scope agenda:write.

mnemosyne_todo_list

Read the backlog back: the lists, and every task with the id you need to change it. Call it before mnemosyne_todo_update, which names tasks by id. A phrase like "delete the task about the invoice" reads perfectly and can still match the wrong task. Filters by list, include_done and limit. The archive is counted rather than listed. Scope todo:read.

mnemosyne_todo_update

Change tasks: edit, complete, move, remove. Every operation names an id from mnemosyne_todo_list. remove archives by default and stays recoverable; permanent: true deletes outright and has to be asked for. The batch applies in order as one save, and each operation reports its own outcome, so a stale id does not sink the ones around it. Scope todo:write.

mnemosyne_todo_categories

Manage the lists themselves: list.create, list.edit to rename or recolour, list.remove. A list that still holds tasks is never removed, and you are told how many are in the way. Move them first: the app will not pick a destination on someone else's behalf. The three original lists can be renamed, never removed. Scope todo:write.

mnemosyne_agenda_list

Read the calendar back: appointments in a time window, each with the id the two tools below require. A repeating event appears once, with its cadence and its next occurrence. An event with nothing left to happen says so rather than showing its original start as a future date. Scope agenda:read.

mnemosyne_agenda_update

Change an appointment: move it, rename it, add or drop a reminder, start or stop it repeating. A field left out is left alone, and null clears it. An unreadable date refuses the change rather than leaving the old one silently in place, and an unknown cadence is refused rather than quietly made a one-off. Scope agenda:write.

mnemosyne_agenda_remove

Remove appointments, by id only. The calendar has no archive: unlike a To-do task, a removed appointment is gone, so the answer names each one by title and start time for the user to check. A repeating appointment goes as a whole series, since the calendar cannot cancel one occurrence. Scope agenda:write.

mnemosyne_resonance_list

List the resonances recorded in the default vault. A resonance is a workspace that tracks one ongoing project.

mnemosyne_position_get

Read where one resonance was left: its phase, and the note written when the work stopped.

mnemosyne_position_update

Record where you left off. It is stored as a DECISION chronicle.

mnemosyne_git_log

Recent commits from the git repository Mnemosyne OS is configured to read (requires monorepo:read scope).

mnemosyne_spine_assignments

How the memories were actually classified: chronicle-to-spine assignments for a vault, newest first, with whole-vault counts per spine and, on request, the taxonomy tree. It is also where the taxon ids come from, so read it instead of guessing a spine_type_filter.

mnemosyne_dream_bridges

The links the Dream State engine found on its own while the machine sat idle, each with its score and an excerpt of both sides, sometimes across two vaults. An empty list is the normal answer and means it has produced none yet, never that the query failed.

Seeing the other agents on this machine

These three read the transcripts coding-agent harnesses already write to disk. No app, no vault, no tokens. That is what makes "check before you commit" cheap enough to actually do.

Tool

What it does

mnemosyne_agent_collisions

Are two agent sessions live on the same project and branch right now? Call it before git add -A, before a commit and before a rebase. One working tree shares one git index, so a commit from one session picks up whatever the other has staged.

mnemosyne_agent_list

The sessions on this machine: conversation name, project, branch, model, last tool, file count, and when a line was last written.

mnemosyne_agent_files

Which files other sessions recently wrote, newest first, with the session each came from. Paths and timestamps only.

No configuration needed. Every shipped connector whose folder exists on this machine is read, and each answer names the folders it actually opened. Override only if your agent writes somewhere unusual:

"env": {
  // Per harness. Absent means "where that agent writes by default".
  "MNEMO_AGENT_SESSIONS": "C:/Users/you/.claude/projects",
  "MNEMO_AGENT_SESSIONS_ANTIGRAVITY": "…",
  "MNEMO_AGENT_SESSIONS_ANTIGRAVITY_IDE": "…",
  // Restrict to a subset. Absent means all of them.
  "MNEMO_AGENT_SOURCES": "claude-code,antigravity"
}

One line you do want, though: ← you. A stdio MCP server is launched with the env its config declares, so the caller's own session id does not arrive on its own. Without it the report cannot mark which line is yours, and it counts one session too many, which is the exact miscount these tools exist to prevent. Pass it through:

"env": {
  "CLAUDE_CODE_SESSION_ID": "${CLAUDE_CODE_SESSION_ID}"
}

If you skip it, or if the variable is not set where Claude Code runs (it then arrives as the literal ${CLAUDE_CODE_SESSION_ID}), the answer says so in one sentence and tells you which of the two happened. It never guesses: an id it cannot place stays neutral rather than becoming "none of these is you", because that would add a phantom session to every warning.

Four limits these tools hold to. A tool that overstates its evidence is worse than no tool.

  • They never say an agent is "working". A crashed agent and an idle one fall equally silent. They report when a line was last seen, and you conclude.

  • They never return content. No message text, no file contents, no tool output. A transcript holds everything that passed in front of an agent for a month. What crosses is metadata.

  • A clean answer is not proof the machine is quiet. It covers the folders it names, and it says which known harnesses were not present. A session whose transcripts live elsewhere does not appear at all.

  • A session it cannot place is reported rather than dropped. Some harnesses record no working directory at all, Antigravity among them, and 91 of 288 sessions measured on one machine carry neither a project nor a branch. Grouping those together would announce collisions that nothing supports, so they are listed separately with the reason.

A file is marked recorded when the harness logged a file-writing tool call, and from a command when a redirection was read out of a shell command that may never have completed. Those are different kinds of fact and are never merged.

Keep the status card honest, without the model remembering

mnemosyne_cockpit_update puts your session on the user's canvas, but only when the model decides to call it. A session that forgets leaves a card saying "working" long after it stopped, and a message the human typed on that card waits for a call that may never come.

The package ships a Claude Code hook that closes both gaps. Install the package so the binary is on your path, then wire it in your own settings. The package never touches them.

npm install -g @mnemosyne_os/mcp
// .claude/settings.json, or ~/.claude/settings.json for every project
{
  "hooks": {
    "SessionStart":     [{ "hooks": [{ "type": "command", "command": "mnemosyne-cockpit-hook" }] }],
    "UserPromptSubmit": [{ "hooks": [{ "type": "command", "command": "mnemosyne-cockpit-hook" }] }],
    "Notification":     [{ "hooks": [{ "type": "command", "command": "mnemosyne-cockpit-hook" }] }],
    "Stop":             [{ "hooks": [{ "type": "command", "command": "mnemosyne-cockpit-hook" }] }],
    "SessionEnd":       [{ "hooks": [{ "type": "command", "command": "mnemosyne-cockpit-hook" }] }]
  }
}

The event name arrives on stdin, so one command serves all five. What each does:

Event

The card

Your mail

SessionStart

appears, working

delivered as context

UserPromptSubmit

working, status is your prompt's first line

delivered as context

Notification

waiting, and the question the harness is putting to you

—

Stop

done, status is the answer's first line

if mail is waiting, the stop is refused and the mail is the reason, so the session reads it instead of ending

SessionEnd

goes away

—

Nothing clears waiting on its own, because answering a question is not a prompt. The card holds it until the turn ends or you type something. The host says how long it has been unconfirmed rather than quietly moving on.

Do not reach for npx here, even though the server line above uses it. npx --package=@mnemosyne_os/mcp mnemosyne-cockpit-hook does work, and it cost 2.1 to 3.7 seconds a run on the machine where the installed binary cost 0.4 to 0.7. UserPromptSubmit fires on every message you send, so that difference is the whole feature.

It cannot cost you a session. Every failure path writes one line to stderr and exits 0. Measured on Windows through the installed binary: 0.40 to 0.58 s with Mnemosyne OS running, 0.47 to 0.66 s with it closed. Most of that is Node starting and the shim npm writes on Windows, since a closed local port refuses at once. A refused stop cannot loop either: the mail is marked delivered when it is handed over, so the next stop finds none and ends normally.

Claude Code only. The event names and the refusal format are Claude Code's hook contract. Cursor and Antigravity have their own and will not fire this one. The three agent-awareness tools above work with every harness; this hook does not.

mnemosyne_memory_query, full parameter reference

{
  query:              string;        // required; be specific, longer is fine
  limit?:             number;        // default 10, capped at 50 server-side
  vault?:             string;        // default: $MNEMO_DEFAULT_VAULT
  spine_type_filter?: string[];      // e.g. ["ARCHITECTURE"], ["GIT","BUGFIX"]
  max_content_chars?: number;        // default 600; trims each result snippet
}

The MCP automatically opts into the semantic ranking branch (Vertex 768D / e5-base) and applies an exact-term boost for identifier-like tokens in your query (codenames, hyphenated tokens, version numbers). The result is a list of chronicles ranked by true semantic relevance, not recency.


At session start
  agent → mnemosyne_position_get("my-project")
        ← phase, last position, what was being worked on

During the session
  agent → mnemosyne_memory_query("auth refactor decisions",
                          spine_type_filter=["ARCHITECTURE","DECISION"])
        ← top 10 chronicles, ranked by relevance

When making a decision worth keeping
  agent → mnemosyne_memory_ingest(
            content="Chose JWT over session cookies because we need stateless
                     workers; trade-off: token revocation needs a denylist.",
            spine_type="DECISION")

At session end
  agent → mnemosyne_position_update("my-project",
                                     position="JWT migration shipped, next:
                                               denylist via Redis",
                                     phase="Phase 12")

Next session
  agent → mnemosyne_position_get("my-project")
        ← Resumes from Phase 12 with full context

Troubleshooting

"Cannot connect to ws://127.0.0.1:7799"

Mnemosyne OS Infinity is not running. Launch it. The MCP retries on every tool call, so once Infinity is up, the next query will succeed.

Agent gets SCOPE_DENIED on a vault

The vault is not in MNEMO_VAULTS. Edit your MCP client config, add the name (uppercased), restart the client.

Which vaults can my agent see?

Ask the agent to run mnemosyne_vault_list. It lists every vault Mnemosyne OS exposes, with its id, name and chronicle count, and it flags the ones outside your MNEMO_VAULTS config. Those return SCOPE_DENIED until you add them.

Tool result is too large for my context window

Use max_content_chars to shrink each snippet. The default is 600. Drop it to 200 for browsing, or raise it to 4000 to read a full file. You can also filter with spine_type_filter to drop noisy types.

My new chronicles do not appear

DocWatch ingests on file save with a small delay. Check the spine: if you wrote a markdown with spine: IDEATIONAL frontmatter, it lands as IDEATIONAL. Query with spine_type_filter=["IDEATIONAL"] to surface it.


Privacy Policy

This server is a bridge, not a service. It has no backend of its own, no account and no hosted endpoint: it opens a WebSocket to 127.0.0.1:7799 on your own machine and relays to the Mnemosyne OS desktop application running there.

What it collects. Nothing. No telemetry, no usage tracking, no analytics, no crash reporting. It holds no identifier for you and never asks for one.

What it stores. Nothing of its own, since it is stateless between calls. Your chronicles live in vaults on your disk, written and managed by the desktop application. The agent-awareness tools read your coding agents' transcript files locally and return metadata only (paths, counts, timestamps), never the text of a message or the contents of a file, and they open no network connection at all.

Who else sees it. Two parties, both of them your choice, and nobody beyond them:

  • Your MCP client. Whatever a tool returns goes to the AI client you connected, Claude or Cursor or another, and travels wherever that client sends it. That is what the server is for. It also means a chronicle you let an agent read leaves your machine when your client is a cloud model. Narrow MNEMO_VAULTS to the domains a given agent should reach. A vault left out is refused, including vaults that exist on the machine.

  • The desktop application, for whatever you configured there yourself: a cloud model, a cloud embedder. Those calls are the application's, made with your own keys. This server neither makes them nor sees them.

The server itself shares with no one, sells nothing and rents nothing.

How long it is kept. By this server, not at all. In the application, for as long as you keep it: memory is deleted where it is made, in the app and by you. Note that mnemosyne_memory_ingest writes a permanent chronicle. It is the one call here the agent cannot undo afterwards.

Contact. Privacy questions go to dev@mnemosyne-os.com, XPACEGEMS LLC, 2932 NW 72 Ave, Miami, FL 33122, USA. Full policy: https://mnemosyne-os.io/confidentialite. Bugs and security reports: https://github.com/Mnemosyne-OS/Mnemosyne-Neural-OS/issues.


The @mnemosyne_os packages

All of them live under one npm organization: npmjs.com/org/mnemosyne_os

Package

What it is

@mnemosyne_os/sdk

Build a Layer 2 app: a Node or browser process talking to the local WebSocket surface

@mnemosyne_os/create-app

npm create @mnemosyne_os/app scaffolds that Layer 2 app in one command

@mnemosyne_os/cartridge-sdk

Build an in-app cartridge: a sandboxed iframe widget rendered on the canvas

@mnemosyne_os/mcp (you are here)

MCP server: plug Claude, Cursor or any MCP agent into the vaults

@mnemosyne_os/design-sdk

Skin the OS with JSON alone, no TypeScript

@mnemosyne_os/public-contracts

The shared types and Zod schemas. No business logic

@mnemosyne_os/agent-transcripts

Read what coding agents already write on disk: the connector format and the interpreter

@mnemosyne_os/affine-reader

Read a local AFFiNE workspace and render its documents to Markdown

@mnemosyne_os/forge

CLI: scaffold, list chronicles, import and export

@mnemosyne_os/sync

The name of the P2P layer to come. A placeholder today, not the library


Where Mnemosyne OS lives

Published by XPACEGEMS LLC. Its official addresses:


License

MIT © Tony Trochet / XPACEGEMS LLC


The OS your code talks to

Mnemosyne OS Infinity Edition · download · mnemosyne-os.io · mnemosyne-os.com

Available Tools

25 tools
mnemosyne_aboutA
Read-only

Re-read the server instructions: the governance tenet, the vault protection model (NORMAL and MAXIMUM, mixableWith, isolated sandbox vaults), the spine model, and the rules for an agent working on someone else's memory. Most clients show this text on connect. Call it if yours did not, or to read it again.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

readOnlyHint already covers the non-mutating safety profile. The description adds useful behavioral context: the tool re-reads client-visible server instructions on connect and enumerates the content topic areas. It doesn't specify exact output shape, but for an informational no-parameter tool this is a minor gap rather than a serious omission.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two efficient sentences front-load the action ('Re-read the server instructions') and then provide a compact, useful list of the contained topics plus a concrete call condition. No filler or repetition of the schema/annotations.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a parameterless read-only about/info tool with no output schema, the description is complete: it says what the tool returns conceptually (server instruction text), what topics it covers, and when to invoke it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has no properties, so the zero-parameter baseline applies. There is nothing for the description to add about argument semantics, and the description doesn't invent any.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb and resource: 'Re-read the server instructions,' and then enumerates exactly what those instructions cover (governance tenet, vault protection model, spine model, memory-work rules). This clearly separates it from the operational sibling tools like mnemosyne_memory_query or mnemosyne_vault_list.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives a concrete condition for calling: 'Most clients show this text on connect. Call it if yours did not, or to read it again.' That is an explicit when-to-use rule, and it implies the negative case (if it was already shown, no call is needed).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

mnemosyne_agenda_addA

Put appointments or deadlines into the user's calendar, the Agenda widget on their canvas. Use it when a conversation names a date and time to remember: a meeting, a deadline, "add this to my calendar". The app routes the write through the widget's own store, so what you file is what the user sees. Works with the app closed on a dev install (the headless daemon reads the file); an npm install has no daemon and needs the app running. To read the calendar back, change or remove an appointment, see mnemosyne_agenda_list, mnemosyne_agenda_update and mnemosyne_agenda_remove.

ParametersJSON Schema
NameRequiredDescriptionDefault
eventsYesOne or more appointments to add.

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the annotations, the description discloses useful behavioral context: writes are routed through the widget's own store, so what is filed is what the user sees, and it explains the daemon difference between dev installs and npm installs. It does not contradict annotations and adds meaningful operational detail, though it omits error/return behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well-structured: purpose first, then usage trigger, then the operational caveat, then sibling routing. Every sentence earns its place, and there is no filler or redundant restatement of schema details.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given that the schema richly documents parameters and the description covers purpose, usage, storage behavior, runtime requirements, and alternatives, the tool is almost fully specified. The main gap is that there is no output schema and the description does not state what a successful call returns or how failures surface.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema already covers 100% of parameters with detailed descriptions, including the ISO 8601 timezone caveat and the meaning of each optional field. The description adds no parameter-level detail beyond framing appointments and deadlines, so the schema carries the load and the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('Put appointments or deadlines') and the resource ('the user's calendar, the Agenda widget on their canvas'). It also distinguishes itself from the sibling read/update/remove tools by explicitly directing readers to mnemosyne_agenda_list, mnemosyne_agenda_update, and mnemosyne_agenda_remove for those operations.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives an explicit trigger: 'Use it when a conversation names a date and time to remember', with concrete examples like meetings, deadlines, and 'add this to my calendar'. It also names alternatives for reading, changing, and removing appointments, giving clear routing guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

mnemosyne_agenda_listA
Read-only

Read the user's calendar back: the appointments in a time window, each with the id you need to change or remove it. Call this before mnemosyne_agenda_update and mnemosyne_agenda_remove, which both name appointments by id. The calendar has no archive and a removal cannot be undone. A repeating event appears once, with its cadence and the date of its next occurrence. Times in the answer are the user's machine local time. Works with the app closed on a dev install (the headless daemon reads the file); an npm install has no daemon and needs the app running. Scope agenda:read.

ParametersJSON Schema
NameRequiredDescriptionDefault
toNoISO 8601 date-time (or epoch ms). Only appointments with an occurrence at or before this. Omitted = no upper bound, which for a repeating event means it always counts.
fromNoISO 8601 date-time (or epoch ms). Only appointments with an occurrence at or after this. Omitted = now.
limitNoMaximum appointments returned (default and ceiling: 200).
include_pastNoInclude appointments with no occurrence left. Default false. Those ones come back with no "next" date rather than with their original start dressed up as a future one.

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare readOnlyHint=true. The description adds substantial behavior beyond that: the calendar has no archive, removals cannot be undone, repeating events appear once with cadence and next occurrence, times are machine local, and daemon availability differs between dev and npm installs. No contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but every sentence earns its place: core purpose and id relevance come first, then crucial behavioral caveats and environment requirements. No filler or repetition of schema content.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description compensates by explaining key return characteristics: each appointment includes its id, repeating events show cadence and next occurrence, and times are local. It also covers operational prerequisites and safety implications, making the tool complete to invoke correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all four parameters including defaults and behavior. The description adds context about time windows and returned ids but no new parameter-specific meaning, so the baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'Read the user's calendar back: the appointments in a time window, each with the id you need to change or remove it.' It clearly differentiates this list tool from sibling mutation tools by explaining that it supplies the ids needed by mnemosyne_agenda_update and mnemosyne_agenda_remove.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly tells the agent when to call this tool: 'Call this before mnemosyne_agenda_update and mnemosyne_agenda_remove.' It also gives environment-specific guidance (dev install with headless daemon vs npm install needing the app running) and scopes access with 'Scope agenda:read.'

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

mnemosyne_agenda_removeA
Destructive

Remove appointments from the user's calendar. Read them with mnemosyne_agenda_list first and pass the ids. This tool matches by id only, because "remove my meetings on Thursday" is how an agent removes the wrong Thursday. There is no archive: unlike a To-do task, a removed appointment is gone, so the answer names each one it removed by title and start time for the user to check. A repeating appointment is removed as a whole series, since the calendar cannot cancel a single occurrence. Works with the app closed on a dev install (the headless daemon reads the file); an npm install has no daemon and needs the app running. Scope agenda:write.

ParametersJSON Schema
NameRequiredDescriptionDefault
event_idsYesThe appointment ids to remove, exactly as mnemosyne_agenda_list printed them in brackets. Up to 50.

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the destructiveHint annotation, the description details real consequences: no archive, removal is permanent, repeating appointments are removed as a whole series, and the response names each removed item. It also discloses environment-dependent daemon behavior, adding significant value.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with the core purpose and each sentence carries important caveat or usage information. It is longer than minimal, but every sentence earns its place; no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Despite lacking an output schema, the description explains what the response will name. It covers destructive semantics, repeating events, the id prerequisite, and the daemon/app runtime condition, making it fully actionable for an agent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema already covers the single parameter at 100%. The description adds useful extra semantics: ids must be exactly as mnemosyne_agenda_list printed them in brackets hire up to 50. This goes beyond the schema's generic wording.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Remove appointments from the user's calendar.' It clearly distinguishes this tool from siblings by emphasizing id-only matching and contrast with archive-able To-do tasks.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives explicit procedural guidance: read appointments with mnemosyne_agenda_list first and pass the printed ids. It also warns against phrase-based removal. It does not spell out all exclusions or alternatives, but the context is clear enough for an agent to use it correctly.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

mnemosyne_agenda_updateA
Destructive

Change an appointment already in the user's calendar: move it, rename it, add or drop a reminder, start or stop it repeating. It names appointments by the id from mnemosyne_agenda_list, so call that first. A field you leave out is left alone, and passing null clears it. A start or end time that cannot be read refuses the change rather than leaving the old one quietly in place. A cadence outside daily, weekly, monthly and yearly is refused rather than turned into a one-off. Works with the app closed on a dev install (the headless daemon reads the file); an npm install has no daemon and needs the app running. Scope agenda:write.

ParametersJSON Schema
NameRequiredDescriptionDefault
changesYesOne entry per appointment. Up to 50 per call.

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate destructiveHint=true, but the description adds substantial behavioral context beyond that: it explains failure modes (refuses unreadable start/end times, refuses invalid cadence), the effect of null on recurrence (drops series end), and environment-specific behavior (works on dev install via daemon, needs app running on npm install). This goes well beyond the annotation flags and gives the agent a clear picture of side effects and constraints.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, dense paragraph that front-loads the core purpose and then covers prerequisites, semantics, failure handling, and environmental notes. Every sentence carries operational weight; there is no filler. While it is longer than strictly minimal, the complexity of the tool justifies the length, and it is organized logically.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation tool with destructiveHint=true and no output schema, the description is remarkably complete: it covers the required id source, the patch semantics, failure conditions, recurrence behavior, and the dev/npm install difference. An agent can correctly invoke this tool for any intended change without additional clarification. The only minor omission is the exact response shape, but that is not needed given the absence of an output schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, with each parameter having a detailed description (e.g., start time timezone semantics, null behavior for clearable fields). The description adds global semantics over the schema: 'A field you leave out is left alone, and passing null clears it' – a rule not per-parameter. It also clarifies the batch nature (up to 50 per call). This enhances the schema's guidance without redundancy.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a clear verb ('Change') and resource ('an appointment already in the user's calendar'), then enumerates the specific actions (move, rename, add/drop reminder, start/stop repeating). It distinguishes itself from siblings like mnemosyne_agenda_add and mnemosyne_agenda_remove by focusing on modification, and it explicitly references the id source (mnemosyne_agenda_list), making its scope unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly states the prerequisite: 'It names appointments by the id from mnemosyne_agenda_list, so call that first.' It also explains the semantics of omitted fields (left alone) and null (clears), which tells the agent how to construct change requests. However, it does not explicitly contrast with add/remove or state when not to use this tool, though the sibling names make that inference easy.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

mnemosyne_agent_collisionsA
Read-only

Check whether two agent sessions are live in the same git working tree and branch right now, across every installed harness. Call it before git add -A, before a commit, and before a rebase: one working tree shares one git index, so a commit from one session picks up whatever the other has staged. It answers from transcript files on disk and needs neither Mnemosyne OS nor a token. Each recorded directory is resolved to its working tree first, because one cd into a subfolder would make two sessions in one repository look like two projects. A clean answer covers what is readable: a session whose transcripts live elsewhere does not appear at all, and sessions whose harness records no directory are listed separately as unplaceable.

ParametersJSON Schema
NameRequiredDescriptionDefault
projectNoRestrict to one working tree, matched exactly (separators and case are normalised). Pass the output of `git rev-parse --show-toplevel`. Omit to check every project on the machine.

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The readOnlyHint is consistent and the description adds substantial non-obvious behavior: it reads transcript files from disk, needs neither Mnemosyne OS nor a token, resolves each directory to a working tree, and explicitly discloses that unreadable sessions are omitted and unplaceable sessions are listed separately. This gives an agent a clear model of what the tool can and cannot detect.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but every sentence earns its place: purpose, triggering scenarios, mechanism, and limitations. The critical before-commit guidance is front-loaded, and the later caveats about transcript readability and unplaceable sessions are important enough to justify the length.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers when to use it, how it works, its data source, and its limitations, and the schema documents the only parameter. The one gap is the absence of output/response shape details, especially since there is no output schema; an agent cannot confidently predict the exact format of a collision report. This is minor but prevents a 5.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The single optional `project` parameter is already fully described in the schema, including exact matching, path normalisation, passing `git rev-parse --show-toplevel`, and omission meaning 'check every project'. The description adds context about working-tree resolution but does not need to repeat the parameter details. Baseline 3 is appropriate given the 100% schema coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Opens with a specific verb and scope: 'Check whether two agent sessions are live in the same git working tree and branch right now, across every installed harness.' This is a distinct operation and clearly not the same as the sibling memory, todo, or agent tools. The title 'Check agent collisions' reinforces the same purpose without needing the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly instructs when to call it: 'Call it before `git add -A`, before a commit, and before a rebase'. It also explains why via the shared git index, which helps an agent decide to insert this check into Git workflows. No near-sibling alternative is needed because the use case is unique.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

mnemosyne_agent_filesA
Read-only

List the files other agent sessions wrote or edited recently, newest first, with the session each one came from. Paths and timestamps only. Each entry says how it is known. recorded means the harness logged a file-writing tool call. from a command means a redirection was read out of a shell command the session ran, which may never have completed. It spans every installed harness and names the agent each line came from. Use it to see what another session has already touched before you edit the same area.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMax files to return (default 40, max 200).
projectNoKeep only files from sessions whose project path contains this string.
containsNoKeep only paths containing this string. A folder name works as well as a file name.

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only provide readOnlyHint=true. The description goes well beyond that by explaining output shape ('Paths and timestamps only'), provenance labels ('recorded' vs 'from a command'), the reliability caveat ('may never have completed'), and scope ('spans every installed harness'). This is rich behavioral context that helps the agent interpret results safely.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the core purpose and ordering, then packs necessary caveats (paths/timestamps only, provenance explanation, unreliability warning) into a compact paragraph. Every sentence earns its place; no filler or repetition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only list tool with three optional params and no output schema, the description explains what is returned ('paths and timestamps', session, agent, how it is known) and why it is useful. Combined with the fully documented schema, an agent has everything it needs to call and interpret correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% (limit, project, contains are all documented with types and constraints), so the baseline is 3. The tool description adds no additional parameter semantics beyond what the schema already says, which is acceptable.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('List the files other agent sessions wrote or edited recently'), gives ordering ('newest first'), and names the exact use case ('see what another session has already touched'). It is clearly distinguished from siblings by targeting agent-written files, which no other mnemosyne_* tool covers.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly says when to use it ('Use it to see what another session has already touched before you edit the same area'). It does not name alternatives or state when not to use it, but the context is clear enough for an agent to select it over unrelated siblings like mnemosyne_memory_query or mnemosyne_agent_list.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

mnemosyne_agent_listA
Read-only

List the other coding-agent sessions on this machine, read from the transcript files their harnesses already write to disk. Metadata only: conversation name, project, git branch, model, last tool, how many files were touched, and when a line was last written. It reads every harness installed here, not just your own, so you can see a session from a different agent working in your repository. Use it before you touch shared state. It reports when a line was last seen and you draw the conclusion: a crashed session and an idle one fall equally silent, so nothing here can tell you an agent is working. Works with the Mnemosyne OS app closed.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoHow many transcripts to open PER HARNESS, newest first (default 40, max 200). Capping the merged total instead would let a chatty agent push a quiet one off the end, and the quiet one is the session you did not know about.
projectNoKeep only sessions whose project path contains this string. Pass the repository folder name to scope the answer to the repo you are working in.
live_onlyNoOnly sessions that wrote a line recently (see the window printed in the answer). Default false, which lists the most recent sessions whether or not they moved lately.

TDQS

A4.3/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The readOnlyHint=true annotation is confirmed and enriched by the description's 'Metadata only' and 'reads' language. More importantly, the description discloses non-obvious behavior beyond the annotation: it 'reads every harness installed here,' it works 'with the Mnemosyne OS app closed,' and it exposes the critical caveat that 'a crashed session and an idle one fall equally silent, so nothing here can tell you an agent is working' — telling the agent it must draw the liveness conclusion itself, not the tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is longer than minimal but every sentence earns its place: the core function leads, followed by the metadata list, the cross-harness scope, the usage trigger, the silent-session caveat, and the app-closed note. It is front-loaded with the primary purpose and well-ordered, though the metadata enumeration and caveat sentences could arguably be tightened without losing meaning.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a filterable listing tool with three well-documented parameters, no required params, and no output schema, the description covers the essential ground: return metadata fields, scope, usage timing, and the liveness limitation. One minor gap is that the exact response structure isn't formally specified, but the description's metadata field list substantially compensates for the absent output schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline of 3 applies — the schema already documents limit, project, and live_only clearly, including the per-harness semantics of limit. The description only loosely reinforces this by mentioning 'how many files were touched' and 'when a line was last written,' which map to the live_only window, but it adds no new parameter syntax or format details beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'List the other coding-agent sessions on this machine, read from the transcript files their harnesses already write to disk.' It enumerates exactly what metadata is returned (conversation name, project, git branch, model, last tool, files touched, last-write time) and scopes itself clearly ('Metadata only', 'every harness installed here, not just your own'). This distinguishes it from siblings like mnemosyne_agent_files and mnemosyne_agent_collisions without needing their schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives an explicit when-to-use directive: 'Use it before you touch shared state.' It also explains the value proposition contextually ('you can see a session from a different agent working in your repository'). It does not name sibling alternatives or state when-not-to-use, but the practical trigger is clear and actionable.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

mnemosyne_cockpit_updateA

Update your own status card on the user's canvas, the cockpit. Call it when you start a task ("working", with a short title and status), when you need the user ("waiting"), when you are stuck ("blocked"), and when you finish ("done"). "waiting" and "blocked" make the card pulse and the taskbar flash. Send "waiting" IN THE SAME TURN as the question you ask the user, right before you stop: the harness only knows to say "waiting" for a permission prompt, and a question asked in the chat is otherwise just the end of a turn — the user never sees that you are waiting for them. Your "waiting" survives the harness saying the turn ended. You declare the state; the app prints it next to the time since your last call, so keep calling at real milestones or the card goes quiet. The answer carries any message the user left on your card. Read it and act on it. Needs the app window open. This is a card, so nothing is stored in memory.

ParametersJSON Schema
NameRequiredDescriptionDefault
stateYes"working" = on it; "waiting" = you asked the human something and stopped; "done" = the task is finished; "blocked" = you cannot continue without them; "closed" = this conversation is over (removes the card).
titleNoWhat this conversation is about, in a few words (the card's title). Give it at least once; later calls may omit it.
detailNoUp to 4 short lines under the status (files, a branch, a count). Optional.
statusNoOne line: what you are doing right now, or what you need. 160 characters max.
desktopNoThe desktop this session should put its cards on, by name, exactly as the human wrote it. Send it once when they tell you; the card stays there for the rest of the conversation without you repeating it. The match is exact apart from case, and an unknown or duplicated name is reported back with the names that exist rather than guessed at. Leave it out and the card goes to whichever desktop answers for this project, which is what most sessions want.
sessionNoOnly when the harness publishes no CLAUDE_CODE_SESSION_ID: a stable id for this conversation, reused on every call.
attentionNoAsk for the human's eye even in "working"/"done" (the card pulses, the taskbar flashes). "waiting" and "blocked" ask on their own.

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations are minimal (readOnlyHint=false, destructiveHint=false), so the description carries the burden and does so richly: 'waiting'/'blocked' pulse the card and flash the taskbar, 'waiting' survives the harness ending the turn, the app prints time-since-last-call, the call requires the app window open, and nothing is stored in memory. It also discloses that the response carries the user's message.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose is front-loaded and the guidance is organized logically from states to the 'waiting' caveat to behavioral notes. It is dense and somewhat long, with the 'waiting'/'blocked' behavior touched on in more than one sentence, but for a stateful tool this length is mostly earned.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation tool with no output schema, the description covers prerequisites ('Needs the app window open'), persistence behavior ('nothing is stored in memory', title given once), the return value (the user's message on the card), and the toolkit's rendering behavior. Nothing an agent needs to invoke it correctly appears to be missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all seven parameters (states, title, detail, desktop, session, attention) with enums and defaults. The description restates the state vocabulary and adds usage context for 'waiting' and title persistence, but adds little semantic detail beyond the schema. Baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'Update your own status card on the user's canvas, the cockpit.' It identifies the action clearly and distinguishes it from all sibling tools (agenda, todo, memory, agent tools) since it is the only one that writes to this status card.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It maps each lifecycle moment to a state ('working' at start, 'waiting' when the user is needed, 'blocked' when stuck, 'done' on finish) and explicitly addresses the tricky case: send 'waiting' in the same turn as the question, right before stopping, because the harness only auto-handles permission prompts. This is explicit when-to-use guidance that no sibling conflicts with.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

mnemosyne_dream_bridgesA
Read-only

List the connections the Dream State engine found between memories while the machine was idle. Each bridge links two memories, sometimes from different vaults, with a Dream Bridge Score, a raw cosine similarity, and a short excerpt of both sides. Use it to surface associations the memory made on its own, to audit whether they are insightful or noise, or to seed creative exploration. An empty list is normal and means the engine has not produced bridges yet. It runs on idle time, when the user enables it in Settings.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMax bridges to return, strongest (highest dbs) first (default: 50, server cap: 200).
min_dbsNoOnly bridges with a Dream Bridge Score at or above this value (0-1).
session_idNoRestrict to one dream scan session.
chronicle_idNoOnly bridges touching this chronicle id (either endpoint).

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint=true, which the 'List' framing matches — no contradiction. Beyond the annotation, the description adds meaningful behavioral context: an empty result is a normal state rather than an error, and data only exists if the idle-time engine has run and the user enabled it in Settings. This helps the agent avoid misinterpreting empty responses or retrying unnecessarily.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four sentences, each earning its place: what the tool lists, what a bridge contains, when to use it, and what empty/normal behavior looks like. The core purpose is front-loaded and there is zero filler or redundancy with the structured fields.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema here, the description carries the burden of describing return values and does so adequately (score, cosine similarity, excerpts of both sides). It also covers data availability and empty-result semantics. The only notable gap is not differentiating the tool from the similarly named mnemosyne_resonance_list sibling, which an agent would benefit from knowing about.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema documents all four parameters (limit with default/cap, min_dbs range, session_id, chronicle_id). Per the baseline rule, the description need not repeat them; it adds indirect value by explaining what a bridge contains, which clarifies the meaning of min_dbs, but provides no parameter-specific guidance beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'List the connections the Dream State engine found between memories while the machine was idle.' It further defines what a bridge is (two memories with a Dream Bridge Score, raw cosine similarity, and excerpts), which disambiguates it from siblings like mnemosyne_resonance_list by attributing the data to the idle-time Dream State engine.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit use cases: 'surface associations the memory made on its own, to audit whether they are insightful or noise, or to seed creative exploration.' It also sets the expectation that an empty list is normal and explains the data depends on the user enabling the engine in Settings. However, it never names the closely related sibling mnemosyne_resonance_list or states when that alternative should be preferred, so exclusions are missing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

mnemosyne_git_logA
Read-only

Read recent commits from the git repository this Mnemosyne OS is configured to read. Read-only. Each commit carries an 8-character hash, the subject line, the author and the date, newest first. The path is set on the app side, so this cannot be pointed at another checkout, and it needs the monorepo:read scope. When either is missing the answer names what is missing, rather than returning an empty list that would read as "no commits". Use mnemosyne_memory_query for the reasoning behind a change.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoNumber of commits to return (default: 20)
sinceNoTime range (e.g. "7 days ago", "2024-01-01")30 days ago

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Read-only is already in annotations, but the description adds critical behavior: commit fields, newest-first ordering, config/scope failure reporting instead of empty list. This tells the agent what to expect and how to interpret 'no commits' scenarios.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four dense sentences with the main action first, then behavior, constraints, and routing. No filler or repeated schema information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description compensates by specifying the return shape (hash, subject, author, date, newest first) and failure behaviors. Scope, configuration, and alternative usage are covered, making it complete for a simple read tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3; limit and since are fully documented in the schema. The description adds no extra parameter-level detail beyond that, but it also doesn't need to.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource ('Read recent commits from the git repository') and states the exact output fields and ordering. It distinguishes itself from mnemosyne_memory_query by noting that tool covers reasoning behind a change, so an agent can pick correctly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly routes to mnemosyne_memory_query for the reasoning behind a change, giving a when-not alternative. It also states the tool is fixed to the configured checkout and requires monorepo:read scope, so agents know when it is usable and what failure looks like.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

mnemosyne_memory_askA
Read-only

Ask a question and get a prose answer grounded in the vault, plus the memories it drew on. It runs the full local RAG pipeline: deeper retrieval, lexical fusion and a re-rank. That also makes it the better retriever, so reach for it whenever you need to find something, and read the Sources list even if you ignore the prose. It is at its best on why, who and how questions that span many memories. It takes up to about 30 seconds. The prose is a model rewording of the sources, so quote the sources instead, or fetch them with mnemosyne_memory_query. Check them before you trust the answer.

ParametersJSON Schema
NameRequiredDescriptionDefault
vaultNoVault to reason over, case-insensitive. Default for this deployment: "DEV". A vault this server was not configured for is refused with SCOPE_DENIED.DEV
questionYesA question in plain language, the way you would ask a colleague who knows the material. Be specific.

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, so safety is covered. The description goes beyond by disclosing performance (up to 30 seconds), the nature of the prose (a model rewording, so quote sources instead), and instructing to verify sources before trusting the answer. This adds rich, non-obvious behavioral context that helps the agent set expectations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact yet dense, with each sentence serving a distinct purpose: main action, pipeline details, usage guidance, performance caveat, and source-trust warning. It front-loads the core purpose and avoids fluff, making it efficient for an agent to parse.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex tool with no output schema, it fully describes the output (prose + sources), performance characteristics, and trust considerations. It also explains why to use the Sources list, covering all the information an agent needs to invoke and interpret the result correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% for both parameters, so the schema already documents them. The description adds extra nuance on the question parameter ('plain language, the way you would ask a colleague who knows the material') and clarifies vault behavior, providing value beyond the schema without being redundant.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb (ask), a resource (the vault), and the exact output (prose answer plus sources). It explicitly differentiates itself from siblings like mnemosyne_memory_query by positioning itself as the better retriever, leaving no ambiguity about what it does.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives explicit when-to-use guidance: 'reach for it whenever you need to find something' and identifies its sweet spot ('why, who and how questions that span many memories'). It also names the alternative (mnemosyne_memory_query) for fetching sources, so the agent knows exactly when to choose this over a sibling.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

mnemosyne_memory_ingestA

Write a memory into a vault: a decision, an architecture note, a debug finding, a session summary. It is stored permanently and indexed for every future agent to retrieve. Use it at the end of a meaningful work session, or whenever a decision is reached that someone will want to recall later.

ParametersJSON Schema
NameRequiredDescriptionDefault
vaultNoTarget vault TOKEN: the folder name uppercased, with spaces and hyphens as underscores. The path-shaped id from mnemosyne_vault_list also works. Default for this deployment: "DEV". Scoped tokens, which are configuration rather than a census: DEV, PERSONAL, SOCIAL. Ingest is permanent, so confirm the vault exists with mnemosyne_vault_list before writing anywhere new.DEV
contentYesThe content to store. Markdown works. Write it self-contained, and say WHY the decision was made rather than only what was decided.
spine_typeNoSemantic type of the content. ARCHITECTURE is boosted heavily in code-scope queries: use it for design docs, big-picture decisions and structural choices. DECISION for narrower trade-offs. BUGFIX and DEBUG for incident learnings. SESSION for where you left off. FEATURE for new capabilities. NOTE for the rest.NOTE

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The annotations are thin (readOnlyHint=false, destructiveHint=false), so the description carries the full disclosure burden, and it delivers the single most important behavioral fact: 'stored permanently and indexed for every future agent to retrieve.' This irreversibility warning is then reinforced in the vault parameter ('Ingest is permanent'), giving agents the clear risk cue they need before calling a write tool. No annotation contradiction arises — a permanent additive write is consistent with destructiveHint=false.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, zero filler: the action and acceptable content types are front-loaded, the permanence trait comes second, and the when-to-use guidance closes. Every sentence earns its place and none repeats what annotations or schema restate, making this an ideal length for agent consumption.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-required-parameter write tool with 100% schema coverage, the description plus schema covers purpose, when to use, target vault semantics, content formatting, and the permanence risk almost completely. The only gap is the operation's response shape (no output schema exists, and the description never says whether the ingest returns a confirmation or identifier), which would help an agent verify successful persistence.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and the parameter descriptions are exceptionally rich — the vault parameter explains token formatting, defaults, scoped tokens, and the permanence caveat, while spine_type maps enum values to usage contexts. Per the baseline rule, the schema already does the heavy lifting, and the main tool description adds no parameter-level meaning of its own, so a 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb+resource pair, 'Write a memory into a vault', and enumerates the exact content kinds it accepts (decision, architecture note, debug finding, session summary). The phrase 'stored permanently and indexed for every future agent to retrieve' implicitly differentiates this write tool from the read-only siblings like mnemosyne_memory_query and mnemosyne_memory_ask, so an agent can tell them apart without opening their schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit when-to-use context: 'at the end of a meaningful work session, or whenever a decision is reached that someone will want to recall later.' The vault parameter additionally embeds a prerequisite ('confirm the vault exists with mnemosyne_vault_list before writing anywhere new'). However, it never names read-side or task alternatives (memory_query, memory_ask, todo_add) or states when NOT to use it, so it stops just short of full exclusionary guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

mnemosyne_memory_queryA
Read-only

Search a vault and get the matching memories back word for word: notes, code, decisions, session summaries, commit history. Nothing is rewritten, so use this when you need the source text itself, to quote it or to write documentation from it. Ranking is vector similarity fused with a local BM25 channel, weighted by spine type. To FIND something rather than quote it, prefer mnemosyne_memory_ask, even when you only want its sources. It retrieves deeper and re-ranks, so it catches rare literal terms this tool misses: proper nouns, identifiers, product names. The score is not a confidence value. A hit and a miss come back with similar numbers, so judge the text, not the number beside it.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMax number of results (default: 10, max: 50)
queryYesThe search query. Be specific. Examples: "why we moved off Postgres", "the retry logic in the payment worker", "notes from the March planning session".
vaultNoVault TOKEN to search, case-insensitive. A token is the folder name uppercased, with spaces and hyphens as underscores. The path-shaped id from mnemosyne_vault_list also works. This deployment's default is "DEV". It is also scoped for: PERSONAL, SOCIAL. That list is configuration rather than a census, so call mnemosyne_vault_list for the vaults that exist on this machine. Anything outside the list is refused.DEV
max_content_charsNoPer-memory snippet size in characters (default 600). Each result is cut to this length, with a hint about its full size. Raise it past 2000 when you need whole files, and watch your context window.
spine_type_filterNoOptional whitelist of spine types. Use ["ARCHITECTURE"] for design docs, ["GIT"] for commit history, ["BUGFIX","DEBUG"] for incident knowledge, ["SOURCE_CODE"] for code only. Omit it to get every type, with the SOURCE_CODE scope weighting deciding the ranking.

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With readOnlyHint true, the description still adds critical behavioral context: 'Nothing is rewritten,' the ranking uses vector similarity fused with a local BM25 channel, and the score is not a confidence value because hits and misses return similar numbers. This goes beyond the annotations and helps the agent correctly interpret results. No contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Each sentence earns its place: purpose, when-to-use, ranking mechanics, alternative routing, and score interpretation. The most important usage guidance is front-loaded. Though it is a longer description, it is dense with actionable information and contains no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a five-parameter tool with no output schema, this description is remarkably complete. It explains what results look like ('word for word'), that snippets are truncated with size hints, how ranking works, and how to interpret the score. Combined with the richly described input schema, an agent has enough to select and invoke the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all five parameters thoroughly, including examples for query and vault token semantics. The description itself does not add parameter-level details, but it doesn't need to; the baseline of 3 applies since the schema carries the load. The description's ranking and score caveats are behavioral, not parameter semantics.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'Search a vault and get the matching memories back word for word.' It names exactly what content is returned (notes, code, decisions, session summaries, commit history) and explicitly separates this tool from mnemosyne_memory_ask. This is unambiguous and distinguishes it from siblings without opening their schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit when-to-use guidance: use it when you need source text itself, to quote it, or to write documentation from it. It also gives an explicit alternative: prefer mnemosyne_memory_ask 'to FIND something rather than quote it,' explaining why because that tool re-ranks and catches rare literal terms. This is clear routing guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

mnemosyne_pheme_radarA
Read-only

Read what the user's Pheme radar last found: fresh threads in the subreddits and Hacker News queries they watch. Each thread carries a topic score, and a tier once they have run the Mnemosyne pass. The tier says how much substance they can bring to that thread. The answer leads with when the scan ran, because it is only as fresh as the last time they opened Pheme and pressed Scan. Use it to find threads to draft a reply for; the user posts it. "No radar" means no scan has been projected yet, so ask them to open Pheme and scan. Needs the app running, though Pheme itself can be closed.

ParametersJSON Schema
NameRequiredDescriptionDefault
tierNoOnly threads of this tier. Omit for all (unranked ones included, last).
limitNoHow many threads at most (default 30, max 100).

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description adds substantial behavioral context beyond the readOnlyHint annotation: it explains the freshness of results, the 'No radar' condition, the prerequisite of the app running, and the structure of the answer (leads with scan time). This fully discloses the tool's behavior and is consistent with the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is about 100 words, but every sentence adds value: it explains the tool's purpose, output semantics, freshness, usage, and prerequisites. It is front-loaded with the core purpose and structured logically, though slightly longer than strictly necessary for a simple read operation.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has no output schema and only two optional parameters, the description covers all essential information: what results are returned (threads with scores and tiers), the freshness caveat, the 'No radar' handling, and the runtime prerequisite. An agent can confidently decide when to call it and what to expect.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 100% schema description coverage, the baseline is 3, but the description enriches the tier parameter by explaining what a tier means ('how much substance they can bring to that thread'). It also clarifies the output includes topic scores and tiers, which helps agents decide whether to filter by tier. The limit parameter is well-covered by the schema alone.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb 'Read' and the resource 'Pheme radar', and explains what it returns: fresh threads from subreddits and Hacker News queries. It also distinguishes itself from siblings like mnemosyne_pheme_watch by focusing on reading results rather than setting up watches, making the purpose unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides a clear use case ('Use it to find threads to draft a reply for') and contextual conditions like the need for the app running and the 'No radar' edge case. However, it does not explicitly mention alternatives or when not to use this tool, though the read/watch distinction is implicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

mnemosyne_pheme_watchA

Add a subreddit, a Hacker News search query or a topic to the user's Pheme radar, or take one off. Pheme is their reputation cartridge: it finds fresh threads worth a genuine reply. The lists belong to them. An operation that would empty one is refused. Every operation reports its own outcome: done, already there, not there, refused. The user gets a receipt in Pheme naming this agent and what changed. Use it when a conversation finds a community worth watching. It scans nothing and posts nothing; the user posts. Needs the app running, though Pheme itself can be closed.

ParametersJSON Schema
NameRequiredDescriptionDefault
opsYesThe changes, in order. Up to 50.

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description adds substantial behavioral context beyond the annotations: operations that would empty a list are refused, each operation reports a specific outcome ('done, already there, not there, refused'), the user receives a receipt naming the agent and what changed, and the tool scans/posts nothing. This fully discloses side effects, refusal behavior, and reporting, with no contradiction to annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with the core action and then provides necessary behavioral details. A few sentences are somewhat narrative ('Pheme is their reputation cartridge...'), but each sentence contributes useful context about outcomes, receipts, or prerequisites. It is appropriately sized for a mutation tool with one nested parameter.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-parameter tool with no output schema, the description is complete: it explains the action, domain, side effects, refusal behavior, return-like outcome reporting, and runtime prerequisite. An agent has enough guidance to select and invoke the tool correctly without needing additional context.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, with the schema already documenting the op/kind/value fields and their enums. The description adds domain framing (Pheme radar, lists, outcomes) but does not introduce additional parameter syntax or constraints. Baseline 3 applies because the schema carries the parameter semantics.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific action ('Add... or take one off') with a specific resource (subreddits, Hacker News search queries, topics on the Pheme radar). It also defines the domain concept 'Pheme' and distinguishes the edit-and-remove operation from mere viewing, which likely separates it from sibling mnemosyne_pheme_radar.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives a clear trigger condition: 'Use it when a conversation finds a community worth watching.' It also provides a prerequisite ('Needs the app running, though Pheme itself can be closed'). It does not explicitly name alternative tools or state when not to use it, but the usage context is clear enough.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

mnemosyne_position_getA
Read-only

Read where one resonance was left: the phase, and the free-text note an agent or the cockpit wrote when it stopped. Read-only. It returns the resonance id, when it was saved, the spine type of the memory, and the whole note. An id nothing was ever saved under answers in plain words and points at mnemosyne_position_update, so a blank history and a failed call do not look alike. Call mnemosyne_resonance_list when you do not know the id.

ParametersJSON Schema
NameRequiredDescriptionDefault
resonance_idYesID of the resonance, for example "checkout-rewrite" or "v2-migration".

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the readOnlyHint annotation, the description discloses the exact return fields and the distinctive unknown-id behavior: a plain-words response that names the update tool, so a blank history is distinguishable from a failed call. No contradiction with annotations exists.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three purposeful sentences front-load the operation and read-only nature, then add return details and fallback behavior. Every clause earns its place, with the alternative routing placed at the end.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter read tool with no output schema, the description fully specifies what the response contains and how the one edge case behaves. An agent has everything needed to decide when to call it and to interpret the result correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema fully documents the single required resonance_id parameter with an example, so the schema carries the parameter-meaning burden. The description reinforces that the id selects a prior saved resonance but adds no new parameter syntax or format details.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and object: 'Read where one resonance was left,' then enumerates what is returned. It differentiates itself from siblings by explicitly referencing mnemosyne_position_update and mnemosyne_resonance_list as related alternatives.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly routes the agent to mnemosyne_resonance_list when the id is unknown, and explains that a missing-id response points to mnemosyne_position_update. This gives clear when-to-use and when-not-to-use guidance beyond the read-only positioning.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

mnemosyne_position_updateA

Record where you left off on a resonance: phase, current state, next steps. Call it at the end of a session. It is stored as a DECISION memory in the vault.

ParametersJSON Schema
NameRequiredDescriptionDefault
phaseYesCurrent phase label (e.g. "Phase 52", "v1.1.0 release")
positionYesDescription of the current position / what was done / what is next
resonance_idYesID of the resonance to update

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description discloses a meaningful behavioral trait beyond the annotations: the output is 'stored as a DECISION memory in the vault.' This tells the agent the write is persistent and goes to a specific memory category. The description does not contradict the readOnlyHint=false and destructiveHint=false annotations, and adds value without needing to repeat them.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences, with the core action front-loaded and the usage trigger immediately following. Every sentence earns its place: the first defines purpose and content, the second gives the usage moment and storage semantics. No filler or restatement of the tool name.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a three-parameter write tool with no output schema, the description covers the what, when, and where-it-stored aspects. It does not explicitly state whether the position is overwritten or appended, but the tool name 'position_update' and required resonance_id make the update semantics reasonably clear. The description is adequate for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and each parameter already has a descriptive schema comment. The tool description adds a light mapping by mentioning 'phase, current state, next steps' which maps to phase and position, but it does not materially expand on the schema's existing parameter descriptions. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb ('Record') and the resource ('a resonance') and exactly what is recorded: phase, current state, next steps. It also distinguishes this position-saving action from siblings like mnemosyne_position_get or memory_ingest by framing it as a session-end bookmark, not a fetch or generic memory write.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit guidance on when to call the tool ('Call it at the end of a session'), which provides clear context. It does not name specific alternatives or when not to use it, but the instruction is unambiguous enough for an agent to know the intended trigger.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

mnemosyne_resonance_listA
Read-only

List the resonances recorded in the default vault. A resonance is a workspace that tracks one ongoing project. Each entry carries its id, its last phase, how many minutes ago it moved, and the id of the memory behind it. Read-only. There is no status filter, so you get every resonance the scan matched in one vault, from at most 30 candidates. An empty result is written in words and means no resonance has been recorded yet. Read one in full with mnemosyne_position_get, or write a new one with mnemosyne_position_update.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the readOnlyHint annotation, the description discloses important behavioral details: there is no status filter, results are capped at 30 candidates, and an empty result is 'written in words' rather than returned as an empty array. It also explains what each entry carries (id, last phase, minutes since moved, memory id), which is valuable context given there is no output schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with the tool's purpose and scoping ('List the resonances recorded in the default vault'), then adds concise, meaningful details about result contents, limits, and empty behavior. Every sentence earns its place, and the sibling-routing guidance is compactly placed at the end without padding.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description is complete for a parameterless read-only list tool: it explains the domain concept, default scope, result fields, candidate limit, empty-result representation, and how to follow up via sibling tools. With no output schema and no parameters, the description carries the full burden of telling an agent what will happen and what it will receive, and it does so thoroughly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters and 100% schema description coverage, so the description has no parameter semantics to add. The description instead uses the parameter-free space to clarify the result shape and domain concept, which is appropriate for a no-argument list tool. The baseline of 4 for zero-parameter tools applies here.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'List the resonances recorded in the default vault.' It defines the domain term ('A resonance is a workspace that tracks one ongoing project') and distinguishes itself from siblings by pointing to mnemosyne_position_get for full detail and mnemosyne_position_update for writing. This is precise and clearly separates it from related tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly says there is no status filter, so the tool returns every resonance the scan matched, up to 30 candidates. It also names alternatives for adjacent use cases: 'Read one in full with mnemosyne_position_get, or write a new one with mnemosyne_position_update.' This gives an agent clear guidance on when to use this list tool versus its siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

mnemosyne_spine_assignmentsA
Read-only

See how Mnemosyne classified its memories. It returns the memory-to-spine assignments for one vault, newest first, and the per-spine counts for the whole vault. Use it to check whether memories landed in the right spines, to read a vault's composition at a glance, or to find the taxon ids to pass as spine_type_filter in mnemosyne_memory_query. Set include_taxonomy to also get the global spine tree.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMax assignments returned (default: 25 to stay context-friendly, server cap: 500).
vaultNoVault to inspect, case-insensitive. Default for this deployment: "DEV".DEV
offsetNoPagination offset (default: 0).
spine_typeNoRestrict the assignment page to one spine taxon id (e.g. "DOCUMENT", "GIT"). Counts stay whole-vault.
include_taxonomyNoAlso return the global spine taxonomy tree (natures → sub-spines).

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The readOnlyHint annotation marks this as safe to call, and the description adds behavioral context beyond the annotation: results are newest-first, spine_type restricts only the assignment page while counts remain whole-vault, and include_taxonomy pulls in the global spine tree. No contradiction with the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences with no waste. The core purpose and return shape are front-loaded in sentence one, use cases in sentence two, and the optional taxonomy flag in sentence three. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description carries the burden of describing return values, which it does: assignments, whole-vault counts, and optionally the taxonomy tree. The readOnly annotation covers the safety profile and the schema documents all parameters, so nothing an agent needs to invoke it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3; the description adds real value on top by explaining that spine_type values can be fed back as spine_type_filter into mnemosyne_memory_query, clarifying the cross-tool contract, and by noting the pagination-vs-counts semantic where spine_type scopes the page but not the counts.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('returns') with a precise resource ('memory-to-spine assignments') and clear scope ('one vault, newest first' + 'per-spine counts for the whole vault'). It clearly distinguishes itself from the close sibling mnemosyne_memory_query by stating its output feeds that tool's spine_type_filter parameter rather than querying memories directly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives three concrete use cases: verifying spine placement, reading vault composition, and obtaining taxon ids for use in mnemosyne_memory_query. It references the sibling alternative where the output is consumed. It lacks an explicit 'use X instead when...' exclusion, but the named use cases provide solid routing guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

mnemosyne_todo_addA

Put tasks into the user's To-do backlog, the To-do widget on their canvas, in order and optionally under named steps. Use it when a conversation has settled what to do. Name the destination list in "list": call once without it to get the existing lists back on a LIST_NOT_FOUND answer, or pass create_list: true to make a new one. Always name a list. The app routes the write through the widget's own store, so what you file is what the user sees. Works with the app closed on a dev install (the headless daemon reads the file); an npm install has no daemon and needs the app running. To read the backlog back or change what is in it, see mnemosyne_todo_list, mnemosyne_todo_update and mnemosyne_todo_categories.

ParametersJSON Schema
NameRequiredDescriptionDefault
listNoDisplayed name of the destination list, case-insensitive: "Today", "WIP", or a list the user created. Omitted, it goes to the first original list. An unknown name with create_list false is refused, and the answer gives the names that exist.
colorNoHex colour (#rrggbb) for a list being created. Optional; the host picks the next swatch otherwise.
tasksYesThe tasks, in execution order. Each item is a string, or {"text": string, "group": string, "description": string} where group is the STEP the task belongs to (e.g. "Step 1 · Mockup"); tasks with the same group are shown under one header in the list. One concrete, actionable line each, 300 characters max; longer detail goes in description (4000 max). A longer text is filed shortened and the answer says so.
create_listNoCreate the list named in "list" when it does not exist. Default false, so an agent does not invent lists in someone's backlog without saying so.

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only state readOnlyHint=false and destructiveHint=false; the description adds substantial context beyond that: the write is routed through the widget's own store so the user sees exactly what is filed, dev-install headless daemon behavior vs npm installs needing the app running, and the 300/4000 character truncation outcome with an explicit answer note.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with purpose then the list-naming rule, with the operational caveats (daemon/npm, store routing) placed after the actionable instruction. Dense and mostly waste-free, though the install-mode sentence is somewhat tangential to invoking the tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation tool with no output schema, it covers the full picture an agent needs: required tasks format, list resolution, creation semantics, truncation feedback, and environmental preconditions. Nothing material is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds real meaning: the case-insensitive fallback to the first original list, why create_list defaults to false ('so an agent does not invent lists in someone's backlog without saying so'), and that group maps to a displayed step header. Only the color parameter is left entirely to the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Put tasks into the user's To-do backlog, the To-do widget on their canvas') with scope beyond the name, including ordering and optional steps. It distinguishes itself from the read/update siblings it names explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives the trigger condition ('Use it when a conversation has settled what to do'), a concrete list-naming procedure including the LIST_NOT_FOUND discovery call and create_list: true, and routes the agent to mnemosyne_todo_list/_update/_categories for other operations. This is about as complete as guidance gets.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

mnemosyne_todo_categoriesA
Read-only

Manage the lists of the user's To-do backlog: create one, rename or recolour one, remove an empty one. It is separate from mnemosyne_todo_update because these change the shape of a workspace rather than the work in it. Two refusals are worth knowing before you call. A list that still holds tasks is never removed, and you are told how many are in the way: move them first, since the app will not pick a destination on someone's behalf. The three original lists can be renamed, never removed, because their contents are what make the file readable. Works with the app closed on a dev install (the headless daemon reads the file); an npm install has no daemon and needs the app running. Scope todo:write.

ParametersJSON Schema
NameRequiredDescriptionDefault
opsYesThe changes, applied in order - so you can create a list and rename another in one call.

TDQS

A3.9/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotation Contradiction: annotations declare readOnlyHint=true, yet the description describes mutating operations ('create one, rename or recolour one, remove an empty one') and explicitly states 'Scope todo:write'. The title 'List To-do lists' also implies a read-only listing while the description describes workspace mutation. Because the description's semantic content directly contradicts the annotation's safety signal, the score is forced to 1 despite the description itself being behaviorally rich.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with purpose and every sentence carries load-bearing information: operation set, sibling discriminator, two refusal conditions, special status of original lists, runtime requirement, and scope. It is dense rather than bloated, though it reads as one long paragraph — a bullet list of the refusals would aid scannability. No sentence is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutating tool with one nested parameter, no output schema, and minimal annotations, the description covers the operations, the sibling alternative, both refusal conditions, the immutable original lists, runtime prerequisites, and required scope. The only substantive gap is the readOnlyHint contradiction, which muddies the reliability of the overall signal; otherwise the description is thorough enough for an agent to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the ops array and its fields are already fully documented in the schema (baseline 3). The description adds meaning beyond the schema by explaining refusal behavior at actual call time: the agent is told how many tasks block a removal, must move them first, and the app will not choose a destination on the user's behalf — operational detail the schema lacks. This added value justifies a 4 rather than a 3.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb+resource ('Manage the lists of the user's To-do backlog') and enumerates the exact operation set: create, rename or recolour, remove an empty one. It also distinguishes itself from the most confusable sibling (mnemosyne_todo_update) by naming it directly and explaining the difference: shape of a workspace vs. work in it. An agent can select this tool correctly without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The sibling routing is explicit: mnemosyne_todo_update is named as the alternative, with the selection rule being whether the change alters workspace shape or work content. The two refusals serve as when-not guidance ('A list that still holds tasks is never removed', 'The three original lists can be renamed, never removed'), and the runtime caveat (daemon on dev install vs. app running on npm install) states a prerequisite condition clearly.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

mnemosyne_todo_listA
Read-only

Read the user's To-do backlog back: the lists that exist and the tasks in them, each with the id you need to change it. Call this before mnemosyne_todo_update, which names tasks by id. A phrase like "delete the task about the invoice" reads perfectly and can still match the wrong task. This is also the way to answer "what is on my plate", or to check whether something is already filed before you add a duplicate. Works with the app closed on a dev install (the headless daemon reads the file); an npm install has no daemon and needs the app running. Scope todo:read.

ParametersJSON Schema
NameRequiredDescriptionDefault
listNoOnly this list, by displayed name (case-insensitive) or by the key shown in brackets. Omitted = every list. A name that matches nothing returns NO tasks rather than silently widening to all of them.
limitNoMaximum tasks returned (default and ceiling: 300). The answer always says how many matched and how many were cut.
include_doneNoInclude tasks already checked off. Default false: the open ones are what almost every question is actually about. Tasks moved to the ARCHIVE are never listed, only counted.

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already provide readOnlyHint=true, so the description is not required to restate that. It adds valuable behavioral context: the difference between dev install (headless daemon reads the file) and npm install (needs app running), and that tasks moved to ARCHIVE are never listed. This goes beyond the annotation and materially affects how an agent plans a call.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with the core purpose and the critical note about IDs, then covers usage and environment context. It is a bit long but every sentence earns its place (the environment note, the scope line, the archive behavior). While it could be tightened, it remains well-structured and readable.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's read-only nature, optional parameters, and lack of output schema, the description covers all the essential operational details: the return includes IDs needed for updates, the duplicate-check use case, the behavior when no list matches, the archive counting behavior, and the environment prerequisites. Nothing critical is missing for an agent to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already explains all three parameters (list, limit, include_done) in detail, including edge cases like omitted list returning every list and unmatched names returning no tasks. The description adds no additional parameter semantics beyond what the schema provides, so the baseline of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Read the user's To-do backlog back' and immediately explains what it returns (lists, tasks, and the IDs needed for changes). It explicitly contrasts itself with mnemosyne_todo_update, which names tasks by id, so an agent can distinguish them without inspecting schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives explicit when-to-use instructions ('Call this before mnemosyne_todo_update') and also positions it as the way to answer 'what is on my plate' or check for duplicates before adding. It mentions the alternative (todo_update) but does not explicitly state when not to use it, so it falls just short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

mnemosyne_todo_updateA
Destructive

Change the user's To-do backlog: edit a task, tick it off, move it to another list, or take it out. Every operation names a task by the id from mnemosyne_todo_list, so call that first. Removing a task archives it by default, and it can be restored. Pass permanent: true only when the user asked for it to be deleted outright. The whole batch is applied in order as one save, and each operation reports its own outcome, so a stale id does not sink the ones around it. Works with the app closed on a dev install (the headless daemon reads the file); an npm install has no daemon and needs the app running. Scope todo:write.

ParametersJSON Schema
NameRequiredDescriptionDefault
opsYesThe changes, applied in order. Up to 100 per call.

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the annotations (readOnlyHint=false, destructiveHint=true), the description discloses that removal archives by default, that archived tasks can be restored, that the batch is applied in order as one save with per-operation outcome reporting, that a stale id won't sink the rest, and that daemon availability depends on the install type. This is far more transparency than the annotations alone provide.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Every sentence earns its place: purpose, prerequisite, archive behavior, permanent-delete guardrail, batch semantics, failure isolation, environment note, and scope. The most decision-relevant information is front-loaded, and there is no filler or repetition of schema text.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is complex, has no output schema, and carries a destructive annotation, yet the description covers invocation prerequisites, side effects, ordering, failure behavior, and environment constraints. The only gap is that the exact response shape is not specified, though 'each operation reports its own outcome' provides a useful partial contract.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema already covers 100% of parameters with detailed descriptions, so the baseline is 3. The tool description adds meaningful operation-level semantics: ids come from mnemosyne_todo_list, batch order matters, stale ids are isolated, and permanent deletion requires explicit user intent. This compensates above baseline without duplicating the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a clear verb and resource: 'Change the user's To-do backlog' and then enumerates the concrete operations (edit, complete, move, remove). It also differentiates from sibling tools by requiring task ids from mnemosyne_todo_listable, making the scope unmistakable.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives explicit and practically important guidance: call mnemosyne_todo_list first, use permanent: true only when the human explicitly asks for outright deletion, and note the environment-dependent daemon behavior. It does not explicitly name an alternative tool such as mnemosyne_todo_add for the 'when not to use' case, so it stops just short of full exclusion guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

mnemosyne_vault_listA
Read-only

List the memory vaults this Mnemosyne OS exposes, each with its token, display name and memory count. Call it first when you are unsure which vault to work against, or when the user names a store you have not seen. Pass the bold token as the vault argument of the other tools. The id line is a local path and the other tools do not take it. Vaults this server was not configured for are flagged here and refused until the user adds them.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the readOnlyHint annotation, the description adds useful behavioral details: vaults have a local-path `id` that other tools reject, unconfigured vaults are flagged and refused until the user adds them, and the response includes token/display-name/memory-count. This prevents the agent from misusing the wrong identifier.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Every sentence earns its place: it names the resource, tells when to call it, explains how to use the output, warns about the non-useful `id` field, and notes unconfigured-vault behavior. The description is front-loaded and free of filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Even without an output schema, the description explains what the response will contain, how to use that response, and the key caveat about unconfigured vaults. For a simple read-only list tool with no parameters, this is complete enough for an agent to call it correctly and act on the results.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, so the description does not need to explain parameter meaning. The baseline of 4 applies here, and the description even adds downstream guidance about how the returned token is used as a parameter in other tools.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('List') and resource ('memory vaults') and states what each listed vault contains: token, display name, and memory count. It clearly distinguishes this tool from sibling list tools like mnemosyne_agent_list or mnemosyne_todo_list by targeting vaults specifically.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives explicit conditions for use: call it first when unsure which vault to work against, or when the user names a store not yet seen. It also tells the agent how to consume the result by passing the token as the `vault` argument to other tools, and clarifies that the `id` line should not be passed.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 1 tool updatev1.7.0-infinity
    • Changedmnemosyne_todo_add2 fields changed
      • changedInput schema / properties / tasks / description
        Previous value: -"The tasks, in execution order. Each item is a string, or {\"text\": string, \"group\": string} where group is the STEP the task belongs to (e.g. \"Step 1 · Mockup\"); tasks with the same group are shown under one header in the list. One concrete, actionable line each; 300 characters max."New value: +"The tasks, in execution order. Each item is a string, or {\"text\": string, \"group\": string, \"description\": string} where group is the STEP the task belongs to (e.g. \"Step 1 · Mockup\"); tasks with the same group are shown under one header in the list. One concrete, actionable line each, 300 characters max; longer detail goes in description (4000 max). A longer text is filed shortened and the answer says so."
      • changedInput schema / properties / tasks / items / anyOf
        Previous value: -[
        -  {
        -    "type": "string"
        -  },
        -  {
        -    "properties": {
        -      "group": {
        -        "type": "string"
        -      },
        -      "text": {
        -        "type": "string"
        -      }
        -    },
        -    "required": [
        -      "text"
        -    ],
        -    "type": "object"
        -  }
        -]New value: +[
        +  {
        +    "type": "string"
        +  },
        +  {
        +    "properties": {
        +      "description": {
        +        "type": "string"
        +      },
        +      "group": {
        +        "type": "string"
        +      },
        +      "text": {
        +        "type": "string"
        +      }
        +    },
        +    "required": [
        +      "text"
        +    ],
        +    "type": "object"
        +  }
        +]
  2. 7 tool updatesv1.6.0-infinity
    • Changedmnemosyne_cockpit_update1 field changed
      • addedInput schema / properties / desktop
        Added value: +{
        +  "description": "The desktop this session should put its cards on, by name, exactly as the human wrote it. Send it once when they tell you; the card stays there for the rest of the conversation without you repeating it. The match is exact apart from case, and an unknown or duplicated name is reported back with the names that exist rather than guessed at. Leave it out and the card goes to whichever desktop answers for this project, which is what most sessions want.",
        +  "type": "string"
        +}
    • Changedmnemosyne_memory_ask2 fields changed
      • changedInput schema / properties / question / description
        Previous value: -"A natural-language question, as you would ask a knowledgeable colleague. Be specific."New value: +"A question in plain language, the way you would ask a colleague who knows the material. Be specific."
      • changedInput schema / properties / vault / description
        Previous value: -"Vault to reason over (case-insensitive). Default for this deployment: \"DEV\". A vault not declared to this MCP is refused with SCOPE_DENIED."New value: +"Vault to reason over, case-insensitive. Default for this deployment: \"DEV\". A vault this server was not configured for is refused with SCOPE_DENIED."
    • Changedmnemosyne_memory_ingest3 fields changed
      • changedInput schema / properties / content / description
        Previous value: -"Content to persist (markdown supported). Be self-contained: include WHY the decision was made, not just WHAT."New value: +"The content to store. Markdown works. Write it self-contained, and say WHY the decision was made rather than only what was decided."
      • changedInput schema / properties / spine_type / description
        Previous value: -"Semantic type of the content. ARCHITECTURE is heavily boosted (×1.40) in SOURCE_CODE scope queries. Use it for design docs, big-picture decisions, structural choices. DECISION for narrower trade-offs. BUGFIX/DEBUG for incident learnings. SESSION for \"here is where I left off\". FEATURE for new capabilities. NOTE for everything else."New value: +"Semantic type of the content. ARCHITECTURE is boosted heavily in code-scope queries: use it for design docs, big-picture decisions and structural choices. DECISION for narrower trade-offs. BUGFIX and DEBUG for incident learnings. SESSION for where you left off. FEATURE for new capabilities. NOTE for the rest."
      • changedInput schema / properties / vault / description
        Previous value: -"Target vault TOKEN: the folder name uppercased, spaces and hyphens as underscores (e.g. MNEMOSYNE_OS). The path-shaped `id` from mnemosyne_vault_list is also accepted and normalized. Default for this deployment: \"DEV\". Tokens this MCP is SCOPED for: a config list, not a census, DEV, PERSONAL, SOCIAL. Ingest is PERMANENT, so confirm the vault EXISTS with mnemosyne_vault_list before writing anywhere you have not written before."New value: +"Target vault TOKEN: the folder name uppercased, with spaces and hyphens as underscores. The path-shaped id from mnemosyne_vault_list also works. Default for this deployment: \"DEV\". Scoped tokens, which are configuration rather than a census: DEV, PERSONAL, SOCIAL. Ingest is permanent, so confirm the vault exists with mnemosyne_vault_list before writing anywhere new."
    • Changedmnemosyne_memory_query4 fields changed
      • changedInput schema / properties / max_content_chars / description
        Previous value: -"Per-chronicle content snippet size in chars (default: 600). Each result is truncated to this length with a hint about total size. Raise to 2000+ when you genuinely need full file content, but be aware results stack up against your context window."New value: +"Per-memory snippet size in characters (default 600). Each result is cut to this length, with a hint about its full size. Raise it past 2000 when you need whole files, and watch your context window."
      • changedInput schema / properties / query / description
        Previous value: -"The search query. Be specific. Examples: \"Phase 51 auto-poll implementation\", \"SDK authentication bug\", \"why did we choose dual-vector dimensions\"."New value: +"The search query. Be specific. Examples: \"why we moved off Postgres\", \"the retry logic in the payment worker\", \"notes from the March planning session\"."
      • changedInput schema / properties / spine_type_filter / description
        Previous value: -"Optional whitelist of spineTypes: restricts results to those types only. Use [\"ARCHITECTURE\"] to surface design docs over code, [\"GIT\"] for commit history, [\"BUGFIX\",\"DEBUG\"] for incident knowledge, [\"SOURCE_CODE\"] to force code-only. Without this, all types are returned (the SOURCE_CODE scope weighting decides ranking)."New value: +"Optional whitelist of spine types. Use [\"ARCHITECTURE\"] for design docs, [\"GIT\"] for commit history, [\"BUGFIX\",\"DEBUG\"] for incident knowledge, [\"SOURCE_CODE\"] for code only. Omit it to get every type, with the SOURCE_CODE scope weighting deciding the ranking."
      • changedInput schema / properties / vault / description
        Previous value: -"Vault TOKEN to query (case-insensitive; the folder name uppercased, spaces and hyphens as underscores). The path-shaped `id` from mnemosyne_vault_list is also accepted and normalized. Mnemosyne OS exposes one vault per tracked folder. This deployment's default is \"DEV\". Tokens this MCP is SCOPED for (a config list, not a census, so some may not be mounted on this machine): PERSONAL, SOCIAL. Call mnemosyne_vault_list for the vaults that actually exist. Anything outside the scoped list is refused."New value: +"Vault TOKEN to search, case-insensitive. A token is the folder name uppercased, with spaces and hyphens as underscores. The path-shaped id from mnemosyne_vault_list also works. This deployment's default is \"DEV\". It is also scoped for: PERSONAL, SOCIAL. That list is configuration rather than a census, so call mnemosyne_vault_list for the vaults that exist on this machine. Anything outside the list is refused."
    • Changedmnemosyne_position_get1 field changed
      • changedInput schema / properties / resonance_id / description
        Previous value: -"ID of the resonance (e.g. \"agent-cockpit\", \"mnemosync-p2p\")"New value: +"ID of the resonance, for example \"checkout-rewrite\" or \"v2-migration\"."
    • Changedmnemosyne_spine_assignments1 field changed
      • changedInput schema / properties / vault / description
        Previous value: -"Vault to inspect (case-insensitive). Default for this deployment: \"DEV\"."New value: +"Vault to inspect, case-insensitive. Default for this deployment: \"DEV\"."
    • Changedmnemosyne_todo_add2 fields changed
      • changedInput schema / properties / create_list / description
        Previous value: -"Create the list named in \"list\" when it does not exist. Default false: an agent must not invent lists in someone's backlog without saying so."New value: +"Create the list named in \"list\" when it does not exist. Default false, so an agent does not invent lists in someone's backlog without saying so."
      • changedInput schema / properties / list / description
        Previous value: -"Displayed name of the destination list (case-insensitive), e.g. \"En cours\", \"WIP\", or a list the human created. Omitted = the first original list. Unknown name + create_list false = refused with the names that exist."New value: +"Displayed name of the destination list, case-insensitive: \"Today\", \"WIP\", or a list the user created. Omitted, it goes to the first original list. An unknown name with create_list false is refused, and the answer gives the names that exist."
  3. 19 tool updatesv1.5.1-infinity
    • Changedmnemosyne_agenda_add1 field changed
      • changedInput schema / properties / events / items / properties / start / description
        Previous value: -"ISO 8601 date-time, e.g. \"2026-09-10T14:00:00\". No timezone offset = read as the HUMAN'S OWN machine local time, never UTC — do not add a \"Z\" unless you mean UTC."New value: +"ISO 8601 date-time, e.g. \"2026-09-10T14:00:00\". No timezone offset = read as the HUMAN'S OWN machine local time, never UTC. Do not add a \"Z\" unless you mean UTC."
    • Addedmnemosyne_agent_list
    • Removedmnemosyne_agents
    • Removedmnemosyne_ask
    • Removedmnemosyne_get_position
    • Removedmnemosyne_ingest
    • Addedmnemosyne_memory_ask
    • Addedmnemosyne_memory_ingest
    • Addedmnemosyne_memory_query
    • Addedmnemosyne_position_get
    • Addedmnemosyne_position_update
    • Removedmnemosyne_query
    • Addedmnemosyne_resonance_list
    • Removedmnemosyne_resonances
    • Addedmnemosyne_todo_categories
    • Removedmnemosyne_todo_lists
    • Removedmnemosyne_update_position
    • Addedmnemosyne_vault_list
    • Removedmnemosyne_vaults
  4. 25 tool updatesv1.10.0
    • First observedmnemosyne_about
    • First observedmnemosyne_agenda_add
    • First observedmnemosyne_agenda_list
    • First observedmnemosyne_agenda_remove
    • First observedmnemosyne_agenda_update
    • First observedmnemosyne_agent_collisions
    • First observedmnemosyne_agent_files
    • First observedmnemosyne_agents
    • First observedmnemosyne_ask
    • First observedmnemosyne_cockpit_update
    • First observedmnemosyne_dream_bridges
    • First observedmnemosyne_get_position
    • First observedmnemosyne_git_log
    • First observedmnemosyne_ingest
    • First observedmnemosyne_pheme_radar
    • First observedmnemosyne_pheme_watch
    • First observedmnemosyne_query
    • First observedmnemosyne_resonances
    • First observedmnemosyne_spine_assignments
    • First observedmnemosyne_todo_add
    • First observedmnemosyne_todo_list
    • First observedmnemosyne_todo_lists
    • First observedmnemosyne_todo_update
    • First observedmnemosyne_update_position
    • First observedmnemosyne_vaults

TDQS

A4.2/5.0

Scored across 25 tools

Disambiguation4/5

Domains are cleanly partitioned by prefix (agenda, todo, memory, agent, pheme), and the one genuinely overlapping pair — memory_query vs memory_ask — is explicitly disambiguated (quote the source vs find/re-rank). The main wrinkle is the resonance domain: the entity is read via resonance_list but written/read via position_get and position_update, so the vocabulary shifts mid-domain. Otherwise each tool has a clearly distinct action+resource.

Naming Consistency4/5

Consistent mnemosyne_ prefix with snake_case domain_action naming (agenda_add/list/update/remove, todo_add/list/update, agent_list/collisions/files). Minor deviations: position_* tools stand in for the resonance domain instead of resonance_get/resonance_update, and mnemosyne_about and mnemosyne_git_log carry no resource-action pair. Still readable and predictable overall.

Tool Count4/5

25 tools is at the heavy end, but they distribute across ~10 well-scoped sub-domains (calendar, todo, memory, vault, resonance, agents, Pheme, cockpit, git, diagnostics), and each tool earns its place with no redundant variants. It sits just under the 'too many' threshold but never feels padded.

Completeness4/5

Core lifecycles are complete: calendar and todo both have add/list/update/remove (todo even splits list-shape management into todo_categories), memory has ingest/query/ask, and agents/Pheme/vault have read and write sides. Gaps are minor: no memory update or delete (memories are permanent by design), no vault creation, and git is read-only (log only). An agent can work around all of these.

Maintenance

ActivityActive
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    A
    maintenance
    Local-first, file-based memory layer for AI agents — one shared Markdown vault across Claude, Codex, Gemini, Cursor and any MCP client. Provides read/write memory tools with an audit trail, per-agent trust levels, and Git sync; no cloud and no lock-in.
    2
    MIT
  • A
    license
    Not graded
    quality
    A
    maintenance
    Local-first, source-grounded memory for AI agents, with citations, bitemporal history, review-gated corrections, and MCP tools for search and recall.
    40 PyPI
    3
    Apache 2.0
  • A
    license
    Not graded
    quality
    C
    maintenance
    Provides a local-first, source-cited memory layer for AI agents, with MCP tools to search, read, explain sources, and propose/apply memory updates.
    70 npm
    13
    Apache 2.0