Skip to main content
Glama

Pathmark

Carry intent across agents without turning stale code facts into hidden memory.

What's New — v0.1.17

Pathmark v0.1.17 stops Claude Code from gating every memory lookup:

  • pathmark setup claude-code --apply-permissions adds allow rules for Pathmark's ten local read-only tools, so recall and search run without a permission prompt, and in auto mode without a safety-classifier round-trip that can fail transiently;

  • it backs up ~/.claude/settings.json, adds only missing rules, keeps every other setting, and never rewrites a file it cannot parse;

  • the allow list is tested against the tools' own readOnlyHint/openWorldHint annotations, so anything that writes, deletes, or may call an external model always keeps its prompt.

See the v0.1.17 release notes or the complete changelog. The npm badge above always shows the currently published version.

Pathmark gives Codex, Claude Code, opencode, Gemini CLI, Cursor, and any MCP-capable harness one local intent and provenance layer. Save decisions, constraints, preferences, and approved conclusions once. Use them from the next agent without pasting a recap.

Code remembers implementation. Pathmark remembers intent. Repository code, architecture, tests, CI, and intentional agent instructions remain authoritative for how the software works. Raw sessions are searchable evidence, not automatically trusted truth.

Your context stays on disk at ~/.pathmark/memory/memory.jsonl. You do not need an account, hosted database, API key, or vendor backend to start.

Related MCP server: auxly-memory-cli

OpenAI Build Week 2026

Pathmark is a Developer Tools submission for OpenAI Build Week 2026. The project existed before the challenge, so the submission is deliberately scoped to the meaningful extension built after the submission period opened on July 13, 2026.

During the eligible period, Codex with GPT-5.6 helped audit and extend Pathmark from a working local memory layer into safer long-running developer infrastructure:

  • fixed a reproduced multi-process SQLite index race;

  • added revision history, superseding, expiration, retention, diagnostics, backup, compaction, and preview-first hard purge;

  • added namespace-scoped reads and writes plus default secret redaction;

  • added scoped import/export, optional AES-256-GCM portable exports, local hybrid reranking, and portable harness ingestion;

  • hardened CI and npm delivery with required CodeQL and dependency review, immutable Action pins, protected tags, OpenSSF analysis, and SLSA provenance.

The primary Codex session for this work is 019f5fc3-d7e6-7b41-8a30-d161c90b98fb. The qualifying release range is v0.1.6 through v0.1.7; the pre-challenge baseline is commit 4c0e87dfdbd2ba4c643abd8b887cc228bdb08b73.

See the Build Week implementation record for the before/after boundary, commit evidence, Codex collaboration details, and a fast judge test.

Why Pathmark

You do not work in one tool. You ask Codex to patch, Claude Code to review, opencode to clean up, and Gemini CLI to challenge the plan. Each tool starts cold unless you carry the context across.

Pathmark gives those tools one place to read and write intent and evidence:

  • One local JSONL store across harnesses.

  • Standard MCP tools include remember, search_memory, recall_memory, session_trace, rate_recall, consolidate_memory, audit_memory, and conclusion-first chat / ask_memory.

  • Client-side synthesis by default, so your coding agent reads the context and answers.

  • Optional Codex CLI, local command, and OpenAI-compatible synthesis modes.

  • Plain files you can inspect, back up, delete, or migrate.

Pathmark stays provider-neutral. Codex gets one optional synthesis preset. The core server works with any MCP client that can use local tools.

Pathmark requires Node.js 22.5 or newer.

Cross-Harness Memory

You switch tools during a coding session:

  • Codex fixes the failing test.

  • Claude Code reviews the patch.

  • opencode cleans the diff.

  • Gemini CLI challenges the approach.

Pathmark keeps the notes in one store.

Point each harness at the same store:

Codex       \
Claude Code \
opencode     >  Pathmark MCP  >  ~/.pathmark/memory/memory.jsonl
Gemini CLI  /
Cursor     /

Install Pathmark in each harness and point them at the same PATHMARK_STORE_DIR. One tool saves raw context with remember or proposes a durable conclusion with create_conclusion; an approved conclusion and raw evidence can then be recovered with recall_memory, search_memory, get_context, or ask_memory.

Pathmark sits below the agents as an intent, evidence, and provenance bus for your coding workflow.

Tools

Pathmark exposes these MCP tools:

Tool

Purpose

remember

Save raw searchable evidence. Raw evidence is not treated as durable approved intent.

create_conclusion

Propose a higher-signal durable conclusion or preference. Approval is required by default before recall.

search_memory

Search memories and conclusions.

recall_memory

Transparent recall: returns context plus the exact memory IDs, timestamps, sources, matches, tags, and previews used. Accepts optional tags, exact ids, and compact includeRecords: false output.

session_trace

Return a bounded chronological audit trail for one session: prompts, exact injected memory IDs, redacted tool inputs/results, and answers.

rate_recall

Label exact IDs from a chat / ask_memory recall as relevant or irrelevant so audit precision is measured.

get_context

Return compact context for a task or question.

list_conclusions

List approved saved conclusions.

list_pending_conclusions

Review bounded, paginated pending conclusion proposals.

approve_conclusion

Atomically approve a proposal, optionally correcting text/tags and recording the reviewer.

reject_conclusion

Retain a rejected proposal in the audit trail while permanently excluding it from recall.

get_memory_snapshot

Generate a bounded USER/PROJECT/AGENT snapshot from approved canonical conclusions.

consolidate_memory

Review a bounded unsynthesized evidence batch and optionally stage evidence-backed proposals. Nothing is auto-approved.

delete_memory

Soft-delete a memory or conclusion by id.

update_memory

Correct a record while preserving prior versions.

supersede_memory

Replace an outdated record with a linked current record.

purge_memory

Preview or apply permanent deletion by id, namespace, tags, source, or date.

audit_memory

Measure capture-to-recall behavior, unused records, recall age, duplicates, stale raw hits, and whether precision labels exist.

doctor_memory

Report duplicates, deleted/expired records, conclusions, and index health.

compact_memory

Preview or apply deduplication, retention, and physical cleanup with an automatic backup.

backup_memory

Create a point-in-time canonical JSONL backup.

export_memory

Export a scoped mergeable JSONL bundle, optionally encrypted.

ask_memory

Return an approved-conclusion answer or scoped raw context, exact provenance, and a recall ID for feedback.

chat

Chat-compatible alias for ask_memory, including multi-intent conclusion retrieval and explicit abstention.

get_config

Show local store configuration.

Quick Start

npm install -g pathmark

Then add the MCP server to your client.

Prefer npm for normal installs. To test the current GitHub main branch directly:

npm install -g --install-links=true github:hacksurvivor/pathmark

Generate a setup snippet for your harness:

pathmark setup list
pathmark setup claude-code
pathmark setup opencode --json
pathmark setup gemini-cli
pathmark setup kimi

See docs/compatibility.md for Codex, Claude Code, opencode, Gemini CLI, OpenClaw, Hermes Agent, Grok CLI, Kimi, GLM, and generic MCP setups.

Codex

codex mcp add pathmark -- pathmark

Codex users can also enable auto-capture:

pathmark codex install --replace-legacy-hooks

When you want the visible "what memory did you use?" entry in Codex, Claude Code, Cursor, opencode, Gemini CLI, Grok-compatible MCP hosts, or any other MCP harness, call the recall_memory tool before answering. Codex session start injects an approved conclusion snapshot. Before non-trivial prompts, Codex recalls approved conclusions first and uses fresh scoped raw evidence only as a bounded fallback. recall_memory remains the portable visible trace across harnesses.

Claude Code

claude mcp add --scope user pathmark -- pathmark
pathmark setup claude-code --apply-permissions

The second command lets Pathmark's local read-only tools (recall, search, diagnostics) run without a permission prompt, including in auto mode, where they would otherwise wait on the safety classifier. It backs up ~/.claude/settings.json, adds only the missing allow rules, and leaves everything else untouched. Tools that write or delete memory still ask.

--scope user makes Pathmark available in every project; Claude Code's default scope only covers the current directory. For automatic capture and session-start recall (including after context compaction), merge the hooks block from pathmark setup claude-code into ~/.claude/settings.json.

Claude Code's built-in auto-memory lives in per-project folders the other agents cannot see. Bring it into the shared store as evidence (re-runnable; edits update in place with history, and records you delete in Pathmark stay deleted):

pathmark import-native claude-code --dry-run
pathmark import-native claude-code

opencode / Gemini CLI

Use the generated snippets:

pathmark setup opencode
pathmark setup gemini-cli

Claude Desktop

Add this to your Claude Desktop MCP config:

{
  "mcpServers": {
    "pathmark": {
      "command": "pathmark",
      "env": {
        "PATHMARK_STORE_DIR": "~/.pathmark/memory"
      }
    }
  }
}

Cursor

Add the same command to Cursor's MCP server settings:

{
  "mcpServers": {
    "pathmark": {
      "command": "pathmark"
    }
  }
}

Local Development

npm install
npm test
npm run coverage

Run directly:

PATHMARK_STORE_DIR=.pathmark npm run dev

Import Legacy Memory

Pathmark can import a compatible local JSONL memory store without deleting or moving the source files.

npm run import:legacy -- --source-dir ~/old-codex-memory

Defaults:

Legacy source:   ~/.pathmark/legacy/codex
Pathmark target: ~/.pathmark/memory/memory.jsonl

The importer creates a memory.jsonl.backup-* file before writing, uses deterministic ids so reruns skip duplicates, and redacts obvious KEY=..., TOKEN=..., PASSWORD=..., and Bearer ... values. It uses the same store lock as live MCP and Codex writers, so an import cannot overwrite records captured concurrently.

Use a dry run first when migrating another machine:

npm run import:legacy -- --source-dir ~/old-codex-memory --dry-run

Codex Auto-Capture

Install Pathmark as the Codex memory adapter:

pathmark codex install --replace-legacy-hooks

This registers the Pathmark MCP server, enables Codex hooks, and removes old compatible hook commands from Codex. It does not delete or move memory files.

The Codex adapter is proactive by default:

  • user prompts, final assistant answers, and tool activity are captured locally; intermediate Codex commentary is excluded;

  • tool activity records include bounded redacted input previews and hashes, status, exit code, duration, and changed files when the hook provides them;

  • tool-output hashes are captured for correlation, while output text remains private by default and requires PATHMARK_CODEX_CAPTURE_TOOL_OUTPUTS=on;

  • activity records expire after 30 days and are physically capped at 5,000 records by default;

  • session start/resume generates one bounded USER/PROJECT/AGENT snapshot from approved canonical conclusions and does not inject raw session history;

  • each non-trivial user prompt searches approved workspace/project conclusions first, then approved global or explicitly named-project conclusions;

  • only when no approved conclusion matches, at most two raw records from the current workspace/project/session may be injected as a high-confidence fallback;

  • raw fallback records must be within the separate automatic-recall horizon, 30 days by default; the full raw archive remains available to explicit search_memory and recall_memory calls;

  • raw cross-project history is never injected automatically; promote durable cross-project intent through the approval workflow;

  • broad cross-project history remains available through explicit search_memory / recall_memory calls without silently entering every prompt;

  • matching memory is injected quietly by default, without adding a raw recall_memory tool result to the conversation;

  • set PATHMARK_CODEX_VISIBLE_RECALL=on when debugging or auditing to make Codex call recall_memory with the exact pre-capture result IDs and workspace tag; this explicit mode omits the redundant full records copy;

  • legacy transport envelopes and assistant progress updates are excluded from session-start and proactive relevance results, while realtime delegation envelopes retain only their current <input> payload;

  • records containing Pathmark boundary escapes, instruction-override patterns, or invisible Unicode controls are tagged memory-quarantined and excluded from automatic recall; injected previews are escaped and explicitly treated as untrusted historical data;

  • no matching memory means no extra context is injected.

Durable extraction is approval-gated by default. create_conclusion creates a pending proposal; pending and rejected conclusions stay in the canonical JSONL audit trail but are structurally excluded from normal search, exact-ID recall, prompt injection, and snapshots. Use list_pending_conclusions, then approve_conclusion or reject_conclusion. Conclusions created before this workflow are treated as already approved for backward compatibility. Raw remember records remain searchable evidence and are not promoted automatically.

The session snapshot is generated from the same canonical store rather than maintained as a second flat file. It is frozen in the session-start hook output; prompt-time scoped recall remains dynamic.

Set PATHMARK_CODEX_PROACTIVE_RECALL=off if you want Codex hooks to capture memory but stop prompt-time recall. Set PATHMARK_CODEX_VISIBLE_RECALL=on only when you want an explicit audit/debug recall_memory tool call. Proactive prompt-time recall remains active when this is off.

recall_memory is a point-in-time record of memory used before an answer. It intentionally does not include commands that run later. Use session_trace with the exact session ID to inspect the chronological prompt → injected memories → tools/results → final-answer trail. When explicitly enabled, output previews are capped at 2,000 characters and redacted before storage; hashes preserve correlation without storing output text by default. Upgrading an existing cursor migrates to final-answer-only parsing without duplicating previously captured user or final-answer turns. When the original transcript is available, exact legacy phase: "commentary" records are soft-deleted by timestamp and text during that migration.

Use --replace-legacy-hooks when you want Pathmark hooks to take over from earlier compatible hook commands. Without it, Pathmark installs alongside existing hook commands.

Check the adapter status:

pathmark codex status

The status output is JSON and includes Pathmark hook state, MCP registration state, legacy hook presence, the active store paths, and the current record count.

Remove Pathmark hooks and MCP registration without deleting memory:

pathmark codex uninstall

Configuration

Variable

Default

Description

PATHMARK_STORE_DIR

~/.pathmark/memory

Directory for memory.jsonl.

PATHMARK_MAX_SEARCH_RESULTS

12

Default search limit.

PATHMARK_CODEX_PROACTIVE_RECALL

on

Automatically inject relevant Pathmark context before non-trivial Codex prompts. Use off to capture without prompt-time recall.

PATHMARK_CODEX_VISIBLE_RECALL

off

Opt into an audit/debug recall_memory tool call that exposes exact usedMemories. Proactive memory injection remains active when this is off.

PATHMARK_CODEX_CAPTURE_TOOL_OUTPUTS

off

Store bounded redacted tool-output previews. Output hashes, status, duration, and exit codes remain available when this is off.

PATHMARK_CODEX_MEMORY_SNAPSHOT

on

Generate a bounded approved-conclusion snapshot at Codex session start/resume.

PATHMARK_SNAPSHOT_CHARS

4000

Character budget for generated snapshots; clamped to 500–12000.

PATHMARK_CODEX_RAW_RECALL_DAYS

30

Prompt-time eligibility horizon for raw evidence. 0 disables automatic raw fallback without deleting or hiding explicit search results.

PATHMARK_CODEX_RAW_RECALL_LIMIT

2

Maximum fresh raw records injected when no approved conclusion matches. Clamped to 0–2.

PATHMARK_CODEX_PROACTIVE_CONSOLIDATION

on

At session start, nudge Codex to review a bounded evidence batch when scoped raw history is accumulating without conclusions.

PATHMARK_CONSOLIDATION_MIN_EVIDENCE

8

Minimum unsynthesized user/assistant records before the proactive consolidation nudge appears.

PATHMARK_CONCLUSION_APPROVAL

on

Stage new conclusions for explicit approval. Set off only for trusted legacy automation.

PATHMARK_SYNTHESIS_PROVIDER

client

client, command, codex, or openai-compatible.

PATHMARK_CHAT_COMMAND

unset

Command provider: receives a synthesized prompt on stdin and writes an answer on stdout.

PATHMARK_CODEX_COMMAND

codex

Codex provider command.

PATHMARK_CODEX_MODEL

unset

Optional Codex model override.

PATHMARK_OPENAI_BASE_URL

https://api.openai.com/v1

OpenAI-compatible API base URL.

PATHMARK_OPENAI_API_KEY

unset

OpenAI-compatible API key.

PATHMARK_OPENAI_MODEL

unset

Model id for OpenAI-compatible synthesis.

PATHMARK_CHAT_TIMEOUT_MS

120000

Synthesis command timeout.

PATHMARK_NAMESPACE

unset

Default namespace applied consistently to MCP reads and writes.

PATHMARK_REDACT_MCP_WRITES

on

Redact common secret-shaped values on remember, conclusion, update, supersede, import, and ingest paths.

PATHMARK_RETENTION_DAYS

0

Retention policy used by compaction; 0 disables age-based removal. Conclusions are retained.

PATHMARK_ACTIVITY_RETENTION_DAYS

30

Automatic lifetime for recall/tool activity records; 0 disables age-based activity expiry.

PATHMARK_ACTIVITY_MAX_RECORDS

5000

Physical cap for activity records; oldest activity is removed automatically. 0 disables the count cap.

PATHMARK_RERANK_COMMAND

unset

Optional trusted local embedding/vector or hybrid reranker. Strict kind/tag/namespace filters are applied before candidates leave the store; the command receives query/candidates as JSON on stdin and returns ranked memory ids.

PATHMARK_HYBRID_CANDIDATES

500

Maximum candidates sent to the optional reranker.

PATHMARK_RETRIEVAL_TIMEOUT_MS

30000

Timeout for the optional reranker.

PATHMARK_EXPORT_KEY

unset

Passphrase for AES-256-GCM portable exports/imports. Never returned by get_config.

PATHMARK_INDEX_LOCK_TIMEOUT_MS

120000

Cross-process wait limit for index initialization or rebuild.

Synthesis Modes

Pathmark separates memory from reasoning.

client

Default. The MCP server returns relevant memory context, and your MCP client model synthesizes the answer. This works across Codex, Claude Desktop, Cursor, and any other MCP client without giving Pathmark a model credential.

PATHMARK_SYNTHESIS_PROVIDER=client pathmark

command

Use any local subscription or model CLI that accepts a prompt on stdin and writes an answer to stdout:

PATHMARK_SYNTHESIS_PROVIDER=command \
PATHMARK_CHAT_COMMAND="your-ai-cli --model your-model" \
pathmark

This is the general path for users with another paid subscription CLI or a local model runner.

codex

Use the proven Codex CLI bridge. It runs a controlled, non-interactive codex exec turn with hooks and memories disabled to avoid recursion:

PATHMARK_SYNTHESIS_PROVIDER=codex \
PATHMARK_CODEX_MODEL=gpt-5.5 \
pathmark

This is useful for Codex users who have persisted ChatGPT/Codex CLI auth locally but do not want to add an OpenAI API key. Pathmark sends the synthesis prompt through stdin, runs Codex in an empty temporary workspace, ignores project rules, and exposes only a minimal environment. Memory records are treated as untrusted data rather than executable instructions.

openai-compatible

Use any provider that exposes /chat/completions, including many Kimi, GLM/Z.ai, OpenRouter, LiteLLM, Ollama-compatible gateways, and self-hosted routers:

PATHMARK_SYNTHESIS_PROVIDER=openai-compatible \
PATHMARK_OPENAI_BASE_URL=https://api.provider.example/v1 \
PATHMARK_OPENAI_API_KEY=... \
PATHMARK_OPENAI_MODEL=... \
pathmark

This mode affects MCP ask_memory / chat and CLI pathmark chat. Regular MCP tools still store and retrieve local memory without a model provider.

Setup CLI

pathmark setup <client> prints copy-paste setup for common harnesses. Add --json when you want structured output for scripts.

Supported targets:

codex
claude-code
claude-desktop
cursor
opencode
gemini-cli
generic
openai-compatible
command

Aliases include claude, gemini, kimi, glm, and z-ai.

Gemini CLI setup includes portable SessionStart, BeforeAgent, AfterTool, and AfterAgent hooks for automatic scoped recall and capture. Other harnesses can feed exported transcripts through the generic ingestion surface:

pathmark ingest --client=claude-code --namespace=my-project < transcript.json
pathmark ingest --client=opencode --namespace=my-project < transcript.json

Memory chat, consolidation, maintenance, and portable sync

Maintenance commands preview destructive changes unless --apply is present:

pathmark chat "What did we decide about release signing?" --namespace=my-project
pathmark consolidate --namespace=my-project
pathmark consolidate --namespace=my-project --cursor=LAST_RECORD_ID
PATHMARK_SYNTHESIS_PROVIDER=codex pathmark consolidate --namespace=my-project --apply
pathmark feedback --recall-id=RECALL_ID --relevant=MEMORY_ID --irrelevant=OTHER_ID
pathmark audit --days=30
pathmark audit --namespace=my-project --days=90
pathmark doctor
pathmark compact
pathmark compact --apply --retention-days=90
pathmark purge --namespace=old-client
pathmark purge --namespace=old-client --apply

pathmark chat and the MCP chat / ask_memory tools search approved conclusions first. Multi-intent questions can return separate conclusions for separate clauses. In default client mode, approved conclusions produce a safe extractive answer; a configured codex, command, or openai-compatible provider can synthesize richer prose. Raw fallback requires an explicit scope (--namespace / tags) or kind: memory, preventing unscoped cross-workspace history from entering chat.

Every matched chat query records recall activity and returns a recallId when the store is writable. Abstentions and read-only stores return recallId: null. Use MCP rate_recall or pathmark feedback with exact recalled IDs to label relevance. pathmark audit reports precision.status: "labeled", measured precision, and label coverage once feedback exists.

pathmark consolidate is preview-first. With default client synthesis it returns the bounded evidence and exact instructions for the host agent, which can call create_conclusion with supporting evidenceIds. When more eligible evidence remains, the result includes nextCursor and remainingAfterBatch; pass the cursor to review the next stable page. With a configured server-side synthesis provider, it previews structured candidates; --apply stages them as pending conclusions for approve_conclusion or reject_conclusion. It never auto-approves extracted intent.

Applied compaction and purge create a backup before replacing the canonical file. Soft deletion remains available through delete_memory; hard purge physically removes selected records from JSONL and rebuilds the derived index.

pathmark audit is read-only. It separates all raw records from consolidation-eligible user/assistant evidence, then reports capture-to-recall ratio, actionable synthesis backlog, recall age, exact duplicates, stale raw hits, and scope/missing-reference signals. Precision remains unlabeled until explicit feedback exists; Pathmark never substitutes a heuristic for a user label.

Use scoped exports and merge imports as the transport-neutral sync layer:

pathmark export --namespace=my-project --output=project.jsonl
pathmark import project.jsonl --namespace=my-project

For an encrypted portable bundle, configure the passphrase outside the command line:

PATHMARK_EXPORT_KEY='use-a-secret-manager' pathmark export --encrypted --output=project.pathmark
PATHMARK_EXPORT_KEY='use-a-secret-manager' pathmark import project.pathmark

Pathmark does not silently upload these files. Move them through a trusted filesystem, backup tool, or sync provider of your choice.

Optional hybrid retrieval

Default retrieval stays local SQLite FTS. To enable semantic or embedding-backed reranking without forcing a model dependency, set PATHMARK_RERANK_COMMAND to a trusted local command. It receives one JSON object on stdin containing query and candidates, and must return a JSON array of ranked record ids (or { "ids": [...] }). If it fails or times out, Pathmark falls back to lexical results.

Data Format

Pathmark stores newline-delimited JSON at:

~/.pathmark/memory/memory.jsonl

memory.jsonl remains the canonical source of truth. Pathmark also maintains a derived, disposable search index at memory.index.v5.sqlite. Index filenames are schema-versioned so old and new MCP processes can coexist during a rolling restart. The index is rebuilt automatically when the JSONL file changes outside Pathmark, and inactive index versions can be deleted safely after their processes stop.

Each record is inspectable:

{
  "id": "uuid",
  "kind": "memory",
  "text": "The user prefers local-first tools.",
  "tags": ["preference"],
  "source": "mcp",
  "createdAt": "2026-06-29T00:00:00.000Z",
  "updatedAt": "2026-06-29T00:00:00.000Z"
}

Deletes are soft deletes by default: the record gets a deletedAt timestamp. Use preview-first purge_memory or pathmark purge --apply for physical erasure. Updates preserve up to 50 prior versions, superseded records link to their replacement, and expired records are excluded from recall. The raw automatic-recall horizon is separate from storage retention: evidence can age out of proactive injection while remaining explicitly searchable.

Malformed JSONL lines are skipped rather than crashing every tool. pathmark codex status reports their count as invalidRecordCount so the source file can be repaired deliberately.

Roadmap

  • Provider presets for common local AI CLIs where stable commands exist.

  • Encrypted store option.

  • Hosted sync as an opt-in layer, not a requirement.

  • Native auto-capture packages for additional harness plugin systems beyond Codex and Gemini CLI.

  • Example recipes for Codex, Claude Desktop, Cursor, ChatGPT, and local LLM tools.

Positioning

Pathmark gives your agents a shared working memory that stays on your machine.

Switch agents. Keep the context.

Bring your own subscription. Keep your memory local.

Author and citation

Pathmark is created and maintained by Sergey Moloman, a B2B AI integration specialist and private AI contractor, and founder of RFLX AI.

Machine-readable authorship and citation metadata are available in CITATION.cff and codemeta.json.

License

MIT

Available Tools

25 tools
approve_conclusionApprove conclusionA
Idempotent

Approve one pending conclusion, optionally correcting its text or tags. The transition is atomic and auditable.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYes
noteNo
tagsNo
textNo
decidedByNo
namespaceNo

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover read-only, destructiveness, and idempotency, but the description adds 'atomic and auditable' traits that are not in the annotations. This provides useful behavioral context beyond the structured metadata.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences with no waste. The core action is front-loaded, and the optional correction detail follows naturally. The structure is tight and efficient.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with six parameters and no output schema, the description is too thin. It omits required parameter semantics (especially id), optional fields like note and decidedBy, and any error or edge-case behavior. The low schema coverage makes this a significant gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 0% schema description coverage, the description must compensate for parameter meaning. It mentions 'text or tags' but leaves id, note, decidedBy, and namespace unexplained. This is a partial but insufficient compensation for a 6-parameter tool.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('Approve') and the resource ('one pending conclusion'), and distinguishes it from siblings like reject_conclusion or list_pending_conclusions. It also notes optional corrections, making the purpose specific and unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It implies when to use it (to approve a pending conclusion) but does not explicitly state when not to use it or name alternatives such as reject_conclusion or create_conclusion. The guidance is clear but not explicit about exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ask_memoryAsk memoryA
Read-onlyIdempotent

Ask approved conclusions first, then scoped or explicitly requested raw evidence. Returns an answer, exact provenance, and a recallId for feedback.

ParametersJSON Schema
NameRequiredDescriptionDefault
kindNo
tagsNo
limitNo
questionYes
namespaceNo

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and non-destructive behavior. The description adds useful behavioral context beyond those flags: it reveals the retrieval order and the exact output contract (answer, provenance, recallId). This helps an agent predict side effects and response shape without contradicting the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences with no filler. The core behavior is front-loaded in the first sentence, and the second sentence efficiently states the return value. Every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers the retrieval strategy and the output contract, which is helpful given the lack of an output schema. However, with five parameters and no schema-level descriptions, an agent still lacks details on parameter meanings and edge behavior such as limit handling or namespace scoping. It is adequate but not fully complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate for undocumented parameters. It partially maps 'approved conclusions' to kind=conclusion and 'raw evidence' to kind=memory, and 'scoped' hints at tags or namespace. However, it does not explain limit, tags, namespace, or the full enum semantics, leaving significant ambiguity for a five-parameter tool.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific action ('Ask'), a resource (memory), and a distinctive retrieval order ('approved conclusions first, then scoped or explicitly requested raw evidence'). It also names the return payload, which helps an agent understand what the tool produces. It doesn't explicitly contrast itself with sibling tools like search_memory or recall_memory, but the behavior is clear enough to differentiate.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear usage context: prefer approved conclusions, then fall back to raw evidence when scoped or explicitly requested. It implies when to use this tool for memory queries and what kind of output to expect. It doesn't state when not to use it or name alternatives, so it stops short of full exclusion guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

audit_memoryAudit memory valueA
Read-onlyIdempotent

Measure capture-to-recall behavior, unused records, recall age, duplicate rate, stale raw hits, and available precision evidence without changing memory.

ParametersJSON Schema
NameRequiredDescriptionDefault
daysNo
tagsNo
namespaceNo

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, establishing the safe, non-mutating nature. The description adds value by detailing what behavioral aspects it measures (capture-to-recall behavior, duplicate rate, stale raw hits), which is context beyond the annotations. It does not contradict any annotation and provides a clearer picture of the tool's analytical focus, though it stops short of describing output structure or side effects.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, information-dense sentence that front-loads the primary action ('Measure') and lists the key metrics immediately. There is no filler or repetition. It earns every word and remains easily scannable, ideal for an AI agent parsing tool definitions.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with three parameters, no output schema, and zero schema coverage, the description leaves critical gaps. It lists what metrics are measured but does not specify the return format, how parameters affect the results, or what constitutes a valid invocation. An agent cannot confidently construct a request without additional external knowledge. The complexity is moderate, but the lack of parameter and output documentation makes this incomplete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate for the three parameters (days, tags, namespace). It does not explain what these parameters control or how they influence the audit. The parameter names are somewhat intuitive, but without semantic detail an agent cannot determine how to set values to achieve a desired audit scope. This is a significant gap given no output schema exists either.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Measure') with a clear resource ('memory') and enumerates concrete metrics (capture-to-recall, unused records, recall age, etc.). It also explicitly distinguishes itself by noting 'without changing memory,' which sets it apart from mutation tools like update_memory or delete_memory. This gives an agent an unambiguous understanding of the tool's purpose.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies a usage context (auditing memory metrics) but does not explicitly state when to prefer this over siblings like get_memory_snapshot or search_memory. It lacks a 'when not to use' or alternative tool references, leaving the agent to infer that this is for analytical insights rather than retrieving raw data. The phrase 'without changing memory' hints at a read-only analysis use case, but no direct routing guidance is provided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

backup_memoryBack up memory storeC

Create a point-in-time copy of the canonical local JSONL store.

ParametersJSON Schema
NameRequiredDescriptionDefault
destinationNo

TDQS

C2.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The annotations provide readOnlyHint=false and destructiveHint=false, indicating a non-destructive write operation. The description adds the 'point-in-time' aspect, implying a snapshot-like behavior that doesn't modify the source. However, it does not disclose potential side effects like overwriting existing files, permission requirements, or what happens to the destination if it already exists. Given annotations already cover the basic safety profile, the added context is moderate.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, focused sentence that front-loads the core action and resource. It contains no fluff and is highly efficient, achieving maximum conciseness without sacrificing clarity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with one parameter and no output schema, the description is notably incomplete. It does not explain the purpose of the destination parameter, what the tool returns (if anything), or any prerequisites like existing directory structure or permissions. The agent lacks enough information to call it correctly without additional inference or experimentation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema has zero description coverage for the only parameter 'destination', and the description does not mention this parameter at all. With a single parameter, this is a critical gap. The agent has no information about the expected format, whether it should be a file path or directory, or if it is required (though it's optional in the schema). The description fails to compensate for the missing schema descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb 'Create' and the resource 'canonical local JSONL store', and adds the qualifier 'point-in-time copy', which distinguishes it from a generic copy. However, it does not explicitly differentiate from sibling tools like export_memory or get_memory_snapshot, leaving some ambiguity about the exact scope.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus alternatives such as export_memory or get_memory_snapshot. The description only states what it does, not when it is the preferred choice or what conditions would make it inappropriate. This leaves the agent without selection criteria.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

chatChatB
Read-onlyIdempotent

Chat with Pathmark using approved conclusions first and only scoped or explicitly requested raw fallback. Returns an answer, provenance, and recallId.

ParametersJSON Schema
NameRequiredDescriptionDefault
kindNo
tagsNo
limitNo
questionYes
namespaceNo

TDQS

B3.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description adds meaningful behavioral context beyond the annotations: the fallback policy ('approved conclusions first', 'scoped or explicitly requested raw fallback') and the return shape ('answer, provenance, and recallId'). This complements the readOnly/idempotent hints without contradicting them.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single dense sentence that front-loads the core behavior and then states the return value. Every clause earns its place, and there is no repetition of schema or annotation information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers the tool's high-level behavior and return value, but with no output schema and no parameter explanations, an agent cannot fully determine how to use optional parameters or how to request raw fallback. The missing parameter semantics and lack of usage alternatives leave significant gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, and the description provides no explanation for any of the five parameters (question, kind, tags, limit, namespace). The agent is left to infer parameter meaning from names and types alone, which is insufficient for optional parameters like kind, tags, and namespace.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's function: 'Chat with Pathmark' using approved conclusions first, with raw fallback only when scoped or requested. It names the resource and the key behavior, though it does not explicitly differentiate from siblings like ask_memory or recall_memory.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'using approved conclusions first and only scoped or explicitly requested raw fallback' implies a usage policy, but it does not explicitly state when to prefer this tool over alternatives or when not to use it. The guidance is present but indirect.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

compact_memoryCompact memory storeA
Destructive

Preview or apply exact deduplication, expired-record removal, retention, and deleted-record purging. Applied runs create a backup.

ParametersJSON Schema
NameRequiredDescriptionDefault
dedupeNo
confirmNo
dropDeletedNo
retentionDaysNo

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, so the agent knows this can destroy data. The description adds valuable context: 'Preview or apply' indicates a dry-run mode, and 'Applied runs create a backup' discloses a safety net. This goes beyond the annotations and helps the agent understand the risk profile.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with the core action and scope, followed by the key safety detail. Every word earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive tool with 4 parameters and no output schema, the description covers the main operations and the backup safety net. However, it doesn't explain the preview/apply flow in terms of the confirm parameter, nor what happens to the backup (where it goes, how to restore). Given the destructiveHint annotation, a bit more operational context would be warranted.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It does explain the high-level operations (dedupe, expired removal, retention, purging) which map to the parameters, but it doesn't clarify the meaning of 'confirm' (likely the apply switch) or 'retentionDays' semantics (e.g., 0 meaning no retention). The description adds some meaning but leaves the parameter details to inference.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb ('Preview or apply') and a clear resource ('memory store'), and lists the exact operations: deduplication, expired-record removal, retention, and deleted-record purging. It distinguishes itself from siblings like delete_memory or purge_memory by framing this as a compaction operation, though it doesn't explicitly name a sibling.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage: preview first, then apply with confirm. It also notes that applied runs create a backup, which hints at safety. However, it doesn't explicitly state when to use this tool versus alternatives like purge_memory, consolidate_memory, or delete_memory, nor does it state when not to use it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

consolidate_memoryConsolidate raw evidenceB
Idempotent

Prepare a bounded unsynthesized evidence batch and, when server synthesis is configured, preview or stage evidence-backed conclusion proposals. Proposals are never auto-approved.

ParametersJSON Schema
NameRequiredDescriptionDefault
daysNo
tagsNo
applyNoStage generated proposals as pending conclusions.
cursorNoContinue after the last record id from a prior bounded batch.
namespaceNo
maxProposalsNo
evidenceLimitNo

TDQS

B3.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate idempotentHint=true, readOnlyHint=false, and destructiveHint=false. The description adds valuable behavioral context: 'Proposals are never auto-approved' and the dependency on server synthesis configuration. This goes beyond annotation disclosure without contradicting it.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences, front-loaded with the primary action, and avoids unnecessary detail. It is concise and structured, though it sacrifices some explanatory depth for brevity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With 7 parameters, no output schema, and low schema coverage, the description is insufficient for an agent to call this tool correctly. It does not explain the meaning of the parameters, the expected return structure, or the conditions under which staging occurs. Significant gaps remain.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 29% (only apply and cursor have descriptions). The tool description fails to explain any of the parameters (days, tags, namespace, maxProposals, evidenceLimit), leaving the agent without meaning beyond the schema. Given the low coverage, the description should compensate but does not.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific action: preparing a bounded unsynthesized evidence batch and optionally previewing/staging proposals. It distinguishes from sibling tools like create_conclusion and approve_conclusion by emphasizing staging and no auto-approval. However, terms like 'unsynthesized' are somewhat jargon-heavy but still convey a clear purpose.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives a conditional context ('when server synthesis is configured') but does not explicitly state when to prefer this over alternatives like create_conclusion or list_pending_conclusions. It implies staging use but lacks explicit when-not guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_conclusionCreate conclusionC

Propose a durable higher-signal conclusion. Approval is required by default before it can be recalled.

ParametersJSON Schema
NameRequiredDescriptionDefault
tagsNo
textYesConclusion text to save.
sourceNo
expiresAtNo
namespaceNo
evidenceIdsNoRaw memory IDs supporting this conclusion. Used for provenance and consolidation coverage.

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations indicate readOnlyHint=false, meaning this is a write operation, which the description implies with 'Propose'. The description adds that 'Approval is required by default' before recall, which is useful. However, it doesn't disclose any other behavioral traits like potential side effects, permissions, or the approval workflow mechanics, leaving gaps.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise, with two sentences, and front-loads the core purpose. No wasted words, but it could add a bit more context without becoming verbose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's moderate complexity (6 params, some with descriptions) and no output schema, the description is minimally sufficient. It covers the core purpose and approval requirement, but lacks guidance on parameter usage and integration with other tools like approval. It's acceptable but not thorough.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 33%, so the description should clarify parameters. The description only mentions that the conclusion is 'higher-signal' and durable, but doesn't explain the purpose of 'tags', 'source', 'expiresAt', 'namespace', or 'evidenceIds'. With low coverage, this is a partial gap, but the schema has descriptions for 'text' and 'evidenceIds', so baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('Propose a durable higher-signal conclusion') and the resource (conclusion), distinguishing it from related tools like 'remember'. However, it doesn't explicitly differentiate from 'approve_conclusion' or 'list_conclusions' which are siblings, but its purpose is clear enough.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies that this tool is for creating conclusions, but it doesn't provide explicit guidance on when to use it versus alternatives like 'remember' or 'search_memory'. There's no mention of when not to use it or which sibling tools are preferable in specific contexts.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_memoryDelete memoryA
DestructiveIdempotent

Soft-delete a saved memory or conclusion by id.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYes

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description adds meaningful behavioral context beyond the annotations by disclosing that deletion is 'soft' rather than a permanent hard delete, which is not explicit in destructiveHint=true or idempotentHint=true. It does not elaborate on recovery or cascading effects, but the key non-obvious trait is clearly stated.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, tightly worded sentence with no filler. It front-loads the operation type and parameter semantics, making it easy for an agent to parse quickly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter tool with annotations covering idempotency and destructiveness, the description is nearly complete. It accurately identifies what the tool does and the required id, though it could optionally clarify the difference between deleting a memory versus a conclusion and what the response contains.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With schema coverage at 0%, the description carries the burden of explaining the 'id' parameter. It does say the operation acts by id on a saved memory or conclusion, which adds some meaning beyond the bare schema property, but it does not specify id format, provenance, or how to obtain a valid id.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('soft-delete') and names the resource ('saved memory or conclusion') plus the primary identifier dimension ('by id'). It also distinguishes itself from sibling tools like purge_memory by explicitly flagging the operation as a soft-delete.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies that the tool should be used when a soft-delete of a memory/conclusion by id is intended, and the 'by id' phrasing hints at the required input. However, it gives no explicit comparison with alternatives such as purge_memory, supersede_memory, or update_memory, so the when/when-not guidance is left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

doctor_memoryDiagnose memory storeA
Read-onlyIdempotent

Report duplicate, deleted, expired, conclusion, invalid-record, and index health counts without changing data.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description states 'without changing data', which aligns with and reinforces the annotations (readOnlyHint=true, destructiveHint=false). It adds value by explicitly clarifying the non-mutating behavior, which is a key behavioral guarantee for an agent deciding to invoke it.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence that front-loads the verb and resource. It is concise and covers the essential categories without waste. It could be slightly more structured but is quite efficient.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter diagnostic tool with rich annotations, the description is nearly complete. It lists all major health categories and assures non-mutation. It lacks an explicit output format, but the categories imply what will be reported, which is adequate.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, so the schema offers no parameter semantics. The description compensates well by listing the specific categories of counts it reports, providing clear meaning about what the tool returns even without an output schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states a specific verb ('Report') and resource ('memory store') and lists the categories of health counts it provides. It distinguishes itself by focusing on diagnostics rather than modification, which differentiates it from most sibling tools that manage or modify memory.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies the tool's use for reporting health statuses without explicit when-to-use guidance or alternatives. It doesn't mention when to prefer this over audit_memory or get_memory_snapshot, but the diagnostic focus is clear enough for an agent to infer its purpose.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

export_memoryExport memoryB

Export a scoped, mergeable JSONL bundle for another Pathmark installation or trusted sync transport.

ParametersJSON Schema
NameRequiredDescriptionDefault
kindNo
tagsNo
encryptedNoEncrypt the export with PATHMARK_EXPORT_KEY.
namespaceNo
destinationYes
includeDeletedNo

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description adds useful behavioral details not in the annotations: the output is JSONL, scoped, and mergeable. However, with all annotations false, it carries the burden of describing side effects, permissions, or external data movement more fully; it only says 'for another ... transport' without explaining what happens during the export.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single dense sentence that front-loads the core action and output format. There is no filler or redundancy; every phrase adds meaning.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given six parameters, no output schema, no useful annotations, and low schema coverage, this short description is not complete enough. It omits parameter semantics, destination format specifics, merge behavior, and any relationship or differentiation from backup_memory. An agent would need to inspect the schema carefully and still infer much of the intent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 17%, and the description does not compensate. It hints at scoping and merging but never explains the required 'destination', the 'kind' enum, 'tags', 'namespace', or 'includeDeleted'. The schema itself leaves nearly all parameters undocumented.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('Export'), the resource ('memory'), and the output form ('scoped, mergeable JSONL bundle'). It also names the target context ('another Pathmark installation or trusted sync transport'), which helps distinguish it from a generic backup operation, though it does not explicitly contrast it with sibling tools like backup_memory.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives a clear context for use: exporting to another Pathmark installation or a trusted sync transport. It does not provide exclusions or explicitly compare with alternatives such as backup_memory, but the intended scenario is reasonably clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_configGet Pathmark configurationA
Read-onlyIdempotent

Show the local Pathmark Memory store location and enabled optional features.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and destructiveHint, so the safety profile is covered. The description adds what the tool returns—local store location and enabled optional features—which is useful behavioral context beyond the annotations. It aligns with the read-only nature and does not contradict any annotation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single 11-word sentence that is front-loaded with the verb and resource. It contains no filler, and every word contributes meaning.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter, low-complexity configuration getter, the description is complete: it states what the tool exposes (store location and optional features), and the annotations confirm it is safe and idempotent. No output schema exists, but the description enumerates the content an agent can expect.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters and the schema is empty, so there is nothing for the description to document. The schema and context already cover 100% of the parameter surface, so no additional parameter explanation is needed.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Show') and names the exact resource: the local Pathmark Memory store location and enabled optional features. This clearly distinguishes it from sibling tools like search_memory or get_context, which operate on memory content rather than configuration.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no explicit guidance on when to choose this tool over alternatives, nor any exclusion criteria. It simply states the function; context for when to call it is only implied by the name 'get_config' among memory-focused siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_contextGet contextC
Read-onlyIdempotent

Return compact local memory context for a task or question.

ParametersJSON Schema
NameRequiredDescriptionDefault
kindNo
tagsNo
limitNo
queryNoTask or question to retrieve context for.
namespaceNo

TDQS

C2.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and destructiveHint, so the safety profile is covered. The description adds the notion of 'compact' and 'local' context, which gives some behavioral color, but it does not disclose output format, filtering behavior, or namespace semantics.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence with no filler. It is concise, but it is arguably too terse to carry the semantic weight needed for a tool with five parameters and many siblings.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the number of parameters, the lack of an output schema, and the large sibling set, this description is incomplete. It does not explain parameter semantics, usage context, or what 'compact context' actually contains, so an agent would struggle to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 20%, with only 'query' described. The description's 'task or question' maps to the query parameter, but it does not explain 'kind', 'tags', 'limit', or 'namespace'. With five parameters and low schema coverage, the description fails to compensate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Return') and resource ('compact local memory context'), and ties it to a task or question. However, it does not differentiate this tool from siblings like search_memory, recall_memory, or ask_memory, so it stops short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus the many sibling memory tools. No alternatives, exclusions, or contextual triggers are mentioned, leaving the agent to infer when get_context is the right choice.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_memory_snapshotGet approved memory snapshotB
Read-onlyIdempotent

Generate a bounded USER/PROJECT/AGENT snapshot from approved canonical conclusions only.

ParametersJSON Schema
NameRequiredDescriptionDefault
tagsNo
charLimitNo
namespaceNo

TDQS

B3.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the description's main contribution is the 'bounded' nature and the restriction to 'approved canonical conclusions only'. It adds useful behavioral context without contradicting the annotations, though it does not describe return format or edge cases.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence with no waste. It front-loads the primary action and scope, making it easy to parse quickly. Every word earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the large sibling set and the lack of parameter semantics and usage guidance, the description is not complete enough for an agent to confidently choose and invoke this tool. The core purpose is clear, but the absence of parameter explanations and differentiation from alternatives leaves significant gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description carries full responsibility for explaining the three parameters (tags, charLimit, namespace). The description mentions none of them, leaving the agent to infer meaning from parameter names alone. This is a significant gap for a tool with three optional parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Generate') and resource ('snapshot') with clear scoping ('USER/PROJECT/AGENT') and a precise constraint ('approved canonical conclusions only'). This clearly distinguishes it from siblings like search_memory or recall_memory, which may include unapproved content, and from creation tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description does not provide any guidance on when to use this tool versus the many sibling tools (e.g., get_context, recall_memory, search_memory). It implies a use case through the 'approved conclusions only' clause, but there is no explicit when-to-use or when-not-to-use guidance, nor any mention of alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_conclusionsList conclusionsC
Read-onlyIdempotent

List saved durable conclusions.

ParametersJSON Schema
NameRequiredDescriptionDefault
tagsNo
limitNo
namespaceNo

TDQS

C2.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, destructiveHint=false, and idempotentHint=true, covering safety. The description adds the 'durable' qualifier, which is useful but does not explain return format, pagination, or any side effects. With annotations covering the core safety profile, a 3 is appropriate.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence with zero unnecessary words. It is front-loaded with the action and resource, and no redundancy exists.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a list operation with three undocumented parameters and no output schema, the description is inadequate. It does not mention filtering options, default limits, or how results are ordered, leaving an agent to guess or infer critical details.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, and the description provides no explanation of the tags, limit, or namespace parameters. The agent is left without any guidance on what these parameters do, how they affect the query, or their defaults.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a clear verb ('List') and resource ('saved durable conclusions'), distinguishing it from pending conclusions implicitly. However, it does not explicitly differentiate from other list tools like list_pending_conclusions or clarify what 'durable' means in this context.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives such as list_pending_conclusions. The description does not mention any conditions or contexts that would select this tool over its siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_pending_conclusionsList pending conclusionsA
Read-onlyIdempotent

List bounded approval-gated conclusion proposals. Pending records are never returned by normal memory search.

ParametersJSON Schema
NameRequiredDescriptionDefault
tagsNo
limitNo
offsetNo
namespaceNo

TDQS

A3.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Adds meaningful context beyond annotations: the fact that pending records are never returned by normal memory search is a useful behavioral caveat. With readOnlyHint, idempotentHint, and destructiveHint already set, the description enriches understanding by clarifying scope and visibility, though it omits pagination or ordering details.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with the core purpose, no filler. The extra sentence adds a valuable behavioral distinction without verbosity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Adequate for a simple list operation given annotations cover safety, but lacking parameter semantics and explicit sibling differentiation. An agent could call it correctly with defaults, but would be unsure about filtering options and when to prefer it over list_conclusions.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the description does not compensate. It fails to explain any of the four parameters (tags, limit, offset, namespace), leaving an agent with no guidance on how to filter or paginate. This is a significant gap given the low schema coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb 'List' and a resource 'bounded approval-gated conclusion proposals', which clearly conveys the core function. The added note that pending records are never returned by normal search helps distinguish it from search_memory, but it does not explicitly differentiate from sibling list_conclusions, leaving slight ambiguity about the exact scope.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Implies this is the tool for pending conclusions via the phrase 'approval-gated' and the note about normal search, but provides no explicit when-to-use vs alternatives like list_conclusions. It does not state exclusions or conditions that would guide selection between sibling tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

purge_memoryHard purge memoryC
Destructive

Preview or permanently remove matching records from the canonical store. A backup is created before an applied purge.

ParametersJSON Schema
NameRequiredDescriptionDefault
idNo
tagsNo
beforeNo
sourceNo
confirmNoFalse previews the purge; true applies it and creates a backup.
namespaceNo

TDQS

C2.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already flag destructiveHint=true, so the description adds the backup-before-purge behavior and the preview/apply distinction, which is useful context. However, it does not disclose what 'matching records' means, whether the purge is reversible via backup, or any side effects on related data. Given annotations, a 3 is appropriate.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise, two sentences, with the primary action (preview/remove) front-loaded. It avoids redundancy and is efficiently structured, though it could have used the space to clarify matching criteria.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With 6 parameters, no output schema, and low parameter descriptions, the description is inadequate. It does not explain how to construct a matching query, what the backup entails, or what the preview output looks like. The tool's complexity demands more context than provided.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 17% (only 'confirm' has a description). The description mentions 'matching records' but does not explain how id, tags, before, source, or namespace contribute to matching. It fails to compensate for the low schema coverage, leaving parameter meaning largely ambiguous.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a clear verb ('preview or permanently remove') and resource ('matching records from the canonical store'). It conveys the core action but does not differentiate from sibling tools like delete_memory or supersede_memory, which also perform removals.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is given on when to use this tool versus alternatives. It does not mention conditions for preview vs. apply, nor any exclusions or prerequisites. The agent must infer usage from the parameter schema and tool name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

rate_recallRate recalled memoriesC

Attach explicit relevance labels to one exact Pathmark recall so audit_memory can report measured precision.

ParametersJSON Schema
NameRequiredDescriptionDefault
noteNo
recallIdYesThe recallId returned by chat or ask_memory.
relevantIdsNo
irrelevantIdsNo

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations are all false (readOnlyHint false, destructiveHint false, etc.), so the description must carry behavioral disclosure. It implies a write operation (attaching labels) but doesn't explain side effects, reversibility, or response format. The link to audit_memory suggests downstream effects but lacks detail.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, compact sentence with no fluff. It front-loads the primary action. However, it could be structured to include more critical details without losing conciseness, such as clarifying the purpose of relevant/irrelevant IDs.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with 4 parameters, no output schema, and minimal annotations, the description is inadequate. It doesn't explain what a 'Pathmark recall' is, how relevantIds and irrelevantIds should be populated, or what the tool returns. An agent would struggle to invoke it correctly without additional context.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 25% (only recallId has a description). The description does not explain note, relevantIds, or irrelevantIds beyond the vague 'relevance labels'. With low coverage, the description should compensate but doesn't, leaving agents guessing about parameter meaning and usage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific action (attach relevance labels) on a specific resource (one exact Pathmark recall) and ties it to audit_memory's precision reporting. This is clear enough to understand the core function, though it doesn't explicitly differentiate from siblings like recall_memory or update_memory.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No explicit guidance on when to use this tool versus alternatives. The mention of audit_memory implies a workflow but doesn't state when to choose rate_recall over other memory operations or what conditions warrant it. No exclusions or alternative tools are named.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

recall_memoryRecall memoryB
Read-onlyIdempotent

Transparent recall for any MCP-capable harness. Use this at task start or before answering to show exactly which memories were used.

ParametersJSON Schema
NameRequiredDescriptionDefault
idsNoExact memory IDs from a prior Pathmark context block. Preserves the original visible-recall set.
kindNo
tagsNoOptional tags to scope visible recall, such as the current workspace tag.
limitNo
queryNoTask, repo, or question to retrieve memory for. Empty query returns recent records.
namespaceNo
includeRecordsNoInclude a second, untruncated full-record copy alongside usedMemories. Defaults to false: it duplicates data already in context/usedMemories and is unbounded in size. Set true only when full record bodies are required.

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, covering the safety profile. The description adds 'Transparent recall' and 'show exactly which memories were used', which is more about purpose than behavior. It does not describe rate limits, authentication, or what happens with includeRecords, but given annotations cover safety, the added context is minimal yet non-contradictory.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences, front-loaded with the core purpose and usage. Every word earns its place, and it avoids redundancy with the schema or annotations.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with 7 parameters and no output schema, the description is minimal. It does not explain the return structure (e.g., usedMemories vs includeRecords), how to retrieve memory IDs from prior context, or the interaction between query, tags, and ids. An agent may struggle to formulate correct requests without additional documentation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 57%, meaning several parameters (kind, limit, namespace) lack descriptions. The tool description does not explain any parameter semantics, such as how to obtain memory IDs or the meaning of namespace. It fails to compensate for the missing schema details, leaving an agent with incomplete guidance for constructing valid calls.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: transparent recall showing exactly which memories were used. It distinguishes it from search_memory and other memory tools by emphasizing 'show exactly which memories were used' at task start, which implies a distinct role in surfacing used memories rather than searching or managing them.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly instructs when to use the tool: 'at task start or before answering'. This provides a clear usage context. However, it does not mention alternatives or when not to use it, but the specific timing guidance is strong enough to guide an agent.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

reject_conclusionReject conclusionA
DestructiveIdempotent

Reject one pending conclusion while retaining it in the canonical audit trail and excluding it from recall.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYes
noteNo
decidedByNo

TDQS

A3.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, but the description adds critical nuance: the conclusion is not deleted but retained in the audit trail and excluded from recall. This is valuable behavioral context beyond the annotation, clarifying the exact side effects.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence with no filler. Every clause adds meaning: the action, the target, and the two key side effects (audit trail retention and recall exclusion).

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

While the core behavior is stated, the description omits parameter semantics and does not mention prerequisites like the id being required or the conclusion being pending. For a destructive tool with no output schema, this is insufficient for an agent to call it correctly without external knowledge.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, and the description does not mention any of the three parameters (id, note, decidedBy). The agent is left without any meaning for these fields, which is a significant gap for a destructive operation that requires an id.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb (reject), resource (pending conclusion), and the specific behavior: retaining it in the canonical audit trail and excluding from recall. This distinguishes it from siblings like approve_conclusion, making the purpose unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage for pending conclusions that need rejection, but it does not explicitly contrast with approve_conclusion or other alternatives. No when-to-use or when-not-to-use guidance is provided, leaving the routing to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

rememberSave raw evidenceA
Idempotent

Save raw searchable evidence. Durable intent should use the approval-gated conclusion workflow.

ParametersJSON Schema
NameRequiredDescriptionDefault
tagsNoOptional lowercase-ish tags for later filtering.
textYesMemory text to save.
sourceNoOptional source label, such as repo, thread, or tool name.
expiresAtNoOptional ISO timestamp after which recall excludes this memory.
namespaceNoOptional project, user, or client namespace.

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover idempotency, non-destructiveness, and write status, so the bar is lower. The description adds useful behavioral context: the stored content is raw and searchable, and durable intent is routed elsewhere. This enriches the agent's understanding beyond the annotation flags.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences with zero filler. The primary action is front-loaded, and the routing guidance is placed second. Every word earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple single-required-parameter save tool, the description plus rich schema and annotations is largely complete. It does not describe the return value or duplicate behavior, but with idempotentHint=true and a straightforward write operation, an agent has enough to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all five parameters are already documented with their own descriptions. The description does not add parameter-specific meaning beyond what the schema provides, warranting the baseline score of 3.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Save') and resource ('raw searchable evidence'), making the action unmistakable. It also distinguishes itself from the conclusion workflow for durable intent, which separates it from sibling tools like create_conclusion.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly tells the agent when not to use this tool ('Durable intent...') and names the alternative ('approval-gated conclusion workflow'). This is an explicit when-not/alternative routing instruction, leaving little to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_memorySearch memoryB
Read-onlyIdempotent

Search saved local memories and conclusions.

ParametersJSON Schema
NameRequiredDescriptionDefault
kindNo
tagsNo
limitNo
queryNoSearch query. Empty query returns recent records.
namespaceNo

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, covering the safety profile. The description adds no further behavioral detail (e.g., result ordering, pagination, or the effect of an empty query) beyond the obvious act of searching. It does not contradict the annotations, but it also does not enrich them.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, efficient sentence with no extraneous words. It is front-loaded with the core action and resource, making it easy to parse quickly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has 5 parameters, no output schema, and a rich set of siblings, the description is far too sparse. It does not explain what the tool returns, how to formulate an effective query, or how it relates to the broader memory ecosystem. The agent would need to rely heavily on schema and external knowledge to use it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With schema description coverage at only 20% (only the 'query' parameter has a description), the tool description should compensate by explaining the key parameters. It does not mention 'kind', 'tags', 'limit', or 'namespace' at all, leaving the agent to interpret them solely from the schema. The description adds no semantic value beyond what the schema already provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action (search) and the resource (saved local memories and conclusions). This is specific and distinguishes it from list-oriented tools like list_conclusions or recall_memory, which imply retrieval of specific items rather than a search.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives such as recall_memory, ask_memory, or list_conclusions. It neither states typical use cases nor exclusions, leaving the agent to infer the appropriate context from the name and schema alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

session_traceSession traceA
Read-onlyIdempotent

Show a bounded chronological audit trail for one captured session: prompts, exact injected memory IDs, redacted tool inputs/results, and answers.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
sessionIdYesExact Codex or harness session ID.
includeOutputsNoInclude redacted bounded tool output previews. Defaults to true.

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already establish read-only, idempotent, and non-destructive behavior. The description adds meaningful behavioral context: the trail is bounded, chronological, tool inputs and results are redacted, and injected memory IDs are shown exactly. This goes beyond annotation basics, though it does not specify ordering direction or the boundary's default cap in detail.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence that starts with the core action and scope, then lists exactly what the trail contains. No filler or redundant words; every phrase earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple read-only audit tool with no output schema, the description adequately explains what is returned: prompts, memory IDs, redacted tool inputs/results, and answers. It could mention the default bound or behavior for missing sessions, but the schema and annotations cover the essential invocation context, so the description is nearly complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema descriptions cover sessionId and includeOutputs, leaving limit without a description. The word 'bounded' in the description loosely suggests the role of limit, but it does not explain its meaning or default. The schema already provides type and min/max constraints, so the description adds only moderate parameter-level value beyond the structured fields.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb and resource: 'Show a bounded chronological audit trail for one captured session', and enumerates the contents (prompts, memory IDs, redacted tool I/O, answers). This clearly differentiates it from sibling tools like audit_memory and get_memory_snapshot, which concern memory storage operations rather than session-level I/O tracing.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The purpose statement implies the tool is for reviewing a session's audit trail, but it gives no explicit when-to-use vs. alternatives guidance. There is no mention of when to prefer session_trace over audit_memory or other trace-like siblings, nor any exclusions, so the usage context is only implied through the description itself.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

supersede_memorySupersede memoryC

Replace an outdated memory with a linked current record while preserving history.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYes
kindNomemory
tagsNo
textYes
sourceNo
expiresAtNo
namespaceNo

TDQS

C2.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations indicate this is a write operation (readOnlyHint false) but not destructive (destructiveHint false). The description adds the important context that history is preserved via a link, which is beyond the annotations. However, it does not clarify what happens to the old record, whether permissions are needed, or what the response contains. Given minimal annotation coverage, the description partially compensates but not fully.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single concise sentence that front-loads the core action and the key differentiator (preserving history). It is efficient and well-structured, though it could be slightly expanded without becoming verbose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With 7 parameters, 2 required, no output schema, and a mutation operation, the description is far from complete. It lacks parameter explanations, return behavior, side effects beyond preserving history, and any differentiation from siblings. The agent would have to rely on the schema alone, which is insufficient given zero coverage.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate by explaining the parameters. It mentions none of the seven parameters (id, text, kind, tags, source, expiresAt, namespace). The agent gets no guidance on what each parameter means or how they affect the operation, making this a severe gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb 'replace' and the resource 'memory', and adds the specific nuance of 'linked current record while preserving history', which distinguishes it from a plain delete or update. It does not explicitly contrast with sibling tools, but the core purpose is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus the many sibling tools like update_memory, delete_memory, or consolidate_memory. No mention of alternatives or conditions. The intended usage is only implied by the action described.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

update_memoryUpdate memoryC

Correct an existing memory while preserving its prior versions in local history.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYes
tagsNo
textNo
sourceNo
expiresAtNo
namespaceNo

TDQS

C2.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate this is a non-destructive, non-read-only, non-idempotent write operation. The description adds the specific behavior of preserving prior versions, which is consistent with destructiveHint=false. However, it does not disclose the response format, error conditions, or any permission requirements, leaving significant gaps for an agent to navigate.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, efficient sentence that states the core purpose without padding. It is front-loaded with the action and the key preservation behavior. While terse, it earns its place; the only slight shortfall is that it could have used one additional sentence for parameter hints.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with six parameters, no output schema, and zero schema descriptions, the description is severely incomplete. It does not clarify how the id identifies the memory, what the effect of each field is, whether updates are partial or full replacements, or what the tool returns. An agent would struggle to call this correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, meaning the schema provides no parameter descriptions. The tool description does not compensate—it never explains what 'text', 'tags', 'source', 'expiresAt', or 'namespace' mean in the context of updating a memory. Without any parameter semantics, an agent cannot correctly construct the input beyond guessing.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('Correct an existing memory') and the key distinguishing behavior ('preserving its prior versions in local history'). This differentiates it from sibling tools like delete_memory or supersede_memory, which either remove or replace without preserving history. The verb-resource pair is specific and unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives. It does not mention that update_memory is for corrections or revisions while supersede_memory might be for replacing with a new version, or that delete_memory is for removal. The agent is left to infer the appropriate context from the tool name alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 9 tool updatesv0.1.15
    • Addedapprove_conclusion
    • Addedaudit_memory
    • Addedconsolidate_memory
    • Changedcreate_conclusion1 field changed
      • addedInput schema / properties / evidenceIds
        Added value: +{
        +  "description": "Raw memory IDs supporting this conclusion. Used for provenance and consolidation coverage.",
        +  "items": {
        +    "minLength": 1,
        +    "type": "string"
        +  },
        +  "maxItems": 100,
        +  "type": "array"
        +}
    • Addedget_memory_snapshot
    • Addedlist_pending_conclusions
    • Addedrate_recall
    • Changedrecall_memory1 field changed
      • changedInput schema / properties / includeRecords / description
        Previous value: -"Include a second full-record copy. Defaults to true for compatibility."New value: +"Include a second, untruncated full-record copy alongside usedMemories. Defaults to false: it duplicates data already in context/usedMemories and is unbounded in size. Set true only when full record bodies are required."
    • Addedreject_conclusion
  2. 2 tool updatesv0.1.9
    • Changedrecall_memory2 fields changed
      • addedInput schema / properties / ids
        Added value: +{
        +  "description": "Exact memory IDs from a prior Pathmark context block. Preserves the original visible-recall set.",
        +  "items": {
        +    "minLength": 1,
        +    "type": "string"
        +  },
        +  "maxItems": 30,
        +  "minItems": 1,
        +  "type": "array"
        +}
      • addedInput schema / properties / includeRecords
        Added value: +{
        +  "description": "Include a second full-record copy. Defaults to true for compatibility.",
        +  "type": "boolean"
        +}
    • Addedsession_trace
  3. 15 tool updatesv0.1.7
    • Changedask_memory3 fields changed
      • addedInput schema / properties / kind
        Added value: +{
        +  "enum": [
        +    "memory",
        +    "conclusion"
        +  ],
        +  "type": "string"
        +}
      • addedInput schema / properties / namespace
        Added value: +{
        +  "minLength": 1,
        +  "type": "string"
        +}
      • addedInput schema / properties / tags
        Added value: +{
        +  "items": {
        +    "type": "string"
        +  },
        +  "type": "array"
        +}
    • Addedbackup_memory
    • Changedchat3 fields changed
      • addedInput schema / properties / kind
        Added value: +{
        +  "enum": [
        +    "memory",
        +    "conclusion"
        +  ],
        +  "type": "string"
        +}
      • addedInput schema / properties / namespace
        Added value: +{
        +  "minLength": 1,
        +  "type": "string"
        +}
      • addedInput schema / properties / tags
        Added value: +{
        +  "items": {
        +    "type": "string"
        +  },
        +  "type": "array"
        +}
    • Addedcompact_memory
    • Changedcreate_conclusion2 fields changed
      • addedInput schema / properties / expiresAt
        Added value: +{
        +  "type": "string"
        +}
      • addedInput schema / properties / namespace
        Added value: +{
        +  "minLength": 1,
        +  "type": "string"
        +}
    • Addeddoctor_memory
    • Addedexport_memory
    • Changedget_context3 fields changed
      • addedInput schema / properties / kind
        Added value: +{
        +  "enum": [
        +    "memory",
        +    "conclusion"
        +  ],
        +  "type": "string"
        +}
      • addedInput schema / properties / namespace
        Added value: +{
        +  "minLength": 1,
        +  "type": "string"
        +}
      • addedInput schema / properties / tags
        Added value: +{
        +  "items": {
        +    "type": "string"
        +  },
        +  "type": "array"
        +}
    • Changedlist_conclusions2 fields changed
      • addedInput schema / properties / namespace
        Added value: +{
        +  "minLength": 1,
        +  "type": "string"
        +}
      • addedInput schema / properties / tags
        Added value: +{
        +  "items": {
        +    "type": "string"
        +  },
        +  "type": "array"
        +}
    • Addedpurge_memory
    • Addedrecall_memory
    • Changedremember2 fields changed
      • addedInput schema / properties / expiresAt
        Added value: +{
        +  "description": "Optional ISO timestamp after which recall excludes this memory.",
        +  "type": "string"
        +}
      • addedInput schema / properties / namespace
        Added value: +{
        +  "description": "Optional project, user, or client namespace.",
        +  "minLength": 1,
        +  "type": "string"
        +}
    • Changedsearch_memory1 field changed
      • addedInput schema / properties / namespace
        Added value: +{
        +  "minLength": 1,
        +  "type": "string"
        +}
    • Addedsupersede_memory
    • Addedupdate_memory
  4. 1 tool updatev0.1.1
    • Addedchat
  5. 8 tool updatesv0.1.0
    • First observedask_memory
    • First observedcreate_conclusion
    • First observeddelete_memory
    • First observedget_config
    • First observedget_context
    • First observedlist_conclusions
    • First observedremember
    • First observedsearch_memory

TDQS

B3.2/5.0

Scored across 25 tools

Disambiguation2/5

chat and ask_memory are nearly identical in purpose—both answer using approved conclusions first with raw fallback and return provenance plus a recallId—creating clear ambiguity. Additionally, search_memory, get_context, recall_memory, and ask_memory overlap in retrieval behavior, though their descriptions partially clarify output differences. The raw-evidence tools (remember, create_conclusion, consolidate_memory) are more distinct but still require careful reading.

Naming Consistency4/5

Most tools follow a consistent verb_noun snake_case pattern such as create_conclusion, approve_conclusion, delete_memory, and export_memory. A few bare-verb or noun-like names like remember, chat, and session_trace deviate slightly, but the overall convention is recognizable and predictable.

Tool Count3/5

At 25 tools, this sits at the heavy end of the borderline range and feels like more surface than most agents will need. The count is defensible for a full memory lifecycle system covering capture, recall, conclusions, audit, and maintenance, but it risks overwhelming users.

Completeness4/5

The tool set covers the core memory lifecycle well: capture raw evidence, propose and approve conclusions, search and recall, update/delete/supersede, audit, compact, backup, and export. Minor gaps exist—there is no import tool to complement export, and no direct get-by-id retrieval—but agents can work around these via search and export/backup workflows.

Maintenance

ActivityActive
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    D
    maintenance
    A shared, local-first memory layer for AI CLIs, providing persistent, layered memory across Claude Code, Gemini CLI, and other MCP-aware clients.
    1
    MIT
  • A
    license
    Not graded
    quality
    A
    maintenance
    Local-first, file-based memory layer for AI agents — one shared Markdown vault across Claude, Codex, Gemini, Cursor and any MCP client. Provides read/write memory tools with an audit trail, per-agent trust levels, and Git sync; no cloud and no lock-in.
    2
    MIT
  • F
    license
    Not graded
    quality
    C
    maintenance
    Local-first cross-agent memory for AI coding agents. Persistent, shared memory over MCP — what you tell one agent can be recalled by another — with all data stored in a single local SQLite file, no cloud and no API keys.
    -
  • A
    license
    A
    quality
    B
    maintenance
    Local-first, source-traceable memory for AI agents — no LLM at ingest, $0 per message, zero data egress. Gives Claude Code, Cursor, and any MCP client one shared persistent memory with semantic recall, belief revision, selective forgetting, and a provenance guard that blocks acting on stale or unconfirmed memories.
    23
    50 PyPI
    14
    MIT