Skip to main content
Glama

structured-memory-mcp

An MCP server that gives an LLM a persistent, structured brain it reads at session start and writes to as it works, so you never re-explain yourself.

Architecture reference implementation. Python, Postgres + pgvector, about 20 MCP tools, no LLM calls inside the server. Everything in the demo is synthetic.


The problem

Every conversation with an LLM starts cold. The usual fixes all fail in a predictable way:

Fix

Why it breaks

Paste chat history into context

Cost and noise grow forever. Old, superseded statements sit next to current ones with equal authority.

A bigger context window

Same pile, bigger. Nothing says which line is true now.

Plain vector search over everything

Great at "what's similar", blind to "what's current", and it always returns something, even when the honest answer is "nothing".

What's missing isn't capacity. It's structure: knowing which kind of fact you're holding (a rule, a current state, a historical event, a hard number), which one wins when they disagree, and when to say "I don't know."

This repo is one answer: split memory into layers with different jobs, different write rules and different staleness rules, serve them over MCP, and teach the model the protocol on connect.

Related MCP server: @brainfeather/mcp

Architecture

flowchart TB
    subgraph Client["MCP client (any host)"]
        Model["LLM"]
    end

    subgraph Server["MCP server — stdio, or http + bearer token"]
        Tools["tool surface (20 tools)"]
        Service["service layer: validate → repository → format"]
        Repo["repository: the only SQL"]
        Tools --> Service --> Repo
    end

    subgraph DB["Postgres + pgvector"]
        subgraph Present["Present tense: state, overwritten in place"]
            direction LR
            Constraints["constraints<br/>standing rules"]
            Records["records<br/>hard data"]
            Pins["pins<br/>state per lane"]
            Wiki["wiki<br/>compiled truth"]
            Threads["threads<br/>open loops"]
        end
        subgraph Past["Past tense: the trail, append-only"]
            Memory["vector memory<br/>hybrid search + honesty floor"]
        end
        subgraph Meta["Meta"]
            Friction["friction<br/>model failure log"]
        end
    end

    Model -- "session start: get_session_continuation()" --> Tools
    Model -- "mid-session: read / write" --> Tools
    Model -- "session close: set_pin, write_wiki, write_memory" --> Tools
    Repo --> DB
    Friction -. "human triage only" .-> Human(("human"))

The seven layers

Layer

Job

Tense

Write rule

constraints

Standing rules, injected every session

present

Added/retired only when the user says so. Never deleted.

pins

One "where things stand + what's next" slot per lane (work, personal, school)

present

Rewritten in place at session close. State, not history.

threads

Open loops tracked below the lane level; quiet ones surface at session start

present

Status changes record only what the user said. Silence never moves a thread.

wiki

One compiled page per entity; CORE pages are always-on operating instructions

present

Rewritten in place with a version bump. Never appended, never dated snapshots.

vector memory

What happened: dated facts and decisions. Hybrid search (vector cosine + BM25, rank-fused) with an honesty floor

past

Append-only, idempotent by key. One fact per write.

records

Hard structured data: invoices, deadlines, measurements

present

Upsert on (kind, title, date). Beats recall when they disagree.

friction

The model's own failure log, evidence-gated

meta

Written by the model, read and acted on by a human only.

Details: docs/architecture.md.

Trust order, in one line

constraints > records > pins > wiki > vector memory > the model's own recall: and if a lower layer is newer than a higher state layer and disagrees, the answer is contested: the model says so and asks, instead of silently picking. Executable version in brain/trust.py; rationale and staleness rules in docs/trust-and-staleness.md.

Session lifecycle

  1. Start: get_session_continuation() returns, deterministically and with no vector search: hard rules → one pin per lane → QUIET THREADS → full text of CORE wiki pages → wiki index.

  2. During: topic-specific history from read_memory / walk_trail; hard data in log_record; the model touches threads, sets constraints, and logs its own friction when it has a concrete artifact.

  3. Close: set_pin for the lane, write_wiki for changed entities, write_memory for durable facts.

Full protocols: docs/protocols.md. Honest failure modes and what the design does about each: docs/failure-modes.md.

Run it locally (about 2 minutes, no API keys)

Prerequisites: Docker, Python 3.11+.

git clone <this-repo> && cd structured-memory-mcp
make setup     # writes .env with a generated DB password (gitignored), creates .venv
make demo      # starts Postgres+pgvector, migrates, seeds a fictional user, runs the walkthrough
make test      # unit tests + integration tests against the live database

make demo plays a scripted session for a fictional freelance analyst (every name, client and number is invented): the session-start payload, a quiet-thread check-in, a search that hits, a search that correctly finds nothing, both trust-order conflicts, the friction evidence gate, and a session close.

An excerpt of real output:

QUIET THREADS — my picture of these is stale; ask to refresh it (silence is missing data, never a verdict):
  Half-marathon training (personal) — quiet 12d (cadence 7d) · Paused after a knee tweak; ...
  Kestrel dashboard handoff (work) — quiet 6d (cadence 3d) · Handoff doc and walkthrough video ...

5. SEARCH THAT SHOULD FIND NOTHING — the honesty floor
No relevant memories found for this query.

6. TRUST ORDER 1 — a record beats a recalled memory
   winner:  record -> $1,450
   flag:    memory disagrees with record; record wins — flag the memory for correction.

7. TRUST ORDER 2 — a NEWER memory contradicts an older pin
   contested: True  (no silent winner)
   flag:      memory is newer than pin and disagrees — the pin may be stale; ask before asserting.

Connect it to an MCP client

After make up migrate seed (or make demo):

# Claude Code
claude mcp add structured-memory -- /absolute/path/to/structured-memory-mcp/.venv/bin/python \
  /absolute/path/to/structured-memory-mcp/server.py

Any MCP host that can spawn a stdio server works: point it at the venv's Python and server.py. The server reads its secrets by name from .env; nothing is baked into the config.

Configuration

All by environment variable; the names are in .env.example. Highlights:

Variable

Purpose

POSTGRES_PASSWORD / DATABASE_URL

Database access (make setup generates a password)

EMBED_PROVIDER

local-hash (default; offline, demo-grade) · voyage · ollama

VOYAGE_API_KEY

Only for EMBED_PROVIDER=voyage (pip install '.[voyage]')

MIN_SIMILARITY

Honesty-floor override (per-provider defaults in brain/config.py)

TRANSPORT

stdio (default) or http (requires MCP_AUTH_TOKEN)

Embeddings: local-hash is a hashed bag-of-words projection. It tracks lexical overlap, not meaning. It exists so the demo and tests run offline. Use a real embedding model for anything real, and recalibrate the floor when you do (docs/failure-modes.md).

What's in the repo

brain/          the server: layers (pure logic), service, tools, repository, search, trust order
migrations/     schema only. No data rows ship in this repo.
demo/           synthetic data for a fictional user + the scripted walkthrough
tests/          pure-logic unit tests, plus integration + MCP-protocol tests (need the DB)
docs/           architecture, trust & staleness, protocols, failure modes, blog draft

Design choices worth a look: pure formatting/selection logic is separated from SQL so every layer's rules are unit-testable without a database; the server makes no LLM calls (retrieval and continuation are deterministic); the honesty floor and the multiplicative re-rank exist because of specific failures documented in docs/failure-modes.md.

Not included (on purpose)

A web dashboard, OAuth, multi-user isolation, automatic contradiction detection across the store, and background "dreaming" jobs.

Writeup

docs/blog/why-structured-memory.md: why structured memory beats dumping chat history into context.

License

MIT

Available Tools

20 tools
get_session_continuationA

CALL THIS FIRST in every conversation. Returns, in order: hard-rule constraints, one pin per lane, QUIET THREADS (if any), the full text of every CORE wiki page, and the wiki index. Deterministic — no vector search.

channel: optional — return only that lane's pin.

ParametersJSON Schema
NameRequiredDescriptionDefault
channelNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and largely discharges it: it discloses the exact ordered contents of the response and that retrieval is deterministic with no vector search. It omits permission/auth requirements and any note about cost or side effects, but for a read tool this is solid disclosure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the urgency directive, then a compact enumerated return list, then the parameter note. Every sentence carries information; only the capitalized domain jargon (QUIET THREADS, CORE) slightly taxes a first-time reader.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be spelled out, yet the description still summarizes them helpfully. Combined with the strong usage directive and parameter note, it is complete enough for an agent to call correctly, though it lacks any auth or failure-mode context.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate for the single parameter, and it does: 'channel: optional — return only that lane's pin' explains the filter and its opt-in nature. It is terse but sufficient for a one-param tool.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb (Returns) and enumerates the exact resources returned in order (hard-rule constraints, one pin per lane, QUIET THREADS, CORE wiki pages, wiki index), and adds a distinctive trait ('Deterministic — no vector search'). An agent can immediately distinguish this bundle-retrieval tool from siblings like load_wiki or read_memory.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'CALL THIS FIRST in every conversation' is an explicit, unambiguous when-to-use directive, reinforced by the deterministic-vs-vector-search contrast that differentiates it from the search_* siblings. No when-not is stated, but the first-call directive is about as strong as guidance gets.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_constraintsC

List standing rules. Retired rules are history, not truth.

ParametersJSON Schema
NameRequiredDescriptionDefault
include_inactiveNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

C2.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full load. It does disclose one real behavioral trait: retired rules exist and are excluded from the default listing. It says nothing about permissions, result size, ordering, or pagination, which an agent calling a list tool would want.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, front-loaded with the action, with no filler. The aphoristic second sentence is memorable and does work, but is slightly obscure in a way that trades precision for style.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be described. For a single-parameter list tool that is roughly adequate, but the description leaves an agent unsure what a 'constraint' is, how many there are, or how this relates to the surrounding constraint tools.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the parameter 'include_inactive' is undescribed. The sentence 'Retired rules are history, not truth' implies what 'inactive' means and that it is off by default, partially compensating, though the framing is exhortative rather than a clear definition of the flag's effect.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a verb+resource ('List standing rules'), but renames the resource as 'rules' rather than 'constraints' and never defines what a constraint actually is. Against siblings like list_entities, list_threads, and list_records it does not make clear why this list is different beyond the active/retired distinction.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit statement of when to call this tool versus alternatives such as set_constraint, retire_constraint, or read_memory. The second sentence only hints, obliquely, that retired rules exist and are probably not wanted by default, which is guidance about the parameter rather than about selecting the tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_entitiesC

List every wiki page: folder, name, version, last update, headline.

ParametersJSON Schema
NameRequiredDescriptionDefault
include_archivedNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral disclosure burden. It implies a read operation but does not state read-only safety, permissions, pagination behavior, or how include_archived changes the result set; the word 'every' may even mislead given the parameter's default false.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single front-loaded sentence with no filler, and it efficiently enumerates the key return fields. It is appropriately sized for the surface-level purpose it communicates.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Because there are no annotations and the only parameter is undocumented, the description is not complete enough for confident invocation. The output schema covers return values, but usage context and include_archived semantics are missing, leaving significant gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The sole parameter, include_archived, has 0% schema description coverage and is never mentioned in the description. The description adds no meaning about what the flag controls or when it should be used, leaving the agent unable to invoke the tool with the parameter correctly understood.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a clear verb and resource: 'List every wiki page,' and enumerates the returned fields (folder, name, version, last update, headline). It does not, however, differentiate this listing tool from siblings like load_wiki or clarify whether it is a complete dump versus a filtered view.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no explicit alternatives, and no conditions for choosing list_entities over load_wiki or any other sibling. The agent is left to infer usage from the tool name and description alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_frictionB

The HUMAN's triage read path. open | triaged | fixed | dismissed | "" for all.

ParametersJSON Schema
NameRequiredDescriptionDefault
statusNoopen

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries full behavioral burden. 'Read path' weakly implies a safe read-only operation, but there is no disclosure of pagination, permissions, ordering, or side effects. The description does not contradict anything, but it leaves key behavioral traits unstated.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The definition is very short and front-loads the purpose before the parameter values. It avoids wasted words, though the terse 'HUMAN's triage read path' phrasing is slightly cryptic rather than perfectly clear.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained. The tool is simple with one optional parameter, and the description supplies the filter values. Still, the purpose and usage context are thin enough that an agent must infer some details from the name and siblings.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must explain the single parameter. It lists the valid status values ('open | triaged | fixed | dismissed') and clarifies that an empty string means all, which is exactly the missing enum/semantic detail the schema lacks.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description calls this 'The HUMAN's triage read path,' which implies a read/list operation but never explicitly states that it lists friction items. Combined with the tool name, an agent can infer the purpose, but the description alone is vague about the resource and verb.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'The HUMAN's triage read path' gives an implicit context for when to use this tool, and the status values suggest filtering. However, there is no explicit guidance on when to choose this over siblings like triage_friction, log_friction, or list_records.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_recordsB

List records, newest first, optionally by kind and ISO date range.

ParametersJSON Schema
NameRequiredDescriptionDefault
kindNo
limitNo
to_dateNo
from_dateNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full disclosure burden. It reveals ordering ('newest first'), which is genuinely useful, but says nothing about read-only safety, side effects, permissions, pagination, or what the returned set's bounds are – thin coverage for a no-annotation tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One sentence, front-loaded with the core action and ordering, then the optional filters. Zero waste; every clause carries information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained. However, with no annotations and a 50-default limit that is never referenced, the definition leaves an agent guessing about result-set size and pagination behavior.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the schema only carries titles. The description gives meaning to kind and an ISO format hint for the date range, but 'limit' is entirely undocumented in both schema and description, and valid 'kind' values remain unspecified.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (List) and resource (records), plus two behavioral facts: default ordering (newest first) and the optional filters. It does not distinguish itself from siblings like list_entities, list_constraints, or list_friction, but the purpose is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied by 'optionally by kind and ISO date range' – an agent can infer that calling with no arguments lists everything. There is no explicit statement of when to use this versus log_record or the other list_* siblings, and no exclusions or prerequisites.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_threadsB

List tracked threads with lane, status, last-touched age and cadence.

ParametersJSON Schema
NameRequiredDescriptionDefault
include_closedNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the burden, and it does disclose the returned attribute set (lane, status, age, cadence). However, it says nothing about the read-only nature, pagination, ordering, or how the default closed-thread filtering behaves.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One front-loaded sentence that states the operation and the returned fields with no filler. It is efficient, though arguably too sparse given the undisclosed parameter.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained, and for a simple list tool the bare description is minimally viable. The gaps are the default filtering behavior of include_closed and any guidance on choosing this tool over the other list_* siblings.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% and the single parameter include_closed is never mentioned in the description. The phrase 'tracked threads' gives no clue that closed threads are excluded by default, so the agent must infer the parameter's meaning from the schema name alone.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Specific verb (List) + resource (tracked threads) plus the fields surfaced (lane, status, last-touched age, cadence). It clearly reads as a read operation and is distinguishable from the write siblings (track_thread, touch_thread), though it never names them explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No statement of when to use this versus alternatives such as list_entities, list_constraints or walk_trail, and no conditions or prerequisites. Usage is only inferable from the tool name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

load_wikiA

Load compiled wiki pages by exact name — deterministic. Misses suggest near matches.

ParametersJSON Schema
NameRequiredDescriptionDefault
entitiesYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the burden and does disclose meaningful behavior: exact-match-only semantics, determinism, and near-match suggestions on a miss. It does not cover permissions or listing/ordering behavior, but the failure-mode disclosure is valuable beyond structured fields.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short clauses, front-loaded with the action and its defining constraint, then the failure behavior. No wasted words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be described. For a single-parameter lookup tool, the description covers input semantics and miss handling adequately; only edge details like multi-entity partial-miss behavior are absent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

One parameter ('entities', an array of strings) with 0% schema coverage. The phrase 'by exact name' clarifies that the items are exact page names, which adds meaning, but the array semantics (one vs many, batching, ordering) are left unstated.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Load) and resource (compiled wiki pages) with a key constraint (by exact name). It implicitly contrasts with write_wiki in the sibling set, though it doesn't name a sibling explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'By exact name — deterministic' implies the appropriate usage context and 'Misses suggest near matches' tells the agent what happens on failure, but no alternative tool or precondition is named. Usage is implied rather than stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

log_frictionA

Log YOUR OWN failure in the moment, with a concrete artifact as evidence (the failed result, the user's exact correction). No artifact = do not call. Zero friction is a normal session. Never read this lane to change your behavior.

category: stale_protocol | wrong_tool | bad_retrieval | user_correction | missing_context | tooling_error | other

ParametersJSON Schema
NameRequiredDescriptionDefault
summaryYes
categoryYes
evidenceYes
proposed_fixNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does meaningful work: it states the evidence precondition, warns the lane is not for behavioral self-modification ('Never read this lane to change your behavior'), and implies write-only semantics. It omits permissions, immutability, dedup behavior, and what logging triggers downstream.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Tight and front-loaded: the core trigger and the hard precondition land in the first two sentences, followed by the normal-case note, the warning, and the category values. Slightly telegraphic phrasing ('this lane') costs a little clarity but nothing is padding.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained, and the description covers the behavioral essentials of what/when/when-not. Given a 4-parameter, 0%-coverage schema, more parameter-level detail would round it out, but nothing critical for correct invocation is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It usefully enumerates the allowed category values (stale_protocol, wrong_tool, bad_retrieval, user_correction, missing_context, tooling_error, other), which the schema does not, but leaves summary, evidence, and proposed_fix semantics (length, format, what counts as artifact) undocumented.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource: logging the agent's own failure in the moment with a concrete artifact as evidence. It implicitly contrasts with the read side of the lane ('Never read this lane'), which helps separate it from list_friction/triage_friction, but it never names those siblings or explicitly distinguishes itself from log_record.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives an explicit trigger (a failure in the moment, with artifact) and an explicit exclusion ('No artifact = do not call'), plus a normal-case note ('Zero friction is a normal session'). It stops short of naming alternative tools for the read/triage side, so it is clear context rather than full when/when-not/alternatives guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

log_recordB

Store hard data (invoice, deadline, measurement…). Records beat recall: when one disagrees with a remembered fact, the record wins. Upserts on (kind, title, date).

ParametersJSON Schema
NameRequiredDescriptionDefault
kindYes
titleYes
amountNo
occurred_onYes
payload_jsonNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description must carry the behavioral burden. It usefully discloses the upsert key (kind, title, date), which tells the agent existing records with that key will be overwritten. However, it says nothing about permissions, conflict resolution details, or what happens to non-key fields on upsert.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences with the core purpose front-loaded and examples inline. The value proposition sentence earns its place by clarifying precedence over memory, though it could be trimmed slightly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need no explanation. But for a 5-parameter mutation tool with zero annotation and schema-description coverage, the description covers only three parameters and omits amount/payload_json semantics and error behavior, leaving real gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate for all five parameters. It only references kind, title, and the date via the upsert key, leaving amount and payload_json completely unexplained, which is a significant gap for a 5-param tool.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (store/log) and resource (hard data records) with concrete examples like invoice, deadline, and measurement. It is distinguishable from memory-oriented siblings such as write_memory and log_friction, though it never names them explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The line 'Records beat recall: when one disagrees with a remembered fact, the record wins' implies when this tool should take precedence over memory tools, but it gives no explicit when-to-use or when-not-to-use guidance relative to siblings like write_memory or log_friction.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

read_memoryA

Hybrid search (vector cosine + BM25, rank-fused, importance/recency re-ranked) over the trail. Use for topic-specific HISTORY — not for current state (that's pins/wiki). Says "No relevant memories found" when nothing clears the honesty floor; believe it.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
queryYes
domainNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the behavioral burden and discloses meaningful traits: hybrid search mechanics, rank fusion, importance/recency re-ranking, and the honesty-floor message when nothing relevant is found. It does not explicitly state read-only behavior or permission requirements, but the core search behavior is well covered.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tightly written sentences front-load the search mechanics and scope before moving to usage guidance and the honesty-floor behavior. Every sentence carries useful information with no wasted wording.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers purpose, usage, and behavior well, and an output schema exists so return values need not be described. However, with 0% schema description coverage for three parameters, it remains incomplete regarding how to use query, limit, and domain correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate for three undocumented parameters. It barely implies that the query should be topic-specific but says nothing about limit or domain, leaving most parameter semantics unexplained.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific operation — hybrid search over the trail — and clarifies it is for topic-specific history rather than current state. It distinguishes itself from pins/wiki, but it does not differentiate from the sibling search_memory, which is likely the closest alternative.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives explicit usage guidance: use for topic-specific HISTORY, not for current state (pins/wiki). This is clear context and an exclusion, but it omits the direct alternative for historical memory search, such as search_memory.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

retire_constraintB

Retire (never delete) the ONE active rule matching the fragment.

ParametersJSON Schema
NameRequiredDescriptionDefault
rule_fragmentYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden, and it does disclose two real traits: retirement is a soft operation ('never delete'), and only active rules are eligible. It omits what happens on zero or multiple matches, whether the change is reversible, and any permission requirements, so the disclosure is partial.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler; the parenthetical 'never delete' is placed exactly where it matters and the emphasis on 'ONE' flags the uniqueness constraint efficiently.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return values need not be described, and one parameter keeps the surface small. Still, for a mutation tool with no annotations, the description leaves failure modes (no match, multiple matches) and reversibility undefined.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% and the sole parameter is just typed as a string, so the description is the only source of meaning — and it does add that the value is a 'fragment' matched against rules rather than an exact identifier. It does not explain matching rules, case sensitivity, or ambiguity handling, so it only partially compensates.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Names a specific verb (retire) and its object (the one active rule matching the fragment), which is enough to separate it from set_constraint and list_constraints. It stops short of explicitly naming a sibling or contrasting behavior, so it is clear but not fully differentiated.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No explicit statement of when to use this versus set_constraint or list_constraints, and no prerequisites or preconditions are given. The 'ONE active rule' phrasing implies a match must be unique, but the agent is left to infer this rather than being told.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_memoryC

Broader sweep with the same ranking as read_memory (more results).

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
queryYes
domainNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and delivers little: it discloses that ranking matches read_memory and returns more results, but says nothing about pagination, result caps, ordering guarantees, or whether the memory store is read-only. The one comparative fact is genuinely useful, but the behavioral picture is thin.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single short sentence with no filler, and the differentiating comparison is front-loaded. It is efficient; the brevity is partly a virtue and partly the source of the gaps scored elsewhere.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained, but for a three-parameter search tool with zero schema descriptions and no annotations the definition leaves the agent guessing about parameters and scope. It is not complete enough to invoke confidently.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across three parameters (query, limit, domain), so the description must compensate and largely does not. 'More results' weakly gestures at the limit parameter, but query and domain—including what a domain scopes to—are completely unexplained anywhere.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description positions the tool relative to a sibling ('broader sweep... same ranking as read_memory') and hints at result volume, but never states the verb+resource plainly (e.g., 'search stored memories by query'). An agent must infer the purpose from the name and the read_memory comparison rather than reading it.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'broader sweep ... (more results)' implies this is the wider-coverage alternative to read_memory, which is a usable selection hint. However, it never states an explicit when-to-use/when-not condition, prerequisites, or how broad the sweep actually is.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

set_constraintB

Add a standing rule when the user states one ("from now on…", "never…").

ParametersJSON Schema
NameRequiredDescriptionDefault
ruleYes
categoryNogeneral

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It implies persistence via "standing rule" but says nothing about precedence, conflict handling, whether existing rules are superseded, or whether confirmation is needed. For a state-mutating tool this is a significant disclosure gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single tight sentence with the action front-loaded and the trigger condition immediately after. Nothing is wasted or buried.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The output schema covers return values, so that need not be explained. However, with no annotations and 0% parameter coverage, the description leaves the category argument and the mutation's behavioral consequences undocumented.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% for two parameters, so the description must compensate and does not. The required "rule" string is only loosely implied ("the user states one"), and the "category" parameter with its "general" default is never mentioned.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

"Add a standing rule" names a specific verb and resource, so the agent knows this creates a persistent constraint. It is not explicit about how it differs from family siblings like retire_constraint and list_constraints, but the create/read/delete split is inferable.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives a concrete trigger condition with example phrasings the user might use ("from now on…", "never…"). It stops short of naming alternatives — e.g. that retire_constraint removes a rule — so it is clear context without explicit exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

set_pinA

Session close: set the pin for ONE lane (work | personal | school). Rewrites that lane in place with the full "where things stand + what's next" state — not a delta. Durable facts belong in write_memory; the pin is state, not history.

ParametersJSON Schema
NameRequiredDescriptionDefault
channelYes
contentYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the burden and largely does: it discloses that the write is a full in-place rewrite of one lane rather than a delta or append, which is the key destructive-behavior fact. It stops short of permission requirements or what happens to sibling lanes / an absent pin, so not a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three dense sentences, front-loaded with the trigger and the object, each carrying distinct information (trigger, scope, semantics, sibling routing). No filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained, and the description covers trigger, scope, write semantics, and the memory-vs-state boundary. Minor gaps remain around overwrite edge cases (missing lane, other lanes' fate) that a mutation tool with no annotations could generously include.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate, and it does: it reveals that 'channel' is one of work|personal|school (an enum the schema omits entirely) and that 'content' must be full state, not a delta. It adds real meaning beyond the bare 'string' schema, though it doesn't define channel format for content syntax.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Specific verb+resource ('set the pin for ONE lane') with scope made explicit (single lane, enumerated values). It also distinguishes itself from write_memory, so an agent can route between a state pin and durable memory without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

States the trigger ('Session close') and the exclusion ('Durable facts belong in write_memory'), naming the alternative and the condition that selects it. This is exactly the when/when-not guidance that makes a writer tool selectable.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

touch_threadB

Mark a thread discussed NOW. Set status ONLY if the user ruled on it — status records what they said, never what silence suggested.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYes
noteNo
statusNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are supplied, so the description carries the full burden. It discloses the semantic of touching a thread as 'NOW' and the inferential rule for status, but says nothing about what happens to the note field, idempotency, or whether prior timestamps are overwritten.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences, front-loaded with the action and followed by the constraint. The emphasized status rule is verbose but earns its place as the tool's key behavioral guardrail.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained. Still, for a mutation tool with zero annotations and 0% parameter coverage, the description leaves name/note behavior and mutation side effects unspecified.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It only clarifies status semantics ('records what they said, never what silence suggested'); name and note are completely undocumented in both schema and description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Mark a thread discussed NOW'), which is clear intent, but does not distinguish itself from the sibling track_thread, so an agent can't tell the two apart from the description alone.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides a real usage rule for the status parameter ('Set status ONLY if the user ruled on it'), which is genuinely useful guidance. However, there is no when-to-use-this-vs-alternatives guidance, notably nothing distinguishing it from track_thread.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

track_threadB

Create/update a thread when the user names an ongoing project or open loop. cadence_days: quiet days before it surfaces (1-365; 0 = never surfaces on silence). status: live | parked | dropped | done. wake_date: ISO date for parked threads.

ParametersJSON Schema
NameRequiredDescriptionDefault
laneYes
nameYes
noteNo
statusNolive
wake_dateNo
cadence_daysNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations, so the description carries the full burden. It usefully explains that cadence_days controls surfacing after quiet days and lists status values, but it never clarifies create-vs-update (upsert) semantics, reversibility, or what happens to existing fields on update.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core action, then terse per-parameter notes. Slightly telegraphic ('0 = never surfaces on silence') but every line earns its place with no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return values need not be described. Still, for a mutation tool with no annotations and 0% schema coverage, the description leaves half the parameters and the create/update distinction unexplained.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It defines cadence_days range/semantics, status enum values, and wake_date, covering three of six parameters, but leaves lane, name, and note entirely undefined.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb pair (create/update) and resource (thread), plus a trigger condition distinguishing it from list_threads/touch_thread. It does not explicitly name how it differs from the sibling touch_thread, but a reader can infer it.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The trigger is stated ('when the user names an ongoing project or open loop'), which is clear usage context. However, no alternatives are named and no when-not-to-use guidance is given, so the agent must infer when touch_thread or list_threads is preferable.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

triage_frictionA

Record the HUMAN's spoken verdict (triaged | fixed | dismissed). A model may never mark its own failure fixed.

ParametersJSON Schema
NameRequiredDescriptionDefault
statusYes
friction_idYes
triage_noteNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does disclose the most important behavioral trait: the verdict must originate from a human, and the model cannot self-apply 'fixed'. It omits permissions, idempotency/overwrite behavior, and what happens to an existing triage record, keeping it out of 5 territory.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, both front-loaded with the operative information (what is recorded, then the hard constraint). No filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return values need not be described, and the key human-only constraint is stated. The gap is the two undocumented parameters (friction_id, triage_note) and no indication of the triage workflow relative to log_friction/list_friction.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and no enums are declared in the schema, so the description's enumeration 'triaged | fixed | dismissed' adds genuinely new meaning for the status parameter. However friction_id and triage_note are left entirely undocumented, so it only partially compensates for the coverage gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Record') and resource ('the HUMAN's spoken verdict'), and even enumerates the accepted status values. It is distinguishable from log_friction/list_friction by its triage-verdict framing, though it never names those siblings explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The second sentence gives a sharp usage constraint: a model may never mark its own failure fixed, which tells the agent exactly when it must NOT call this tool unilaterally. It stops short of naming an alternative or describing the human-handoff flow, so it is clear context rather than full when/when-not routing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

walk_trailA

Walk the trail in date order, oldest first, no scoring. For "what happened in March" / "when did X change". ISO dates, inclusive.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
domainNo
to_dateNo
from_dateNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden, and it does disclose meaningful traits: chronological ordering, oldest-first, and explicitly 'no scoring'. However, it says nothing about pagination/limit behavior, permissions, or the read-only nature an agent would want to assume for a mutation-capable API family.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences with no filler. The core behavior (order, no scoring) is front-loaded and the use-case cue follows immediately; every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be described. But with zero annotations and 0% schema description coverage across 4 parameters, the description should do more about limit/domain semantics and the tool's safety profile than it does.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate. It clarifies the two date parameters usefully ('ISO dates, inclusive') and implies chronological ordering, but leaves 'limit' and 'domain' entirely unexplained, so half the parameters remain opaque in both schema and description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('walk the trail') and adds distinguishing behavioral detail ('in date order, oldest first, no scoring'). The term 'trail' is domain jargon and no sibling is named, so an agent must infer how this differs from list_records or search_memory, but the operation is otherwise clear.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides concrete triggering questions ('what happened in March' / 'when did X change') that map the tool to a clear use case. It stops short of naming alternatives or saying when NOT to use it (e.g., vs. list_records or search_memory), so it is context without exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

write_memoryA

Persist one durable fact or decision. One specific fact per call, under 300 words.

domain: free-form short label, e.g. decisions | people | projects | general importance: 9-10 decisions · 7-8 relationship/strategy updates · 5 context · 3 ephemeral idempotency_key: unique slug; the same key twice is a safe no-op entities: wiki page names this fact touches (powers topic->page push)

ParametersJSON Schema
NameRequiredDescriptionDefault
domainNogeneral
contentYes
entitiesNo
importanceNo
idempotency_keyNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and discloses meaningful traits: idempotency_key makes a repeated call a safe no-op, content is limited to 300 words, and entities power a topic-to-page push. It does not cover permissions or overwrite semantics, but the key behavioral details are present.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with the core purpose, then uses compact labeled lines to explain each parameter's semantics. Every sentence earns its place with no repetition or filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no annotations, 5 parameters, 0% schema coverage, and an existing output schema, the description provides enough to invoke the tool correctly: it defines most parameter semantics and important behaviors like idempotency and word limit. It stops short of explicitly naming the content parameter and lacks sibling routing guidance, but is largely complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, and it does for four of five parameters: domain gets examples, importance gets a numeric interpretation, idempotency_key explains safe duplication, and entities explains linkage. The content parameter is only implicitly defined via 'one specific fact per call, under 300 words', which keeps it from a 5.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'Persist one durable fact or decision' and scopes it to one fact per call under 300 words. It does not explicitly differentiate from sibling tools like write_wiki or log_record, so it falls short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied by 'Persist one durable fact or decision' and the importance scale, but there is no explicit when-to-use guidance or comparison to alternatives such as write_wiki or log_record. The domain examples and importance mapping are helpful parameter guidance rather than selection criteria.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

write_wikiB

Create or rewrite a page IN PLACE (synthesized current truth, never an appended log). folder="CORE" marks an always-on operating page injected in full every session.

ParametersJSON Schema
NameRequiredDescriptionDefault
entityYes
folderNo
contentYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations exist, so the description carries the full behavioral burden. It usefully discloses that writes are in-place rewrites (implying prior content is replaced) and that folder=CORE injects a page into every session, which is non-obvious behavior. It omits permissions, reversibility, and the consequence for existing page content.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, front-loaded with the core create/rewrite semantics before the folder special case. Dense and free of filler, though the capitalized 'IN PLACE' and parenthetical add mild visual noise.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be described. For a mutation tool with no annotations and 0% schema coverage, the description covers overwrite semantics and the CORE folder behavior but leaves key gaps: what happens to overwritten content, and how entity/content are expected to be formed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate for all three parameters. It explains the special value 'CORE' for folder, adding real meaning, but leaves entity and content entirely undefined — no format, naming convention, or size expectations.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb pair (create/rewrite) and resource (wiki page), and the parenthetical 'synthesized current truth, never an appended log' sharply defines the write mode, distinguishing it from log-style siblings like log_record and write_memory. It does not explicitly name those siblings, so it stops short of 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The 'never an appended log' clause implies when this tool is appropriate versus append-style tools, and the folder=CORE note signals a special use case. However, no alternative tool is named and no conditions or prerequisites are given for choosing write_wiki over write_memory or set_constraint.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 20 tool updatesv0.1.0
    • First observedget_session_continuation
    • First observedlist_constraints
    • First observedlist_entities
    • First observedlist_friction
    • First observedlist_records
    • First observedlist_threads
    • First observedload_wiki
    • First observedlog_friction
    • First observedlog_record
    • First observedread_memory
    • First observedretire_constraint
    • First observedsearch_memory
    • First observedset_constraint
    • First observedset_pin
    • First observedtouch_thread
    • First observedtrack_thread
    • First observedtriage_friction
    • First observedwalk_trail
    • First observedwrite_memory
    • First observedwrite_wiki

TDQS

B3.4/5.0

Scored across 20 tools

Disambiguation4/5

Most tools target distinct resources and actions, but read_memory and search_memory overlap heavily (both hybrid search the trail), and write_memory vs write_wiki could be confused about where durable information belongs. The detailed descriptions help, but the boundaries are not perfectly crisp.

Naming Consistency5/5

All tool names use snake_case with a consistent verb_noun pattern (load_wiki, retire_constraint, write_memory, list_constraints, track_thread, etc.). The convention is predictable throughout with no mixed styles.

Tool Count3/5

20 tools is heavy for a single MCP server and sits in the borderline range. While each tool covers a distinct facet (memory, wiki, constraints, threads, friction, records, session state), the surface could likely be consolidated to reduce selection overhead.

Completeness4/5

The server covers core lifecycle operations across its subdomains: memory read/write/search/walk, wiki load/write/list, constraints list/set/retire, threads list/track/touch, friction log/list/triage, and records log/list. Missing hard delete/update for memories, records, and wiki pages may be intentional by design, but still leaves minor gaps.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • F
    license
    A
    quality
    C
    maintenance
    Provides an external context buffer so a local model can maintain session memory across long tasks and search and page through large documents without loading them whole into its limited context window.
    15
    1
    -
  • A
    license
    Not graded
    quality
    A
    maintenance
    Enables MCP agents to maintain durable, evidence-aware project knowledge, retrieve precise excerpts on demand, and track decisions, conflicts, and revisions across sessions.
    1
    Apache 2.0
  • A
    license
    Not graded
    quality
    C
    maintenance
    Enables multiple MCP-compatible AI clients to share persistent, versioned project knowledge across sessions with conflict-safe updates, provenance, hybrid retrieval, stale-memory handling, and context-budgeted recall.
    MIT