Skip to main content
Glama

EMET

Enduring Memory & Epistemic Tiers — a multi-agent, multi-access memory server for the Model Context Protocol.

EMET stores agent memory in MongoDB Atlas across six layers, keeps a separate word-for-word transcript tier, and verifies every write by reading it back. It runs as a local stdio server and as a remote HTTP server from the same source, so it works as a bridge: AI models on several hosts — desktop, CLI, browser, phone, another vendor's agent — read and write one shared memory, and the work, checks, and decisions one leaves behind are there for the next. That includes Grok Bot teams and single-session Grok chats connected to the same server: what the bots save is there for a fresh Grok chat to find once it initializes, and the reverse. To connect a single-session chat (a Grok project, or any chat host that takes a remote MCP connector and supports its OAuth sign-in), add your deployed server's URL, https://<your-host>/mcp, as a remote MCP connector there and complete the sign-in with the passphrase you set (EMET_OAUTH_PASSPHRASE; see Connecting clients). A connector that only takes a fixed header or key will not work. Then put the instruction block from docs/BOOTSTRAP.md in the project instructions: the chat reaches for that memory most reliably, unprompted from the first turn, with the block there, or, as a fallback where there is no such field, with it pasted as the first message. The store has one owner; other people can join only through the optional Guest (read access by default) and Member (read and write) access, both off until turned on.

Status: 0.x. The stored document shapes are not frozen yet. Breaking changes to the schema may land before 1.0. Pin a version if you depend on it.


Two ways in

A session that has not initialized can read everything and can write only a new drop (drops/<date>-<slug>.md): its items, wrap-up and handoff. A session with a registered source and the full startup has full access. Every initialized session sees unhandled drops first. A board edit is refused only for a drop that was listed at that session's boot. Drops are capped at 20 per source per day, and they are never deleted.

Related MCP server: KusiScribe Core

Editing working state

The server provisions a blank board at setup. No session can write the board. It is edited with patch_doc and a reason; an added row names its origin. Once any working document exists, it is edited, not rewritten. A change that leaves a document under half the size it had when this session first touched it is refused. Working documents are not retired. restore_doc is the one whole-document recovery, for an owner-listed tag, and only a stored conforming revision whose digest matches.

The name

EMET is אמת — Hebrew for truth. It is also an acronym for what the server does: Enduring Memory & Epistemic Tiers.

The choice is not decoration. In the golem legend the word emet is inscribed to animate the body, and erasing its first letter leaves מת — met, dead. One character is the whole difference.

That is why the verification layer is called aleph. Every write in EMET is read back and compared before it is reported as saved. That read-back is the aleph: a memory system without it still returns answers, still looks healthy, and is quietly a corpse — confidently serving things it never actually stored. Aleph is also the first letter, which matches the order the protocol insists on: verify before retrying, verify the premise before applying the remedy, read back before trusting.

The same image explains what setup is for. A golem is inert until the word is written on it; a fresh model instance is unformed until the record is written into it. emet_initialize is that inscription — it hands the model an identity, the person it serves, where things stood, and the disciplines it is bound by. Nothing mystical is claimed here. It is a literal description of what the first tool call does.

Both halves of the acronym are load-bearing.

Enduring is not a modest claim, and it is not a feature competing with other memory products. It is measured against what a language model is by default: an etch-a-sketch. Every session starts from nothing, and ends by forgetting everything. You explain yourself again, re-establish what you are working on again, re-derive decisions you already made. Enduring names the difference between an assistant that knows you and one meeting you for the first time, again. That is the character of the thing, not a marketing adjective attached to it.

Epistemic Tiers is how that endurance stays trustworthy. Tier 1 holds curated claims; tier 2 holds the verbatim record that justifies them — justified belief and its evidence, kept structurally separate so neither can quietly overwrite the other. Memory that persists but cannot be checked is just confident noise that survives longer.


One memory, every device

This is the part worth understanding before anything else.

Most MCP memory servers run as a local process: your desktop launches it over stdio, and the memory lives wherever that process can reach. That works until you pick up your phone.

EMET ships two entry points over one source tree:

src/index.js

stdio

desktop and CLI, launched by the client

src/http-remote.js

streamable HTTP + OAuth

phone, tablet, browser, another vendor's agent

They are not two projects. Both import the same tools.js, database.js, validation.js and content_analyzer.js — the transport is the only thing that differs. There is no fork to keep in sync, no second copy drifting out of date, no "mobile version" lagging the real one. A fix lands once and both transports have it.

Deploy the HTTP entry point to any host that can run a container — the Dockerfile is included, and docs/DEPLOY.md walks through Railway and Fly — and every device you own reaches the same store. Write a decision from your laptop, ask about it from your phone an hour later, and it is simply there. Add a third host and it joins the same memory rather than starting its own.

The remote endpoint is OAuth-protected: a device is authorized once with a passphrase you set, and the token persists across restarts and redeploys. Leave EMET_SOURCE_TAG unset on a server several hosts share: each host passes its own tag as source on its writes, and you register the tags in EMET_SOURCE_TAGS (see .env.example), so every entry records which machine produced it — with several hosts writing to one store, provenance stops being optional bookkeeping and becomes the thing that lets you trust what you read.

Running cost for a single user is a small MongoDB cluster plus a hobby-tier container host. Both have free tiers adequate to start.


Best practices

Split the work by what each host does best. A single-session chat (for example, a Grok project chat, connected as described at the top) only works when you send it a message. It can't watch EMET, and bots can't message it. So when a bot team needs the chat's help:

  1. The bots leave a task in EMET with everything needed to do it, in a document whose id they give you.

  2. You send the chat one line that names that document, for example: "Initialize EMET, read threads/open-tasks.md, do the task, and save the result to EMET."

  3. The chat saves its draft to EMET.

  4. The bots check it.

Good fits for the chat: long drafting and review done in one sitting. Good fits for the bots: work in the repository, on the server, on a schedule, or in accounts. If the chat runs on a separate plan, its work draws on that plan's allowance, not the bots'.

Fetch every startup page, then open the session with that startup's load_id. Call emet_initialize {} first and fetch every page it names; ask for {"page": "all"} only on a host known to read long results whole. Pass the load_id from those pages to emet_session_open. An open with no load_id, with a page missing, or with a load_id the server has no record of (one from before a restart, or 60 minutes past its last page) is refused, and the reply says what to call. The page check runs on your own startup, and a skipped page is refused before anything is written. See docs/BOOTSTRAP.md and docs/BOTS.md.

Set the source-tag list before more than one host or bot writes. The guards that refuse an unregistered tag are off unless EMET_SOURCE_TAGS is set; with it unset, any caller can write under any tag. Add each host's tag to the list before that host's setup interview. With several hosts writing one store, the tag is what lets you trust who wrote what. See docs/SETUP.md.

Put the instruction block in the project's instructions. Connecting the server gives the assistant the ability to remember; it does not make it use that ability. A host with an instructions field is the strongest integration: paste the block from docs/BOOTSTRAP.md there. Where there is no such field, paste it as the first message of a session. With the block in place, the assistant reaches for memory most reliably from the first turn, instead of answering from what it can infer and saving nothing.

Approve reads standing, and keep an eye on the three writes that replace a document. Reads cost nothing and cannot be wrong destructively, so a reasonable default for most users is standing approval for reads. Additive writes are recoverable, because a correction supersedes what it corrects rather than overwriting it. The three replacing writes archive the prior version, but a wrong one is silently a different document rather than an error. Do not carry this posture to a filesystem or shell server. See docs/PERMISSIONS.md.

Reconnect each host after an upgrade that changes the tool list. Hosts cache the tool list: after a deploy that adds or renames tools, each host keeps offering the old set until you reconnect the connector there. New tools stay invisible until the host sees the new list, and the host may ask again for approvals you had already given. See docs/DEPLOY.md.

Register a bot before its first write, and set both lists in one change. One bot, one tag. In the same deployment, add the tag to EMET_SOURCE_TAGS and append bots/<tag>/HANDOFF.md to EMET_RECORD_DOCS. A registered tag that is not listed as a bot can still write the main records. The hub's tag is registered as a host and is never listed as a bot. See docs/BOTS.md.

Prove the write path before you trust a new deploy. Run emet_status with probe: true. It writes a test document, reads the same value back and removes it; a read alone cannot show a broken write path. A role missing a permission fails there, loudly, instead of in the middle of a session. Then save something from one device and recall it from another. See docs/SETUP.md.

Keep house style in the store, not pasted into every project. A document written with write_doc is read by every host and changed in one place; text pasted into three projects is three copies that will disagree. Delivery preferences belong in bootstrap/USER.md, not in one vendor's own memory, which does not travel to the next host. See docs/BOOTSTRAP.md.


Attribution

EMET is a derivative of CASCADE Enterprise RAM by Jason Glass / C.I.P.S. LLC, published at github.com/For-Sunny/cascade-memory-enterprise, used and redistributed under the MIT license the project adopted at commit acdb5c8c99cdfc29fb44db2cd828d7752be744db (v2.2.2, 2026-02-19). Both copyright lines are retained in LICENSE.

Link unavailable: as of 2026-10-03 that repository no longer resolves (GitHub answers 404), and no moved copy, public mirror or archive of it was found. The address and commit are kept here as the record of where this code came from.

How much is upstream, measured rather than estimated. Line-level comparison of this tree against upstream v2.2.2 at the commit named above.

Method, stated so it can be reproduced: exact, case-sensitive comparison of non-blank lines; an upstream line counts as retained if the same line is present in this tree's file. "How much of upstream survives here" and "how much of this file is upstream's" are different questions with different answers, so both are given. scripts/measure-upstream.mjs regenerates this table against a clone of the commit above (while the upstream repository is unavailable, that needs a copy kept from before) — run it after any change to src/, because a figure nobody can cheaply re-derive is a figure that goes quietly stale.

File

Lines here

Upstream lines retained

of upstream

of this file

src/content_analyzer.js

450

361 of 394

91.6%

93.5%

src/validation.js

1132

765 of 798

95.9%

78.2%

src/index.js

867

541 of 631

85.7%

74.0%

src/tools.js

2625

798 of 1007

79.2%

33.2%

src/database.js

992

353 of 651

54.2%

40.2%

src/http-remote.js

232

—

—

no upstream counterpart

src/oauth-provider.js

357

—

—

no upstream counterpart

src/layers.js, floor.js, aleph.js, session.js, setup.js, gaps.js, provenance.js, redact.js, annotations.js, request-log.js, disciplines.js, metadata.js, env.js

3296

—

—

no upstream counterpart

Upstream's server/decay.js and server/healthcheck.js have no counterpart here and are not carried at all. Of the five files that do have one, the content analyser and the validation layer are substantially Jason's work, and most of his tools.js and index.js survives here as well — the routing patterns, the validators, and the error classes are his. What changed is described below; what did not change is his.

What diverged from upstream

  • Storage: SQLite on a RAM disk → MongoDB. Upstream's ram_disk_manager and server/decay.js temporal-decay engine are not carried forward, and the SQLite/RAM code paths have been removed rather than left dormant.

  • Transport: stdio only → stdio and streamable HTTP with OAuth, from one source tree.

  • Tools: 7 → 28. Three are carried from upstream (recall, query_layer, save_to_layer); remember, get_stats, get_echo_stats and get_status are dropped as duplicates or folded into emet_status. Twenty-five have no upstream equivalent — the document store (read_doc/ read_doc_history/write_doc/list_docs/patch_doc/retire_doc/rename_doc/ validate_doc/accept_nonconformance), the transcript tier (save_transcript/ revise_transcript/query_transcripts — which also reads the pre-2026-07 exchange archive), semantic_recall, revise_memory, corpus_recall, emet_gaps, emet_floor, the setup and session family (emet_initialize, emet_status, emet_setup_complete, emet_session_open, emet_transcript_append, emet_session_close), and the two access doors (emet_invite_create, emet_member_access_create).

  • The write path: upstream validates and inserts. Here every layer write, transcript write and document write passes through one door that scrubs secrets, decides the source tag, fills missing accountability fields with visible markers, digests content and metadata into their own columns, reads the row back, and returns one receipt shape. A test fails the build on any second door. See The door below.

  • Metadata vocabulary: upstream's whitelist relocated any field it did not know under custom. EMET's own fields — supersedes, superseded_by, superseded_at, corrects, corrected_by, valid_until, valid_from, routing, assertion_origin, and the marker booleans — are first-class, and anything that is still relocated is named in the receipt.

  • Charter as data: layers.js carries, beside each layer's prose, its numeric importance band, whether it may be revised, whether provenance is required, its aliases and its expected share, and the write path reads those rather than hardcoding them.

No claim is made that layered memory, entry supersession, or a split curated/verbatim store originate here. Comparable ideas ship in other projects. EMET's aim is that the substrate is not something you have to think about.


Install

Never deployed anything? Let your AI do it: docs/GUIDED-SETUP.md. You sign in to two services when asked and paste one URL; the assistant installs, configures, deploys and verifies the rest. It needs a shell — Claude Code and Cowork have one; Claude Desktop gets one from an MCP server such as Desktop Commander (one command, in the guide).

Full walkthrough: docs/SETUP.md — Atlas, both transports, and the first-run interview. Deploying remotely: docs/DEPLOY.md. Making your assistant actually use it: docs/BOOTSTRAP.md. Running bots alongside your assistant — registering their tags, their own handoffs, the session loop: docs/BOTS.md.

Short version:

npm install
cp .env.example .env      # fill in MONGODB_URI and MONGODB_DB
npm start

Requires Node 20+ and a MongoDB connection. .env is gitignored — never commit it.

Local (stdio) — as an MCP server entry:

{ "mcpServers": { "emet": { "command": "node", "args": ["/path/to/emet/src/index.js"] } } }

Remote (HTTP) — npm run start:http, or deploy the included Dockerfile to any container host. Set PORT, PUBLIC_URL (the OAuth issuer — it must match the deployed URL exactly) and EMET_OAUTH_PASSPHRASE. Then add https://<your-host>/mcp as a custom connector in each client. See docs/DEPLOY.md.

Give every host its own source tag: the host passes it as source on its writes, and EMET_SOURCE_TAGS registers the tags. Leave EMET_SOURCE_TAG unset on a server several hosts share, since a default tag makes every host look the same; set it only where one host uses the server (a local stdio entry, for example). Every write records the host that produced it; a shared default tag is ambiguity rather than attribution.

First run

Say "initialize". That is the whole interface.

emet_initialize is the front door and the only thing anyone needs to know about. On a new store it returns a short interview — what the assistant should be called, how it should carry itself, your name, what you are using this for — which your assistant runs conversationally, then records with emet_setup_complete, which also creates the owner person. On an established store the same call returns the assistant's identity, your profile, and where things were left, so the session opens already knowing you. A read-scoped emet_initialize writes nothing; a write-scoped one may provision a missing board. First-time setup is the owner's: emet_setup_complete is refused under a bot tag, and on a store that is already set up it is refused unless reissue: true is passed on the owner's explicit word.

Requiring a setup command would have been backwards: the one time you need it is the first session, before you have read anything. So the branch happens inside the tool, and you never have to know which case you were in.

Connecting the server gives your assistant the ability to remember; it does not make it use that ability unprompted. docs/BOOTSTRAP.md has a short block to paste into your project instructions, and explains what still works on hosts that have no instructions field at all.

Two behaviours worth knowing. Setup refuses to finish on partial answers rather than inventing an identity for itself. And if the store cannot be reached, initialize says so plainly instead of proceeding as though memory were available — fabricated continuity is worse than an honest gap.

Every session after that: how to start, how to end

Memory is only as good as the two moments at the edges of a conversation. Nothing your assistant learned survives unless it was written to the store, and the moment to write it arrives without warning — a closed tab, a context limit, a phone call. So every session has one word at the start and one at the end.

Start: say "initialize". Your assistant calls emet_initialize, which returns who it is, who you are, where things were left and the operating disciplines, and picks up from there. Before it asks you anything, it names your items from the handoff and state documents: anything left half-finished, what is dated in the next seven days, each open item you own with its next action, and whatever was left undecided. An item not named at the start of a session is easy to lose, so a count or a bare "where do you want to start?" is not treated as a startup. Do this in every new conversation, on every device — the first message, before the question you actually came to ask. A session that skips it is a session that answers from guesswork and saves nothing.

End: say "wrap up" (or goodbye, or that you are opening a new session). Your assistant then writes four things — the handoff, the state document, an episodic entry for the session, and the verbatim transcript — and calls emet_session_close, which checks that each one actually landed and returns state: "complete". Wait for that word before you close the window. If it comes back anything else, something did not save, and the assistant will say which.

Why both, every time: the layers are curated meaning and the transcript is the word-for-word record; the next session reads the first and can reach into the second by timestamp. Close without the ritual and the next session inherits a hole — it cannot see what happened, only that something did, and emet_gaps will count it. Start without it and the assistant is a stranger with your tools. The instruction block in docs/BOOTSTRAP.md makes both of these automatic on hosts that take standing instructions; on hosts that do not, the two words are what you supply.


Design notes

The six layers

Memory is split by what kind of thing an entry is, not by topic. Each entry carries an importance score and optional temporal validity.

Layer

Holds

Importance

Written

identity

Who the assistant is and who it serves. Durable self-definition and the user's standing requirements.

0.9 – 1.0

Rarely, by design

semantic

Facts that stay true regardless of when they were learned. People, systems, definitions, established conclusions.

0.6 – 0.9

When a durable fact is established or corrected

episodic

What happened, and when. Session records, problems solved, progress through long work.

0.4 – 0.7

At session close and at milestones

procedural

How something is done, and decisions that govern future action — with the reasoning.

0.5 – 0.8

When a decision constrains later work

meta

Everything visible only from outside the exchange: how the memory system behaves, and how the assistant behaves toward the user.

0.6 – 0.9

On retrieval failures and wins, protocol drift, and corrections to how the user is worked with

working

Short-lived context, in-flight state, notes for whoever picks the work up.

0.3 – 0.6

When parking state between sessions

The distinction that does most of the work: episodic, semantic and procedural are about the world — what happened, what is true, how something is done. Meta is about the system — looking back at the exchange from outside it. Material belonging in meta feels like nothing else fits it, which is precisely why it gets discarded when meta has no definition.

Every layer works from the first install

Creating six collections is not the same as having six layers. A layer with no stated charter has no answer to "does this belong here?", so nothing routes to it — and an empty collection throws no error and fails no test, so it can stay empty indefinitely without anyone noticing.

In the deployment this design came from, the meta layer sat nearly unused for about a year. The collection existed from day one and the tools accepted writes to it; it was simply never the obvious place to put anything. When it was finally given a definition it filled immediately — which is the proof that the material had been there all along and was being thrown away.

Two things in this repo exist because of that:

  • src/layers.js ships the charter for all six layers in code, and emet_initialize returns it — on a brand-new store as well as an established one. Same reasoning as the operating disciplines: documentation is read by people, and this has to reach the model at the start of every session. What belongs, what does not belong, when to write, when to read, and what the layer's absence looks like.

  • assessLayerHealth() reports a layer that has gone dark. A store cannot know what should have been written, but it can see that one layer took almost nothing while the others filled, and say so in the first payload of the next session. Run against the real history above, it flags meta at 10 entries / 1.8% of the store — and goes quiet once the layer is in genuine use. It reports; it never blocks, and it stays silent on a store too new to have a pattern.

Two tiers, deliberately independent. Layer entries are curated meaning; a separate transcript store holds the word-for-word conversation. When a layer entry is relevant but thin, the transcript holds the depth that was compressed out of it. Summarising at transcript-write time collapses the two tiers into one and destroys the second line of inquiry — so the transcript path stores exactly what it is given. If a payload will not transmit, split it across parts; do not compress it.

How tier 1 reaches tier 2 — two links, one precedence rule.

An entry records the transcript span it came from: session_id plus a range of exchange indices, written at save time as derived_from. Entries consolidated from other entries carry {entries: [id, ...]} instead, because an entry with no transcript origin must be able to say so rather than carry a nearest-timestamp guess dressed as a citation. Both are stored as top-level indexed columns rather than inside the metadata subdocument, so the reverse question — every entry drawn from this transcript — is entirely a query.

Every entry also has a timestamp, so a window around it can always be searched. The two are not redundant copies: the pointer is exact but conditional, the timestamp universal but approximate — it depends on choosing the window, degrades when sessions overlap, and when it is wrong it returns the neighbouring span rather than failing.

Precedence is what keeps them from conflicting: the pointer wins whenever present, the timestamp is the fallback when it is null. Never averaged, never combined, and the timestamp is never preferred merely because it returned something. resolveProvenance() applies the rule so no caller has to decide it.

Keeping both buys something neither gives alone — they can check each other. If an entry's pointer names a session whose span sits nowhere near that entry's own timestamp, something is wrong: a mis-written pointer, a field copied from another entry, an entry attributed to the wrong session. crossCheckProvenance() reports the disagreement. Two independently derived answers that can disagree is redundancy in the engineering sense; one link that can only be trusted is not.

The reverse direction — what did we take from this session? — is not a second structure. source_session_id is indexed, so it is the same data read the other way. A stored reverse index would be a mirror of data that already exists, and mirrors are how forks are born.

Entries written before pointers existed keep a null pointer and resolve by timestamp, and that remains correct for them — superseded, not wrong. They are not backfilled: the only way to invent a pointer for them is to infer it from the timestamp, which is exactly the weaker method the pointer replaces. Old records are not rewritten to look like new ones. Note also that an empty timestamp-window search is not proof that nothing was recorded — a period may hold no transcript at all.

Supersession, not mutation. revise_memory writes a new entry carrying supersedes and marks the parent superseded_by. Nothing is edited or deleted. Superseded entries stop surfacing in normal reads; include_superseded: true restores them for audit. Episodic entries are immutable — events do not get revised.

Corrections, not supersession, for partial fixes. When a later entry corrects only part of an earlier one (one item of an episode, say), superseding would hide the parent's items that still stand. Instead the new entry carries metadata.corrects: <parent id> (same layer; a missing parent is refused before anything is written), and the parent is marked: its corrected_by list gains the new id, its content, digest and liveness are untouched, and its metadata digest is re-attested in the same write. Reads return corrected_by, so a reader of the parent sees that a correction exists. corrected_by is a door marker; a caller cannot set it. Entries corrected before this existed (e.g. episodic #604 correcting #603) carry no link; the fix is a new correction entry naming corrects: 603, not an edit or a backfill.

Verified writes - src/aleph.js. Every write is read back and compared before it is reported as saved. Document writes archive the prior version, then confirm both the sha256 and the version; layer writes confirm the stored content at the id that was allocated. The id is checked by content rather than existence on purpose: ids come from a counters collection, and a counter read from the wrong collection restarts at 1 and silently overwrites real entries - which looks exactly like a successful save.

A verification failure raises AlephError, deliberately distinct from a database error. A database error means the store could not be reached and the write probably did not happen. An AlephError means the call succeeded and the store does not contain what it should. That is the more serious of the two and the one that must never be retried blindly, so it is never flattened into "the save failed".

The aleph answers two questions, at two moments. At the write: was it recorded accurately? The read-back above, before any receipt. At every later read: is it still what was recorded? Each row carries its content and metadata digests, an HMAC signature over both (when EMET_SIGNING_KEY is set; the key stays on the service), and a link in its layer's hash chain (each entry's chain digest covers its own content digest and its predecessor's). Every read recomputes the digests and reports integrity; emet_gaps walks the whole chain for any host and counts a row altered after writing (broken_chain) or removed from mid-chain (chain_discontinuities). Expected count for both: zero.

The door

Truth integrity, accountability and validation are meant to be a guaranteed property of the store, not an option a model can decline to use. A file a model can be told about is a file it can ignore. So the guarantee comes from a property, not a module: exactly one way to write each store, with the checks inside it. saveMemory for the six layers, writeDoc for documents (patch_doc delegates to it), saveTranscript and reviseTranscript for the verbatim tier. tests/write-path.test.js scans src/ for every MongoDB write verb and fails the build on any that sits outside those functions — it found and removed a second document door in setup.js.

Inside the door, in order:

  • Secrets are masked before anything is stored or hashed (src/redact.js): the literal values of secret-bearing environment variables, credentials inside URIs, bearer tokens, JWTs, and key: value pairs whose key names a secret become *****. The mask is visible on purpose, the receipt's redacted reports the number of spans this write changed (text already carrying the mask is left alone and not counted, so a read-modify-write of a document reports zero for the masks it already held), and the same scrubber wraps every rendered log line.

  • The source tag is decided here, not by the transport: the caller's tag, else the environment default marked source_defaulted: true, else null marked source_missing: true.

  • Every missing accountability field becomes a visible marker rather than an absence — provenance_missing, assertion_origin_missing, importance_defaulted, source_unregistered — because absence cannot be queried and "missing" can. A missing field is flagged, not refused, with one exception (2026-09-29): a semantic or procedural entry with no source is refused, because a durable claim that cannot point to where it was said is exactly what the record exists to prevent. The door stamps the source from the session's transcript position when the caller names its session. Importance defaults to the layer's charter midpoint; a value outside the band is kept and warned about. The door never infers an assertion's origin.

  • The routing decision is kept in metadata.routing — whether the layer was chosen by the caller or fell back, and what both routers thought — so routing can be measured later instead of remembered. An unlabelled write is never placed by a router: it lands in working, flagged routing.method: 'fallback', with a receipt warning. Naming the layer is the norm.

  • Content and metadata are digested into their own columns (sha256, metadata_sha256, the latter over canonical JSON), read back and compared before the write is reported. Every read path — recall, query_layer, semantic_recall — recomputes both on the way out and attaches integrity and metadata_integrity: true, false (the row is returned and flagged, never dropped), or 'unattested' for rows written before the column existed. Supersession and the correction mark are the only sanctioned mutations of stored metadata, and both re-attest the parent.

  • One receipt shape on both transports: layer, id, timestamp, verified, sha256, digest_stored, metadata_sha256, metadata_digest_stored, redacted, provenance, source, assertion_origin, importance, routing, warnings[] plus whichever markers apply.

And the record's honesty about its own holes is a query: emet_gaps counts, with the ids behind each count, unsourced assertions, unattributed inferences, ambiguous attribution by tag and by session, self-routed writes, unattested rows, digest mismatches, pointer disagreements against the transcript store, broken supersession, broken correction links, expired-but-live rows, consolidation candidates, and redacted spans by source. It identifies and never fills — where a source is knowable the fix is revise_memory with a real pointer; where it is not, the gap stands as a fact and the count stops growing from that date. emet_initialize carries the counts beside layer_health.

What can be guaranteed this way: every write attested, attributed, scrubbed, and every absence visible. What cannot be guaranteed without inviting fabrication: every write complete. The store does not promise the latter.

Guardrails & Monorails

The door guards each write. Guardrails & Monorails guards the order of a session's writes. Guardrails (the floor, the character, the served instructions) bound a field where the AI moves freely and does its judgment work. The monorail (the session graph, src/session-graph.js) carries it from one field to the next without asking it to choose: open, work, record, board, handoff, close. The car can't derail or leave the beam. Where the flow must fork, a switch at the junction (a small logic unit in EMET's code, tripped only by its condition, like a transistor; no hosted model, no API) routes the session to the right guardrail space. The track throws the switch, never the AI. The stations are where the processing happens, and each station's door opens only when every stop before it has been made. Switches are designed, not yet built. The server checks the track on every governed write, so it is the same for any AI on any device. No host hook is needed. Every skipped step is refused, and the refusal says how to pass. The full practice, the gates and the honest limit are in docs/GUARDRAILS-AND-MONORAILS.md.


Tools

recall · semantic_recall · query_layer · save_to_layer · revise_memory · read_doc · read_doc_history · write_doc · list_docs · patch_doc · retire_doc · rename_doc · save_transcript · revise_transcript · query_transcripts · corpus_recall · emet_gaps · emet_floor · validate_doc · accept_nonconformance · emet_invite_create

Retained revisions — every document write archives the prior copy; read_doc_history lists what is retained for a document and returns any retained revision in full, re-hashed against its stored digest. Retention that nothing could read back was a promise without a witness.

One reader for the second tier — query_transcripts reads the live transcript store and the legacy exchange archive (rows written by earlier hosts, before the transcript store existed) together, in one session shape, each result naming the collection it came from. Nothing is migrated between them.

Document control — validate_doc checks a controlled document against the template that governs it, and the same check runs inside the write path on every document write, where no caller can skip it. A record that fails is marked, never rejected: it lands, carries a nonconformance marker naming what failed, and stays counted at session start until it is corrected by a later revision or dispositioned with accept_nonconformance. The checks are form and continuity — sections present, in order and filled; no placeholder left unfilled; every item on the previous revision still present or carrying a disposition. Whether an answer is true is not something a machine can check, and this one does not pretend to.

Session start — emet_initialize (start here): fetch every page it names, or pass page: "all" on a host with no size limit; then emet_session_open with that startup's load_id. Session ids are dated in the local time zone (EMET_TIMEZONE, default America/New_York); a closed session is never reopened, so a slug whose id is already closed opens as <id>-2, -3, and the result says so. After a complete close a session takes exactly one more append, its closing exchange, within EMET_CLOSE_GRACE_MINUTES (default 5, at most 60) of the close, written into the session's last part, and is then locked; a retry of it writes nothing. Corrections with revise_transcript stay allowed but must keep exactly the same exchange numbers.

During the session — emet_transcript_append after every reply, one exchange at a time. Every write tool also takes session_id and exchange, and returns session_graph.next, the next stop on the monorail (Guardrails & Monorails).

Session end — emet_session_close. Checks whether the session's work actually reached the store — transcript saved and intact, continuity documents updated, an episodic entry written — and reports exactly what is missing. It does not write the artifacts: only the model holds the session's content, and a tool that generated a plausible-looking summary would be manufacturing the false continuity the disciplines exist to prevent. emet_initialize returns the same protocol at session start, which is the only moment early enough to matter.

Setup internals and diagnostics — emet_status (environment, layers, health; probe: true writes and reads back a probe) · emet_setup_complete

Approvals

Every tool declares MCP annotations — readOnlyHint, destructiveHint, idempotentHint, openWorldHint — in one table in src/annotations.js, which is also the complete statement of what EMET does to a store. 11 of the 29 tools cannot modify anything; four can replace the visible content of a document; openWorldHint is false everywhere, because EMET reaches its own database and nothing else — no filesystem, no shell, no outbound network.

A server that declares nothing is treated as if every call might do anything, so a host that asks before acting asks before every call, including the reads that make up most of a session. These hints let a host stop asking about reads without being told to trust writes. They are hints, not enforcement: what is actually auto-approved is a host-side setting. See docs/PERMISSIONS.md.


Tests

npm test

1,222 assertions across 39 suites (counted 2026-09-29). No framework, no database, no credentials — each suite is a plain node script on one shared harness (tests/harness.js), and the runner discovers every *.test.js, so npm test runs on a fresh clone. Per-suite counts are not kept here because they go stale; npm test prints them.

validation.test.js and content_analyzer.test.js began as upstream's own suites, ported unchanged apart from the import path; every one of upstream's assertions still stands, and the additions cover EMET's metadata vocabulary and the shared layer list. They are the original author's assertions about the original author's code, and they are kept passing.

The rest cover this tree. Some are fixtures rather than unit tests — they read the source and fail on a class of drift: write-path.test.js allows no MongoDB write outside the sanctioned doors, requires each door to scrub before it hashes, and requires that nothing in the tree removes from a layer, document or transcript collection; dispatch.test.js requires the declared tool list and both transports' dispatch switches to be the same set, and that a list response can carry a caller-facing warning; annotations.test.js requires every tool to declare its approval hints; atlas-role.test.js requires the database role to grant no remove on the record. The others test behaviour — among them aleph.test.js (read-back, digests, signatures, the chain walk), session-graph.test.js (the monorail's gates and lab results), receipt.test.js, provenance-stamp.test.js, session.test.js, gaps.test.js and redact.test.js. Their failure cases matter more than their success cases — a verifier that passes something it should have caught converts an undetected problem into a confident assurance. layers.test.js asserts against the store shape that produced the year-long dormant layer, so the check is tested against the failure it was written for rather than a hypothetical one.

Upstream's index.test.js, decay.test.js and integration.test.js are not carried: they target the SQLite backend and the temporal decay engine, neither of which exists here. Running them would produce failures that say nothing about this code.

Benchmarks

EMET's implementation language was evaluated rather than assumed. The same specification was implemented in JavaScript, Python, Go and Rust, deployed identically, and run against live Atlas — with a digest equivalence check proving all four did the same work before any timing was read. JavaScript won, on serial in-process cost, on throughput under concurrency, and on standards provenance as the only candidate carrying an international language standard (ECMA-262 and ISO/IEC 16262).

The exercise's most useful output was not a language choice but a function to optimise: the write-path scrubber was paying nearly full cost on documents containing no secrets, and fixing that saved roughly thirty times what the best available language change could have.

Method, full results, and an explicit account of what was not measured: docs/BENCHMARKS.md. scripts/redact-bench.mjs measures the scrubber alone — run it before and after any change to src/redact.js.


Known gaps

Honest list, because the point of this project is that it behaves predictably:

  • Rows written before 2026-09-05 carry no content digest and rows written before 2026-09-07 carry no metadata digest; reads report them as 'unattested', and they are not backfilled — a digest computed today would attest to today's copy, not to the write. emet_gaps counts the boundary.

  • Entries written before provenance existed keep a null pointer and resolve by timestamp window; emet_gaps counts them as unsourced. They are not backfilled, for the reason given under Provenance.

  • Two content routers are carried from upstream (a substring keyword router and a content analyser). Measured against a 1,175-row corpus both agree with the stored layer well under half the time (41% and 36%), so neither places entries: an unlabelled write lands in working, flagged, with both opinions recorded in routing (keyword_layer, analyzer_layer). Retuning them into something fit to route is possible with those numbers as the baseline; it is not scheduled.

  • consolidation_candidates in emet_gaps is tag overlap, not semantic similarity.

  • The monorail (session graph) refuses every skipped step (enforce only since 2026-10-03; there is no report mode or off switch). Its turn coverage is only as strong as the exchange numbers hosts report, and it never sees the reply text. See Guardrails & Monorails.

  • The scheduled weekly intake writes without opening a session, so the session-order check refuses those writes (gate no_session) until the routine opens a session with its startup's load_id and passes the session_id. Its old source tag, cloud-scheduled-intake, was retired on 2026-10-02 but is still accepted rather than refused: flagged on memory writes, and let through on transcript appends. emet_gaps keeps counting the rows already written under it as ambiguous attribution instead of rewriting history.

  • save_transcript's description does not state that it stores content verbatim; the implementation does.

  • No export / import / restore-and-compare yet.

  • The document store is not prefixed, so one database holds one instance's identity. Run separate databases rather than separate prefixes if you want separate memories.

  • Dynamic Client Registration is still served beside Client ID Metadata Documents; the MCP specification's deprecation window closes around 2027-07. It is scheduled to come off on 2027-06-01.


License

Authorship of the commit history, including a correction recorded as an addition rather than a rewrite: docs/AUTHORSHIP.md.

MIT. See LICENSE — both the original C.I.P.S. LLC notice and the Dreamforge Studio LLC notice must travel with any copy.

Available Tools

29 tools
accept_nonconformanceA

Record the written disposition for a nonconformance that will not be corrected. A marker is never cleared silently: it leaves either because a later revision passes, or because someone accepts it in writing here. The marker is RETAINED with the acceptance beside it, and the acceptance covers only the revision it names - the next write is validated fresh, and accepted markers stay counted at session start, separately from open ones. The reason must be at least 10 characters: it is what an audit reads later.

ParametersJSON Schema
NameRequiredDescriptionDefault
doc_idYesDocument carrying the open marker.
reasonYesWhy this nonconformance is accepted rather than corrected. At least 10 characters.
sourceNoSource tag - which host produced this write. One name on every write since 2026-09-27.
exchangeNoThe exchange (turn) this write belongs to, counting from 1. A write for exchange n while exchange n-1 was never appended is refused by the session graph until n-1 is appended.
session_idNoTranscript session id.
accepted_byNoAlias of `source` (the older name, kept one release). Prefer `source`.

TDQS

A4.3/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare write/non-idempotent/non-destructive; the description adds far more: the marker is RETAINED rather than cleared, the acceptance covers only the named revision so the next write is validated fresh, and accepted markers are counted separately from open ones at session start. It also states the audit-facing durability of the reason.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core action, then layered with genuinely useful lifecycle behavior. Every sentence carries weight, though the closing constraint note slightly overlaps the schema and keeps it from being maximally tight.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 6-param write tool with full schema coverage and no output schema, the description covers purpose, lifecycle, persistence, and validation semantics well. It stops short of naming sibling alternatives, but an agent has what it needs to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so doc_id, reason, source, exchange, etc. are already documented. The description reinforces the ≥10-character reason constraint and its audit purpose, but adds no syntax or format detail beyond the schema. Baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: 'Record the written disposition for a nonconformance that will not be corrected.' It makes the scope unmistakable (only for nonconformances deliberately left uncorrected) and clearly separates this action from the alternative path where a revision later passes.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives clear selection context: use it when a marker will not be corrected, since it otherwise leaves only via a passing revision. However, it never names the sibling tools (e.g. validate_doc or write_doc) that represent the other path, so the routing is implied rather than explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

corpus_recallA
Read-onlyIdempotent

Semantic search over an optional supplementary knowledge corpus that the user has loaded separately from memory. Returns nothing unless a corpus has been populated; this is not where session memory lives.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMax results (default 5, max 25)
queryYesNatural-language query (raw text; concept match, phrased however the experience is phrased)
corpusNoOptional: restrict to one corpus name

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already establish readOnly, idempotent, and non-destructive behavior, so the description's job is to add context. It does so by disclosing the empty-result condition ('returns nothing unless a corpus has been populated') and the separate-corpus scope, both of which are useful traits not visible in the annotations or schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences with the primary action front-loaded and the scoping caveat immediately after. No filler, though the second sentence's second clause slightly overlaps the first sentence's 'loaded separately from memory' idea.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 3-parameter, no-output-schema read tool with full schema coverage, the description supplies the key precondition and scope boundary. The only real gap is the absence of explicit routing guidance relative to recall/semantic_recall.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description adds no syntax, format, or default details for query, limit, or corpus beyond what the schema already documents.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: semantic search over a supplementary knowledge corpus. It also explicitly rules out session memory, which helps separate it from the recall/semantic_recall siblings, but it never names those siblings directly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives a precondition (a corpus must have been populated) and implicitly signals usage via the contrast with session memory, but there is no explicit statement of when to choose this over recall, semantic_recall, or query_layer.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

emet_floorA
Read-onlyIdempotent

Read the floor: the standing description of how this assistant engages with the person using it. It ships in the server code, so setup cannot overwrite it and no install can lose it - the character document holds the user's modifiers on top of it. emet_initialize already returns the whole floor at session start; call this to read it again without re-running initialize. Use compact: true for the short form partway through a long session, when the start of the conversation is far behind. Read-only, and re-reading is not the same as being bound by it.

ParametersJSON Schema
NameRequiredDescriptionDefault
compactNoReturn the short re-assertion form instead of the whole floor.
sectionNoReturn one section only, e.g. 'neurodivergent_support'. Call with no arguments to see the section names.

TDQS

A4.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint and destructiveHint=false, so the safety profile is covered. The description adds genuinely new context beyond that: the floor ships in server code, setup cannot overwrite it, installs cannot lose it, and the character document layers user modifiers on top. The closing note that re-reading is not the same as being bound by it also characterizes the operation's limits.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the definition, then provenance, then the initialize relationship, then the compact tip. Every sentence carries information, though the server-code/character-document provenance sentence is denser than strictly needed for invocation.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-required-parameter, read-only tool with no output schema and full schema coverage, the description supplies everything needed: what the floor is, why it persists, how it relates to initialize, and when to use the compact form. No meaningful gap remains.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3 — both compact and section are already documented in the schema. The description goes one step further by supplying the situational trigger for compact (long session, distant conversation start), which the schema's terse 'short re-assertion form' does not convey.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Read the floor') and then defines what the floor is — the standing description of how the assistant engages with the user. It explicitly distinguishes itself from the sibling emet_initialize, noting that initialize already returns the whole floor at session start, so an agent can route correctly without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Names the alternative (emet_initialize) and the exact condition for choosing this tool instead: re-reading without re-running initialize. It also gives concrete when-to-use guidance for the compact parameter ('partway through a long session, when the start of the conversation is far behind').

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

emet_gapsA
Read-onlyIdempotent

The record's honesty about its own holes, as a query. Read-only; identifies and never fills. Counts, with the ids behind them, for each gap class: unsourced assertions, unattributed inferences, ambiguous attribution (by tag and by session), uncertain routing, unattestable rows (pre-2026-09-05 boundary), digest mismatches (the alarm - expect zero), pointer disagreements against the transcript store, broken supersession, expired-but-live rows, consolidation candidates (tag families; a list for a person), and redacted spans by source. Call it when asked how trustworthy the record is, before a consolidation pass, or after any migration. Closing a gap is a person's decision: revise_memory with a real derived_from where the source is knowable; otherwise the gap stands as a fact.

ParametersJSON Schema
NameRequiredDescriptionDefault
layerNoRestrict to one layer. Default: all six.
limitNoMax ids returned per class (1-500, default 50). Counts are always complete.
sinceNoUnix seconds; only entries with timestamp >= since. Default: the whole store.

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly/idempotent/non-destructive, but the description adds value beyond them: "identifies and never fills," the return expectation "digest mismatches (the alarm - expect zero)," and the note that counts are complete while ids are capped. It stops short of describing the response shape in full.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Dense but front-loaded: the operation and its read-only nature come first, then the enumerated classes, then usage. The list of gap classes is long but each item is distinct and informative. The opening metaphor is a minor flourish that costs a few words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only diagnostic query with no output schema, the description is largely sufficient: it names every gap class, states counts are complete, and explains what closing a gap entails. Only the exact response structure and id-format remain unspecified, which is a minor gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents layer, limit, and since in detail. The description adds no parameter-level syntax or defaults; its references to tags, sessions, and the 2026-09-05 boundary describe gap classes rather than inputs. Baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

After a metaphorical opening ("the record's honesty about its own holes"), it states a concrete verb and resource: "Counts, with the ids behind them, for each gap class," and enumerates those classes (unsourced assertions, digest mismatches, broken supersession, etc.). This distinguishes it from siblings like emet_status or query_layer, though the poetic framing costs some immediate clarity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit triggers are given: "Call it when asked how trustworthy the record is, before a consolidation pass, or after any migration." It also names the follow-up path, routing remediation to revise_memory "with a real derived_from where the source is knowable," and clarifies the tool itself never fills gaps.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

emet_initializeA
Idempotent

START HERE. Call this FIRST in every session, before answering anything that might have prior context, and always when the user opens with 'initialize', 'start', 'begin', 'hello' or similar. It is the only setup step a user should ever have to know about. A large startup comes in pages: read page_notice first and call again with the page it names until the last page, before replying. If this memory store is new it returns a short interview to run conversationally; if it is established it returns who you are, how you speak (character, when the store keeps one - authoritative on voice for the whole session), who you are serving, and where things were left. The user never needs to know which case applied.

ParametersJSON Schema
NameRequiredDescriptionDefault
pageNoThe startup is served in pages when it is larger than one result (each page fits under 20,000 bytes). Omit for page 1; every page but the last says which page to ask for next - call again with that page, before replying, until a page says it is the last. "all" returns the whole startup in one result, for a host with no size limit.
load_idNoThe load_id your startup returned in `startup_page.load_id`: every emet_initialize result carries one (each page of a paged startup, the whole startup with page "all" or when it fits one result, and a store that is not set up yet). Pass it with every later page so all pages come from the same startup, and pass the same load_id to emet_session_open so its startup page check runs on this startup.
recent_limitNoHow many recent context entries to return (default 5).

TDQS

A4.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare idempotent, non-destructive, and readOnlyHint=false; the description adds real behavioral detail beyond that - paginated startup, the read-page_notice/recall loop, load_id threading, and the conditional new-store interview vs. established-store profile. It doesn't explain why readOnlyHint is false despite describing a largely read-oriented setup, a minor gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the imperative call-to-action and ordered logically (when to call, then paging, then outcomes). It is dense and long, but each sentence carries procedural meaning, so little is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description carries the burden and does: it explains paged returns, the load_id contract, both new-store and established-store outcomes, and the sequencing with emet_session_open. An agent has everything needed to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description goes further by explaining the paging workflow ('read page_notice first and call again'), the role of load_id across pages and into emet_session_open, and the defaulting behavior around recent context - value beyond the raw field descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific action (start/set up the session for a memory store) with front-loaded imperative 'START HERE. Call this FIRST,' and distinguishes its role from siblings like emet_session_open by naming the setup step it owns. An agent immediately knows this is the session-initiation tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly says when to call (first in every session, when the user opens with 'initialize'/'start'/'begin'/'hello', or before any answer relying on prior context) and how to sequence it with emet_session_open and the paging loop. Both the trigger conditions and the alternative/next-step tool are named.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

emet_invite_createA

EMET Guest Access: create a one-time code for a guest. They paste it on the connect screen instead of the passphrase. Default access is read-only. Only the owner can create a guest code. EMET Member Access is a different door.

ParametersJSON Schema
NameRequiredDescriptionDefault
daysNoDays until the code expires. Default 14.
scopesNoDefault: emet.read only
display_nameYesName the guest will be known by in this store

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare the safe/one-shot profile (readOnlyHint=false, destructiveHint=false, idempotentHint=false). The description adds genuinely new context: the code is one-time, access defaults to read-only, and creation is restricted to the owner — a meaningful permission constraint not captured elsewhere.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four short, front-loaded sentences with no filler. The core action leads, followed by mechanism, constraints, and a routing note to the sibling tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a small three-parameter mutation tool with full schema coverage and no output schema, the description covers purpose, auth restriction, defaults, and sibling disambiguation. Only the returned artifact (the actual code) is left implicit, a minor gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the schema already documents display_name, days (default 14), and scopes (default emet.read). The description's 'default access is read-only' merely restates the scopes default, adding no syntax or meaning beyond the schema; baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('create a one-time code for a guest') and explains the mechanism (pasted on the connect screen instead of the passphrase). It also explicitly distinguishes itself from the sibling emet_member_access_create by calling member access 'a different door'.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Clear context for when this applies (guest access via a code rather than a passphrase) and it names the adjacent alternative (EMET Member Access). It stops short of spelling out the selection criteria between guest vs. member access beyond the metaphor.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

emet_member_access_createA

EMET Member Access: create a one-time code for a new person with read and write access. They paste it on the connect screen instead of the passphrase. This is not EMET Guest Access, and it never makes them the owner. Only the owner can create it. Off unless EMET_MEMBER_ACCESS=1.

ParametersJSON Schema
NameRequiredDescriptionDefault
daysNoDays until the code expires. Default 14.
display_nameYesName this person will be known by in this store. A name never selects an existing person.

TDQS

A4.1/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only cover the generic safety profile (readOnly=false, destructive=false, idempotent=false, openWorld=false). The description adds substantive behavior beyond them: creator authorization ('Only the owner can create it'), the fact that it does not confer ownership, the redeem-on-connect-screen flow, single-use nature of the code, and a feature-flag prerequisite.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The purpose is front-loaded in the first sentence, followed by disambiguation and constraints in compact form. The telegraphic fragments ('Only the owner can create it. Off unless EMET_MEMBER_ACCESS=1.') are slightly clipped but each carries real information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description should indicate what the call yields; it implies a generated code but does not state the return shape or what happens when the tool is disabled or the caller lacks owner rights. Otherwise it covers prerequisites and scope well for a simple two-parameter tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both parameters (display_name, days) are already fully documented in the schema, including the important note that a name never selects an existing person. The description adds no additional parameter-level detail, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a specific verb and resource ('create a one-time code for a new person with read and write access') and explicitly frames the scope against 'EMET Guest Access' and ownership. However, it never distinguishes itself from the actual sibling tool 'emet_invite_create', which is the closest candidate for confusion, so sibling differentiation is incomplete.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It states clear conditions: only the owner can call it, the recipient uses the code on the connect screen instead of the passphrase, and the tool is inactive unless EMET_MEMBER_ACCESS=1. What is missing is an explicit routing statement against alternatives such as emet_invite_create or the guest-access path.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

emet_session_closeA

Call this when the session is ending - the user says goodbye, asks you to wrap up, mentions restarting or opening a new session, or you are about to run out of room. An acknowledgement of the turn just completed - thanks, great, perfect - is not a session end; if it is ambiguous, ask rather than close. Checks whether this session's work has actually been written to the store and reports exactly what is missing. It does not write anything itself. RECEIPTS ARE REQUIRED: pass receipts naming what you wrote and read back for each close step - working documents, the episodic session record id, the board and its version, the transcript part ids and highest exchange index as query_transcripts returned them, and the handoff version (written last). Each receipt is compared with the store; omitted or mismatched receipts, a handoff that was not the last write, or a transcript no layer entry cites (derived_from) return state 'blocked'. Each missing line carries a scope: touch only what it names. THREE blocked closes on one session return state 'escalate': stop retrying, say the speak lines, and hand the user the list - the counter does not reset on a changed receipt, only with a new session. Say every line of the returned speak to the user, in order - a wrap that leaves it out is a failed close. Never tell the user their session is saved unless this returns state 'complete'.

ParametersJSON Schema
NameRequiredDescriptionDefault
sourceNoYour host's source tag: the registered tag this session was written under. The close is never refused for it; an unregistered tag, one that did not write this session, or none (with no server default) gets a warning that the close is not attributed.
exchangeNoThe exchange (turn) this write belongs to, counting from 1. A write for exchange n while exchange n-1 was never appended is refused by the session graph until n-1 is appended.
receiptsNoRequired for state 'complete'. One entry per close step, taken from what the tools returned this session - not from recall.
session_idNoThe session_id used (or to be used) for this session's transcript, e.g. '2026-09-03_short-description'. Without it the transcript cannot be checked; receipts.transcript_session_id is used when this is omitted.
thread_docsNoAdditional document ids this session should have updated, e.g. ['threads/some-project.md'].
session_startNoUnix seconds when the session began. Defaults to 12 hours ago, which makes the answer approximate.

TDQS

A4.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

It discloses behavior far beyond the annotations: the receipt-comparison protocol, that omitted/mismatched receipts, a non-last handoff, or an uncited transcript yield state 'blocked'; that three blocked closes yield 'escalate' and that the counter persists across receipt changes; and that every 'speak' line must be relayed. One mild tension: readOnlyHint=false implies side effects while the text says it 'does not write anything itself', though the persistent blocked-close counter plausibly accounts for the non-idempotent annotation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The text is long but dense, and the most consequential rules are front-loaded (when to call, then what it does, then RECEIPTS ARE REQUIRED) with emphasis markers for the blocked/escalate branches. A few clauses, such as the repeated warnings about relaying speak lines, could be tightened, but almost every sentence carries an operative rule.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Despite a complex 6-parameter schema with a nested receipts object and no output schema, the description covers the return states ('complete', 'blocked', 'escalate'), the per-line 'scope' field, and the required 'speak' output, so an agent knows both how to call it and how to act on the result. Nothing essential is left unspecified.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds real semantics the schema cannot: receipts are required for state 'complete', each receipt is compared against the store, and each missing line carries a 'scope' telling the agent to touch only what it names. It also stresses receipts must come from tool returns, not recall, which is meaningful usage guidance for the nested object.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a precise verb+resource: it 'Checks whether this session's work has actually been written to the store and reports exactly what is missing' and explicitly disclaims writing ('It does not write anything itself'). This distinguishes it cleanly from sibling write tools like save_to_layer, write_doc, and emet_transcript_append, and from the sibling emet_status.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives explicit triggering conditions ('user says goodbye, asks you to wrap up, mentions restarting or opening a new session, or you are about to run out of room') and an explicit exclusion ('An acknowledgement of the turn just completed - thanks, great, perfect - is not a session end'), plus a tie-breaker ('if it is ambiguous, ask rather than close'). This is exactly the when/when-not guidance the dimension rewards.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

emet_session_openA

Open a transcript session. Returns session_id, dated in the configured local time zone (EMET_TIMEZONE, default America/New_York). A slug whose id is already CLOSED is never reopened: the session opens as -2, -3, ... and the result's notice says so. Pass load_id from your startup pages (startup_page.load_id on every emet_initialize page) so the startup page check runs on your own startup; an open without load_id, or with a page of that startup never fetched, is refused (the refusal says how to pass). After every reply, call emet_transcript_append with that id and the one exchange that just happened. Do not dump the whole chat at close.

ParametersJSON Schema
NameRequiredDescriptionDefault
slugNoShort name for the session, e.g. grok-init
load_idNoRequired for the startup page check (by policy; the schema leaves it optional so older hosts can still open): the load_id from `startup_page` on your emet_initialize pages - page 1 returns it, every page carries it, and the last page names it. An open without load_id, or with a page of that load known not fetched, is refused before anything is written (gate startup_pages); the refusal names the exact calls that pass.

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes well beyond the annotations: discloses the timezone (EMET_TIMEZONE, default America/New_York), the CLOSED-slug collision rule (session opens as <id>-2, -3 with a `notice`), and the refusal path for a missing/unmatched load_id. These are non-obvious behavioral traits that the readOnlyHint/not-destructiveHint/not-idempotentHint annotations cannot convey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core action and return values, then the load_id gate and workflow. Dense with information, though a couple of sentences are long and slightly clause-heavy; little is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, and the description covers the return (session_id, timezone, `notice`) plus the prerequisite gate and the required follow-up call. An agent has everything needed to invoke it correctly and continue the session workflow.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds real meaning: it explains why load_id is optional in the schema yet required by policy, and clarifies slug collision semantics. It gives more semantic context than the schema alone.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ("Open a transcript session") and immediately specifies what it returns (session_id) and the scoping behavior. An agent can distinguish it from emet_session_close and emet_transcript_append without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly instructs to pass load_id from startup_page and states the refusal condition when it is missing or unfetched, naming the gate. It also prescribes the follow-up workflow ("After every reply, call emet_transcript_append with that id") and warns against dumping the whole chat, which routes the agent correctly.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

emet_setup_completeA
Destructive

First-time setup, the owner's only: record the user's interview answers as bootstrap/IDENTITY.md, bootstrap/CHARACTER.md and bootstrap/USER.md and as an identity-layer entry, and create the owner person. Every write is read-back verified. Refused under a bot tag. On a store that is already set up it is refused unless reissue is true (the owner's explicit word): a re-run writes new versions of the user's identity documents; prior versions are archived, never destroyed.

ParametersJSON Schema
NameRequiredDescriptionDefault
sourceNoYour source tag. Setup is refused under a bot tag.
answersYesAnswers keyed exactly as the interview questions returned by emet_initialize.
reissueNoOnly on the owner's explicit word: re-run setup on a store that is already set up. Without it, a re-run is refused and nothing is written.

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already flag destructiveHint=true and idempotentHint=false; the description adds substance beyond them: every write is read-back verified, bot-tag refusal, and the key safety nuance that on re-run prior versions are archived, never destroyed. That directly informs the agent about irreversibility and preconditions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the scoping statement ('First-time setup, the owner's only') and packs multiple constraints into a tight block. Slightly dense and repeats the bot-tag refusal that also appears in the schema, but no sentence is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive, bot-refusing, non-idempotent mutation with a nested answers object and no output schema, the description covers preconditions, re-run semantics, and archival behavior well. It does not state what the tool returns, which is the only notable gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents source, answers, and reissue. The description adds context (owner-only, explicit-word requirement for reissue, refusal otherwise) but no parameter syntax or format detail beyond what the schema provides. Baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Names a specific effect (record interview answers into bootstrap/IDENTITY.md, CHARACTER.md, USER.md plus an identity-layer entry, and create the owner person) with a clear actor/scope ('the owner's only'). This is unmistakably distinct from siblings like emet_initialize, emet_status, and write_doc.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit when-to-use ('First-time setup'), when-not ('Refused under a bot tag'; refused on an already-set-up store unless reissue is true), and the escape hatch (reissue = owner's explicit word). Conditions for selecting the behavior are fully spelled out.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

emet_statusA
Idempotent

Diagnostic in one call (replaces emet_setup_status, get_status and emet_verify_connection, 2026-09-27). Reports environment, database, collections, vector index, whether identity and user documents exist, and per-layer counts and health. Pass probe: true to also prove the write path - connect, write a probe document, read the same value back, remove it; a read alone cannot reveal a broken write path. Prefer emet_initialize for a normal session start.

ParametersJSON Schema
NameRequiredDescriptionDefault
probeNoAlso write and read back a probe document (default false). Use after a deploy or a configuration change.

TDQS

A4.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With readOnlyHint=false already declared, the description explains why: probe:true writes a document, reads it back, then removes it, so the agent understands the write and its cleanup rather than assuming a pure read. That closes the apparent gap between a 'status' tool and a non-read-only annotation. It does not cover auth requirements or failure modes.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three front-loaded sentences with no filler; the diagnostic scope comes first, then the probe semantics, then the routing hint. The parenthetical list of replaced tools is the only slightly bulky element, but it earns its place as a migration cue.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, so the description carries the burden of describing the return surface - and it does, enumerating environment, database, collections, index, identity/user documents, and counts. Combined with the probe explanation and the emet_initialize route, nothing needed to call it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so baseline would be 3, but the description adds real value beyond the schema by explaining the rationale - 'a read alone cannot reveal a broken write path' - which tells the agent why probe exists, not just what it does.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific purpose - a one-call diagnostic - and enumerates exactly what it reports (environment, database, collections, vector index, identity/user docs, per-layer counts and health). It also names the siblings it replaces and distinguishes itself from emet_initialize, so an agent can tell it apart without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly routes the agent: 'Prefer emet_initialize for a normal session start' names the alternative and its condition, and the probe parameter is scoped to 'after a deploy or a configuration change'. Both when-to-use and when-not-to-use are stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

emet_transcript_appendA

Append ONE exchange to the session transcript. Call after every reply. Idempotent on exchange_index. Rolls to a new part after 8 exchanges. After a complete close the session takes exactly ONE more append - its closing exchange, numbered next, within EMET_CLOSE_GRACE_MINUTES (default 5) of the close; it is written into the session's last part whatever session_id it names, and the result's closing_exchange says it was accepted and that the session is now locked. A retry of that exchange number is answered as already saved and writes nothing. exchange_index, when given, must be a whole number of 1 or more (a number or its text). Any other append to a closed session is refused and nothing is written; that exchange belongs in the next session you open the usual way. THE FIRST APPEND OF A SESSION IS CHECKED: its assistant_message must carry the measured machine line - a machine name from the install notes plus the clock reading, as the shell returned them - or the words no local shell MCP. Without one of those the append is refused and the transcript does not start.

ParametersJSON Schema
NameRequiredDescriptionDefault
sourceNoSource tag - which host produced this write. One name on every write since 2026-09-27.
channelNoAlias of `source` (the older name, kept one release). Prefer `source`.
session_idYes
user_messageYes
session_startNo
exchange_indexNo
assistant_messageYes

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes far beyond the annotations: part rolling after 8 exchanges, the single closing-exchange exception after close, the grace window, the lock state, refusal semantics, and the first-append machine-line gate. The only blemish is the claim "Idempotent on exchange_index" sitting against idempotentHint=false; it is defensible as scoped idempotency (a repeated index writes nothing) rather than a safety-relevant contradiction, but the two signals are not obviously reconciled for the agent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the core purpose and routing rule in the first two sentences, then layers edge cases in descending priority. It is long and occasionally dense (the closing-exchange paragraph is a single run-on chain), and "whatever session_id it names" is a minor detour, but almost every sentence carries operational weight.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex, stateful, multi-constraint write tool with no output schema, the description covers the full lifecycle: normal append, part rollover, close grace append, retry resolution, and refusal. It even describes the observable result field (`closing_exchange`) that confirms acceptance, so an agent can act and verify without further documentation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 29% across 7 parameters, so the description must compensate. It does add real meaning for exchange_index (whole number of 1 or more, number or its textual form) and for assistant_message on the first append (must carry the measured machine line or "no local shell MCP"), but source, channel, session_start, user_message and the closing-exchange assistant_message content are left to the schema or undocumented.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a precise verb and resource ("Append ONE exchange to the session transcript") and even emphasizes the unit of work with capitalised "ONE", which separates it from bulk siblings like save_transcript and revise_transcript. An agent knows exactly what this writes and at what granularity without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly says "Call after every reply", then spells out the when-not cases: appends to a closed session are refused and "that exchange belongs in the next session you open the usual way", with a named exception window (EMET_CLOSE_GRACE_MINUTES). This is textbook when/when-not guidance with the alternative route named.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_docsA
Read-onlyIdempotent

List continuity documents (metadata only; optional doc_id prefix filter, e.g. 'threads/'). Documents marked obsolete by retire_doc are RETAINED but hidden here by default - pass include_retired to see them, each flagged with its retirement date and reason.

ParametersJSON Schema
NameRequiredDescriptionDefault
prefixNo
include_retiredNoInclude documents marked obsolete. Default false.

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly/idempotent/non-destructive, so safety is covered. The description adds genuinely non-obvious behavior: retired docs are retained but hidden by default, and when shown are flagged with retirement date and reason. This is useful context beyond the annotations, though it doesn't describe pagination or ordering.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences with no filler. The core purpose and the default-vs-override behavior are front-loaded, and every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description carries the return-value burden and does note 'metadata only' plus the retirement flagging. For a simple two-parameter read tool this is nearly complete; only ordering/pagination is left unstated.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 50% (prefix has no schema description), and the description compensates by explaining prefix as a doc_id prefix filter with an example ('threads/'). include_retired's semantics are reinforced in both places, so the description meaningfully closes the coverage gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ("List continuity documents") and adds a clear scope qualifier (metadata only). It is instantly distinguishable from siblings like read_doc, write_doc, and retire_doc, which are single-document operations.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explains the key usage condition: retired documents are hidden by default and you pass include_retired to see them. It does not, however, explicitly name an alternative tool (e.g. read_doc) or state when this listing is preferable over corpus_recall/semantic_recall, so the routing guidance is good but not exhaustive.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

patch_docA
Destructive

Replace one exact substring in a document (str_replace semantics) without resending the whole file. THE PREFERRED WAY TO MAKE SMALL EDITS to large docs like STATE.md. old_str must match EXACTLY ONCE - zero or multiple matches are rejected rather than guessed. Pass new_str as an empty string to delete. Same versioning, archiving, and sha256 read-back as write_doc.

ParametersJSON Schema
NameRequiredDescriptionDefault
editsNoSeveral replacements in one revision. Each old_str must match the current text exactly once. Applied together, so a failed edit writes nothing.
doc_idYes
originNoRequired when a board edit adds a row id. Names the session, chat or drop the row came from.
reasonYesRequired, at least 10 characters. A removed or closed board row says why.
sourceNoSource tag - which host produced this write. One name on every write since 2026-09-27.
new_strNoReplacement text. Empty string deletes the matched text.
old_strNoExact text to replace. Must occur exactly once in the current text. Omit when `edits` is set.
exchangeNoThe exchange (turn) this write belongs to, counting from 1. A write for exchange n while exchange n-1 was never appended is refused by the session graph until n-1 is appended.
session_idNoTranscript session id. The floor is measured from the size when this session first touched the document.
updated_byNoAlias of `source` (the older name, kept one release). Prefer `source`.
expected_versionNoOptional optimistic-concurrency guard.

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true and non-idempotent, so the mutation profile is known. The description adds genuinely useful behavior beyond that: exact-once match enforcement with ambiguous matches rejected rather than guessed, empty new_str as delete, and parity with write_doc's versioning/archiving/sha256 read-back. It does not discuss permissions or what a failed multi-edit leaves behind (that detail lives in the schema).

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four tight sentences, front-loaded with the operation and its mechanism, then preference guidance, then the critical exact-match constraint, then delete behavior and parity guarantees. No filler and each sentence carries distinct information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For an 11-parameter mutation tool with no output schema, the description covers the core mental model (substring replacement, ambiguity rejection, delete, versioning parity) and the schema carries the remaining parameters at 91% coverage. It omits any narrative about the session/exchange graph constraints and batch-edit atomicity, which are only in schema text, but nothing essential to invoking it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 91%, so the schema already documents nearly every parameter including the edits array, origin, reason, exchange and expected_version. The description only restates old_str/new_str semantics already present in the schema, adding no syntax, format, or constraint detail beyond it. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Names a specific verb+resource with its mechanism ('Replace one exact substring ... str_replace semantics') and explicitly distinguishes itself from the full-file alternative ('without resending the whole file'). An agent can tell this apart from write_doc and revise_memory from the first sentence alone.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear selection context: 'THE PREFERRED WAY TO MAKE SMALL EDITS to large docs like STATE.md' implicitly routes full rewrites to write_doc. It stops short of naming write_doc explicitly or stating when-not to use patching (e.g. multi-region rewrites), so it is clear but not fully exclusionary.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

query_layerC
Read-onlyIdempotent

Query a specific memory layer with structured filters

ParametersJSON Schema
NameRequiredDescriptionDefault
layerYesMemory layer to query
optionsNoQuery options with structured filters

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false and openWorldHint=false, so the safety profile is covered. The description adds essentially nothing behavioral beyond 'structured filters' — no note on superseded/expired handling, layer semantics, or result shape — even though the schema's include_superseded text shows there is non-trivial behavior worth surfacing.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler or redundancy. It is efficient, though its brevity borders on under-specification for a tool with a nested options object and six-valued enum.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With annotations covering safety and a fully documented schema, nothing critical is missing to invoke the tool, but for a layered retrieval tool surrounded by recall siblings the description leaves the agent without routing or layer-semantics context. Adequate, not complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, including detailed enum, filter, ordering, and include_superseded semantics, so the schema does the heavy lifting and the baseline is 3. The description adds no extra parameter meaning beyond what the schema already documents.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (query), a specific resource (a memory layer), and the mechanism (structured filters), which is much clearer than a restated title. However, it does nothing to distinguish itself from the many sibling recall tools (recall, semantic_recall, corpus_recall), so an agent cannot tell from the description alone which one to pick for a layered/filtered query.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit when-to-use guidance, no mention of alternatives, and no conditions that select this tool over the recall-family siblings. The phrase 'structured filters' weakly implies the use case (filter-driven retrieval), but nothing tells the agent when filtering is preferred over semantic recall.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

query_transcriptsA
Read-onlyIdempotent

Retrieve session transcripts - the second tier, word for word. Reads the live store written by save_transcript AND the legacy archive (exchanges written by earlier hosts, June 2025 to June 2026), grouped into the same session shape; every result carries source naming the collection it came from. Nothing is migrated between them. Use it when a layer entry is relevant but thin: the transcript around its timestamp holds the depth the layer compressed out.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMaximum number of sessions to return (default: 50, max: 500)
session_idNoOptional: Filter by specific session ID (e.g. '2026-04-16_001')
timestamp_endYesUnix timestamp - end of time range (filters by session_start)
timestamp_startYesUnix timestamp - start of time range (filters by session_start)

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly, idempotent, non-destructive, and closed-world, so the safety profile is covered. The description adds real value beyond them: it discloses that reads span two collections (live store + legacy archive June 2025–June 2026), that nothing is migrated between them, and that each result carries a `source` field naming its origin.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose is front-loaded in the first clause, followed by provenance detail and the usage trigger. Four sentences that each carry information; only the 'second tier, word for word' phrasing is slightly cryptic, but it earns its place by contrasting with the compressed layer.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description usefully explains the return shape (sessions grouped, each carrying `source`) and the two backing stores. Annotations handle safety; parameters are fully covered by the schema. Adequate for a read-only query tool, though it could note ordering or pagination behavior.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all four parameters (limit, session_id, timestamp_start/end) are already documented in the schema. The description adds no syntax or filtering nuance beyond that, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Retrieve session transcripts') and clarifies it reads both the live store and the legacy archive. This distinguishes it from query_layer/recall by framing it as the 'second tier' depth source, though the tier jargon is not self-evident without sibling context.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives a concrete when-to-use condition ('when a layer entry is relevant but thin: the transcript around its timestamp holds the depth the layer compressed out'), implicitly positioning it against query_layer. No explicit exclusions or negative guidance, but the selection condition is clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

read_docA
Read-onlyIdempotent

Read a continuity document from the versioned document store. A document longer than one page (18,000 UTF-8 bytes by default) comes back in pages: the result then carries page with next_offset - call again with that offset (and expected_version set to the first page's version) until next_offset is null, and do not act on a document you have only part of. If the document changes between pages the read is refused: restart from offset 0. sha256 and size (characters; size_bytes in UTF-8 bytes) always describe the WHOLE document.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoOptional: page size in bytes as the host receives it (default 18,000): a page ends before its JSON-escaped text plus 1,500 bytes kept for the result envelope would pass it, and never splits a character. 0 returns the whole document in one result.
doc_idYesDocument id, e.g. 'STATE.md', 'HANDOFF.md', 'threads/pcb-laser-board.md', 'bootstrap/IDENTITY.md'
offsetNoOptional: UTF-8 byte offset to start from (default 0). Use the previous page's `page.next_offset`.
expected_sha256NoOptional: the `sha256` your first page returned - the same pin as expected_version, and the only pin for a shipped guide (guides have no version).
expected_versionNoOptional: the `version` your first page returned. Pass it with every later page; if the document changed in between, the read is refused (restart from offset 0) so pages of two versions are never joined.

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly/idempotent/non-destructive, but the description adds substantial behavior beyond them: paging rules, refusal-and-restart semantics when a document changes between pages, and the guarantee that `sha256`/`size` describe the whole document rather than the page. These are exactly the traits an agent needs to avoid joining pages of two versions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the core action, then layers paging, pinning, and invariants in dense sentences with no filler. Every sentence carries operational information an agent must act on.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, so the description must describe return values, and it does: `page` with `next_offset`, `version`, `sha256`, `size`, and `size_bytes`. Combined with the paging protocol, an agent can call and consume the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds cross-parameter meaning the schema does not: `expected_version` must be the first page's version, and `expected_sha256` is the same pin and the only pin for shipped guides. That interaction logic goes beyond per-field schema text.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Read a continuity document from the versioned document store'), which clearly separates it from write_doc, list_docs, and read_doc_history. The pagination framing further distinguishes it from sibling read tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit operational guidance: call again with `next_offset` until it is null, pin `expected_version` to the first page's version, and do not act on a partial document. It does not, however, say when to prefer this over siblings like corpus_recall, semantic_recall, or read_doc_history, so routing among read tools is left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

read_doc_historyA
Read-onlyIdempotent

Read a RETAINED revision of a document, or list the revisions retained for it. Every write archives the prior copy (nothing is deleted); this is the reader for that archive. Without a version: the list of retained revisions with their digests and dates. With a version: that revision's full content, re-hashed against its stored digest. Read-only. Use it for an audit question - what did this document say at version N - or to compare a document with what it replaced.

ParametersJSON Schema
NameRequiredDescriptionDefault
doc_idYesDocument id, e.g. 'CHARTER.md', 'STATE.md'
versionNoOptional: the revision to read. Omit to list what is retained.

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover read-only, idempotent, and non-destructive, yet the description adds substantial context: every write archives the prior copy, nothing is deleted, and a returned revision is re-hashed against its stored digest. This is genuine behavioral disclosure beyond the structured hints.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core action, then the archive rationale, then the two parameter behaviors and use cases. Information-dense, though the middle clauses are slightly long and could be trimmed without losing meaning.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description carries the return-value burden and does so: it describes both the list form (digests and dates) and the single-revision form (full content, re-hashed). Combined with the archive semantics, an agent has everything needed to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds return-shape meaning per parameter value: omitting version yields the list of revisions with digests and dates, supplying version yields that revision's full re-hashed content. That is more than the schema's terse per-field notes.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (read/list) plus resource (retained revision) and scope ('Every write archives the prior copy... this is the reader for that archive'), clearly separating it from the sibling read_doc which reads the current copy. An agent can distinguish it without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit use cases are given: 'audit question - what did this document say at version N' and 'compare a document with what it replaced.' This is clear context for when to reach for it. It does not name read_doc explicitly as the alternative for current content, so it stops short of a full when/when-not statement.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

recallB
Read-onlyIdempotent

Search and retrieve memories across layers by keyword

ParametersJSON Schema
NameRequiredDescriptionDefault
layerNoOptional: Search only in specific layer
limitNoMaximum number of results to return (default: 10)
queryYesSearch query to match against memory content

TDQS

B3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare the full safety profile (readOnlyHint=true, idempotentHint=true, destructiveHint=false, openWorldHint=false), so the description is not the primary carrier of behavior here. It adds some value by disclosing the retrieval mechanism (keyword matching) and cross-layer scope, but says nothing about result volume, ordering, or pagination behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single short sentence with no filler, and the core verb and scope are front-loaded. Nothing in it is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple three-parameter read tool with no output schema and safety fully covered by annotations, the description is minimally sufficient. It falls short on the one thing that actually matters in this crowded sibling set: telling the agent when to pick recall over query_layer, semantic_recall, or corpus_recall.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all three parameters (query, layer enum, limit) are already documented, and the baseline is 3. The description adds only marginal meaning beyond the schema – 'by keyword' hints at matching semantics and 'across layers' corresponds to the layer filter, but no format or syntax detail is added.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a clear verb+resource ('search and retrieve memories') plus the mechanism ('by keyword', 'across layers'), so the basic purpose is legible. However, the sibling set contains several closely overlapping retrieval tools (corpus_recall, query_layer, semantic_recall), and the description offers no differentiation between them, leaving an agent unable to tell which retriever applies.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance and no mention of alternatives, despite obvious candidates such as query_layer and semantic_recall. The only implicit signal is that the tool is keyword-based, which the agent must infer as a routing criterion on its own.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

rename_docA

Reissue a continuity document under a new doc_id. The destination is written first through the normal write path (same redaction, digest, versioning and sha256 read-back), and only then is the old id marked obsolete with retire_doc. NOTHING IS DELETED: the old id keeps its content and stays readable at its own id, flagged retired with the reason 'Reissued as ', and hidden from list_docs. No tombstone document is needed, because the retired original IS the tombstone and it still carries the real content. Refuses if the destination already exists, because a rename onto an occupied id is a merge decision and merges are not guessed.

ParametersJSON Schema
NameRequiredDescriptionDefault
new_idYesDestination document id. Must not already exist.
old_idYesExisting document id to reissue from. It is retired, not removed.
reasonYesWhy the document is reissued. Required, at least 10 characters.
sourceNoSource tag - which host produced this write. One name on every write since 2026-09-27.
exchangeNoThe exchange (turn) this write belongs to, counting from 1. A write for exchange n while exchange n-1 was never appended is refused by the session graph until n-1 is appended.
session_idNoTranscript session id.
updated_byNoAlias of `source` (the older name, kept one release). Prefer `source`.

TDQS

A4.3/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Far exceeds the annotations: it specifies the write ordering (destination written first, then retire_doc), that NOTHING IS DELETED, that the old id stays readable and is flagged retired with a reason, that it is hidden from list_docs, and the merge-refusal rule. The destructiveHint=false annotation is corroborated and richly explained rather than merely restated.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the core action and ordering before elaborating on the retire semantics and refusal rule; every sentence carries information. It runs slightly long, and the tombstone sentence is somewhat redundant with the preceding NOTHING IS DELETED clause.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description covers the mutation's effects, ordering, and refusal behavior thoroughly for a 7-parameter write tool. It leaves the source/exchange/session_id write-path parameters to the schema, which is reasonable given their full descriptions there.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the schema already documents all seven parameters. The description reinforces new_id/old_id semantics (destination must not exist; old id is retired, not removed) but adds no syntax or format detail beyond the schema, which is the expected baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a precise verb+resource ('Reissue a continuity document under a new doc_id') and immediately frames it against the related retire_doc operation. An agent can distinguish it from write_doc, retire_doc, and patch_doc without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explains the operational context clearly (reissue an existing doc under a new id) and names an explicit refusal condition (destination already exists, because that would be a merge decision). It does not spell out when to prefer this over, say, patch_doc or write_doc, but the context is unambiguous.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

restore_docA
Destructive

Write a retained revision back as the current document. The retained row is copied through the normal write path, with a reason. Nothing is deleted.

ParametersJSON Schema
NameRequiredDescriptionDefault
doc_idYesDocument id to restore.
reasonYesWhy this revision is restored. At least 10 characters.
sourceYesSource tag. Must be in EMET_RESTORE_TAGS. A missing source is refused.
versionYesRetained revision number to write back.
exchangeNoThe exchange (turn) this write belongs to, counting from 1. A write for exchange n while exchange n-1 was never appended is refused by the session graph until n-1 is appended.
session_idNoTranscript session id. Required on Path 2.

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true and idempotentHint=false, so the safety profile is known. The description adds real context beyond that: the retained row goes through the normal write path and no rows are deleted, which usefully qualifies what 'destructive' means here. It still doesn't state permission requirements or failure behavior beyond what the schema carries.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, front-loaded with the action, and the clarifying 'Nothing is deleted' earns its place by tempering the destructive hint. No filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation tool with annotations and a fully documented six-parameter schema, the description covers the essential mechanism. It lacks any note on permissions, partial-failure behavior, or what the caller gets back, but with no output schema and rich annotations those gaps are minor.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so every parameter (including the source-tag allowlist and exchange ordering rule) is already documented in the schema. The description alludes to 'a reason' and 'retained revision' but adds no format or constraint detail beyond the schema. Baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: writing a retained revision back as the current document. An agent understands the operation immediately, but the description never distinguishes it from siblings like patch_doc, write_doc, or read_doc_history.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied by the mechanics described (restoring a retained revision), but there is no explicit when-to-use guidance, no exclusions, and no mention of the sibling tools that also write document content. The agent must infer that this is the revert path rather than a normal write.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

retire_docA
Idempotent

Mark a continuity document OBSOLETE. Nothing is deleted: the document keeps its id, its full content, its version and its digest, and read_doc still returns it - flagged retired, with the reason - so it can never be silently lost. What changes is visibility: list_docs stops returning it unless include_retired is true, so retired documents leave the working view and the startup scans. This is document control as ISO 7.5.3 describes it - issue the revision, mark the prior copy obsolete, retain it, prevent its unintended use - and it is the document-store counterpart of revise_memory superseding a layer entry. THE STORE HAS NO DELETE, BY DESIGN. Use for genuine litter: tombstones, merged staging documents, test scratch. Reversible: write_doc on the same id brings it back into the working view.

ParametersJSON Schema
NameRequiredDescriptionDefault
doc_idYesDocument id to mark obsolete.
reasonYesWhy it is obsolete. Required, at least 10 characters. Drops only.
sourceNoSource tag - which host produced this write. One name on every write since 2026-09-27.
exchangeNoThe exchange (turn) this write belongs to, counting from 1. A write for exchange n while exchange n-1 was never appended is refused by the session graph until n-1 is appended.
retired_byNoAlias of `source` (the older name, kept one release). Prefer `source`.
session_idNoTranscript session id.

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=false and idempotentHint=true, but the description goes far beyond them: it states nothing is deleted, the id/content/version/digest are retained, read_doc still returns it flagged retired, and list_docs hides it unless include_retired is true. It also discloses reversibility via write_doc and that the store has no delete by design.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core action, and each sentence carries information (retention, visibility, reversibility). The ISO 7.5.3 exposition and the revise_memory analogy are slightly discursive but still earn their place by anchoring intent.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a non-destructive mutation with annotations covering the safety profile, the description supplies everything an agent needs: what changes, what is preserved, effect on list_docs and startup scans, and how to undo it. No output schema is needed since the visibility behavior is fully explained.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so every parameter including doc_id, reason, source, exchange and retired_by is already documented in the schema. The description adds no format or syntax detail about parameters, so the baseline 3 for a fully-covered schema applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a precise verb+resource ("Mark a continuity document OBSOLETE") and immediately scopes what that means versus deletion. It distinguishes itself from siblings by naming read_doc, list_docs, write_doc and revise_memory and describing how each relates to a retired document.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

"Use for genuine litter: tombstones, merged staging documents, test scratch" gives clear positive guidance, and the revise_memory comparison situates it among alternatives. It stops short of explicit when-not-to-use exclusions beyond the implicit "not for live documents," so it falls just shy of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

revise_memoryA

Supersede an existing entry with a revised one (supersession, not mutation). Creates a NEW entry carrying metadata.supersedes= and marks the target entry metadata.superseded_by=. Superseded entries STILL SURFACE in recall / query_layer / semantic_recall, marked with superseded_by and is_superseded, so the revision is read alongside what it replaced (query_layer options.include_superseded=false gives the live-only view). Episodic revision is prohibited - events are immutable.

ParametersJSON Schema
NameRequiredDescriptionDefault
layerYesLayer of the entry being revised (episodic is prohibited)
contentYesFull revised content - must stand alone without the superseded entry
exchangeNoThe exchange (turn) this write belongs to, counting from 1. A write for exchange n while exchange n-1 was never appended is refused by the session graph until n-1 is appended.
metadataNoMetadata for the new entry (supersedes is set automatically; include source tag)
target_idYesID of the entry this revision supersedes
session_idNoYour transcript session id from emet_session_open. The session graph reads this session's state before the write and returns `session_graph.next` - the next step. A governed write with no session is refused (open one with emet_session_open and load_id).

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotation contradiction: destructiveHint=false is actively supported by the disclosure that superseded entries still surface in recall/query_layer/semantic_recall marked with superseded_by and is_superseded, while idempotentHint=false aligns with each revision minting a new id. The description goes well beyond the annotations by disclosing the exact write mechanics (new entry carrying metadata.supersedes, target marked metadata.superseded_by), the read-visibility consequence, and the immutability constraint on episodic.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core semantics in the first sentence, then layers in write mechanics, read visibility, and the prohibition in tight, non-redundant sentences. Every sentence carries distinct information an agent needs.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a six-parameter, non-idempotent write with no output schema, the description is complete: it explains the effective outcome (a new id plus a superseded target), how superseded content remains visible to readers, and the immutability carve-out. Nothing an agent needs to invoke this safely is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and the schema already documents layer, content, exchange, metadata, target_id and session_id in detail (including the session-graph refusal behavior). The description adds the constraint 'episodic is prohibited' echoed from the schema, but contributes no additional syntax or format meaning to the parameters, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb+resource and immediately disambiguates the operation: 'Supersede an existing entry with a revised one (supersession, not mutation).' It distinguishes this from mutation-style updates and from creation by naming the mechanics (new entry + metadata.supersedes), so an agent can tell it apart from save_to_layer or patch-style siblings without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It clearly states the condition for use (revising an existing entry rather than creating one) and an explicit exclusion ('Episodic revision is prohibited - events are immutable'), plus the read-side alternative (query_layer options.include_superseded=false for the live-only view). It stops short of explicitly naming the sibling to use for a fresh write (e.g., save_to_layer), so this is clear context rather than full when/when-not/alternatives routing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

revise_transcriptA

Supersede a stored transcript part with a corrected one. Nothing is edited or deleted: a NEW document is written under the same session_id with an incremented revision and a pointer to its parent, and the parent is marked superseded. Use when a saved part has a mispaired, missing or wrong exchange. The reason parameter is required - an unexplained correction is a silent rewrite. To ADD a new part, use save_transcript instead. Allowed on a closed session (a correction deletes nothing), but there a revision may only correct the exchanges the part already holds - never add new ones - and must carry exactly the same exchange numbers (none dropped or repeated; numbers sent as text are stored as numbers).

ParametersJSON Schema
NameRequiredDescriptionDefault
reasonYesWhat was wrong and how this differs. Required, minimum 10 characters. This is the audit record of the correction.
sourceNoSource tag - which host produced this write. Defaults to the parent's.
channelNoAlias of `source` (the older name, kept one release). Prefer `source`.
exchangeNoThe exchange (turn) this write belongs to, counting from 1. A write for exchange n while exchange n-1 was never appended is refused by the session graph until n-1 is appended.
exchangesYesThe COMPLETE corrected exchange list for this part, verbatim. A revision replaces the whole part.
session_idYesThe session_id of the part to correct, e.g. '2026-09-04_x_verbatim_part14'. Must already exist.
session_endNoOptional. Defaults to the parent's.
session_startNoOptional. Defaults to the parent's.

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare non-read-only, non-destructive, non-idempotent, and the description explains why: nothing is edited or deleted, a new document is written and the parent marked superseded. It adds behavioral facts beyond the annotations, including the mandatory reason ('an unexplained correction is a silent rewrite') and the closed-session exchange-number rules. No contradiction with the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Dense but well front-loaded: the mechanism comes first, then when-to-use, then the closed-session caveat. Each sentence carries distinct information, though the final clause is long and could be split for readability.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For an 8-parameter mutation tool with no output schema, the description covers replacement semantics, the required audit reason, sibling routing, and edge-case behavior on closed sessions. An agent has everything needed to invoke it correctly; return values are not its burden.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3; the description still adds meaning by requiring the reason and explaining the closed-session constraint that exchanges must match exactly and text numbers are stored as numbers. It does not add syntax detail for source/channel/session_start/session_end beyond the schema, keeping it at 4 rather than 5.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States the specific verb and resource ('Supersede a stored transcript part with a corrected one') and immediately clarifies the mechanism (new document, incremented revision, parent superseded). It distinguishes itself from the sibling save_transcript by contrast. An agent can identify the operation without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly names the trigger ('when a saved part has a mispaired, missing or wrong exchange') and the alternative ('To ADD a new part, use save_transcript instead'). It further qualifies usage on a closed session with concrete constraints, leaving nothing to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

save_to_layerA

Save memory to a specific layer with full control over metadata. PASS session_id - your transcript session from emet_session_open - on every save made inside a session: the server stamps derived_from with the exchange in progress. Semantic and procedural entries are REFUSED without a source (session_id, or metadata.derived_from); nothing is written when refused.

ParametersJSON Schema
NameRequiredDescriptionDefault
layerYesTarget memory layer
contentYesMemory content to save
exchangeNoThe exchange (turn) this write belongs to, counting from 1. A write for exchange n while exchange n-1 was never appended is refused by the session graph until n-1 is appended.
metadataNoFull metadata control
session_idNoYour transcript session id (from emet_session_open). The server reads that session's counter and stamps derived_from {session_id, exchange_start, exchange_end} with the exchange in progress. Ignored when metadata.derived_from is given. Required in practice for semantic and procedural saves made in a session.

TDQS

A4.2/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations cover the safety profile (readOnly=false, destructive=false, idempotent=false), and the description adds substantial behavior beyond that: the server stamps derived_from with the in-progress exchange, semantic/procedural entries are REFUSED without a source, and nothing is written when refused. That refusal atomicity is exactly the kind of consequence an agent must know before calling.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The purpose is front-loaded in the first clause, followed by the critical session_id directive and the refusal consequence. Three dense sentences, all earning their place, though the emphasis-shift to the refusal rule makes it slightly less cleanly ordered.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation tool with no output schema but full annotation coverage, the description supplies the key behavioral facts an agent needs: source requirement, refusal semantics, and derived_from stamping. Remaining gaps (return shape, interaction with revise_memory) are minor given the annotations and detailed schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3; the description earns a bump by tying session_id to a concrete outcome (derived_from is stamped) and to the refusal rule, reinforcing the interaction between session_id and metadata.derived_from. It does not add format detail beyond the already-rich schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('save memory to a specific layer') with an explicit scope ('full control over metadata'), which is clearly distinct from siblings like save_transcript or revise_memory. It does not name an alternative directly, so it lands at a strong 4 rather than a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives an explicit operational rule: pass session_id on every save made inside a session, and semantic/procedural entries are refused without a source. This is clear usage context, but it never contrasts with sibling tools or states when to prefer this over revise_memory, so it stops short of 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

save_transcriptA

FALLBACK. The primary path is emet_session_open then emet_transcript_append after every reply; use this only where a host cannot append per reply, or to load a Claude Code CLI session from its JSONL file. Saves the full word-for-word dialog of a session - full transcripts only, no summaries. Pass exchanges (chat hosts) or jsonl_path (CLI). Refused for a session that was closed complete.

ParametersJSON Schema
NameRequiredDescriptionDefault
sourceNoSource tag - which host produced this write. One name on every write since 2026-09-27.
channelNoAlias of `source` (the older name, kept one release). Prefer `source`.
exchangeNoThe exchange (turn) this write belongs to, counting from 1. A write for exchange n while exchange n-1 was never appended is refused by the session graph until n-1 is appended.
exchangesNoClaude.ai Project mode: full array of exchange objects for this session
jsonl_pathNoCLI mode: absolute path to Claude Code session JSONL file. Tool reads and parses it automatically.
session_idYesUnique session identifier, e.g. 'claude_project_2026-05-23_001'
session_endNoUnix timestamp — session end (optional, defaults to now)
session_startYesUnix timestamp — session start

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, destructiveHint=false, idempotentHint=false and closed-world, so the safety profile is covered. The description adds real behavioral context beyond that: the tool is a non-preferred fallback, it accepts two distinct input modes, and it is refused for a session closed as complete — a failure mode the annotations do not convey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four dense sentences with the fallback status front-loaded and no filler; every clause carries routing, mode or refusal information. It is packed tightly rather than bloated, though the compression makes it slightly telegraphic.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Covers the key call decisions for an 8-parameter, 2-required, non-idempotent write: when to use it, which parameter family matches which host, and a refusal condition. It does not describe what a successful write returns, but no output schema exists and the primary selection/invocation needs are met.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description exceeds it by mapping parameters to operating modes ('Pass exchanges (chat hosts) or jsonl_path (CLI)'), which tells the agent which grouping of parameters applies to its host before it reads the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a concrete verb+resource ('Saves the full word-for-word dialog of a session') and immediately bounds the scope ('full transcripts only, no summaries'). It also distinguishes itself from the sibling emet_transcript_append by naming it as the primary path this tool falls back from.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Front-loads 'FALLBACK' and spells out the preferred alternative (emet_session_open + emet_transcript_append after every reply) and the exact conditions that select this tool instead (host cannot append per reply, or loading a CLI JSONL session). It also names a when-not case: refused for a session already closed complete.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

semantic_recallA
Read-onlyIdempotent

Semantic (meaning-based) memory search across all six memory layers via Atlas Vector Search autoEmbed. Finds conceptually related memories even without keyword overlap. Complements keyword-based recall/query_layer.

ParametersJSON Schema
NameRequiredDescriptionDefault
layerNoOptional: restrict to one layer
limitNoMax results (default 10)
queryYesNatural-language query - embedded automatically server-side (voyage-4)

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false, and openWorldHint=false, so safety is covered. The description adds real value beyond them by disclosing the retrieval mechanism (vector search, server-side autoEmbed) and scope (all six layers), which explains result behavior. It does not describe ranking or result shape, but annotations and the limit param carry much of that burden.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences, front-loaded with the core purpose and mechanism, with no filler. The final sentence routing to keyword-based siblings is efficiently placed at the end, though the phrasing could be marginally tighter.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only search tool with full annotation coverage, a fully documented schema, and no output schema, the description supplies enough: what it does, how it differs from keyword search, and the embedding mechanism. Return format and pagination with limit are not covered, but that is a minor gap given the structured data available.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so baseline is 3. The description reinforces that search spans all six layers, implicitly mapping to the layer enum's default, but adds no syntax or format detail beyond the schema's own per-parameter descriptions (including the voyage-4 embedding note).

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (semantic search) and resource (memories across six layers), and names the mechanism (Atlas Vector Search autoEmbed) plus the distinguishing property (finds conceptually related memories without keyword overlap). It explicitly contrasts itself with the keyword-based siblings recall/query_layer, so an agent can differentiate without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description names the alternatives (keyword-based recall/query_layer) and states the condition that selects this tool (when there is no keyword overlap but conceptual relation is expected). It stops short of explicit when-not/exclusion phrasing, but the routing signal is clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

validate_docA
Read-onlyIdempotent

Check a controlled document against the template that governs it, WITHOUT writing anything. The same check the write path runs on every document write - this is the freestanding form, for auditing records written before validation existed, for counting what adoption will mark before it marks anything, and for checking new content before it is stored. Pass doc_id to check a stored document (the previous revision is pulled from the archive so the carry-forward check runs). Pass content, with template_id or a doc_id the map covers, to check text that is not stored. Checks FORM and CONTINUITY only: sections present, in order and filled; no placeholder left unfilled; NOT COMPLETED carrying a reason; a board's counts matching its rows; every row on the previous revision still present or dispositioned. It cannot check whether an answer is true, or whether a 'none' states its scope honestly - those stay with the author and the reader.

ParametersJSON Schema
NameRequiredDescriptionDefault
doc_idNoStored document to check. Alone, it is read from the store and compared against its own previous revision.
contentNoText to check instead of a stored document. Requires template_id, or a doc_id the map covers.
template_idNoTemplate to check against, e.g. 'templates/BOARD.md'. Read from the copy shipped with this release, never the store copy.

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly/idempotent/non-destructive, yet the description adds substantial behavior: no writes occur, the previous revision is pulled from the archive so the carry-forward check runs, and the check scope is bounded to FORM and CONTINUITY. It additionally discloses what the tool cannot do (verify truth or honest scope of a 'none'), which is exactly the kind of limit an agent needs.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the purpose and the no-write guarantee, then use cases, then parameter behavior, then scope limits. Dense and mostly waste-free, though the middle sentence packs three use cases and the boundary paragraph is long relative to the rest.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 3-optional-param, no-output-schema tool, the description is nearly sufficient: it explains inputs, the archive comparison, and the check's boundaries. It does not describe the shape of the result (e.g., that a report/count of findings is returned), leaving the response format to inference.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds real meaning: doc_id pulls the prior revision from the archive for carry-forward comparison, content requires template_id or a covered doc_id, and template_id is read from the shipped copy rather than the store copy. These interaction rules go beyond the field-level schema text.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Specific verb ('Check') plus resource ('controlled document against the template that governs it') and an explicit scoping qualifier ('WITHOUT writing anything'). It distinguishes itself from the write path it mirrors and from doc-mutating siblings, so an agent can identify it without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Enumerates three concrete use cases (auditing pre-validation records, counting adoption impact, pre-write checks of unstored content) and notes it is the freestanding counterpart of the write-path check, which implicitly routes agents away from write_doc/patch_doc. It lacks an explicit 'do not use this for X' exclusion or a named alternative tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

write_docA
Destructive

Write or update a continuity document in the document store (versioned, prior copy archived to documents_history and marked obsolete, sha256 read-back verified). RECORDS - the handoff (HANDOFF.md) and any id in EMET_RECORD_DOCS - are superseded, never amended: patch_doc is refused on them, and a new revision is written here WHOLE with supersedes set to the version currently in force (ISO 9001 7.5.3 / 13485 4.2.5).

ParametersJSON Schema
NameRequiredDescriptionDefault
doc_idYes
reasonNoRequired on a record revision (at least 10 characters). Stored as edit_reason.
sourceNoSource tag - which host produced this write. One name on every write since 2026-09-27.
contentYes
exchangeNoThe exchange (turn) this write belongs to, counting from 1. A write for exchange n while exchange n-1 was never appended is refused by the session graph until n-1 is appended.
session_idNoTranscript session id. The floor is measured from the size when this session first touched the document.
supersedesNoRequired when writing a new revision of a RECORD (the handoff): the version currently in force, which this revision supersedes. Ignored for ordinary documents.
updated_byNoAlias of `source` (the older name, kept one release). Prefer `source`.

TDQS

A4.3/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare destructiveHint=true and idempotentHint=false; the description adds substantial behavioral detail beyond them - the prior copy is archived to documents_history and marked obsolete, the write is sha256 read-back verified, and record writes must carry supersedes pointing at the version in force. That is exactly the kind of consequence information an agent needs before a mutation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose is front-loaded in the first clause, followed by behavioral guarantees and the record rule. The parentheticals (documents_history, sha256, ISO references) are dense but each carries real information; the single long sentence is slightly heavy but not padded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation tool with annotations present and no output schema, the description covers the versioning, archival and record-supersession workflow well. Remaining gaps (reason minimum length, session_id floor semantics, exchange ordering) are left to the schema, which is acceptable given 75% coverage.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 75% and the schema already documents reason, source, exchange, session_id and supersedes in detail. The description restates the supersedes requirement rather than adding format or syntax beyond the schema, so with the schema doing the heavy lifting this sits at the baseline 3.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (write or update) and resource (continuity document in the document store), and immediately distinguishes itself from patch_doc by naming the record case where patching is refused. An agent can tell it apart from read_doc, retire_doc, rename_doc and patch_doc without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives explicit routing for RECORDS: those are superseded, never amended, and patch_doc is refused on them, so write_doc is the required path. However, it never states the inverse condition for ordinary documents (i.e. when a small edit should go through patch_doc instead), so the alternative-usage guidance is one-sided rather than complete.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 14 tool updatesv0.2.2
    • Changedaccept_nonconformance2 fields changed
      • changedInput schema / properties / exchange / description
        Previous value: -"The exchange (turn) this write belongs to, counting from 1. A write for exchange n while exchange n-1 was never appended is flagged by the session graph."New value: +"The exchange (turn) this write belongs to, counting from 1. A write for exchange n while exchange n-1 was never appended is refused by the session graph until n-1 is appended."
      • changedInput schema / properties / session_id / description
        Previous value: -"Your transcript session id from emet_session_open. The session graph reads this session's state before the write and returns `session_graph.next` - the next step."New value: +"Transcript session id."
    • Changedemet_initialize2 fields changed
      • addedInput schema / properties / load_id
        Added value: +{
        +  "description": "The load_id your startup returned in `startup_page.load_id`: every emet_initialize result carries one (each page of a paged startup, the whole startup with page \"all\" or when it fits one result, and a store that is not set up yet). Pass it with every later page so all pages come from the same startup, and pass the same load_id to emet_session_open so its startup page check runs on this startup.",
        +  "type": "string"
        +}
      • addedInput schema / properties / page
        Added value: +{
        +  "description": "The startup is served in pages when it is larger than one result (each page fits under 20,000 bytes). Omit for page 1; every page but the last says which page to ask for next - call again with that page, before replying, until a page says it is the last. \"all\" returns the whole startup in one result, for a host with no size limit."
        +}
    • Changedemet_session_close2 fields changed
      • changedInput schema / properties / exchange / description
        Previous value: -"The exchange (turn) this write belongs to, counting from 1. A write for exchange n while exchange n-1 was never appended is flagged by the session graph."New value: +"The exchange (turn) this write belongs to, counting from 1. A write for exchange n while exchange n-1 was never appended is refused by the session graph until n-1 is appended."
      • addedInput schema / properties / source
        Added value: +{
        +  "description": "Your host's source tag: the registered tag this session was written under. The close is never refused for it; an unregistered tag, one that did not write this session, or none (with no server default) gets a warning that the close is not attributed.",
        +  "type": "string"
        +}
    • Changedemet_session_open1 field changed
      • addedInput schema / properties / load_id
        Added value: +{
        +  "description": "Required for the startup page check (by policy; the schema leaves it optional so older hosts can still open): the load_id from `startup_page` on your emet_initialize pages - page 1 returns it, every page carries it, and the last page names it. An open without load_id, or with a page of that load known not fetched, is refused before anything is written (gate startup_pages); the refusal names the exact calls that pass.",
        +  "type": "string"
        +}
    • Changedpatch_doc7 fields changed
      • addedInput schema / properties / edits
        Added value: +{
        +  "description": "Several replacements in one revision. Each old_str must match the current text exactly once. Applied together, so a failed edit writes nothing.",
        +  "items": {
        +    "type": "object"
        +  },
        +  "type": "array"
        +}
      • changedInput schema / properties / exchange / description
        Previous value: -"The exchange (turn) this write belongs to, counting from 1. A write for exchange n while exchange n-1 was never appended is flagged by the session graph."New value: +"The exchange (turn) this write belongs to, counting from 1. A write for exchange n while exchange n-1 was never appended is refused by the session graph until n-1 is appended."
      • changedInput schema / properties / old_str / description
        Previous value: -"Exact text to replace. Must occur exactly once - include surrounding context to disambiguate."New value: +"Exact text to replace. Must occur exactly once in the current text. Omit when `edits` is set."
      • addedInput schema / properties / origin
        Added value: +{
        +  "description": "Required when a board edit adds a row id. Names the session, chat or drop the row came from.",
        +  "type": "string"
        +}
      • addedInput schema / properties / reason
        Added value: +{
        +  "description": "Required, at least 10 characters. A removed or closed board row says why.",
        +  "type": "string"
        +}
      • changedInput schema / properties / session_id / description
        Previous value: -"Your transcript session id from emet_session_open. The session graph reads this session's state before the write and returns `session_graph.next` - the next step."New value: +"Transcript session id. The floor is measured from the size when this session first touched the document."
      • changedInput schema / required
        Previous value: -[
        -  "doc_id",
        -  "old_str",
        -  "new_str"
        -]New value: +[
        +  "doc_id",
        +  "reason"
        +]
    • Changedread_doc4 fields changed
      • addedInput schema / properties / expected_sha256
        Added value: +{
        +  "description": "Optional: the `sha256` your first page returned - the same pin as expected_version, and the only pin for a shipped guide (guides have no version).",
        +  "type": "string"
        +}
      • addedInput schema / properties / expected_version
        Added value: +{
        +  "description": "Optional: the `version` your first page returned. Pass it with every later page; if the document changed in between, the read is refused (restart from offset 0) so pages of two versions are never joined.",
        +  "type": "number"
        +}
      • changedInput schema / properties / limit / description
        Previous value: -"Optional: page size in characters (default 40,000). 0 returns the whole document in one result."New value: +"Optional: page size in bytes as the host receives it (default 18,000): a page ends before its JSON-escaped text plus 1,500 bytes kept for the result envelope would pass it, and never splits a character. 0 returns the whole document in one result."
      • changedInput schema / properties / offset / description
        Previous value: -"Optional: character offset to start from (default 0). Use the previous page's `page.next_offset`."New value: +"Optional: UTF-8 byte offset to start from (default 0). Use the previous page's `page.next_offset`."
    • Changedrename_doc4 fields changed
      • changedInput schema / properties / exchange / description
        Previous value: -"The exchange (turn) this write belongs to, counting from 1. A write for exchange n while exchange n-1 was never appended is flagged by the session graph."New value: +"The exchange (turn) this write belongs to, counting from 1. A write for exchange n while exchange n-1 was never appended is refused by the session graph until n-1 is appended."
      • addedInput schema / properties / reason
        Added value: +{
        +  "description": "Why the document is reissued. Required, at least 10 characters.",
        +  "type": "string"
        +}
      • changedInput schema / properties / session_id / description
        Previous value: -"Your transcript session id from emet_session_open. The session graph reads this session's state before the write and returns `session_graph.next` - the next step."New value: +"Transcript session id."
      • changedInput schema / required
        Previous value: -[
        -  "old_id",
        -  "new_id"
        -]New value: +[
        +  "old_id",
        +  "new_id",
        +  "reason"
        +]
    • Addedrestore_doc
    • Changedretire_doc4 fields changed
      • changedInput schema / properties / exchange / description
        Previous value: -"The exchange (turn) this write belongs to, counting from 1. A write for exchange n while exchange n-1 was never appended is flagged by the session graph."New value: +"The exchange (turn) this write belongs to, counting from 1. A write for exchange n while exchange n-1 was never appended is refused by the session graph until n-1 is appended."
      • changedInput schema / properties / reason / description
        Previous value: -"Why it is obsolete, e.g. 'merged into threads/income-transition.md'. Recorded on the document and shown by read_doc."New value: +"Why it is obsolete. Required, at least 10 characters. Drops only."
      • changedInput schema / properties / session_id / description
        Previous value: -"Your transcript session id from emet_session_open. The session graph reads this session's state before the write and returns `session_graph.next` - the next step."New value: +"Transcript session id."
      • changedInput schema / required
        Previous value: -[
        -  "doc_id"
        -]New value: +[
        +  "doc_id",
        +  "reason"
        +]
    • Changedrevise_memory2 fields changed
      • changedInput schema / properties / exchange / description
        Previous value: -"The exchange (turn) this write belongs to, counting from 1. A write for exchange n while exchange n-1 was never appended is flagged by the session graph."New value: +"The exchange (turn) this write belongs to, counting from 1. A write for exchange n while exchange n-1 was never appended is refused by the session graph until n-1 is appended."
      • changedInput schema / properties / session_id / description
        Previous value: -"Your transcript session id from emet_session_open. The session graph reads this session's state before the write and returns `session_graph.next` - the next step."New value: +"Your transcript session id from emet_session_open. The session graph reads this session's state before the write and returns `session_graph.next` - the next step. A governed write with no session is refused (open one with emet_session_open and load_id)."
    • Changedrevise_transcript1 field changed
      • changedInput schema / properties / exchange / description
        Previous value: -"The exchange (turn) this write belongs to, counting from 1. A write for exchange n while exchange n-1 was never appended is flagged by the session graph."New value: +"The exchange (turn) this write belongs to, counting from 1. A write for exchange n while exchange n-1 was never appended is refused by the session graph until n-1 is appended."
    • Changedsave_to_layer1 field changed
      • changedInput schema / properties / exchange / description
        Previous value: -"The exchange (turn) this write belongs to, counting from 1. A write for exchange n while exchange n-1 was never appended is flagged by the session graph."New value: +"The exchange (turn) this write belongs to, counting from 1. A write for exchange n while exchange n-1 was never appended is refused by the session graph until n-1 is appended."
    • Changedsave_transcript1 field changed
      • changedInput schema / properties / exchange / description
        Previous value: -"The exchange (turn) this write belongs to, counting from 1. A write for exchange n while exchange n-1 was never appended is flagged by the session graph."New value: +"The exchange (turn) this write belongs to, counting from 1. A write for exchange n while exchange n-1 was never appended is refused by the session graph until n-1 is appended."
    • Changedwrite_doc3 fields changed
      • changedInput schema / properties / exchange / description
        Previous value: -"The exchange (turn) this write belongs to, counting from 1. A write for exchange n while exchange n-1 was never appended is flagged by the session graph."New value: +"The exchange (turn) this write belongs to, counting from 1. A write for exchange n while exchange n-1 was never appended is refused by the session graph until n-1 is appended."
      • addedInput schema / properties / reason
        Added value: +{
        +  "description": "Required on a record revision (at least 10 characters). Stored as edit_reason.",
        +  "type": "string"
        +}
      • changedInput schema / properties / session_id / description
        Previous value: -"Your transcript session id from emet_session_open. The session graph reads this session's state before the write and returns `session_graph.next` - the next step."New value: +"Transcript session id. The floor is measured from the size when this session first touched the document."
  2. 28 tool updatesv0.1.0
    • First observedaccept_nonconformance
    • First observedcorpus_recall
    • First observedemet_floor
    • First observedemet_gaps
    • First observedemet_initialize
    • First observedemet_invite_create
    • First observedemet_member_access_create
    • First observedemet_session_close
    • First observedemet_session_open
    • First observedemet_setup_complete
    • First observedemet_status
    • First observedemet_transcript_append
    • First observedlist_docs
    • First observedpatch_doc
    • First observedquery_layer
    • First observedquery_transcripts
    • First observedread_doc
    • First observedread_doc_history
    • First observedrecall
    • First observedrename_doc
    • First observedretire_doc
    • First observedrevise_memory
    • First observedrevise_transcript
    • First observedsave_to_layer
    • First observedsave_transcript
    • First observedsemantic_recall
    • First observedvalidate_doc
    • First observedwrite_doc

TDQS

A3.6/5.0

Scored across 29 tools

Disambiguation4/5

Tools target distinct resources and actions, and descriptions explicitly differentiate overlapping clusters such as recall vs semantic_recall vs query_layer, and emet_transcript_append vs save_transcript. However, multiple retrieval and transcript tools still share similar purposes, so an agent must read carefully to avoid misselection.

Naming Consistency3/5

All names use snake_case, but the set mixes prefixed tools (emet_status, emet_session_open) with bare tools (query_layer, write_doc) and mixes verb_noun with noun_verb patterns (corpus_recall, semantic_recall). It remains readable, but the convention is not consistent.

Tool Count2/5

At 29 tools, the surface is heavy and exceeds the typical 3–15 range. While the domain is broad, several specialized tools (emet_floor, emet_status, emet_setup_complete, access creation) could likely be consolidated or parameterized.

Completeness4/5

The set covers memory layers, transcripts, versioned documents, setup, diagnostics, and access creation comprehensively. Minor gaps remain: no access revocation/list, no explicit memory deletion or redaction tool, and no corpus population tool.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    A
    maintenance
    Enables agents to read and write a durable, append-only memory with verified provenance, supporting search, record capture, and context preparation across sessions and machines.
    MIT
  • A
    license
    Not graded
    quality
    C
    maintenance
    Enables autonomous multi-agent workflows to capture, preserve, and retrieve immutable transcripts and atomic memory cards through a local MCP server, providing tools for forensic search, topic mapping, and structured fact access without losing context fidelity.
    1
    MIT
  • A
    license
    Not graded
    quality
    C
    maintenance
    Enables AI agents to store and recall an append-only, hash-chained memory with citation-backed semantic search, fact verification, contradiction detection, and replay/audit capabilities through local MCP tools.
    1
    AGPL 3.0